i
i
“K27443˙FM” — 2018/7/17 — 16:16 — page 2 — #2
i
i
Fundamentals of Probability
With Stochastic Processes
Fourth Edition
i
i
i
i
i
i
“K27443˙FM” — 2018/7/17 — 16:16 — page 3 — #3
i
i
i
i
i
i
i
i
“K27443˙FM” — 2018/7/17 — 16:16 — page 4 — #4
i
i
Fundamentals of Probability
With Stochastic Processes
Fourth Edition
Saeed Ghahramani
Western New England University
Springfield, Massachusetts, USA
i
i
i
i
i
i
“K27443˙FM” — 2018/7/17 — 16:16 — page 6 — #6
i
i
CRC Press
Taylor & Francis Group
6000 Broken Sound Parkway NW, Suite 300
Boca Raton, FL 33487-2742
© 2019 by Taylor & Francis Group, LLC
CRC Press is an imprint of Taylor & Francis Group, an Informa business
No claim to original U.S. Government works
Printed on acid-free paper
Version Date: 20180712
International Standard Book Number-13: 978-1-498-75509-2 (Hardback)
This book contains information obtained from authentic and highly regarded sources. Reasonable efforts have been
made to publish reliable data and information, but the author and publisher cannot assume responsibility for the validity
of all materials or the consequences of their use. The authors and publishers have attempted to trace the copyright
holders of all material reproduced in this publication and apologize to copyright holders if permission to publish in this
form has not been obtained. If any copyright material has not been acknowledged please write and let us know so we may
rectify in any future reprint.
Except as permitted under U.S. Copyright Law, no part of this book may be reprinted, reproduced, transmitted, or utilized
in any form by any electronic, mechanical, or other means, now known or hereafter invented, including photocopying,
microfilming, and recording, or in any information storage or retrieval system, without written permission from the
publishers.
For permission to photocopy or use material electronically from this work, please access www.copyright.com
(http://www.copyright.com/) or contact the Copyright Clearance Center, Inc. (CCC), 222 Rosewood Drive, Danvers,
MA 01923, 978-750-8400. CCC is a not-for-profit organization that provides licenses and registration for a variety of
users. For organizations that have been granted a photocopy license by the CCC, a separate system of payment has been
arranged.
Trademark Notice: Product or corporate names may be trademarks or registered trademarks, and are used only for
identification and explanation without intent to infringe.
Visit the Taylor & Francis Web site at
http://www.taylorandfrancis.com
and the CRC Press Web site at
http://www.crcpress.com
i
i
i
i
✐
✐
“K27443” — 2018/7/13 — 14:41 — page v — #1
✐
✐
To Lili, Adam, and Andrew
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page vi — #2
✐
✐
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page vii — #3
✐
✐
C ontents
◮ Preface
xiii
◮ 1 Axioms of Probability
1
1.1
1.2
1.3
1.4
1.5
1.6
1.7
1.8
Introduction
1
Sample Space and Events
3
Axioms of Probability
12
Basic Theorems
18
Continuity of Probability Function
27
Probabilities 0 and 1
29
Random Selection of Points from Intervals
What Is Simulation?
35
Chapter 1 Summary
37
Review Problems
39
Self-Test on Chapter 1
42
◮ 2 Combinatorial Methods
2.1
2.2
2.3
2.4
2.5
Introduction
45
Counting Principle
45
Number of Subsets of a Set
Tree Diagrams
50
Permutations
55
Combinations
62
Stirling’s Formula
80
Chapter 2 Summary
81
Review Problems
82
Self-Test on Chapter 2
85
30
45
50
vii
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page viii — #4
✐
✐
viii
Contents
◮ 3 Conditional Probability and Independence
3.1
3.2
3.3
3.4
3.5
Conditional Probability
87
Reduction of Sample Space
The Multiplication Rule
97
Law of Total Probability
101
Bayes’ Formula
111
Independence
120
Chapter 3 Summary
137
Review Problems
139
Self-Test on Chapter 3
142
◮ 4
4.1
4.2
4.3
4.4
4.5
4.6
87
91
Distribution Functions and
Discrete Random Variables
Random Variables
145
Distribution Functions
149
Discrete Random Variables
158
Expectations of Discrete Random Variables
164
Variances and Moments of Discrete Random Variables
Moments
183
Standardized Random Variables
187
Chapter 4 Summary
188
Review Problems
190
Self-Test on Chapter 4
192
145
178
◮ 5 Special Discrete Distributions
5.1
5.2
5.3
Bernoulli and Binomial Random Variables
195
Expectations and Variances of Binomial Random Variables
Poisson Random Variable
209
Poisson as an Approximation to Binomial
209
Poisson Process
213
Other Discrete Random Variables
222
Geometric Random Variable
222
Negative Binomial Random Variable
225
Hypergeometric Random Variable
227
Chapter 5 Summary
236
Review Problems
238
Self-Test on Chapter 5
240
195
202
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page ix — #5
✐
✐
Contents
◮ 6 Continuous Random Variables
6.1
6.2
6.3
7.3
7.4
7.5
7.6
Uniform Random Variable
279
Normal Random Variable
285
Correction for Continuity
288
Exponential Random Variables
299
Gamma Distribution
306
Beta Distribution
311
Survival Analysis and Hazard Function
Chapter 7 Summary
322
Review Problems
325
Self-Test on Chapter 7
326
8.2
8.3
8.4
279
317
◮ 8 Bivariate Distributions
8.1
243
Probability Density Functions
243
Density Function of a Function of a Random Variable
253
Expectations and Variances
259
Expectations of Continuous Random Variables
259
Variances of Continuous Random Variables
264
Chapter 6 Summary
272
Review Problems
274
Self-Test on Chapter 6
275
◮ 7 Special Continuous Distributions
7.1
7.2
ix
329
Joint Distributions of Two Random Variables
329
Joint Probability Mass Functions
329
Joint Probability Density Functions
333
Independent Random Variables
348
Independence of Discrete Random Variables
349
Independence of Continuous Random Variables
351
Conditional Distributions
360
Conditional Distributions: Discrete Case
361
Conditional Distributions: Continuous Case
366
Transformations of Two Random Variables
374
Chapter 8 Summary
383
Review Problems
386
Self-Test on Chapter 8
389
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page x — #6
✐
✐
x
Contents
◮ 9 Multivariate Distributions
9.1
9.2
9.3
Joint Distributions of n > 2 Random Variables
Joint Probability Mass Functions
393
Joint Probability Density Functions
401
Random Sample
406
Order Statistics
411
Multinomial Distributions
417
Chapter 9 Summary
422
Review Problems
425
Self-Test on Chapter 9
426
393
393
◮ 10 More Expectations and Variances
10.1
10.2
10.3
10.4
10.5
Expected Values of Sums of Random Variables
Covariance
439
Correlation
450
Conditioning on Random Variables
455
Bivariate Normal Distribution
470
Chapter 10 Summary
475
Review Problems
478
Self-Test on Chapter 10
480
◮ 11
429
429
Sums of Independent Random
Variables and Limit Theorems
483
11.1
11.2
11.3
Moment-Generating Functions
483
Sums of Independent Random Variables
494
Markov and Chebyshev Inequalities
501
Chebyshev’s Inequality and Sample Mean
505
11.4 Laws of Large Numbers
511
11.5 Central Limit Theorem
520
Chapter 11 Summary
529
Review Problems
531
Self-Test on Chapter 11
533
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page xi — #7
✐
✐
Contents
◮ 12
537
Stochastic Processes
Introduction
537
More on Poisson Processes
538
What Is a Queueing System?
548
PASTA: Poisson Arrivals See Time Average
12.3 Markov Chains
552
Classifications of States of a Markov Chain
Absorption Probability
568
Period
570
Steady-State Probabilities
572
12.4 Continuous-Time Markov Chains
582
Steady-State Probabilities
587
Birth and Death Processes
589
Chapter 12 Summary
599
Review Problems
603
Self-Test on Chapter 12
605
xi
12.1
12.2
549
560
◮ Appendix Tables
607
◮ Answers to Odd-Numbered Exercises
611
◮ Index
623
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page xii — #8
✐
✐
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page xiii — #9
✐
✐
Preface
This one- or two-term basic probability text is written for majors in mathematics, physical sciences, engineering, statistics, actuarial science, business and finance, operations research, and
computer science. It can also be used by students who have completed a basic calculus course.
Our aim is to present probability in a natural way: through interesting and instructive examples
and exercises that motivate the theory, definitions, theorems, and methodology. Examples and
exercises have been carefully designed to arouse curiosity and hence encourage the students to
delve into the theory with enthusiasm.
In this edition, there is an ample number of insurance related probability examples and
exercises. These along with hundreds of other types of general indeterministic problems, which
are also subjects of actuarial studies, make the book a perfect source for actuarial students.
Authors are usually faced with two opposing impulses. One is a tendency to put too much
into the book, because everything is important and everything has to be said the author’s way!
On the other hand, authors must also keep in mind a clear definition of the focus, the level,
and the audience for the book, thereby choosing carefully what should be “in” and what “out.”
Hopefully, this book is a resolution of the tension generated by these opposing forces.
Since the publication of the third edition of this book, I have received much additional
correspondence and feedback from faculty and students in this country and abroad. The comments, discussions, recommendations, and reviews helped me to improve the book in many
ways. All detected errors were corrected, and the text has been fine-tuned for accuracy. More
explanations and clarifying comments have been added to certain sections. More insightful and
better solutions are given for a number of problems and exercises.
In this fourth edition, I have added a self-quiz at the end of each section and a self-test
at the end of each chapter so that the students can quiz and test themselves. This edition is
also accompanied with a rich companion website. To further students’ knowledge and proficiency in probability, I have written the Companion Website for Fundamentals of Probability
with Stochastic Processes, Fourth Edition, which shall henceforth be referred to as the companion website. The companion website includes additional examples, topics, and applications for
more in-depth studies. It also includes complete solutions to all self-tests and self-quiz problems so that students can assess their own work and can monitor their own success in learning
this challenging field of study. This companion website is available to all instructors and students who use the fourth edition of the book. I have also written an Instructor’s Solutions
Manual that gives detailed solutions to virtually all of the 1600 exercises of the book and all
exercises of the companion website. The third useful resource for the fourth edition of the book
is the excellent Test Bank, which is authored by Dr. Joshua Stangle.
xiii
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page xiv — #10
✐
✐
xiv
Preface
The Solutions Manual and the Test Bank, which are also included in the companion website, are only available to the instructors of the book and are password protected. To access the
companion website, one should follow the URL,
http://www.crcpress.com/cw/ghahramani.
To distinguish between the material in the book and the material in the companion website, in a given chapter of the book, numbers are used to specify sections, theorems, examples,
figures, etc., while in the companion website, following the chapter number, sections are identified with capital letters, and examples, theorems, figures, etc., are identified by lower case
letters. For example, in the book, Section 5.3, Example 7.2, and Figure 8.5 refer to the third
section of Chapter 5, second example of Chapter 7, and fifth figure of Chapter 8, respectively.
However, in the companion website, for example, Section 5B, Example 7a, and Figure 8d refer
to the second section of Companion for Chapter 5, first example of Companion for Chapter 7,
and fourth figure of Companion for Chapter 8, respectively.
Instructors should enjoy the versatility of this text and its companion website. They can
choose their favorite examples and exercises from a collection of 2,096 and, if necessary, omit
some sections and/or theorems to teach at an appropriate level. In particular, I have added 538
new examples and exercises to this edition, almost all of which are of applied nature and in
realistic contexts.
Exercises for most sections are divided into two categories: A and B. Those in category
A are routine, and those in category B are challenging. However, not all exercises in category
B are uniformly challenging. Some of those exercises are included because students find them
somewhat difficult.
I have tried to maintain an approach that is mathematically rigorous and, at the same time,
closely matches the historical development of probability. Whenever appropriate, I include
historical remarks, and also include discussions of a number of probability problems published
in journals such as Mathematics Magazine and American Mathematical Monthly. These are
interesting and instructive problems that deserve discussion in classrooms.
Chapter 13 of the companion website concerns computer simulation. That chapter is divided into several sections, presenting algorithms that are used to find approximate solutions
to complicated probabilistic problems. These sections can be discussed independently when
relevant materials from earlier chapters are being taught, or they can be discussed concurrently,
toward the end of the semester. Although I believe that the emphasis should remain on concepts, methodology, and the mathematics of the subject, I also think that students should be
asked to read the material on simulation and perhaps do some projects. Computer simulation is
an excellent means to acquire insight into the nature of a problem, its functions, its magnitude,
and the characteristics of the solution.
Other Important Features of the Book Include:
•
The historical roots and applications of many of the theorems and definitions are presented in detail, accompanied by suitable examples or counterexamples.
•
As much as possible, examples and exercises for each section do not refer to exercises
in other chapters or sections—a style that often frustrates students and instructors.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page xv — #11
✐
✐
Preface
xv
•
Whenever a new concept is introduced, its relationship to preceding concepts and theorems is explained.
•
Although the usual analytic proofs are given, simple probabilistic arguments are presented to promote deeper understanding of the subject.
•
The book begins with discussions on probability and its definition, rather than with combinatorics. I believe that combinatorics should be taught after students have learned the
preliminary concepts of probability. The advantage of this approach is that the need for
methods of counting will occur naturally to students, and the connection between the
two areas becomes clear from the beginning. Moreover, combinatorics becomes more
interesting and enjoyable.
•
Students beginning their study of probability have a tendency to think that sample spaces
always have a finite number of sample points. To minimize this proclivity, the concept
of random selection of a point from an interval is introduced in Chapter 1 and applied
where appropriate throughout the book. Moreover, since the basis of simulating indeterministic problems is selection of random points from (0, 1), in order to understand
simulations, students need to be thoroughly familiar with that concept.
•
Often, when we think of a collection of events, we have a tendency to think about them
in either temporal or logical sequence. So, if, for example, a sequence of events A1 ,
A2 , . . . , An occurs in time or in some logical order, we can usually immediately
write
down the probabilities P (A1 ), P A2 | A1 , . . ., P An | A1 A2 · · · An−1 without
much computation. However, we may be interested in probabilities of the intersection
of events, or probabilities of events unconditional on the rest, or probabilities of earlier
events, given later events. These three questions motivated the need for the multiplication
rule, the law of total probability, and Bayes’ theorem. I have given the multiplication rule
a section of its own so that each of these fundamental uses of conditional probability
would have its full share of attention and coverage.
•
Borel’s normal number theorem is discussed in Chapter 11, and a version of a famous
set that is not an event is presented in Chapter 1.
•
The concepts of expectation and variance are introduced early, because important concepts should be defined and used as soon as possible. One benefit of this practice is that,
when random variables such as Poisson and normal are studied, the associated parameters will be understood immediately rather than remaining ambiguous until expectation
and variance are introduced. Therefore, from the beginning, students will develop a natural feeling about such parameters.
•
Special attention is paid to the Poisson distribution; it is made clear that this distribution is frequently applicable, for two reasons: first, because it approximates the binomial
distribution and, second, it is the mathematical model for an enormous class of phenomena. The comprehensive presentation of the Poisson process and its applications can be
understood by junior- and senior-level students.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page xvi — #12
✐
✐
xvi
Preface
•
Students often have difficulties understanding functions or quantities such as the density
function of a continuous random variable
R and the formula for mathematical expectation.
For example, they may wonder why xf (x) dx is the appropriate definition for E(X)
and why correction for continuity is necessary. I have explained the reason behind such
definitions, theorems, and concepts, and have demonstrated why they are the natural
extensions of discrete cases.
•
The first six chapters include many examples and exercises concerning selection of random points from intervals. Consequently, in Chapter 7, when discussing uniform random
variables, I have been able to calculate the distribution and (by differentiation) the density function of X , a random point from an interval (a, b). In this way the concept of a
uniform random variable and the definition of its density function are readily motivated.
•
In Chapters 7 and 8 the usefulness of uniform densities is shown by using many examples. In particular, applications of uniform density in geometric probability theory are
emphasized.
•
Normal density, arguably the most important density function, is readily motivated by
De Moivre’s theorem. In Section 7.2, I introduce the standard normal density, the elementary version of the central limit theorem, and the normal density just as they were
developed historically. Experience shows this to be a good pedagogical approach. When
teaching this approach, the normal density becomes natural and does not look like a
strange function appearing out of the blue.
•
Exponential random variables naturally occur as times between consecutive events of
Poisson processes. The time of occurrence of the nth event of a Poisson process has a
gamma distribution. For these reasons I have motivated exponential and gamma distributions by Poisson processes. In this way we can obtain many examples of exponential
and gamma random variables from the abundant examples of Poisson processes already
known. Another advantage is that it helps us visualize memoryless random variables by
looking at the interevent times of Poisson processes.
•
Joint distributions and conditioning are often trouble areas for students. A detailed explanation and many applications concerning these concepts and techniques make these
materials somewhat easier for students to understand.
•
A section on pattern appearance is presented in Section 10B of Companion for Chapter
10. Even though the method discussed in this section is intuitive and probabilistic, it
should help the students understand such paradoxical-looking results as the following.
On the average, it takes almost twice as many flips of a fair coin to obtain a sequence of
five successive heads as it does to obtain a tail followed by four heads.
•
If a fair coin is tossed a very large number of times, the general perception is that heads
occurs as often as tails. In Section 11B of Companion for Chapter 11, I have explained
what is meant by “heads occurs as often as tails.”
•
To study the risk or rate of “failure,” per unit of time of “lifetimes” that have already
survived a certain length of time, I have included a section, Survival Analysis and Hazard
Functions, in Chapter 7.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page xvii — #13
✐
✐
Preface
xvii
•
Chapter 12 covers more in-depth material on Poisson processes. It also presents the basics of Markov chains and continuous-time Markov chains. The Companion for Chapter
12 includes a complete section on Brownian motion. The topics are covered in some
depth. Therefore, the current edition and its companion website combined have enough
material for a second course in probability as well. The level of difficulty of the chapter
on stochastic processes is consistent with the rest of the book. I believe the explanations
in the book and its companion website make some challenging material more easily accessible to undergraduate and beginning graduate students. We assume only calculus as
a prerequisite. Throughout the chapter, as examples, certain important results from such
areas as queuing theory, random walks, branching processes, superposition of Poisson
processes, and compound Poisson processes are discussed. I have also explained what
the famous theorem, PASTA, Poisson Arrivals See Time Average, states. In short, the
chapter on stochastic processes is laying the foundation on which students’ further pure
and applied probability studies and work can build.
•
Some practical, meaningful, nontrivial, and relevant applications of probability and
stochastic processes in finance, economics, and actuarial sciences are presented.
•
Ever since 1853, when Gregor Johann Mendel (1822–1884) began his breeding experiments with the garden pea Pisum sativum, probability has played an important role in the
understanding of the principles of heredity. In the companion website, I have included
sections and examples on genetics to demonstrate the extent of that role.
•
For random sums of random variables, I have discussed Wald’s equation in Chapter 10
and its analogous case for variance. Certain applications of Wald’s equation have been
discussed in the exercises, as well as in Chapter 12, Stochastic Processes.
•
In this book and its companion website, in mathematical discussions of a topic, there is
no distinction between normal and boldface text.
•
The answers to the odd-numbered exercises of the book are included at the end of the
book.
Sample Syllabi
For a one-term course on probability, instructors have been able to omit many sections without
difficulty. The book is designed for students with different levels of ability, and a variety of
probability courses, applied and/or pure, can be taught using this book. A typical one-semester
course on probability would cover Chapters 1 and 2; Sections 3.1–3.5; Chapters 4, 5, 6; Sections 7.1–7.4; Sections 8.1–8.3; Section 9.1; Sections 10.1–10.3; and Chapter 11.
A follow-up course on introductory stochastic processes, or on a more advanced probability
would cover the remaining material in the book with an emphasis on Sections 8.4, 9.2–9.3, 10.4
and, especially, the entire Chapter 12 and Section 12B of Companion for Chapter 12.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page xviii — #14
✐
✐
xviii
Preface
A course on discrete probability would cover Sections 1.1–1.5; Chapters 2, 3, 4, and 5;
The subsections Joint Probability Mass Functions, Independence of Discrete Random Variables, and Conditional Distributions: Discrete Case, from Chapter 8; the subsection Joint Probability Mass Functions, from Chapter 9; Section 9.3; selected discrete topics from Chapters 10
and 11; and Section 12.3.
Acknowledgments
While writing the manuscript, many people helped me either directly or indirectly. Lili, my
beloved wife, deserves an accolade for her patience and encouragement; as do my wonderful
children.
According to Ecclesiastes 12:12, “of the making of books, there is no end.” Improvements
and advancement to different levels of excellence cannot possibly be achieved without the help,
criticism, suggestions, and recommendations of others. I have been blessed with so many colleagues, friends, and students who have contributed to the improvement of this textbook. One
reason I like writing books is the pleasure of receiving so many suggestions and so much help,
support, and encouragement from colleagues and students all over the world. My experience
from writing the four editions of this book indicates that collaboration and camaraderie in the
scientific community is truly overwhelming.
For the fourth edition of this book, its solutions manual, and its companion website, my
brother, Dr. Soroush Ghahramani, Professor and Chair of the Department of Architecture at
West Valley College in Saratoga, California, using AutoCad, with utmost patience and meticulosity, sketched each and every one of the figures. As a result, the illustrations are very accurate
and clear. I am most indebted to my brother for his hard work and diligence.
For the fourth edition, I wrote many new AMS-LATEX files. My assistants at Western New
England University, Jody Levesque and Dale-Marie Dahlke, with utmost patience, keen eyes,
positive attitude, and eagerness typed these hand-written files into the computer. My great colleague and friend, Professor Ann Kizanis, who is known for being a perfectionist and for having
eagle eyes, read, very carefully, these new files and made many helpful suggestions. Admiring
the various editions of this book, Ann made passionate pleas that I put aside my other projects
and write a new update of this book. Her urging is one of the reasons I wrote this fourth edition.
It gives me a distinct pleasure to thank Ann, Jody, and Dale-Marie for their enthusiastic help.
For the first two editions of the book, my colleagues and friends at Towson University
read or taught from various revisions of the text and offered useful advice. In particular, I
am grateful to Professors Mostafa Aminzadeh, Raouf Boules, James P. Coughlin, Ohoe Kim,
Martha Siegel, Houshang Sohrab, and my late dear friends Sayeed Kayvan and Jerome Cohen.
I want to thank my colleagues Professors Coughlin and Sohrab, especially, for their kindness
and the generosity with which they spent their time carefully reading the entire text every time
it was revised.
Professor Jay Devore from California Polytechnic Institute—San Luis Obispo, made excellent comments that improved the manuscript substantially for the first edition. From Boston
University, Professor Mark E. Glickman’s careful review and insightful suggestions and ideas
helped me in writing the second edition. I was very lucky to receive thorough reviews of the
third edition from Professor James Kuelbs of University of Wisconsin, Madison, Professor
Robert Smits of New Mexico State University, and Ms. Ellen Gundlach from Purdue University. The thoughtful suggestions and ideas of these colleagues improved the third edition of
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page xix — #15
✐
✐
Preface
xix
this book in several ways. I am most grateful to Drs. Devore, Glickman, Kuelbs, Smits, and
Ms. Gundlach.
I am also most grateful to the following professors for their valuable suggestions and constructive criticisms of the current edition of the book: David Busekist, Southeastern Louisiana
University; Christopher Fallaize, University of Nottingham; Trent Gaugler, Lafayette College;
Prakash Gorroochurn, Columbia University; Jose Guardiola, Texas A&M University, Corpus
Christi; Vera Ioudina, Texas State University, San Marcos; Syed Kirmani, University of Northern Iowa; Vasile Lauric, Florida A&M University; Joshua Stangle, University of Wisconsin,
Superior; and Wayne Tarrant, Rose-Hulman Institute of Technology.
Special thanks are due to CRC’s visionary editor, David Grubbs, for his leadership, encouragement, and assistance in seeing this new effort through. I am in awe of the organizational,
time management, and synthesizing skills of Marlena Sullivan, the development editor of Taylor & Francis. It was a true pleasure working with her during the publication of this book. I
would like to thank Marlena for her remarkable help in this undertaking.
Last, but not least, I want to express my gratitude for all the technical help I received, for
17 years, from my dear friend and colleague Professor Howard Kaplon of Towson University,
and all technical help I regularly receive from Bill Landry, my good friend and colleague at
Western New England University.
Saeed Ghahramani
sghahram@wne.edu
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page xx — #16
✐
✐
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 1 — #17
✐
✐
Chapter 1
A xioms of Probability
1.1
INTRODUCTION
In search of natural laws that govern a phenomenon, science often faces “events” that may
or may not occur. The event of disintegration of a given atom of radium is one such example
because, in any given time interval, such an atom may or may not disintegrate. The event of
finding no defect during inspection of a microwave oven is another example, since an inspector
may or may not find defects in the microwave oven. The event that an orbital satellite in space
is at a certain position is a third example. In any experiment, an event that may or may not
occur is called random. If the occurrence of an event is inevitable, it is called certain, and if
it can never occur, it is called impossible. For example, the event that an object travels faster
than light is impossible, and the event that in a thunderstorm flashes of lightning precede any
thunder echoes is certain.
Knowing that an event is random determines only that the existing conditions under which
the experiment is being performed do not guarantee its occurrence. Therefore, the knowledge
obtained from randomness itself is hardly decisive. It is highly desirable to determine quantitatively the exact value, or an estimate, of the chance of the occurrence of a random event. The
theory of probability has emerged from attempts to deal with this problem. In many different
fields of science and technology, it has been observed that, under a long series of experiments,
the proportion of the time that an event occurs may appear to approach a constant. It is these
constants that probability theory (and statistics) aims at predicting and describing as quantitative measures of the chance of occurrence of events. For example, if a fair coin is tossed
repeatedly, the proportion of the heads approaches 1/2. Hence probability theory postulates
that the number 1/2 be assigned to the event of getting heads in a toss of a fair coin.
Historically, from the dawn of civilization, humans have been interested in games of chance
and gambling. However, the advent of probability as a mathematical discipline is relatively
recent. Ancient Egyptians, about 3500 B.C., were using astragali, a four-sided die-shaped bone
found in the heels of some animals, to play a game now called hounds and jackals. The ordinary
six-sided die was made about 1600 B.C. and since then has been used in all kinds of games. The
ordinary deck of playing cards, probably the most popular tool in games and gambling, is much
more recent than dice. Although it is not known where and when dice originated, there are reasons to believe that they were invented in China sometime between the seventh and tenth centuries. Clearly, through gambling and games of chance people have gained intuitive ideas about
the frequency of occurrence of certain events and, hence, about probabilities. But surprisingly,
studies of the chances of events were not begun until the fifteenth century. The Italian scholars
Luca Paccioli (1445–1514), Niccolò Tartaglia (1499–1557), Girolamo Cardano (1501–1576),
1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 2 — #18
✐
✐
2
Chapter 1
Axioms of Probability
and especially Galileo Galilei (1564–1642) were among the first prominent mathematicians
who calculated probabilities concerning many different games of chance. They also tried to
construct a mathematical foundation for probability. Cardano even published a handbook on
gambling, with sections discussing methods of cheating. Nevertheless, real progress started in
France in 1654, when Blaise Pascal (1623–1662) and Pierre de Fermat (1601–1665) exchanged
several letters in which they discussed general methods for the calculation of probabilities.
In 1655, the Dutch scholar Christian Huygens (1629–1695) joined them. In 1657 Huygens
published the first book on probability, De Ratiocinates in Aleae Ludo (On Calculations in
Games of Chance). This book marked the birth of probability. Scholars who read it realized
that they had encountered an important theory. Discussions of solved and unsolved problems
and these new ideas generated readers interested in this challenging new field.
After the work of Pascal, Fermat, and Huygens, the book written by James Bernoulli
(1654–1705) and published in 1713 and that by Abraham de Moivre (1667–1754) in 1730
were major breakthroughs. In the eighteenth century, studies by Pierre-Simon Laplace (1749–
1827), Siméon Denis Poisson (1781–1840), and Karl Friedrich Gauss (1777–1855) expanded
the growth of probability and its applications very rapidly and in many different directions. In
the nineteenth century, prominent Russian mathematicians Pafnuty Chebyshev (1821–1894),
Andrei Markov (1856–1922), and Aleksandr Lyapunov (1857–1918) advanced the works of
Laplace, De Moivre, and Bernoulli considerably. By the early twentieth century, probability
was already a developed theory, but its foundation was not firm. A major goal was to put it
on firm mathematical grounds. Until then, among other interpretations perhaps the relative
frequency interpretation of probability was the most satisfactory. According to this interpretation, to define p, the probability of the occurrence of an event A of an experiment, we study a
series of sequential or simultaneous performances of the experiment and observe that the proportion of times that A occurs approaches a constant. Then we count n(A), the number of times
that A occurs during n performances of the experiment, and we define p = limn→∞ n(A)/n.
This definition is mathematically problematic and cannot be the basis of a rigorous probability
theory. Some of the difficulties that this definition creates are as follows:
1.
In practice, limn→∞ n(A)/n cannot be computed since it is impossible to repeat an experiment infinitely many times. Moreover, if for a large n, n(A)/n is taken as an approximation for the probability of A, there is no way to analyze the error.
2.
There is no reason to believe that the limit of n(A)/n, as n → ∞, exists. Also, if the
existence of this limit is accepted as an axiom, many dilemmas arise that cannot be solved.
For example, there is no reason to believe that, in a different series of experiments and
for the same event A, this ratio approaches the same limit. Hence the uniqueness of the
probability of the event A is not guaranteed.
3.
By this definition, probabilities that are based on our personal belief and knowledge are not
justifiable. Thus statements such as the following would be meaningless.
•
•
•
•
The probability that the price of oil will be raised in the next six months is 60%.
The probability that the 50,000th decimal figure of the number π is 7 exceeds 10%.
The probability that it will snow next Christmas is 30%.
The probability that Mozart was poisoned by Salieri is 18%.
In 1900, at the International Congress of Mathematicians in Paris, David Hilbert (1862–
1943) proposed 23 problems whose solutions were, in his opinion, crucial to the advancement
of mathematics. One of these problems was the axiomatic treatment of the theory of probability. In his lecture, Hilbert quoted Weierstrass, who had said, “The final object, always to be kept
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 3 — #19
✐
✐
Section 1.2
Sample Space and Events
3
in mind, is to arrive at a correct understanding of the foundations of the science.” Hilbert added
that a thorough understanding of special theories of a science is necessary for successful treatment of its foundation. Probability had reached that point and was studied enough to warrant
the creation of a firm mathematical foundation. Some work toward this goal had been done by
Émile Borel (1871–1956), Serge Bernstein (1880–1968), and Richard von Mises (1883–1953),
but it was not until 1933 that Andrei Kolmogorov (1903–1987), a prominent Russian mathematician, successfully axiomatized the theory of probability. In Kolmogorov’s work, which
is now universally accepted, three self-evident and indisputable properties of probability (discussed later) are taken as axioms, and the entire theory of probability is developed and rigorously based on these axioms. In particular, the existence of a constant p, as the limit of the
proportion of the number of times that the event A occurs when the number of experiments
increases to ∞, in some sense, is shown. Subjective probabilities based on our personal knowledge, feelings, and beliefs may also be modeled and studied by this axiomatic approach.
In this book we study the mathematics of probability based on the axiomatic approach.
Since in this approach the concepts of sample space and event play a central role, we now
explain these concepts in detail.
1.2
SAMPLE SPACE AND EVENTS
If the outcome of an experiment is not certain but all of its possible outcomes are predictable
in advance, then the set of all these possible outcomes is called the sample space of the experiment and is usually denoted by S . Therefore, the sample space of an experiment consists of
all possible outcomes of the experiment. These outcomes are sometimes called sample points,
or simply points, of the sample space. In the language of probability, certain subsets of S are
referred to as events. So events are sets of points of the sample space. Some examples follow.
Example 1.1 For the experiment of tossing a coin once, the sample space S consists of two
points (outcomes), “heads” (H) and “tails” (T). Thus S = {H, T}. Example 1.2 Suppose that an experiment consists of two steps. First a coin is flipped. If the
outcome is tails, a die is tossed. If the outcome is heads, the coin is flipped again. The sample
space of this experiment is S = {T1, T2, T3, T4, T5, T6, HT, HH}. For this experiment, the
event of heads in the first flip of the coin is E = {HT, HH}, and the event of an odd outcome
when the die is tossed is F = {T1, T3, T5}. Example 1.3 Consider measuring the lifetime of a light bulb. Since any nonnegative real
number can be considered as the lifetime of the light bulb (in hours), the sample space is
S = {x : x ≥ 0}. In this experiment, E = {x : x ≥ 100} is the event that the light bulb
lasts at least 100 hours, F = {x : x ≤ 1000} is the event that it lasts at most 1000 hours, and
G = {505.5} is the event that it lasts exactly 505.5 hours. Example 1.4 Suppose that a study is being done on all families with one, two, or three
children. Let the outcomes of the study be the genders of the children in descending order of
their ages. Then
S = b, g, bg, gb, bb, gg, bbb, bgb, bbg, bgg, ggg, gbg, ggb, gbb .
Here the outcome b means that the child is a boy, and g means that it is a girl. The events
F = {b, bg, bb, bbb, bgb, bbg, bgg} and G = {gg, bgg, gbg, ggb} represent families where the
eldest child is a boy and families with exactly two girls, respectively. ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 4 — #20
✐
✐
4
Chapter 1
Axioms of Probability
Example 1.5 A bus with a capacity of 34 passengers stops at a station some time between
11:00 A.M. and 11:40 A.M. every day. The sample space of the experiment, consisting of counting
the number of passengers on the bus and measuring the arrival time of the bus, is
n
2o
(1.1)
S = (i, t) : 0 ≤ i ≤ 34, 11 ≤ t ≤ 11 ,
3
where i represents the number of passengers and t the arrival time of the bus in hours and
fractions of hours. The subset of S defined by F = (27, t) : 11 13 < t < 11 23 is the event
that the bus arrives between 11:20 A.M. and 11:40 A.M. with 27 passengers. Remark 1.1 Different manifestations of outcomes of an experiment might lead to different
representations for the sample space of the same experiment. For instance, in Example 1.5,
the outcome that the bus arrives at t with i passengers is represented by (i, t), where t is
expressed in hours and fractions of hours. By this representation, (1.1) is the sample space of
the experiment. Now if the same outcome is denoted by (i, t), where t is the number of minutes
after 11 A.M. that the bus arrives, then the sample space takes the form
S1 = (i, t) : 0 ≤ i ≤ 34, 0 ≤ t ≤ 40 .
To the outcome that
the bus arrives at 11:20 A.M. with 31 passengers, in S the corresponding
point is 31, 11 13 , while in S1 it is (31, 20). Example 1.6 (Round-Off Error) Suppose that each time Jay charges an item to his credit
card, he will round the amount to the nearest dollar in his records. Therefore, the round-off
error, which is the true value charged minus the amount recorded, is random, with the sample
space
S = 0, 0.01, 0.02, . . . , 0.49, −0.50, −0.49, . . . , −0.01 ,
where we have assumed that for any integer dollar amount a, Jay rounds a.50 to a + 1. The
event of rounding off at most 3 cents in a random charge is given by
0, 0.01, 0.02, 0.03, −0.01, −0.02, −0.03 . If the outcome of an experiment belongs to an event E, we say that the event E has
occurred. For example, if we draw two cards from an ordinary deck of 52 cards and observe that one is a spade and the other a heart, all of the events {sh}, {sh, dd}, {cc, dh, sh},
{hc, sh, ss, hh}, and {cc, hh, sh, dd} have occurred because sh, the outcome of the experiment, belongs to all of them. However, none of the events {dh, sc}, {dd}, {ss, hh, cc}, and
{hd, hc, dc, sc, sd} has occurred because sh does not belong to any of them.
In the study of probability theory the relations between different events of an experiment
play a central role. In the remainder of this section we study these relations. In all of the following definitions the events belong to a fixed sample space S .
Subset
Equality
An event E is said to be a subset of the event F if, whenever E occurs, F
also occurs. This means that all of the sample points of E are contained in
F . Hence considering E and F solely as two sets, E is a subset of F in the
usual set-theoretic sense: that is, E ⊆ F .
Events E and F are said to be equal if the occurrence of E implies the
occurrence of F, and vice versa; that is, if E ⊆ F and F ⊆ E, hence
E = F.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 5 — #21
✐
✐
Section 1.2
Sample Space and Events
5
Intersection
An event is called the intersection of two events E and F if it occurs only
whenever E and F occur simultaneously. In the language of sets this event is
denoted by EF or E ∩F because it is the set containing exactly the common
points of E and F .
Union
An event is called the union of two events E and F if it occurs whenever at
least one of them occurs. This event is E ∪ F since all of its points are in E
or F or both.
Complement
An event is called the complement of the event E if it only occurs whenever
E does not occur. The complement of E is denoted by E c .
Difference
An event is called the difference of two events E and F if it occurs whenever
E occurs but F does not. The difference of the events E and F is denoted
by E − F . It is clear that E c = S − E and E − F = EF c .
Certainty
An event is called certain if its occurrence is inevitable. Thus the sample
space is a certain event.
Impossibility
An event is called impossible if there is certainty in its nonoccurrence.
Therefore, the empty set ∅, which is S c , is an impossible event.
Mutually Exclusiveness If the joint occurrence of two events E and F is impossible, we
say that E and F are mutually exclusive. Thus E and F are mutually exclusive if the occurrence of E precludes the occurrence of F, and vice versa.
Since the event representing the joint occurrence of E and F is EF, their
intersection, E and F, are mutually exclusive if EF = ∅. A set of events
{E1 , E2 , . . .} is called mutually exclusive if the joint occurrence of any two
of them is impossible, that is, if ∀i 6= j, Ei Ej = ∅. Thus {E1 , E2 , . . .} is
mutually exclusive if and only if every pair of them is mutually exclusive.
Sn
Tn
S∞
T∞
The events i=1 Ei , i=1 Ei , i=1 Ei , and i=1 Ei are defined in a way similar
Sn to
E1 ∪ E2 and E1 ∩ E2 . For example, if {E1 , E2 , . . . , En } is a set of events, T
by i=1 Ei
n
we mean the event in which at least one of the events Ei , 1 ≤ i ≤ n, occurs. By i=1 Ei we
mean an event that occurs only when all of the events Ei , 1 ≤ i ≤ n, occur.
Example 1.7 At a busy international airport, arriving planes land on a first-come,
first-served basis. Let
E = there are at least five planes waiting to land,
F = there are at most three planes waiting to land,
H = there are exactly two planes waiting to land.
Then
1.
E c is the event that at most four planes are waiting to land.
2.
F c is the event that at least four planes are waiting to land.
3.
E is a subset of F c ; that is, if E occurs, then F c occurs. Therefore, EF c = E.
4.
H is a subset of F ; that is, if H occurs, then F occurs. Therefore, F H = H .
5.
E and F are mutually exclusive; that is, EF = ∅. E and H are also mutually exclusive
since EH = ∅.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 6 — #22
✐
✐
6
Chapter 1
6.
Axioms of Probability
F H c is the event that the number of planes waiting to land is zero, one, or three.
Sometimes Venn diagrams are used to represent the relations among events of a sample
space. The sample space S of the experiment is usually shown as a large rectangle and, inside
S, circles or other geometrical objects are drawn to indicate the events of interest. Figure 1.1
presents Venn diagrams for EF, E ∪ F, E c , and (E c G) ∪ F . The shaded regions are the
indicated events.
S
S
E
E
F
E
EF
S
F
F
S
E
F
E
G
Ec
Figure 1.1
(E c G)
F
Venn diagrams of the events specified.
Example 1.8 Figure 1.2 shows an electric circuit in which, at any random time, each of the
switches located at 1, 2, 3, 4, and 5 is either closed or open. For 1 ≤ i ≤ 5, let Ei be the event
that, at a random time, the switch at location i is closed. In terms of Ei ’s, describe the event
that a signal fed to the input is transmitted to the output.
Solution: The event that a signal fed to the input is transmitted to the output is
F = E1 E5 ∪ E1 E3 E4 ∪ E2 E4 ∪ E2 E3 E5 . ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 7 — #23
✐
✐
Section 1.2
Figure 1.2
Sample Space and Events
7
Electric circuit of Example 1.8.
Unions, intersections, and complementations satisfy many useful relations between events.
A few of these relations are as follows:
(E c )c = E,
E ∪ E c = S,
ES = E,
and EE c = ∅.
Commutative laws:
E ∪ F = F ∪ E,
EF = F E.
Associative laws:
E ∪ (F ∪ G) = (E ∪ F ) ∪ G,
E(F G) = (EF )G.
Distributive laws:
(EF ) ∪ H = (E ∪ H)(F ∪ H),
(E ∪ F )H = (EH) ∪ (F H).
De Morgan’s first law:
(E ∪ F )c = E c F c ,
n
[
Ei
n
\
Ei
i=1
c
=
c
=
n
\
Eic ,
i=1
∞
[
Ei
∞
\
Ei
i=1
c
=
c
=
∞
\
Eic .
i=1
De Morgan’s second law:
(EF )c = E c ∪ F c ,
i=1
n
[
Eic ,
i=1
i=1
∞
[
Eic .
i=1
Another useful relation between E and F, two arbitrary events of a sample space S, is
E = EF ∪ EF c .
This equality readily follows from E = ES and distributivity:
E = ES = E(F ∪ F c ) = EF ∪ EF c .
These and similar identities are usually proved by the elementwise method. The idea is to
show that the events on both sides of the equation are formed of the same sample points. To
use this method, we prove set inclusion in both directions. That is, sample points belonging to
the event on the left also belong to the event on the right, and vice versa. An example follows.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 8 — #24
✐
✐
8
Chapter 1
Axioms of Probability
Example 1.9 Prove De Morgan’s first law: For E and F, two events of a sample space S,
(E ∪ F )c = E c F c .
Proof: First we show that (E ∪ F )c ⊆ E c F c ; then we prove the reverse inclusion E c F c ⊆
(E ∪ F )c . To show that (E ∪ F )c ⊆ E c F c , let x be an outcome that belongs to (E ∪ F )c .
Then x does not belong to E ∪ F, meaning that x is neither in E nor in F . So x belongs to both
E c and F c and hence to E c F c . To prove the reverse inclusion, let x ∈ E c F c . Then x ∈ E c
and x ∈ F c , implying that x 6∈ E and x 6∈ F . Therefore, x 6∈ E ∪ F and thus x ∈ (E ∪ F )c .
Note that Venn diagrams are an excellent way to give intuitive justification for the validity
of relations or to create counterexamples and show invalidity of relations. However, they are
not appropriate to prove relations. This is because of the large number of cases that must be
considered (particularly if more than two events are involved). For example, suppose that by
means of Venn diagrams, we want to prove the identity (EF )c = E c ∪ F c . First we must
draw appropriate representations for all possible ways that E and F can be related: cases such
as EF = ∅, EF 6= ∅, E = F, E = ∅, F = S, and so on. Then in each particular case we
should find the regions that represent (EF )c and E c ∪ F c and observe that they are the same.
Even if these two sets have different representations in only one case, the identity would be
false.
EXERCISES
A
1.
From the letters of the word MISSISSIPPI, a letter is chosen at random. What is a sample
space for this experiment? What is the event that the outcome is a vowel?
2.
Last month, an insurance company sold 57 life insurance policies. Define a sample space
for the number of claims that the company will receive from the beneficiaries of this
group within the next 15 years. What is the event that the company receives at least 3
but no more than 8 claims?
3.
In the experiment of flipping a coin three times, what does the event
E = {THH, HTH, HHT, HHH}
represent?
4.
5.
In theexperiment of tossing two dice, what dothe following events represent?
E = (1, 3), (2, 6), (3, 1), (6, 2) and F = (1, 5), (2, 4), (3, 3), (4, 2), (5, 1) .
A deck of six cards consists of three black cards numbered 1, 2, 3, and three red cards
numbered 1, 2, 3. First, Vann draws a card at random and without replacement. Then
Paul draws a card at random and without replacement from the remaining cards. Let A
be the event that Paul’s card has a larger number than Vann’s card. Let B be the event
that Vann’s card has a larger number than Paul’s card.
(a)
Are A and B mutually exclusive?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 9 — #25
✐
✐
Section 1.2
(b)
Sample Space and Events
9
Are A and B complements of one another?
6.
A box contains three red and five blue balls. Define a sample space for the experiment
of recording the colors of three balls that are drawn from the box, one by one, with
replacement.
7.
Define a sample space for the experiment of choosing a number from the interval (0, 20).
Describe the event that such a number is an integer.
8.
Define a sample space for the experiment of putting three different books on a shelf
in random order. If two of these three books are a two-volume dictionary, describe the
event that these volumes stand in increasing order side-by-side (i.e., volume I precedes
volume II).
9.
Two dice are rolled. Let E be the event that the sum of the outcomes is odd and F be
the event of at least one 1. Interpret the events EF, E c F, and E c F c .
10.
Define a sample space for the experiment of drawing two coins from a purse that contains two quarters, three nickels, one dime, and four pennies. For the same experiment
describe the following events:
(a)
drawing 26 cents;
(b)
drawing more than 9 but less than 25 cents;
(c)
drawing 29 cents.
11.
A telephone call from a certain person is received some time between 7:00 A.M. and 9:10
A.M. every day. Define a sample space for this phenomenon, and describe the event that
the call arrives within 15 minutes of the hour.
12.
A device that has two components fails if at least one of its components breaks down.
The device is observed at a random time. Let Ai , 1 ≤ i ≤ 2, denote the outcome that
the ith component is operative at the random time. In terms of Ai ’s, (a) define a sample
space for the status of the system at the random time; (b) describe the event that the
device is not operative at that random time.
13.
When flipping a coin more than once, what experiment has a sample space defined by
the following?
S = {TT, HTT, THTT, HHTT, HHHTT, HTHTT, THHTT, . . . }.
14.
Let E, F, and G be three events; explain the meaning of the relations E ∪ F ∪ G = G
and EF G = G.
15.
A limousine that carries passengers from an airport to three different hotels just left the
airport with two passengers. Describe the sample space of the stops and the event that
both of the passengers get off at the same hotel.
16.
A psychologist, who is interested in human complexion, in a study of hues and shades,
shows her subjects three pieces of wood colored almond, lemon, and flax, respectively.
She then asks them to identify their favorite colors. Define a sample space for the answer
given by a random subject.
17.
For a saw blade manufacturer’s products, the global demand, per month, for band saws is
between 30 and 36 thousands; for reciprocating saws, it is between 28 and 33 thousands;
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 10 — #26
✐
✐
10
Chapter 1
Axioms of Probability
and for hole saws, it is between 300 and 600 thousands. Define a sample space for the
demands for these three types of saws by this manufacturer in a random month. Describe
the event that, in such a month, the demand for band saws is at least 33 thousands and
for hole saws, it is less than 435 thousands. In this exercise, by “between two numbers
a and b,” we mean inclusive of both a and b.
18.
An insurance company sells a joint life insurance policy to Alexia and her husband, Roy.
Define a sample space for the death or survival of this couple in five years. What is the
event that at that time only one of them lives?
19.
Find the simplest possible expression for the following events.
(a)
(b)
(E ∪ F )(F ∪ G).
(E ∪ F )(E c ∪ F )(E ∪ F c ).
20.
At a certain university, every year eight to 12 professors are granted University Merit
Awards. This year among the nominated faculty are Drs. Jones, Smith, and Brown. Let
A, B, and C denote the events, respectively, that these professors will be given awards.
In terms of A, B, and C, find an expression for the event that the award goes to (a) only
Dr. Jones; (b) at least one of the three; (c) none of the three; (d) exactly two of them;
(e) exactly one of them; (f) Drs. Jones or Smith but not both.
21.
A device that has n, n > 1, components fails to operate if at least one of its components
breaks down. The device is observed at a random time. Let Ai , 1 < i ≤ n, denote the
event that the ith component is operative at the random time. In terms of Ai ’s, describe
the event that the device is operative at that random time.
22.
Prove that the event B is impossible if and only if for every event A,
A = (B ∩ Ac ) ∪ (B c ∩ A).
23.
Let E, F, and G be three events. Determine which of the following statements are correct and which are incorrect. Justify your answers.
(a)
(b)
(c)
(d)
(E − EF ) ∪ F = E ∪ F .
F c G ∪ E c G = G(F ∪ E)c .
(E ∪ F )c G = E c F c G.
EF ∪ EG ∪ F G ⊂ E ∪ F ∪ G.
24.
Travis picks up darts to shoot toward an 18′′ -diameter dartboard aiming at the bullseye
on the board. Describe a sample space for the point at which a dart hits the board.
25.
For the experiment of flipping a coin until a heads occurs, (a) describe the sample space;
(b) describe the event that it takes an odd number of flips until a heads occurs.
26.
In an experiment, cards are drawn, one by one, at random and successively from an
ordinary deck of 52 cards. Let An be the event that no face card or ace appears on the
first n − 1 drawings, and the nth draw is an ace. In terms of An ’s, find an expression
for the event that an ace appears before a face card, (a) if the cards are drawn with
replacement; (b) if they are drawn without replacement.
27.
A point is chosen at random from the interval (−1, 1). Let E1 be the event that it falls in
the interval (−1/2, 1/2), E2 be the event that it falls in the interval (−1/4, 1/4) and, in
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 11 — #27
✐
✐
Section 1.2
Sample Space and Events
11
general,
< ∞, Ei be the event that the point is in the interval (−1/2i , 1/2i ).
S∞for 1 ≤ i T
∞
Find i=1 Ei and i=1 Ei .
B
28.
Prove De Morgan’s second law, (AB)c = Ac ∪ B c , (a) by elementwise proof; (b) by
applying De Morgan’s first law to Ac and B c .
29.
Let A and B be two events. Prove the following relations by the elementwise method.
(a)
(b)
30.
31.
(A − AB) ∪ B = A ∪ B .
(A ∪ B) − AB = AB c ∪ Ac B .
Let {An }∞
n=1 be a sequence of events. Prove that for every event B,
S∞
S∞
(a) B
i=1 Ai =
i=1 BAi .
T∞
S T∞
(b) B
i=1 Ai =
i=1 (B ∪ Ai ).
Define a sample space for the experiment of putting in a random order seven different
books on a shelf. If three of these seven books are a three-volume dictionary, describe the
event that these volumes stand in increasing order side by side (i.e., volume I precedes
volume II and volume II precedes volume III).
32.
In a mathematics department of 31 voting faculty members, there are three candidates
running for the chair position. The voting procedure adopted by the department is
approval voting, in which the eligible voters can vote for as many candidates as they
wish. The candidate with the maximum votes will win. Let Aij , 1 ≤ i ≤ 31, 1 ≤ j ≤ 3,
be the event that, in the next election, the ith faculty member votes for the j th candidate.
In terms of Aij ’s, describe the sample space for the next election of this department.
33.
Let {A1 , A2 , A3 , . . .} be a sequence of events. Find an expression for the event that
infinitely many of the Ai ’s occur.
34.
Let {A1 , A2 , A3 , . . .} be a sequence of events of a sample space S . Find S
a sequence
n
{B
,
B
,
B
,
.
.
.}
of
mutually
exclusive
events
such
that
for
all
n
≥
1,
1
2
3
i=1 Ai =
Sn
B
.
i
i=1
Self-Quiz on Section 1.2
Time allotted: 20 Minutes
Each problem is worth 2.5 points.
1.
Jody, Ann, Bill, and Karl line up in a random order to get their photo taken. Describe the
event that, on the line, males and females alternate.
2.
In a large hospital, there are 100 patients scheduled to have heart bypass surgery. Let
Ei , 1 ≤ i ≤ 100, be the event that the ith patient lives through the postoperative
period of the surgery. In terms of Ei ’s, (a) describe the event that all patients survive the
critical postoperative period; (b) describe the event that exactly one patient dies during
that period.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 12 — #28
✐
✐
12
Chapter 1 Axioms of Probability
3.
Find the simplest possible expression for the event (E ∪ F )(F ∪ G)(EG ∪ F c ).
4.
Consider the system shown by the diagram of Figure 1.3, consisting of 7 components
denoted by 1, 2, . . . , 7. Suppose that each component is either functioning or not functioning, with no other capabilities. Suppose that the system itself also has two performance capabilities, functioning and not functioning. For 1 ≤ i ≤ 7, let Ei be the event
that component i is functioning. In terms of Ei ’s, describe the event that a signal fed to
the input is transmitted to the output.
2
4
1
7
3
Figure 1.3
1.3
5
6
A diagram for the system of Self-Quiz Problem 4.
AXIOMS OF PROBABILITY
In mathematics, the goals of researchers are to obtain new results and prove their correctness,
create simple proofs for already established results, discover or create connections between
different fields of mathematics, construct and solve mathematical models for real-world problems, and so on. To discover new results, mathematicians use trial and error, instinct and inspired guessing, inductive analysis, studies of special cases, and other methods. But when a
new result is discovered, its validity remains subject to skepticism until it is rigorously proven.
Sometimes attempts to prove a result fail and contradictory examples are found. Such examples that invalidate a result are called counterexamples. No mathematical proposition is settled
unless it is either proven or refuted by a counterexample. If a result is false, a counterexample
exists to refute it. Similarly, if a result is valid, a proof must be found for its validity, although
in some cases it might take years, decades, or even centuries to find it.
Proofs in probability theory (and virtually any other theory) are done in the framework of
the axiomatic method. By this method, if we want to convince any rational person, say Sonya,
that a statement L1 is correct, we will show her how L1 can be deduced logically from another
statement L2 that might be acceptable to her. However, if Sonya does not accept L2 , we should
demonstrate how L2 can be deduced logically from a simpler statement L3 . If she disputes
L3 , we must continue this process until, somewhere along the way we reach a statement that,
without further justification, is acceptable to her. This statement will then become the basis
of our argument. Its existence is necessary since otherwise the process continues ad infinitum
without any conclusions. Therefore, in the axiomatic method, first we adopt certain simple,
indisputable, and consistent statements without justifications. These are axioms or postulates.
Then we agree on how and when one statement is a logical consequence of another one and,
finally, using the terms that are already clearly understood, axioms and definitions, we obtain
new results. New results found in this manner are called theorems. Theorems are statements
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 13 — #29
✐
✐
Section 1.3
Axioms of Probability
13
that can be proved. Upon establishment, they are used for discovery of new theorems, and the
process continues and a theory evolves.
In this book, our approach is based on the axiomatic method. There are three axioms upon
which probability theory is based and, except for them, everything else needs to be proved. We
will now explain these axioms.
Definition 1.1 (Probability Axioms) Let S be the sample space of a random phenomenon. Suppose that to each event A of S, a number denoted by P (A) is associated with A.
If P satisfies the following axioms, then it is called a probability and the number P (A) is said
to be the probability of A.
Axiom 1
P (A) ≥ 0.
Axiom 2
P (S) = 1.
Axiom 3
If {A1 , A2 , A3 , . . .} is a sequence of mutually exclusive events (i.e., the joint
occurrence of every pair of them is impossible: Ai Aj = ∅ when i 6= j ), then
P
∞
[
i=1
∞
X
Ai =
P (Ai ).
i=1
Note that the axioms of probability are a set of rules that must be satisfied before S and P
can be considered a probability model.
Axiom 1 states that the probability of the occurrence of an event is always nonnegative. Axiom 2 guarantees that the probability of the occurrence of the event S that is certain is 1. Axiom
3 states that for a sequence of mutually exclusive events the probability of the occurrence of at
least one of them is equal to the sum of their probabilities.
Axiom 2 is merely a convenience to make things definite. It would be equally reasonable
to have P (S) = 100 and interpret probabilities as percentages (which we frequently do).
Let S be the sample space of an experiment. Let A and B be events of S . We say that A
and B are equally likely if P (A) = P (B). Let ω1 and ω2 be sample points of S . We say
that ω1 and ω2 are equally likely if the events {ω1 } and {ω2 } are equally likely, that is, if
P ({ω1 }) = P ({ω2 }).
We will now prove some immediate implications of the axioms of probability.
Theorem 1.1
The probability of the empty set ∅ is 0. That is, P (∅) = 0.
Proof: Let A1 = S and Ai = ∅ for i ≥ 2; then A1 , A2 , A3 , . . . is a sequence of mutually
exclusive events. Thus, by Axiom 3,
P (S) = P
∞
[
i=1
implying that
P∞
∞
∞
X
X
Ai =
P (Ai ) = P (S) +
P (∅),
i=1
i=2 P (∅) = 0. This is possible only if P (∅) = 0.
i=2
Axiom 3 is stated for a countably infinite collection of mutually exclusive events. For this
reason, it is also called the axiom of countable additivity. We will now show that the same
property holds for a finite collection of mutually exclusive events as well. That is, P also
satisfies finite additivity.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 14 — #30
✐
✐
14
Chapter 1 Axioms of Probability
Theorem 1.2
Let {A1 , A2 , . . . , An } be a mutually exclusive set of events. Then
P
n
[
i=1
n
X
P (Ai ).
Ai =
i=1
Proof: For i > n, let Ai = ∅. Then A1 , A2 , A3 , . . . is a sequence of mutually exclusive
events. From Axiom 3 and Theorem 1.1 we get
P
n
[
i=1
∞
∞
[
X
P (Ai )
Ai = P
Ai =
i=1
i=1
=
=
n
X
i=1
n
X
P (Ai ) +
∞
X
i=n+1
P (Ai ) =
n
X
P (Ai ) +
i=1
∞
X
P (∅)
i=n+1
P (Ai ). i=1
Letting n = 2, Theorem 1.2 implies that if A1 and A2 are mutually exclusive, then
P (A1 ∪ A2 ) = P (A1 ) + P (A2 ).
(1.2)
The intuitive meaning of (1.2) is that if an experiment can be repeated indefinitely, then for two
mutually exclusive events A1 and A2 , the proportion of times that A1 ∪ A2 occurs is equal to
the sum of the proportion of times that A1 occurs and the proportion of times that A2 occurs.
For example, for the experiment of tossing a fair die, S = {1, 2, 3, 4, 5, 6} is the sample space.
Let A1 be the event that the outcome is 6, and A2 be the event that the outcome is odd. Then
A1 = {6} and A2 = {1, 3, 5}. Since all sample points are equally likely to occur (by the
definition of a fair die) and the number of sample points of A1 is 1/6 of the number of sample
points of S, we expect that P (A1 ) = 1/6. Similarly, we expect that P (A2 ) = 3/6. Now
A1 A2 = ∅ implies that the number of sample points of A1 ∪ A2 is (1/6 + 3/6)th of the
number of sample points of S . Hence we should expect that P (A1 ∪ A2 ) = 1/6 + 3/6, which
is the same as P (A1 )+P (A2 ). This and many other examples suggest that (1.2) is a reasonable
relation to be taken as Axiom 3. However, if we do this, difficulties arise when a sample space
contains infinitely many sample points, that is, when the number of possible outcomes of an
experiment is not finite. For example, in successive throws of a die let An be the event
S∞that the
first 6 occurs on the nth throw. Then we would be unable to find the probability of n=1 An ,
which represents the event that eventually a 6 occurs. For this reason, Axiom 3, which is the
infinite analog of (1.2), is required. It by no means contradicts our intuitive ideas of probability,
and one of its great advantages is that Theorems 1.1 and 1.2 are its immediate consequences.
A significant implication of (1.2) is that for any event A, P (A) ≤ 1. To see this, note that
P (A ∪ Ac ) = P (A) + P (Ac ).
Now, by Axiom 2,
P (A ∪ Ac ) = P (S) = 1;
therefore, P (A) + P (Ac ) = 1. This and Axiom 1 imply that P (A) ≤ 1. Hence
The probability of the occurrence of an event is always some number between 0 and 1. That is,
0 ≤ P (A) ≤ 1.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 15 — #31
✐
✐
Section 1.3
Axioms of Probability
15
⋆ Remark 1.2† Let S be the sample space of an experiment. The set of all subsets of
S is denoted by P(S) and is called the power set of S . Since the aim of probability theory
is to associate a number between 0 and 1 to every subset of the sample space, probability is
a function P from P(S) to [0, 1]. However, in theory, there is one exception to this: If the
sample space S is not countable, not all of the elements of P(S) are events. There are elements
of P(S) that are not (in a sense defined in more advanced probability texts) measurable. These
elements are not events (see Example 1.22). In other words, it is a curious mathematical fact
that the Kolmogorov axioms are inconsistent with the notion that every subset of every sample
space has a probability. Since in real-world problems we are only dealing with those elements
of P(S) that are measurable, we are not concerned with these exceptions. We must also add
that, in general, if the domain of a function is a collection of sets, it is called a set function.
Hence probability is a real-valued, nonnegative, countably additive set function. Example 1.10 A coin is called unbiased or fair if, whenever it is flipped, the probability of
obtaining heads equals that of obtaining tails. Suppose that in an experiment an unbiased coin
is flipped. The sample space of such an experiment is S = {T, H}. Since the events {H} and
{T} are equally likely to occur, P ({T}) = P ({H}), and since they are mutually exclusive,
P {T, H} = P {T} + P {H} .
Hence Axioms 2 and 3 imply that
1 = P (S) = P {H, T} = P {H} + P {T} = P {H} + P {H} = 2P {H} .
This gives that P {H} = 1/2 and P {T} = 1/2. Now suppose that an experiment consists
of flipping a biased coin where the outcome of tails is twice as likely as heads; then P {T} =
2P {H} . Hence
1 = P (S) = P {H, T} = P {H} + P {T} = P {H} + 2P {H} = 3P {H} .
This shows that P {H} = 1/3; thus P {T} = 2/3. Example 1.11 Sharon has baked five loaves of bread that are identical except that one
of them is underweight. Sharon’s husband chooses one of these loaves at random. Let Bi ,
1 ≤ i ≤ 5, be the event that he chooses the ith loaf. Since all five loaves are equally likely to
be drawn, we have
P {B1 } = P {B2 } = P {B3 } = P {B4 } = P {B5 } .
But the events {B1 }, {B2 }, {B3 }, {B4 }, and {B5 } are mutually exclusive, and the sample
space is S = {B1 , B2 , B3 , B4 , B5 }. Therefore, by Axioms 2 and 3,
1 = P (S) = P {B1 } + P {B2 } + P {B3 } + P {B4 } + P {B5 } = 5 · P {B1 } .
This gives P {B1 } = 1/5 and hence P {Bi } = 1/5, 1 ≤ i ≤ 5. Therefore, the
probability that Sharon’s husband chooses the underweight loaf is 1/5. From Examples 1.10 and 1.11 it should be clear that if a sample space contains N points
that are equally likely to occur, then the probability of each outcome (sample point) is 1/N .
†
Throughout the book, items that are optional and can be skipped are identified by ⋆’s.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 16 — #32
✐
✐
16
Chapter 1 Axioms of Probability
In general, this can be shown as follows. Let S = {s1 , s2 , . . . , sN } be the sample space of an
experiment; then, if all of the sample points are equally likely to occur, we have
P {s1 } = P {s2 } = · · · = P {sN } .
But P (S) = 1, and the events {s1 }, {s2 }, . . . , {sN } are mutually exclusive. Therefore,
1 = P (S) = P {s1 , s2 , . . . , sN }
= P {s1 } + P {s2 } + · · · + P {sN } = N P {s1 } .
This shows that P {s1 } = 1/N . Thus P {si } = 1/N for 1 ≤ i ≤ N .
One simple consequence of the axioms of probability is that, if the sample space of an
experiment contains N points that are all equally likely to occur, then the probability of the
occurrence of any event A is equal to the number of points of A, say N (A), divided by N .
Historically, until the introduction of the axiomatic method by A. N. Kolmogorov in 1933, this
fact was taken as the definition of the probability of A. It is now called the classical definition
of probability. The following theorem, which shows that the classical definition is a simple
result of the axiomatic approach, is also an important tool for the computation of probabilities
of events for experiments with finite sample spaces.
Theorem 1.3 Let S be the sample space of an experiment. If S has N points that are all
equally likely to occur, then for any event A of S,
P (A) =
N (A)
,
N
where N (A) is the number of points of A.
Proof: Let S = {s1 , s2 , . . . , sN }, where each si is an
outcome (a sample point) of the
experiment. Since the outcomes are equiprobable, P {si } = 1/N for all i, 1 ≤ i ≤ N . Now
let A = {si1 , si2 , . . . , siN (A) }, where sij ∈ S for all ij . Since {si1 }, {si2 }, . . . , {siN (A) } are
mutually exclusive, Axiom 3 implies that
P (A) = P {si1 , si2 , . . . , siN (A) }
= P {si1 } + P {si2 } + · · · + P {siN (A) }
=
1
1
N (A)
1
+
+ ··· +
. =
N
N
N
N
|
{z
}
N (A) terms
Example 1.12 Let S be the sample space of flipping a fair coin three times and A be the
event of at least two heads; then
S = HHH, HTH, HHT, HTT, THH, THT, TTH, TTT
and A = {HHH, HTH, HHT, THH}. So N = 8 and N (A) = 4. Therefore, the probability
of at least two heads in flipping a fair coin three times is N (A)/N = 4/8 = 1/2. Example 1.13 An elevator with two passengers stops at the second, third, and fourth floors.
If it is equally likely that a passenger gets off at any of the three floors, what is the probability
that the passengers get off at different floors?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 17 — #33
✐
✐
Section 1.3
Axioms of Probability
17
Solution: Let a and b denote the two passengers and a2 b4 mean that a gets off at the second
floor and b gets off at the fourth floor, with similar representations for other cases. Let A be the
event that the passengers get off at different floors. Then
S = a2 b2 , a 2 b3 , a 2 b4 , a 3 b2 , a 3 b3 , a 3 b4 , a 4 b2 , a 4 b3 , a 4 b4
and A = {a2 b3 , a2 b4 , a3 b2 , a3 b4 , a4 b2 , a4 b3 }. So N = 9 and N (A) = 6. Therefore, the
desired probability is N (A)/N = 6/9 = 2/3. Example 1.14 A number is selected at random from the set of integers 1, 2, . . . , 1000 .
What is the probability that the number is divisible by 3?
Solution: Here the sample space contains 1000 points, so N = 1000. Let A be the set of all
numbers between 1 and 1000 that are divisible by 3. Then A = {3m : 1 ≤ m ≤ 333}. So
N (A) = 333. Therefore, the probability that a random natural number between 1 and 1000 is
divisible by 3 is equal to 333/1000. Example 1.15 A number is selected at random from the set {1, 2, . . . , N }. What is the
probability that the number is divisible by k, 1 ≤ k ≤ N ?
Solution: Here the sample
space contains N points. Let A be the event that the outcome is
divisible by k . Then A = km : 1 ≤ m ≤ [N/k] , where [N/k] is the greatest integer less
than or equal to N/k (to compute [N/k], just divide N by k and round down). So N (A) =
[N/k] and P (A) = [N/k]/N. Remark 1.3 As explained in Remark 1.1, different manifestations of outcomes of an experiment may lead to different representations for the sample space of the same experiment. Because of this, different sample points of a representation might not have the same probability of
occurrence. For example, suppose that a study is being done on families with three children. Let
the outcomes of thestudy be the number of girls and the number of boys
in a randomly selected
family. Then S = bbb, bgg, bgb, bbg, ggb, gbg, gbb, ggg and Ω = bbb, bbg, bgg, ggg are
both reasonable sample spaces for the genders of the children of the family. In S, for example,
bgg means that the first child of the family is a boy, the second child is a girl, and the third
child is also a girl. In Ω, bgg means that the family has one boy and two girls. Therefore, in
S all sample points occur with the same probability, namely,
probabili 1/8. In Ω, however,
ties associated to the sample points are not equal: P {bbb} = P {ggg} = 1/8 whereas
P {bbg} = P {bgg} = 3/8. Finally, we should note that for finite sample spaces, if nonnegative probabilities are assigned to sample points so that they sum to 1, then the probability axioms hold. Let
S = {w1 , w2 , . . . , wn }
be a sample space. Let p1 , p2 , . . . , pn ben nonnegative real numbers with
P be defined on subsets of S by P {wi } = pi , and
P {wi1 , wi2 , . . . , wiℓ } = pi1 + pi2 + · · · + piℓ .
Pn
i=1 pi = 1. Let
It is straightforward to verify that P satisfies the probability axioms. Hence P defines a probability on the sample space S .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 18 — #34
✐
✐
18
Chapter 1
1.4
BASIC THEOREMS
Theorem 1.4
Axioms of Probability
For any event A, P (Ac ) = 1 − P (A).
Proof: Since AAc = ∅, A and Ac are mutually exclusive. Thus
P (A ∪ Ac ) = P (A) + P (Ac ).
But A ∪ Ac = S and P (S) = 1, so
1 = P (S) = P (A ∪ Ac ) = P (A) + P (Ac ).
Therefore, P (Ac ) = 1 − P (A).
This theorem states that the probability of nonoccurrence
of the event A is 1 minus the
probability of its occurrence. For example, consider S = (i, j) : 1 ≤ i ≤ 6, 1 ≤ j ≤ 6 ,
the sample
space of tossing two fair dice. If A is the event of getting a sum of 4, then
A = (1, 3), (2, 2), (3, 1) and P (A) = 3/36. Theorem 1.4 states that the probability of
Ac , the event of not getting a sum of 4, which is harder to count, is 1 − 3/36 = 33/36.
As another example, consider the experiment of selecting a random number from the set
{1, 2, 3, . . . , 1000}. By Example 1.14, the probability that the number selected is divisible
by 3 is 333/1000. Thus by Theorem 1.4, the probability that it is not divisible by 3, a quantity
harder to find directly, is 1 − 333/1000 = 667/1000.
Theorem 1.5
If A ⊆ B, then
P (B − A) = P (BAc ) = P (B) − P (A).
Proof: A ⊆ B implies that B = (B − A) ∪ A (see Figure 1.4). But (B − A)A =
∅.
So the events B − A and A are mutually exclusive, and P (B) = P (B − A) ∪ A =
P (B − A) + P (A). This gives P (B − A) = P (B) − P (A). Corollary
If A ⊆ B, then P (A) ≤ P (B).
Proof: By Theorem 1.5, P (B − A) = P (B) − P (A). Since P (B − A) ≥ 0, we have that
P (B) − P (A) ≥ 0. Hence P (B) ≥ P (A). B
A
Figure 1.4
B_A
A ⊆ B implies that B = (B − A) ∪ A.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 19 — #35
✐
✐
Section 1.4 Basic Theorems
19
This corollary says that, for instance, it is less likely that a computer has one defect than
it has at least one defect. Note that in Theorem 1.5, the condition of A ⊆ B is necessary. The
relation P (B − A) = P (B) − P (A) is not true in general. For example, in rolling a fair die,
let B = {1, 2} and A = {3, 4, 5}, then B − A = {1, 2}. Therefore, P (B − A) = 1/3,
P (B) = 1/3, and P (A) = 1/2. Hence P (B − A) 6= P (B) − P (A).
Theorem 1.6
P (A ∪ B) = P (A) + P (B) − P (AB).
Proof: A ∪ B = A ∪ (B − AB) (see Figure 1.5) and A(B − AB) = ∅, so A and B − AB
are mutually exclusive events and
P (A ∪ B) = P A ∪ (B − AB) = P (A) + P (B − AB).
(1.3)
Now since AB ⊆ B, Theorem 1.5 implies that
P (B − AB) = P (B) − P (AB).
Therefore, (1.3) gives
P (A ∪ B) = P (A) + P (B) − P (AB). Figure 1.5
The shaded region is B−AB.
Thus A ∪ B = A ∪ (B − AB).
Example 1.16 Suppose that in a community of 400 adults, 300 bike or swim or do both,
160 swim, and 120 swim and bike. What is the probability that an adult, selected at random
from this community, bikes?
Solution: Let A be the event that the person swims and B be the event that he or she bikes;
then P (A ∪ B) = 300/400, P (A) = 160/400, and P (AB) = 120/400. Hence the relation
P (A ∪ B) = P (A) + P (B) − P (AB) implies that
P (B) = P (A ∪ B) + P (AB) − P (A)
300 120 160
260
=
+
−
=
= 0.65. 400 400 400
400
Example 1.17 A number is chosen at random from the set of integers {1, 2, . . . , 1000}.
What is the probability that it is divisible by 3 or 5 (i.e., either 3 or 5 or both)?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 20 — #36
✐
✐
20
Chapter 1
Axioms of Probability
Solution: The number of integers between 1 and N that are divisible by k is computed by
dividing N by k and then rounding down (see Examples 1.15). Therefore, if A is the event that
the outcome is divisible by 3 and B is the event that it is divisible by 5, then P (A) = 333/1000
and P (B) = 200/1000. Now AB is the event that the outcome is divisible by both 3 and 5.
Since a number is divisible by 3 and 5 if and only if it is divisible by 15 (3 and 5 are prime
numbers), P (AB) = 66/1000 (divide 1000 by 15 and round down to get 66). Thus the desired
probability is computed as follows:
P (A ∪ B) = P (A) + P (B) − P (AB)
200
66
467
333
+
−
=
= 0.467. =
1000 1000 1000
1000
Theorem 1.6 gives a formula to calculate the probability that at least one of A and B occurs.
We may also calculate the probability that at least one of the events A1 , A2 , A3 , . . . , and An
occurs. For three events A1 , A2 , and A3 ,
P (A1 ∪ A2 ∪ A3 ) = P (A1 ) + P (A2 ) + P (A3 ) − P (A1 A2 )
− P (A1 A3 ) − P (A2 A3 ) + P (A1 A2 A3 ).
For four events,
P (A1 ∪ A2 ∪ A3 ∪ A4 ) = P (A1 ) + P (A2 ) + P (A3 ) + P (A4 ) − P (A1 A2 )
− P (A1 A3 ) − P (A1 A4 ) − P (A2 A3 ) − P (A2 A4 )
− P (A3 A4 ) + P (A1 A2 A3 ) + P (A1 A2 A4 )
+ P (A1 A3 A4 ) + P (A2 A3 A4 ) − P (A1 A2 A3 A4 ).
We now explain a procedure, which will enable us to find P (A1 ∪ A2 ∪ · · · ∪ An ), the
probability that at least one of the events A1 , A2 , · · · , An occurs.
Inclusion-Exclusion Principle To calculate P (A1 ∪ A2 ∪ · · · ∪ An ), first find all of the
possible intersections of events from A1 , A2 , . . . , An and calculate their probabilities. Then
add the probabilities of those intersections that are formed of an odd number of events, and
subtract the probabilities of those formed of an even number of events.
The following formula is an expression for the principle of inclusion-exclusion. It follows by
induction. (For an intuitive proof, see Example 2.33.)
P
n
[
i=1
n
n−1
n
n−2
X
X X
X n−1
X
Ai =
P (Ai ) −
P (Ai Aj ) +
i=1
i=1 j=i+1
− · · · + (−1)n−1 P (A1 A2 · · · An ).
n
X
P (Ai Aj Ak )
i=1 j=i+1 k=j+1
Example 1.18 Suppose that 25% of the population of a city read newspaper A, 20% read
newspaper B, 13% read C, 10% read both A and B, 8% read both A and C, 5% read B and C,
and 4% read all three. If a person from this city is selected at random, what is the probability
that he or she does not read any of these newspapers?
Solution: Let E, F, and G be the events that the person reads A, B, and C, respectively. The
event that the person reads at least one of the newspapers A, B, or C is E ∪ F ∪ G. Therefore,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 21 — #37
✐
✐
Section 1.4 Basic Theorems
21
1 − P (E ∪ F ∪ G) is the probability that he or she reads none of them. Since
P (E ∪ F ∪ G) = P (E) + P (F ) + P (G) − P (EF ) − P (EG)
− P (F G) + P (EF G)
= 0.25 + 0.20 + 0.13 − 0.10 − 0.08 − 0.05 + 0.04 = 0.39,
the desired probability equals 1 − 0.39 = 0.61.
Example 1.19 Dr. Grossman, an internist, has 520 patients, of which (i) 230 are hypertensive, (ii) 185 are diabetic, (iii) 35 are hypochondriac and diabetic, (iv) 25 are all three, (v) 150
are none, (vi) 140 are only hypertensive, and finally, (vii) 15 are hypertensive and hypochondriac but not diabetic. Find the probability that Dr. Grossman’s next appointment is hypochondriac but neither diabetic nor hypertensive. Assume that appointments are all random. This
implies that even hypochondriacs do not make more visits than others.
Solution: Let T, C, and D denote the events that the next appointment of Dr. Grossman is
hypertensive, hypochondriac, and diabetic, respectively. The Venn diagram of Figure 1.6 shows
that the number of patients with only hypochondria is 30. Therefore, the desired probability is
30/520 ≈ 0.06. Figure 1.6
Theorem 1.7
Venn diagram of Example 1.19.
P (A) = P (AB) + P (AB c ).
Proof: Clearly, A = AS = A(B ∪ B c ) = AB ∪ AB c . Since AB and AB c are mutually
exclusive,
P (A) = P (AB ∪ AB c ) = P (AB) + P (AB c ). Example 1.20 In a community, 32% of the population are male smokers; 27% are female
smokers. What percentage of the population of this community smoke?
Solution: Let A be the event that a randomly selected person from this community smokes.
Let B be the event that the person is male. By Theorem 1.7,
P (A) = P (AB) + P (AB c ) = 0.32 + 0.27 = 0.59.
Therefore, 59% of the population of this community smoke.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 22 — #38
✐
✐
22
Chapter 1
Axioms of Probability
Remark 1.4 In the real world, especially for games of chance, it is common to express
probability in terms of odds. We say that the odds in favor of an event A are r to s if P (A) =
r/(r + s). Similarly, the odds against an event A are r to s if P (A) = s/(r + s). Therefore,
if the odds in favor of an event A are r to s, then the odds against A are s to r . For example, in
drawing a card at random from an ordinary deck of 52 cards, the odds against drawing an ace
are 48 to 4 or, equivalently, 12 to 1. The odds in favor of an ace are 4 to 48 or, equivalently, 1
to 12. If for an event A, P (A) = p, then the odds in favor of A are p to 1 − p. Therefore, for
example, in flipping a fair coin three times, by Example 1.12, the odds in favor of HHH are 1/8
to 7/8 or, equivalently, 1 to 7. EXERCISES
A
1.
Gottfried Wilhelm Leibniz (1646–1716), the German mathematician, philosopher,
statesman, and one of the supreme intellects of the seventeenth century, believed that
in a throw of a pair of fair dice, the probability of obtaining the sum 11 is equal to that
of obtaining the sum 12. Do you agree with Leibniz? Explain.
2.
Show that if A and B are mutually exclusive, then P (A) + P (B) ≤ 1.
3.
Suppose that 33% of the people have O+ blood and 7% have O− . What is the probability
that the next president of the United States has type O blood?
4.
Sixty-eight minutes through Morituri, the 1965 Marlon Brando movie, a group of underground anti-Nazis wanting to escape from a Nazi ship to a nearby island state that
the chances of their succeeding are 15-to-1. Eighty-two minutes through the movie, the
same group refers to this estimate stating that the odds against successfully escaping to
that island are 15-to-1. Are these two statements consistent? Why or why not?
5.
The probability that an earthquake will damage a certain structure during a year is 0.015.
The probability that a hurricane will damage the same structure during a year is 0.025. If
the probability that both an earthquake and a hurricane will damage the structure during
a year is 0.0073, what is the probability that next year the structure will not be damaged
by a hurricane or an earthquake?
6.
In a probability test, for two events E and F of a sample space, Tina’s calculations
resulted in P (E) = 1/4, P (F ) = 1/2, and P (EF ) = 3/8. Is it possible that Tina
made a mistake in her calculations? Why or why not?
7.
For events A and B, suppose that the probability that at least one of them occurs is 0.8
and the probability that both of them occur is 0.3. Find the probability that exactly one
of them occurs.
8.
The admission office of a college admits only applicants whose high school GPA is at
least 3.0 or whose SAT score is 1200 or higher. If 38% of the applicants of this college
have at least a 3.0 GPA, 30% have a SAT of 1200 or higher, and 15% have both, what
percentage of all the applicants are admitted to the college?
9.
Jacqueline, Bonnie, and Tina are the only contestants in an athletic race, where it is not
possible to tie. The probability that Bonnie wins is 2/3 that of Jacqueline winning and
4/3 that of Tina winning. Find the probability of each of these three athletes winning.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 23 — #39
✐
✐
Section 1.4 Basic Theorems
23
10.
A motorcycle insurance company has 7000 policyholders of whom 5000 are under 40.
If 4100 of the policyholders are males and under 40, 1100 are married and under 40, and
550 are married males who are under 40, find the probability that the next motorcycle
policyholder of this company who gets into an accident is a single female who is under
40. Assume that it is equally likely for all policyholders to get into an accident.
11.
From an ordinary deck of 52 cards, we draw cards at random and without replacement.
What is the probability that at least two cards must be drawn to obtain a face card?
12.
Suppose that the probability that a driver is a male, and has at least one motor vehicle
accident during a one-year period, is 0.12. Suppose that the corresponding probability
for a female is 0.06. What is the probability of a randomly selected driver having at least
one accident during the next 12 months?
13.
Suppose that 75% of all investors invest in traditional annuities and 45% of them invest
in the stock market. If 85% invest in the stock market and/or traditional annuities, what
percentage invest in both?
14.
In a horse race, the odds in favor of the first horse winning in an 8-horse race are 2 to 5.
The odds against the second horse winning are 7 to 3. What is the probability that one
of these horses will win?
15.
A professor asks her students to present three methods of generating a random state out
of US states. One of the methods a student, Walter, introduces is to draw a congressman
from the list of 535 voting members of the House of Representatives and then select the
state to which that member belongs. Is the state selected by this method random?
16.
Suppose that in the following table, by the probability of the interval [x, y) we mean the
probability that, in a certain region, a person dies on or after his or her xth birthday but
before his or her y th birthday.
17.
18.
Interval
Probability of the interval
Interval
Probability of the interval
[0, 10)
[10, 20)
[20, 30)
[30, 40)
[40, 50)
0.0634
0.0129
0.0218
0.0317
0.0634
[50, 60)
[60, 70)
[70, 80)
[80, 90)
[90, 100)
[100, ∞)
0.1257
0.2089
0.2710
0.1683
0.0277
0.0052
Based on this table of mortality rates, what is the probability that a baby just born in this
region reaches the age of 60? What is the probability that he or she lives to be at least 20
but dies before 50?
Let S = {ω1 , ω2 , ω3 } be the sample space of an experiment. If P {ω1 , ω2 } = 0.5 and
P {ω1 , ω3 } = 0.7, find P {ω1 } , P {ω2 } , and P {ω3 } .
Excerpt from the TV show The Rockford Files:
Rockford: There are only two doctors in town. The chances of both autopsies
being performed by the same doctor are 50–50.
Reporter: No, that is only for one autopsy. For two autopsies, the chances
are 25–75.
Rockford: You’re right.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 24 — #40
✐
✐
24
Chapter 1
Axioms of Probability
Was Rockford right to agree with the reporter? Explain why or why not?
19.
In a major state university, to study the performance relationship between calculus I and
calculus II grades, the mathematics department reviewed the letter grades of all students
who had taken both of the calculus courses in the last twenty years. For each pair of
letter grades (X, Y ), the department calculated the proportion of all of the students who
had earned X in calculus I and Y in calculus II. The results are given in the following
table.
Calculus II
Calculus I
F
D
C
B
A
F
D
C
B
A
0.00
0.08
0.02
0.01
0.00
0.00
0.05
0.03
0.01
0.01
0.00
0.04
0.18
0.09
0.06
0.00
0.02
0.08
0.07
0.10
0.00
0.01
0.02
0.04
0.08
According to the rules of this university, if a student takes a course several times, only
the last grade will be recorded. Based on this table, find the probability that a randomly
selected student (a) does better in calculus I than in calculus II; (b) does better in calculus
II than in calculus I; (c) earns the same letter grade in both calculus courses.
20.
A company has only one position with three highly qualified applicants: John, Barbara,
and Marty. However, because the company has only a few women employees, Barbara’s
chance to be hired is 20% higher than John’s and 20% higher than Marty’s. Find the
probability that Barbara will be hired.
21.
In a psychiatric hospital, the number of patients with schizophrenia is three times the
number with psychoneurotic reactions, twice the number with alcohol addictions, and 10
times the number with involutional psychotic reaction. If a patient is selected randomly
from the list of all patients with one of these four diseases, what is the probability that
he or she suffers from schizophrenia? Assume that none of these patients has more than
one of these four diseases.
22.
Let A and B be two events. Prove that
P (AB) ≥ P (A) + P (B) − 1.
23.
A card is drawn at random from an ordinary deck of 52 cards. What is the probability
that it is (a) a black ace or a red queen; (b) a face or a black card; (c) neither a heart nor
a queen?
24.
Which of the following statements is true? If a statement is true, prove it. If it is false,
give a counterexample.
(a)
If P (A) + P (B) + P (C) = 1, then the events A, B, and C are mutually
exclusive.
(b)
If P (A ∪ B ∪ C) = 1, then A, B, and C are mutually exclusive events.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 25 — #41
✐
✐
Section 1.4 Basic Theorems
25
25.
Suppose that in the Baltimore metropolitan area 25% of the crimes occur during the day
and 80% of the crimes occur in the city. If only 10% of the crimes occur outside the city
during the day, what percent occur inside the city during the night? What percent occur
outside the city during the night?
26.
Suppose that, in a temperate coniferous forest, 40% of randomly selected quarter-acre
plots have both cedar and cypress trees, 25% have both cedar and redwood trees, 20%
have cypress and redwoods, and 15% have all three. What is the probability that in a
randomly selected quarter-acre plot in this forest we can find exactly two of the three
types of these trees?
27.
Let A, B, and C be three events. Prove that
P (A ∪ B ∪ C)
= P (A) + P (B) + P (C) − P (AB) − P (AC) − P (BC) + P (ABC).
28.
Let A, B, and C be three events. Show that exactly two of these events will occur with
probability
P (AB) + P (AC) + P (BC) − 3P (ABC).
29.
Eleven chairs are numbered 1 through 11. Four girls and seven boys sit on these chairs
at random. What is the probability that chair 5 is occupied by a boy?
30.
A ball is thrown at a square that is divided into n2 identical squares.PThe probability
that
n Pn
the ball hits the square of the ith column and j th row is pij , where i=1 j=1 pij = 1.
In terms of pij ’s, find the probability that the ball hits the j th horizontal strip.
31.
Among 33 students in a class, 17 of them earned A’s on the midterm exam, 14 earned A’s
on the final exam, and 11 did not earn A’s on either examination. What is the probability
that a randomly selected student from this class earned an A on both exams?
32.
From a small town 120 persons were selected at random and asked the following question: Which of the three shampoos, A, B, or C, do you use? The following results were
obtained: 20 use A and C, 10 use A and B but not C, 15 use all three, 30 use only C, 35
use B but not C, 25 use B and C, and 10 use none of the three. If a person is selected at
random from this group, what is the probability that he or she uses (a) only A; (b) only
B; (c) A and B ? (Draw a Venn diagram.)
33.
The coefficients of the quadratic equation x2 + bx + c = 0 are determined by tossing
a fair die twice (the first outcome is b, the second one is c). Find the probability that the
equation has real roots.
34.
Two integers m and n are called relatively prime if 1 is their only common positive
divisor. Thus 8 and 5 are relatively prime, whereas 8 and 6 are not. A number is selected
at random from the set {1, 2, 3, . . . , 63}. Find the probability that it is relatively prime
to 63.
35.
A number is selected randomly from the set {1, 2, . . . , 1000}. What is the probability
that (a) it is divisible by 3 but not by 5; (b) it is divisible neither by 3 nor by 5?
36.
The secretary of a college has calculated that from the students who took calculus,
physics, and chemistry last semester, 78% passed calculus, 80% physics, 84% chemistry,
60% calculus and physics, 65% physics and chemistry, 70% calculus and chemistry, and
55% all three. Show that these numbers are not consistent, and therefore the secretary
has made a mistake.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 26 — #42
✐
✐
26
Chapter 1
Axioms of Probability
B
37.
From an ordinary deck of 52 cards, we draw cards at random and without replacement
until only cards of one suit are left. Find the probability that the cards left are all spades.
38.
A number is selected at random from the set of natural numbers {1, 2, . . . , 1000}. What
is the probability that it is divisible by 4 but neither by 5 nor by 7?
39.
For a Democratic candidate to win an election, she must win districts I, II, and III. Polls
have shown that the probability of winning I and III is 0.55, losing II but not I is 0.34,
and losing II and III but not I is 0.15. Find the probability that this candidate will win all
three districts. (Draw a Venn diagram.)
40.
For events E and F, show that
P E ∪ F + P E ∪ F c + P E c ∪ F + P E c ∪ F c = 3.
41.
42.
Two numbers are successively selected at random and with replacement from the set
{1, 2, . . . , 100}. What is the probability that the first one is greater than the second?
Let A1 , A2 , A3 , . . . be a sequence of events of a sample space. Prove that
P
∞
[
n=1
∞
X
An ≤
P (An ).
n=1
This is called Boole’s inequality.
43.
Let A1 , A2 , A3 , . . . be a sequence of events of an experiment. Prove that
P
∞
\
n=1
Hint:
∞
X
An ≥ 1 −
P (Acn ).
n=1
Use Boole’s inequality, discussed in Exercise 42.
44.
In a certain country, the probability is 49/50 that a randomly selected fighter plane returns from a mission without mishap. Mia argues that this means there is one mission
with a mishap in every 50 consecutive flights. She concludes that if a fighter pilot returns
safely from 49 consecutive missions, he should return home before his fiftieth mission.
Is Mia right? Explain why or why not.
45.
Let P be a probability defined on a sample space S . For events A of S define Q(A) =
2
P (A) and R(A) = P (A)/2. Is Q a probability on S ? Is R a probability on S ? Why
or why not?
In its general case, the following exercise has important applications in coding theory,
telecommunications, and computer science.
46.
(The Hat Problem) A game begins with a team of three players entering a room
one at a time. For each player, a fair coin is tossed. If the outcome is heads, a red hat
is placed on the player’s head, and if it is tails, a blue hat is placed on the player’s
head. The players are allowed to communicate before the game begins to decide on a
strategy. However, no communication is permitted after the game begins. Players cannot
see their own hats. But each player can see the other two players’ hats. Each player is
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 27 — #43
✐
✐
Section 1.5
Continuity of Probability Function
27
given the option to guess the color of his or her hat or to pass. The game ends when
the three players simultaneously make their choices. The team wins if no player’s guess
is incorrect and at least one player’s guess is correct. Obviously, the team’s goal is to
develop a strategy that maximizes the probability of winning. A trivial strategy for the
team would be for two of its players to pass and the third player to guess red or blue as
he or she wishes. This gives the team a 50% chance to win. Can you think of a strategy
that improves the chances of the team winning?
Self-Quiz on Section 1.4
Time allotted: 20 Minutes
1.
Let A and B be mutually exclusive events of an experiment with P (A) = 1/3 and
P (B) = 1/4. What is the probability that neither A occurs nor B ? (3 points)
2.
Zack has two weeks to read a book assigned by his English teacher and two weeks to
write an essay assigned by his philosophy professor. The probability that he completes
both assignments on time is 0.6, and the probability that he completes at least one of
them on time is 0.95. What is the probability that he ends up meeting the deadline only
for one assignment? (3 points)
3.
Suppose that two-thirds of Americans traveling to Europe visit at least one of the three
countries France, England, and Italy. Furthermore, suppose that one-half of them visit
England, one-third visit France, and one-fourth visit Italy. If for each pair of these countries, one-fifth of the travelers visit that pair, what fraction of Americans traveling to
Europe visit all of these three countries? (4 points)
1.5
CONTINUITY OF PROBABILITY FUNCTIONS
Let R denote (here and everywhere else throughout the book) the set of all real numbers. We
know from calculus that a function f : R → R is called continuous at a point c ∈ R if
limx→c f (x) = f (c). It is called continuous on R if it is continuous at all points c ∈ R. We
also know that this definition is equivalent to the sequential criterion f : R → R is continuous
on R if and only if, for every convergent sequence {xn }∞
n=1 in R,
lim f (xn ) = f ( lim xn ).
n→∞
n→∞
(1.4)
This property, in some sense, is shared by the probability function. To explain this, we need to
introduce some definitions. But first recall that probability is a set function from P(S), the set
of all possible events of the sample space S, to [0, 1].
A sequence {En , n ≥ 1} of events of a sample space is called increasing if
E1 ⊆ E2 ⊆ E3 ⊆ · · · ⊆ En ⊆ En+1 · · · ;
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 28 — #44
✐
✐
28
Chapter 1
Axioms of Probability
it is called decreasing if
E1 ⊇ E2 ⊇ E3 ⊇ · · · ⊇ En ⊇ En+1 ⊇ · · · .
For an increasing sequence of events {En , n ≥ 1}, by limn→∞ En we mean the event that at
least one Ei , 1 ≤ i < ∞ occurs. Therefore,
lim En =
n→∞
∞
[
Ei .
i=1
Similarly, for a decreasing sequence of events {En , n ≥ 1}, by limn→∞ En we mean the
event that every Ei occurs. Thus in this case
lim En =
n→∞
∞
\
Ei .
i=1
The following theorem expresses the property of probability function that is analogous to (1.4).
Theorem 1.8
(Continuity of Probability Function)
decreasing sequence of events, {En , n ≥ 1},
For any increasing or
lim P (En ) = P ( lim En ).
n→∞
n→∞
Proof: For the case where {En , n ≥ 1} is increasing, let F1 = E1 , F2 = E2 − E1 ,
F3 = E3 − E2 , . . . , Fn = En − En−1 , . . . . Clearly, {Fi , i ≥ 1} is a mutually exclusive set
of events that satisfies the following relations:
n
[
i=1
∞
[
Fi =
Fi =
i=1
(see Figure 1.7). Hence
P ( lim En ) = P
n→∞
n
[
i=1
∞
[
Ei = En ,
n = 1, 2, 3, . . . ,
Ei
i=1
∞
[
i=1
= lim P
n→∞
Ei = P
n
[
i=1
∞
[
i=1
Fi =
Fi = lim P
n→∞
∞
X
P (Fi ) = lim
i=1
n
[
i=1
n→∞
n
X
P (Fi )
i=1
Ei = lim P (En ),
n→∞
Sn
where the last equality follows since {En , n ≥ 1} is increasing, and hence i=1 Ei = En .
This establishes the theorem for increasing sequences.
c
If {En , n ≥ 1} is decreasing, then En ⊇ En+1 , ∀n, implies that Enc ⊆ En+1
, ∀n.
c
Therefore, the sequence {En , n ≥ 1} is increasing and
\
∞
∞
∞
\
c [
P ( lim En ) = P
Ei = 1 − P
Ei
=1−P
Eic
n→∞
i=1
i=1
i=1
= 1 − P ( lim Enc ) = 1 − lim P (Enc ) = 1 − lim
n→∞
n→∞
n→∞
= 1 − 1 + lim P (En ) = lim P (En ). n→∞
1 − P (En )
n→∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 29 — #45
✐
✐
Section 1.6
Figure 1.7
Probabilities 0 and 1
29
The circular disks are the Ei ’s and the shaded circular annuli
are the Fi ’s, except for F1 , which equals E1 .
Example 1.21 Suppose that some individuals in a population produce offspring of the same
kind. The offspring of the initial population are called second generation, the offspring of the
second generation
are called third
generation, and so on. Furthermore, suppose that with probability exp − (2n2 + 7)/(6n2 ) the entire population completely dies out by the nth generation
before producing any offspring, what is the probability that such a population survives forever?
Solution: Let En denote the event of extinction of the entire population by the nth generation;
then
E1 ⊆ E2 ⊆ E3 ⊆ · · · ⊆ En ⊆ En+1 ⊆ · · ·
because if En occurs, then En+1 also occurs. Hence, by Theorem 1.8,
P {population survives forever} = 1 − P {population eventually dies out}
∞
[
=1−P
Ei = 1 − lim P (En )
i=1
n→∞
2n2 + 7 = 1−e−1/3 . = 1 − lim exp −
n→∞
6n2
1.6
PROBABILITIES 0 AND 1
Events with probabilities 1 and 0 should not be misinterpreted. If E and F are events with
probabilities 1 and 0, respectively, it is not correct to say that E is the sample space S and F is
the empty set ∅. In fact, there are experiments in which there exist infinitely many events each
with probability 1, and infinitely many events each with probability 0. An example follows.
Suppose that an experiment consists of selecting a random point from the interval (0, 1).
Since every point in (0, 1) has a decimal representation such as
0.529387043219721 · · · ,
the experiment is equivalent to picking an endless decimal from (0, 1) at random (note that if a
decimal terminates, all of its digits from some point on are 0). In such an experiment we want
to compute the probability of selecting the point 1/3. In other words, we want to compute the
probability of choosing 0.333333 · · · in a random selection of an endless decimal. Let An be
the event that the selected decimal has 3 as its first n digits; then
A1 ⊃ A2 ⊃ A3 ⊃ A4 ⊃ · · · ⊃ An ⊃ An+1 ⊃ · · · ,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 30 — #46
✐
✐
30
Chapter 1
Axioms of Probability
since the occurrence of An+1 guarantees the occurrence of An . Now P (A1 ) = 1/10 because
there are 10 choices 0, 1, 2, . . . , 9 for the first digit, and we want only one of them, namely
3, to occur. P (A2 ) = 1/100 since there are 100 choices 00, 01, . . . , 09, 10, 11, . . . , 19, 20,
. . . , 99 for the first two digits, and we want only one of them, 33, to occur. P (A3 ) = 1/1000
because there are 1000 choices 000, 001, . . . , 999 for the first three digits, and we want only
n
one
T∞ of them, 333, to occur. Continuing this argument, we have P (An ) = (1/10) . Since
n=1 An = {1/3}, by Theorem 1.8,
∞
1 n
\
is selected = P
An = lim P (An ) = lim
= 0.
P
n→∞
n→∞ 10
3
n=1
1
Note that there is nothing special about the point 1/3. For any other point 0.α1 α2 α3 α4 · · ·
from (0, 1), the same argument could be used to show that the probability of its occurrence is
0 (define An to be the event that the first n digits of the selected decimal are α1 , α2 , . . . , αn ,
respectively, and repeat the same argument). We have shown that in random selection of points
from (0, 1), the probability of the occurrence
of any particular point is 0. Now for t ∈ (0, 1),
let Bt = (0, 1) − {t}. Then P {t} = 0 implies that
P (Bt ) = P {t}c = 1 − P {t} = 1.
Therefore, there are infinitely many events, Bt ’s, each with probability 1 and none equal to the
sample space (0, 1).
1.7
RANDOM SELECTION OF POINTS FROM INTERVALS
In Section 1.6, we showed that the probability of the occurrence of any particular point in
a random selection of points from an interval (a, b) is 0. This implies immediately that if
[α, β] ⊆ (a, b), then the events that the point falls in [α, β], (α, β), [α, β), and (α, β] are
a + b a + b
a+b
and
, b ; since
all equiprobable. Now consider the intervals a,
2
2
2
is the midpoint of (a, b), it is reasonable to assume that
p1 = p2 ,
(1.5)
a + b
where p1 is the probability that the point belongs to a,
and p2 is the probability
2
a + b a + b
that it belongs to
, b . The events that the random point belongs to a,
and
2
2
a + b , b are mutually exclusive and
2
a + b ha + b a,
∪
, b = (a, b);
2
2
therefore,
p1 + p2 = 1.
This relation and (1.5) imply that
p1 = p2 = 1/2.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 31 — #47
✐
✐
Section 1.7
Random Selection of Points from Intervals
31
Hence the probability that a random point selected from (a, b) falls into the interval
ha + b a + b
is 1/2. The probability that it falls into
, b is also 1/2. Note that the
a,
2
2
length of each of these intervals is 1/2 of the length of (a, b). Now consider the intervals
2a + b i 2a + b a + 2b i
a + 2b 2a + b
a + 2b
a,
,
, and
,
, b . Since
and
are the
3
3
3
3
3
3
points that divide the interval (a, b) into three subintervals with equal lengths, we can assume
that
p1 = p2 = p3 ,
(1.6)
2a + b i
,
where p1 , p2 , and p3 are the probabilities that the point falls into a,
3
2a + b a + 2b i
a + 2b ,
, and
, b , respectively. On the other hand, these three
3
3
3
intervals are mutually disjoint and
Hence
2a + b i 2a + b a + 2b i a + 2b a,
∪
,
∪
, b = (a, b).
3
3
3
3
p1 + p2 + p3 = 1.
This relation and (1.6) imply that
p1 = p2 = p3 = 1/3.
Therefore, the probability that a random point selected from (a, b) falls into the interval
2a + b a + 2b i
2a + b i
is 1/3. The probability that it falls into
,
is 1/3, and the
a,
3
3
3
a + 2b probability that it falls into
, b is 1/3. Note that the length of each of these
3
intervals is 1/3 of the length of (a, b). These and other similar observations indicate that the
probability of the event that a random point from (a, b) falls into a subinterval (α, β) is equal
to (β − α)/(b − a).
Note that in this discussion we have assumed that subintervals of equal lengths are
equiprobable. Even though two subintervals of equal lengths may differ by a finite or a countably infinite set (or even a set of measure zero), this assumption is still consistent with our
intuitive understanding of choosing random points from intervals. This is because in such an
experiment, the probability of the occurrence of a finite or countably infinite set (or a set of
measure zero) is 0.
Thus far, we have based our discussion of selecting random points from intervals on our intuitive understanding of this experiment and not on a mathematical definition. Such discussions
are often necessary for the creation of appropriate mathematical meanings for unclear concepts.
The following definition, which is based on our intuitive analysis, gives an exact mathematical
meaning to the experiment of random selection of points from intervals.
Definition 1.2 A point is said to be randomly selected from an interval (a, b) if any two
subintervals of (a, b) that have the same length are equally likely to include the point. The
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 32 — #48
✐
✐
32
Chapter 1
Axioms of Probability
probability associated with the event that the subinterval (α, β) contains the point is defined to
be (β − α)/(b − a).
As explained before, choosing a random number from (0, 1) is equivalent to choosing randomly all the decimal digits of the number successively. Since in practice this is impossible,
choosing exact random points or numbers from (0, 1) or any other interval is only a theoretical matter. Approximate random numbers, however, can be generated by computers. Most of
the computer languages, some scientific computer software, and some calculators are equipped
with subroutines that generate approximate random numbers from intervals. However, since it
is difficult to construct good random number generators, there are computer languages, software programs, and calculators that are equipped with poor random number generator algorithms. An excellent reference for construction of good random number generators is The Art
of Computer Programming, Volume 2, Seminumerical Algorithms, third edition, by Donald
E. Knuth (Addison Wesley, 1998). Simple mechanical tools can also be used to find such approximations. For example, consider a spinner mounted on a wheel of unit circumference (radius 1/2π ). Let A be a point on the perimeter of the wheel. Each time that we flick the spinner,
it stops, pointing toward some point B on the wheel’s circumference. The length of the arc
AB (directed, say, counterclockwise) is an approximate random number between 0 and 1 if the
spinner does not have any “sticky” spots (see Figure 1.8).
Figure 1.8
Spinner, a model to generate random numbers.
Finally, when selecting a random point from an interval (a, b), we may think of an extremely large hypothetical box that contains infinitely many indistinguishable balls. Imagine
that each ball is marked by a number from (a, b), each number of (a, b) is marked on exactly
one ball, and the balls are completely mixed up, so that in a random selection of balls, any
two of them have the same chance of being drawn. With this transcendental model in mind,
choosing a random number from (a, b) is then equivalent to drawing a random ball from such
a box and looking at its number.
⋆ Example 1.22 (A Set that Is Not an Event) Let an experiment consist of selecting
a point at random from the interval [−1, 2]. We will construct a set that is not an event (i.e.,
it is impossible to associate a probability with the set). We begin by defining an equivalence
relation on [0, 1]: x ∼ y if x − y is a rational number. Let Q = {r1 , r2 , . . .} be the set
of rational numbers in [−1, 1]. Clearly, x is equivalent to y if x − y ∈ Q. The fact that
this relation is reflexive, symmetric, and transitive is trivial. Therefore, being an equivalence
relation, it partitions the interval [0, 1] into disjoint equivalence classes (Λα ). These classes are
such that if x and y ∈ Λα for some α, then x − y is rational. However, if x ∈ Λα and y ∈ Λβ ,
and α 6= β, then x − y is irrational. These observations imply that for each α, the equivalence
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 33 — #49
✐
✐
Section 1.7
Random Selection of Points from Intervals
33
S
class Λα is countable. Since α Λα = [0, 1] is uncountable, the number of equivalence classes
is uncountable. Let E be a set consisting of exactly one point from each equivalence class
Λα . The existence of such a set is guaranteed by the Axiom of Choice. We will show, by
contradiction, that E is not an event. Suppose that E is an event, and let p be the probability
associated with E . For each positive integer n, let En = {rn + x : x ∈ E} ⊆ [−1, 2]. For
each rn ∈ Q, En is simply a translation of E . Thus, for all n, the set En is also an event, and
P (En ) = P (E) = p.
S∞We now make two more observations: (1) For n 6= m, En ∩ Em = ∅, (2) [0, 1] ⊂
n=1 En . To prove (1), let t ∈ En ∩ Em . We will show that En = Em . If t ∈ En ∩ Em ,
then for some rn , rm ∈ Q, and x, y ∈ E, we have that t = rn + x = rm + y. That is,
x − y = rm − rn is rational, and x − y belongs to the same equivalence class. Since E has
exactly one point from each equivalence class, we must have x = y, hence rn = rm , hence
En = Em . To prove (2), let x ∈ [0, 1]. Then x ∼ y for some y ∈ E . This implies that x − y
is a rational
S∞ number in Q. That is, for some n, x − y = rn , or x = y + rn , or x ∈ En . Thus
x ∈ n=1 En .
Putting (1) and (2) together, we obtain
∞
[
En ≤ 1,
1/3 = P [0, 1] ≤ P
n=1
or
1/3 ≤
∞
X
P (En ) =
n=1
∞
X
n=1
p ≤ 1.
P∞
This is a contradiction because n=1 p is either 0 or ∞. Hence E is not an event, and we
cannot associate a probability with this set. EXERCISES
A
1.
A bus arrives at a station every day at a random time between 1:00 P.M. and 1:30 P.M.
What is the probability that a person arriving at this station at 1:00 P.M. will have to wait
at least 10 minutes?
2.
Past experience shows that every new book by a certain publisher captures randomly
between 4 and 12% of the market. What is the probability that the next book by this
publisher captures at most 6.35% of the market?
3.
Which of the following statements are true? If a statement is true, prove it. If it is false,
give a counterexample.
(a)
If A is an event with probability 1, then A is the sample space.
(b)
If B is an event with probability 0, then B = ∅.
4.
Let A and B be two events. Show that if P (A) = 1 and P (B) = 1, then
P (AB) = 1.
5.
A point is selected at random from the interval (0, 2000). What is the probability that it
is an integer?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 34 — #50
✐
✐
34
6.
7.
8.
9.
Chapter 1
Axioms of Probability
Suppose that a point is randomly selected from the interval (0, 1). Using Definition 1.2,
show that all numerals are equally likely to appear as the first digit of the decimal representation of the selected point.
1
For an experiment with sample space S = (0, 2), for n ≥ 1, let En = 0, 1 +
n
and P (En ) = (3n + 2)/5n. Find the probability of E = (0, 1]. Note that this is
not the experiment of choosing a point at random from the interval (0, 2) as defined in
Section 1.7.
A point is chosen at random from the interval (−1, 1). Let E1 be the event that it falls
in the interval (−1, 1/3], E2 be the event that it falls in the interval (−1, 1/9], E3 be
the event that it falls in the interval (−1, 1/27] and, in general,
i < ∞, Ei be
S∞ for 1 ≤
T∞
the event that the point is in the interval (−1, 1/3i ]. Find i=1 Ei and i=1 Ei .
For the experiment of choosing a point at random from the interval (0, 1), let
En = (1/2 − 1/2n, 1/2 + 1/2n), n ≥ 1.
T∞
(a) Prove that n=1 En = {1/2}.
(b)
Using part (a) and the continuity of probability function, show that the probability
of selecting 1/2 in a random selection of a point from (0, 1) is 0.
B
10.
Is it possible to define a probability on a countably infinite sample space so that the
outcomes are equally probable?
11.
Let A1 , A2 , . . . , An be n events. Show that if
P (A1 ) = P (A2 ) = · · · = P (An ) = 1,
then P (A1 A2 · · · An ) = 1.
12.
A point is selected at random from the interval (0, 1). What is the probability that it is
rational? What is the probability that it is irrational?
13.
Suppose that a point is randomly selected from the interval (0, 1). Using Definition 1.2,
show that all numerals are equally likely to appear as the nth digit of the decimal representation of the selected point.
P∞
Let {A1 , A2 , A3 , .. .} be a sequenceof events. Prove that if the series n=1 P (An )
T∞ S∞
converges, then P
= 0. This is called the Borel–Cantelli lemma.
m=1
n=m An
P∞
It says that if n=1 P (An ) < ∞, the probability that infinitely many of the An ’s occur
is 0.
S∞
Hint: Let Bm = n=m An and apply Theorem 1.8 to {Bm , m ≥ 1}.
14.
15.
Show that the result of Exercise 11 is not true for an infinite number of events. That is,
show that if {Et : 0 < t < 1} is a collection of events for which P (Et ) = 1, it is not
\
necessarily true that P
Et = 1.
t∈(0,1)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 35 — #51
✐
✐
Section 1.8
16.
What Is Simulation?
35
Let A be the set of rational numbers in (0, 1)
. Since A is countable, it can be written as a
sequence i.e., A = {rn : n = 1, 2, 3, . . .} . Prove that for any ε > 0, A can be covered
by a sequence of open balls whose total length is less than ε. That is, P
∀ε > 0, there exists
∞
a sequence of open intervals (αn , βn ) such that rn ∈ (αn , βn ) and n=1 (βn − αn ) <
ε. This important result explains why in a random selection of points from (0, 1) the
probability of choosing a rational is zero.
Hint: Let αn = rn − ε/2n+2 , βn = rn + ε/2n+2 .
Self-Quiz on Section 1.7
Time allotted: 20 Minutes
Each problem is worth 5 points.
1.
Suppose that for events A and B, P (AB) = 0. Does this imply that A and B are
mutually exclusive? Why or why not?
2.
For the experiment of choosing a point at random from the interval [0, 1], let En =
h1
3
−
1
1
2 i
, n ≥ 1. Applying the Continuity of Probability Function to
, +
n+2 3 n+2
En ’s, show that P (1/3 is selected) = 0.
1.8
WHAT IS SIMULATION?
Solving a scientific or an industrial problem usually involves mathematical analysis and/or
simulation. To perform a simulation, we repeat an experiment a large number of times to assess
the probability of an event or condition occurring. For example, to estimate the probability of at
least one 6 occurring within four rolls of a die, we may do a large number of experiments rolling
a die four times and calculate the number of times that at least one 6 is obtained. Similarly, to
estimate the fraction of time that, in a certain bank all the tellers are busy, we may measure the
lengths of such time intervals over a long period X, add them, and then divide by X . Clearly, in
simulations, the key to reliable answers is to perform the experiment a large number of times or
over a long period of time, whichever is applicable. Since manually this is almost impossible,
simulations are carried out by computers. Only computers can handle millions of operations in
short periods of time.
To simulate a problem that involves random phenomena, generating random numbers from
the interval (0, 1) is essential. In almost every simulation of a probabilistic model,we will need
to select random points from the interval (0, 1). For example, to simulate the experiment of
tossing a fair coin, we draw a random number from (0, 1). If it is in (0, 1/2), we say that the
outcome is heads, and if it is in [1/2, 1), we say that it is tails. Similarly, in the simulation of
die tossing, the outcomes 1, 2, 3, 4, 5, and 6, respectively, correspond to the events that the
random point from (0, 1) is in (0, 1/6), [1/6, 1/3), [1/3, 1/2), [1/2, 2/3), [2/3, 5/6), and
[5/6, 1).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 36 — #52
✐
✐
36
Chapter 1
Axioms of Probability
As discussed in Section 1.7, choosing a random number from a given interval is, in practice,
impossible. In real-world problems, to perform simulation we use pseudorandom numbers
instead. To generate n pseudorandom numbers from a uniform distribution on an interval (a, b),
we take an initial value x0 ∈ (a, b), called the seed, and construct a function ψ so that the
sequence {x1 , x2 , . . . , xn } ⊂ (a, b) obtained recursively from
xi+1 = ψ(xi ),
0 ≤ i ≤ n − 1,
(1.7)
satisfies certain statistical tests for randomness. (Choosing the tests and constructing the function ψ are complicated matters beyond the scope of this book.) The function ψ takes a seed
and generates a sequence of pseudorandom numbers in the interval (a, b). Clearly, in any pseudorandom number generating process, the numbers generated are rounded to a certain number
of decimal places. Therefore, ψ can only generate a finite number of pseudorandom numbers,
which implies that, eventually, some xj will be generated a second time. From that point on, by
(1.7), a pitfall is that the same sequence of numbers that appeared after xj ’s first appearance
will reappear. Beyond that point, numbers are not effectively random. One important aspect of
the construction of ψ is that the second appearance of any of the xj ’s is postponed as long as
possible.
It should be noted that passing certain statistical tests for randomness does not mean that the
sequence {x1 , x2 , . . . , xn } is a randomly selected sequence from (a, b) in the true mathematical sense discussed in Section 1.7. It is quite surprising that there are deterministic real-valued
functions ψ that, for each i, generate an xi+1 that is completely determined by xi , and yet the
sequence {x1 , x2 , . . . , xn } passes certain statistical tests for randomness.
For convenience, throughout this chapter, in practical problems, by random number we
simply mean pseudorandom number. In general, for good choices of ψ, the generated numbers
are usually sufficiently random for practical purposes.
Most of the computer languages and some scientific computer software are equipped with
subroutines that generate random numbers from intervals and from sets of integers. For example, in Mathematica, from Wolfram Research, Inc. (http://www.wolfram.com), the command
Random[ ] selects a random number from (0, 1), Random[Real, {a, b}] chooses a random
number from the interval (a, b), and Random [Integer, {m, m + n}] picks up an integer
from {m, m + 1, m + 2, . . . , m + n} randomly. However, there are computer languages that
are not equipped with a subroutine that generates random numbers, and there are a few that are
equipped with poor algorithms. As mentioned in Section 1.7, even though it is difficult to construct good pseudorandom number generators, an excellent reference for construction of such
numbers is The Art of Computer Programming, Volume 2, 3rd edition, by Donald E. Knuth.
At this point, let us emphasize that the main goal of scientists and engineers is always
to solve a problem mathematically. It is a mathematical solution that is accurate, exact, and
completely reliable. Simulations cannot take the place of a rigorous mathematical solution.
They are widely used (a) to find good estimations for solutions of problems that either cannot
be modeled mathematically or whose mathematical models are too difficult to solve; (b) to get
a better understanding of the behavior of a complicated phenomenon; and/or (c) to obtain a
mathematical solution by acquiring insight into the nature of the problem, its functions, and
the magnitude and characteristics of its solution. Intuitively, it is clear why the results that are
obtained by simulations are good. Theoretically, most of them can be justified by the strong
law of large numbers, which will be discussed in Section 11.4.
Chapter 13 of the Companion Website for this book concerns computer simulation. That
chapter is divided into several sections, presenting algorithms that are used to find approximate
solutions to complicated probabilistic problems. These sections can be studied independently
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 37 — #53
✐
✐
Chapter 1
Summary
37
when relevant materials from earlier chapters are being studied, or they can be learned concurrently.
CHAPTER 1 SUMMARY
◮ If the outcome of an experiment is not certain but all of its possible outcomes are predictable in advance, then the set of all these possible outcomes is called the sample space of the
experiment and is usually denoted by S . These outcomes are sometimes called sample points,
or simply points, of the sample space. Certain subsets of S are referred to as events. So events
are sets of points of the sample space. If the outcome of an experiment belongs to an event E,
we say that the event E has occurred.
◮ In the study of probability theory the relations between different events of an experiment
play a central role. The following are among the most important examples:
De Morgan’s first law:
c
c
c
(E ∪ F ) = E F ,
n
[
Ei
n
\
Ei
i=1
c
=
c
=
n
\
Eic ,
∞
[
Ei
∞
\
Ei
i=1
i=1
c
=
c
=
∞
\
Eic .
i=1
De Morgan’s second law:
(EF )c = E c ∪ F c ,
i=1
n
[
Eic ,
i=1
i=1
∞
[
Eic .
i=1
Another useful relation between E and F, two arbitrary events of a sample space S, is
E = EF ∪ EF c .
◮ If the joint occurrence of two events E and F is impossible, we say that E and F are
mutually exclusive. So E and F, are mutually exclusive if EF = ∅. A set of events
{E1 , E2 , . . .} is called mutually exclusive if the joint occurrence of any two of them is impossible, that is, if every pair of them is mutually exclusive.
◮ Probability Axioms: Let S be the sample space of a random phenomenon. Suppose
that to each event A of S, a number denoted by P (A) is associated with A. If P satisfies
the following axioms, then it is called a probability and the number P (A) is said to be the
probability of A.
Axiom 1
Axiom 2
Axiom 3
P (A) ≥ 0.
P (S) = 1.
If {A1 , A2 , A3 , . . .} is a sequence of mutually exclusive events (i.e., the joint
occurrence of every pair of them is impossible: Ai Aj = ∅ when i 6= j ), then
P
∞
[
i=1
∞
X
Ai =
P (Ai ).
i=1
Some immediate implications of the axioms of probability are the following:
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 38 — #54
✐
✐
38
Chapter 1
Axioms of Probability
• P (∅) = 0.
• Let {A1 , A2 , . . . , An } be a mutually exclusive set of events. Then
P
n
[
i=1
n
X
Ai =
P (Ai ).
i=1
• Let S be the sample space of an experiment. If S has N points that are all equally likely
to occur, then for any event A of S, P (A) = N (A)/N, where N (A) is the number of
points of A.
• For any event A, P (Ac ) = 1 − P (A).
• If A ⊆ B, then P (B − A) = P (BAc ) = P (B) − P (A).
• If A ⊆ B, then P (A) ≤ P (B).
• P (A ∪ B) = P (A) + P (B) − P (AB).
• P (A1 ∪ A2 ∪ A3 ) = P (A1 ) + P (A2 ) + P (A3 ) − P (A1 A2 ) − P (A1 A3 ) − P (A2 A3 )
+ P (A1 A2 A3 ).
• Inclusion-Exclusion Principle To calculate P (A1 ∪A2 ∪ · · · ∪An ), first find all of the
possible intersections of events from A1 , A2 , . . . , An and calculate their probabilities.
Then add the probabilities of those intersections that are formed of an odd number of
events, and subtract the probabilities of those formed of an even number of events. The
following formula is an expression for this principle.
P
n
[
i=1
n
n−1 X
n
n−2 X
n−1
X
X
X
Ai =
P (Ai ) −
P (Ai Aj ) +
i=1
i=1 j=i+1
n
X
P (Ai Aj Ak )
i=1 j=i+1 k=j+1
− · · · + (−1)n−1 P (A1 A2 · · · An ).
• P (A) = P (AB) + P (AB c ).
• (Continuity of Probability Function) For any increasing or decreasing sequence of
events, {En , n ≥ 1},
lim P (En ) = P ( lim En ).
n→∞
n→∞
◮ A point is said to be randomly selected from an interval (a, b) if any two subintervals
of (a, b) that have the same length are equally likely to include the point. The probability associated with the event that the subinterval (α, β) contains the point is defined to be
(β − α)/(b − a).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 39 — #55
✐
✐
Chapter 1
Review Problems
39
REVIEW PROBLEMS
1.
The number of minutes it takes for a certain animal to react to a certain stimulus is a
random number between 2 and 4.3. Find the probability that the reaction time of such
an animal to this stimulus is no longer than 3.25 minutes.
2.
Two dice are rolled. What is the event that the outcomes are consecutive?
3.
From a phone book, a phone number is selected at random. (a) What is the event that the
last digit is an odd number? (b) What is the event that the last digit is divisible by 3?
4.
Let P be the set of all subsets of A = {1, 2}. We choose two distinct sets randomly
from P . Define a sample space for this experiment, and describe the following events:
(a)
The intersection of the sets chosen at random is empty.
(b)
The sets are complements of each other.
(c)
One of the sets contains more elements than the other.
5.
In a tutoring center, there are three computers that can be up or down at any given time.
For 1 ≤ i ≤ 3, let Ei be the event that computer i is up at a random time. In terms of
Ei ’s, describe the event that at least two computers are up.
6.
Aiden just bought a stock for $320. Define a sample space for the price of this stock in
two years. Define the event that he makes money selling this stock at that time.
7.
In a certain experiment, whenever the event A occurs, the event B also occurs. Which
of the following statements is true and why?
(a)
If we know that A has not occurred, we can be sure that B has not occurred as
well.
(b)
If we know that B has not occurred, we can be sure that A has not occurred as
well.
8.
A department store accepts only its own credit card or an American Express card.
Customers not carrying one of these two cards must pay with cash. If 47% of the customers of this store carry American Express, 32% carry the store’s credit card, and 12%
carry both, what percentage of the store customers have no choice but to pay with cash?
9.
Kayla has two cars, an Audi and a Jeep. At a given time, whether these cars are operative
or totaled, due to accidents, is of concern to her insurance company. Define a sample
space for the driving conditions of these cars at a random time. What is the event that at
least one of the two cars is operative at that random time? Assume that a car that is not
totaled is operative.
10.
For a saw blade manufacturer’s products, the global demand, per month, for band saws is
between 30 and 36 thousands; for reciprocating saws, it is between 28 and 33 thousands;
for hole saws, it is between 300 and 600 thousands; and for hacksaws, it is between 500
and 650 thousands. Define a sample space for the demands for these four types of saws
by this manufacturer in a random month. Describe the event that, in such a month, the
demand for reciprocating saws is no more than 28.5 thousands and for hacksaws, it is
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 40 — #56
✐
✐
40
Chapter 1
Axioms of Probability
at least 570 thousands. In this exercise, by “between two numbers a and b”, we mean
inclusive of both a and b.
11.
The following relations are not always true. In each case give an example to refute
them.
(a)
(b)
P (A ∪ B) = P (A) + P (B).
P (AB) = P (A)P (B).
12.
A coin is tossed until, for the first time, the same result appears twice in succession.
Define a sample space for this experiment.
13.
In a midwest town, 80% of households have cable TV, 60% have an internet subscription,
and 90% have at least one of these. What percentage of the households of this town have
both cable TV and an internet subscription?
14.
15.
Let A and B be two events of an experiment
with P (A) = 1/3 and P (B) = 1/4. Find
the maximum value for P A ∪ B .
16.
Let A, B, and C be three events. Prove that
The number of the patients now in a hospital is 63. Of these 37 are male and 20 are for
surgery. If among those who are for surgery 12 are male, how many of the 63 patients
are neither male nor for surgery?
P (A ∪ B ∪ C) ≤ P (A) + P (B) + P (C).
17.
Let A, B, and C be three events. Show that
P (A ∪ B ∪ C) = P (A) + P (B) + P (C)
if and only if P (AB) = P (AC) = P (BC) = 0.
18.
Suppose that 40% of the people in a community drink or serve white wine, 50% drink
or serve red wine, and 70% drink or serve red or white wine. What percentage of the
people in this community drink or serve both red and white wine?
19.
Suppose that, in a temperate coniferous forest, 60% of randomly selected quarter-acre
plots have cedar trees, 45% have cypress trees, 30% have redwoods, 40% have both
cedar and cypress, 25% have cedar and redwoods, 20% have cypress and redwoods,
and 80% of such plots have at least one of these trees. What is the probability that in
a randomly selected quarter-acre plot in this forest we can find all three types of these
trees?
Suppose that P E ∪ F = 0.75 and P E ∪ F c = 0.85. Find P (E).
20.
21.
Anthony, Bob, and Carl, three American race car drivers, will compete in a professional
Trans-Am road race. Past records show that the probability that one of these former
champions wins is 10/13. If Anthony is twice as likely to win as Bob, and Bob is three
times more likely than Carl to win the race, find the respective probabilities of these
three people winning.
22.
Five customers enter a wireless corporate store to buy smartphones. If the probability
that at least two of them purchase an Android smartphone is 0.6 and the probability that
all of them buy non-Android smartphones is 0.17, what is the probability that exactly
one Android smartphone is purchased?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 41 — #57
✐
✐
Chapter 1
23.
Review Problems
41
Answer the following question, asked of Marilyn Vos Savant in the “Ask Marilyn”
column of Parade Magazine, March 3, 1996.
My dad heard this story on the radio. At Duke University, two students
had received A’s in chemistry all semester. But on the night before the
final exam, they were partying in another state and didn’t get back to
Duke until it was over. Their excuse to the professor was that they had a
flat tire, and they asked if they could take a make-up test. The professor
agreed, wrote out a test and sent the two to separate rooms to take it.
The first question (on one side of the paper) was worth 5 points, and they
answered it easily. Then they flipped the paper over and found the second
question, worth 95 points: ‘Which tire was it?’ What was the probability
that both students would say the same thing? My dad and I think it’s 1 in
16. Is that right?
24.
Let A and B be two events. Suppose that P (A), P (B), and P (AB) are given. What is
the probability that neither A nor B will occur?
25.
Let A and B be two events. The event (A − B) ∪ (B − A) is called the symmetric
difference of A and B and is denoted by A ∆ B . Clearly, A ∆ B is the event that exactly
one of the two events A and B occurs. Show that
P (A ∆ B) = P (A) + P (B) − 2P (AB).
26.
27.
Let S =
Suppose that
{ω1 , ω2 , ω3 , . . .} be the sample space of an experiment.
P {ω1 } = 1/8 and, for a constant k, 0 < k < 1, P {ωi+1 } = k · P {ωi }
for i ≥ 1. Find P {ωi } , for i > 1.
A bookstore receives six boxes of books per month on six random days of each month.
Suppose that two of those boxes are from one publisher, two from another publisher, and
the remaining two from a third publisher. Define a sample space for the possible orders
in which the boxes are received in a given month by the bookstore. Describe the event
that the last two boxes of books received last month are from the same publisher.
28.
Mildred, a commuter student studying at Western New England University, reported to
her advisor that there are three traffic lights that she needs to pass while driving from
home to school. She said that, based on her experience, 10% of the time all of the three
traffic lights are green, 35% of the time exactly two of them are green, and 43% of the
time exactly one of them is green. Assuming that Mildred’s observations are correct,
what is the probability that on a random day she encounters (a) no green traffic lights;
(b) at most one non-green (red or yellow) light; (c) at least two non-green lights?
29.
Suppose that in a certain town the number of people with blood type O and blood type
A are approximately the same. The number of people with blood type B is 1/10 of
those with blood type A and twice the number of those with blood type AB. Find the
probability that the next baby born in the town has blood type AB.
30.
A number is selected at random from the set of natural numbers {1, 2, 3, . . . , 1000}.
What is the probability that it is not divisible by 4, 7, or 9?
31.
A number is selected at random from the set {1, 2, 3, . . . , 150}. What is the probability
that it is relatively prime to 150? See Exercise 34, Section 1.4, for the definition of
relatively prime numbers.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 42 — #58
✐
✐
42
Chapter 1
32.
A point is selected randomly
from the interval (0, 2). For n ≥ 2, let En be the event
√
T∞
n
that it is in the interval (0, 3 ). Determine the event n=2 En .
33.
Axioms of Probability
Suppose that each day the price of a stock moves up 1/8 of a point, moves down 1/8 of
a point, or remains unchanged. For i ≥ 1, let Ui and Di be the events that the price of
the stock moves up and down on the ith trading day, respectively. In terms of Ui ’s and
Di ’s, find an expression for the event that the price of the stock
(a)
remains unchanged on the ith trading day;
(b)
moves up every day of the next n trading days;
(c)
remains unchanged on at least one of the next n trading days;
(d)
is the same as today after three trading days;
(e)
does not move down on any of the next n trading days.
34.
A bus traveling from Baltimore to New York breaks down at a random location. What
is the probability that the breakdown occurred after passing through Philadelphia? The
distances from New York and Philadelphia to Baltimore are, respectively, 199 and 96
miles.
35.
The coefficients of the quadratic equation ax2 + bx + c = 0 are determined by tossing
a fair die three times (the first outcome is a, the second one b, and the third one c). Find
the probability that the equation has no real roots.
Self-Test on Chapter 1
Time allotted: 120 Minutes
Each problem is worth 10 points.
1.
To determine who pays for dinner, Crispin, Allison, and Terry each flip an unbiased coin.
The one whose flip has a different face up will pay. If all of the flips land on the same
face, they start all over again. (a) What is the probability that Crispin ends up paying for
dinner; (b) what is the probability that no more than one round of flips is necessary?
2.
Every day, a major ball manufacturer produces at least 800, but no more than 1300 of
each of its five products: baseballs, tennis balls, softballs, basketballs, and soccer balls.
Define a sample space for the production levels of these five types of balls that this
manufacturer manufactures on a random working day. Describe the event that, on this
day, the difference between the numbers of baseballs and soccer balls produced will not
exceed 100.
Hint: Let x1 , x2 , x3 , x4 , and x5 be the production level, in hundreds, for baseballs,
tennis balls, softballs, basketballs, and soccer balls, respectively.
3.
A device that has three components fails if at least one of its components breaks down.
The device is observed at a random time. Let Ai , 1 ≤ i ≤ 3, denote the outcome that
the ith component is operative at the random time. In terms of Ai ’s, (a) define a sample
space for the status of the system at the random time; (b) describe the event that the
device is not operative at that random time.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 43 — #59
✐
✐
Chapter 1
Self-Test Problems
43
4.
Suppose that, in a particular geographical area, of people aged 55 and older, 13.4% suffer
from Parkinson’s disease, 11.3% suffer from Alzheimer’s disease, and 2.26% suffer from
both of these completely separate neurodegenerative illnesses. What percentage of the
people from this age group in this geographic area have neither Parkinson’s disease nor
Alzheimer’s disease?
5.
For an experiment, E and F are two events with P (E) = 0.4 and P (E c F c ) = 0.35.
Calculate P E c F .
6.
Last semester, all freshman students of a college took calculus, biology, and English. If
18% of them received an A in calculus, 10% received an A in both calculus and biology,
13% received an A in both English and calculus, and 7% received an A in all three
courses, what is the probability that a randomly selected freshman from this college, last
semester, received an A in calculus, but not in biology and not in English?
7.
For an experiment with sample space S = (0, 2), for n ≥ 1, let En =
1 − 1/n, 1 + 1/n and P (En ) = (2n + 1)/3n. For this experiment, find the probability that the event {1} occurs. Note that this is not the experiment of choosing a point
at random from the interval (0, 2), as defined in Section 1.7.
8.
Five customers enter a wireless corporate store to purchase smartphones. If the probability that at least three of them purchase an Android smartphone is 0.54, what is the
probability that at most two of them buy such a phone?
Hint: For 0 ≤ i ≤ 5, let AiP
be the event that exactly i of these customers buy an
5
Android smartphone. Note that i=0 P (Ai ) = 1.
9.
10.
In some towns of a country, the harmful substances lead and asbestos fibers find their
way into drinking water. In a study, it was found that, in 13% of those towns, the drinking
water supplies have neither lead nor asbestos fibers, in 32% of them the drinking water
supplies have lead, and in 43% the drinking water supplies have asbestos fibers. In what
percentage of the towns are the drinking water supplies contaminated with exactly one
of these two impurities?
Suppose that on a certain week, only one-third of the travelers visiting Paris took a trip
to the Eiffel Tower’s top, one-half visited the Louvre Museum, and one-third took a tour
of Notre Dame Cathedral. If none of the tourists visited all three sites, and for each pair
of the sites, one-fourth of the travelers visited that pair, find the fraction of this particular
group of tourists who did not visit any of the three sites.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 44 — #60
✐
✐
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 45 — #61
✐
✐
Chapter 2
C ombinatorial Methods
2.1
INTRODUCTION
The study of probability includes many applications, such as games of chance, occupancy and
order problems, and sampling procedures. In some of such applications, we deal with finite
sample spaces in which all sample points are equally likely to occur. Theorem 1.3 shows that,
in such cases, the probability of an event A is evaluated simply by dividing the number of
points of A by the total number of sample points. Therefore, some probability problems can be
solved simply by counting the total number of sample points and the number of ways that an
event can occur.
In this chapter we study a few rules that enable us to count systematically. A branch of
mathematics, combinatorial analysis, deals with methods of counting: a very broad field with
applications in virtually every branch of applied and pure mathematics. Besides probability
and statistics, it is used in information theory, coding and decoding, linear programming, transportation problems, industrial planning, scheduling production, group theory, foundations of
geometry, and other fields. Combinatorial analysis, as a formal branch of mathematics, began
with Tartaglia in the sixteenth century. After Tartaglia, Pascal, Fermat, Chevalier Antoine de
Méré (1607–1684), James Bernoulli, Gottfried Leibniz, and Leonhard Euler (1707–1783) made
contributions to this field. The mathematical development of the twentieth century accelerated
development by combinatorial analysis.
2.2
COUNTING PRINCIPLES
Suppose that there are n routes from town A to town B, and m routes from B to a third town,
C . If we decide to go from A to C via B, then for each route that we choose from A to B, we
have m choices from B to C . Therefore, altogether we have nm choices to go from A to C via
B . This simple example motivates the following principle, which is the basis of this chapter.
Theorem 2.1 (Counting Principle) If the set E contains n elements and the set F
contains m elements, there are nm ways in which we can choose, first, an element of E and
then an element of F .
Proof: Let E = {a1 , a2 , . . . , an } and F = {b1 , b2 , . . . , bm }; then the following rectangular
array, which consists of nm elements, contains all possible ways that we can choose, first, an
45
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 46 — #62
✐
✐
46
Chapter 2
Combinatorial Methods
element of E and then an element of F .
(a1 , b1 ),
(a2 , b1 ),
..
.
(a1 , b2 ), . . . , (a1 , bm )
(a2 , b2 ), . . . , (a2 , bm )
(an , b1 ), (an , b2 ), . . . , (an , bm ) Now suppose that a fourth town, D, is connected to C by ℓ routes. If we decide to go from
A to D, passing through C after B, then for each pair of routes that we choose from A to C,
there are ℓ possibilities from C to D . Therefore, by the counting principle, the total number
of ways we can go from A to D via B and C is the number of ways we can go from A to
C through B times ℓ, that is, nmℓ. This concept motivates a generalization of the counting
principle.
Theorem 2.2 (Generalized Counting Principle)
Let E1 , E2 , . . . , Ek be sets with
n1 , n2 , . . . , nk elements, respectively. Then there are n1 × n2 × n3 × · · · × nk ways in which
we can, first, choose an element of E1 , then an element of E2 , then an element of E3 , . . . , and
finally an element of Ek .
In probability, this theorem is used whenever we want to compute the total number of possible outcomes when k experiments are performed. Suppose that the first experiment has n1
possible outcomes, the second experiment has n2 possible outcomes, . . . , and the k th experiment has nk possible outcomes. If we define Ei to be the set of all possible outcomes of the
ith experiment, then the total number of possible outcomes coincides with the number of ways
that we can, first, choose an element of E1 , then an element of E2 , then an element of E3 , . . . ,
and finally an element of Ek ; that is, n1 × n2 × · · · × nk .
Example 2.1
How many outcomes are there if we throw five dice?
Solution: Let Ei , 1 ≤ i ≤ 5, be the set of all possible outcomes of the ith die. Then Ei =
{1, 2, 3, 4, 5, 6}. The number of the outcomes of throwing five dice equals the number of ways
we can, first, choose an element of E1 , then an element of E2 , . . . , and finally an element of
E5 . Thus we get 6 × 6 × 6 × 6 × 6 = 65 . Remark 2.1 Consider experiments such as flipping a fair coin several times, tossing a number of fair dice, drawing a number of cards from an ordinary deck of 52 cards at random and
with replacement, and drawing a number of balls from an urn at random and with replacement.
In Section 3.5, discussing the concept of independence, we will show that all the possible outcomes in such experiments are equiprobable. Until then, however, in all the problems dealing
with these kinds of experiments, we assume, without explicitly so stating in each case, that the
sample points of the sample space of the experiment under consideration are all equally likely.
Example 2.2
In tossing four fair dice, what is the probability of at least one 3?
Solution: Let A be the event of at least one 3. Then Ac is the event of no 3 in tossing the four
dice. N (Ac ) and N, the number of sample points of Ac and the total number of sample points,
respectively, are given by 5 × 5 × 5 × 5 = 54 and 6 × 6 × 6 × 6 = 64 . Therefore, P (Ac ) =
N (Ac )/N = 54 /64 . Hence P (A) = 1 − P (Ac ) = 1 − 625/1296 = 671/1296 ≈ 0.52.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 47 — #63
✐
✐
Section 2.2
Counting Principle
47
Example 2.3 Virginia wants to give her son, Brian, 14 different baseball cards within a
7-day period. If Virginia gives Brian cards no more than once a day, in how many ways can this
be done?
Solution: Each of the baseball cards can be given on 7 different days. Therefore, in 7 × 7 ×
· · · × 7 = 714 ≈ 6.78 × 1011 ways Virginia can give the cards to Brian. In experiments such as drawing cards from an ordinary deck of 52 cards or drawing balls
from an urn, if after each draw the card or the ball is put into the deck or into the urn, respectively, we say that cards or balls are drawn with replacement. If they are not returned into the
deck or into the urn, we say that they are drawn without replacement.
Example 2.4 A box contains 7 identical balls numbered 1 through 7. Three balls are drawn,
one by one, at random, and with replacement, and their numbers are recorded. What is the
probability that (a) all three outcomes are odd; (b) exactly one outcome is odd?
Solution: (a) Since the draws are with replacement, there are 7 possibilities for each draw.
Thus the sample space has 7 × 7 × 7 = 343 points. In 4 × 4 × 4 = 64 ways, all three outcomes
are odd. That is, each ball drawn has one of the numbers 1, 3, 5, or 7. So the probability that all
three outcomes are odd is 64/343 ≈ 0.187.
(b) In 4 × 3 × 3 = 36 ways, the first outcome is odd and the second and third outcomes are
even (2, 4, or 6). In 3 × 4 × 3 = 36 ways, the first and third outcomes are even, but the second
outcome is odd. Similarly, in 3×3×4 = 36 ways, the first two outcomes are even and the third
one is odd. So the probability that exactly one outcome is odd is (3 × 36)/343 ≈ 0.315. Example 2.5 A box contains 7 identical balls numbered 1 through 7. Three balls are drawn,
one by one, at random, and without replacement, and their numbers are recorded. What is the
probability that (a) all three outcomes are odd; (b) exactly one outcome is odd?
Solution: (a) Since the draws are without replacement, there are 7 possibilities for the first
draw, 6 possibilities for the second draw, and 5 possibilities for the third draw. Thus the sample
space has 7 × 6 × 5 = 210 points. Clearly, in 4 × 3 × 2 = 24 ways, all three outcomes are
odd. That is, each ball drawn has one of the numbers 1, 3, 5, or 7. So the probability that all
three outcomes are odd is 24/210 ≈ 0.114.
(b) In 4 × 3 × 2 = 24 ways, the first outcome is odd, and the second and third outcomes are
even. In 3 × 4 × 2 = 24 ways, the first and third outcomes are even, but the second outcome
is odd. Similarly, in 3 × 2 × 4 = 36 ways, the first two outcomes are even and the third one is
odd. So the probability that exactly one outcome is odd is (3 × 24)/210 ≈ 0.34. Example 2.6 Rose has invited n friends to her birthday party. If they all attend, and each one
shakes hands with everyone else at the party exactly once, what is the number of handshakes?
Solution 1: There are n + 1 people at the party and each of them shakes hands with the other
n people. This is a total of (n + 1)n handshakes, but that is an overcount since it counts “A
shakes hands with B” as one handshake and “B shakes hand with A” as a second. Since each
handshake is counted exactly twice, the actual number of handshakes is (n + 1)n/2.
Solution 2: Suppose that guests arrive one at a time. Rose will shake hands with all the n guests.
The first guest to appear will shake hands with Rose and all the remaining guests. Since we have
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 48 — #64
✐
✐
48
Chapter 2
Combinatorial Methods
already counted his or her handshake with Rose, there will be n − 1 additional handshakes.
The second guest will also shake hands with n − 1 fellow guests and Rose. However, we have
already counted his or her handshakes with Rose and the first guest. So this will add n − 2
additional handshakes. Similarly, the third guest will add n − 3 additional handshakes, and so
on. Therefore, the total number of handshakes will be n + (n − 1) + (n − 2) + · · · + 3 + 2 + 1.
Comparing solutions 1 and 2, we have the well-known relation
1 + 2 + 3 · · · + (n − 2) + (n − 1) + n =
n(n + 1)
. 2
Example 2.7 At a state university in Maryland, there is hardly enough space for students to
park their cars in their own lots. Jack, a student who parks in the faculty parking lot every day,
noticed that none of the last 10 tickets he got was issued on a Monday or on a Friday. Is it wise
for Jack to conclude that the campus police do not patrol the faculty parking lot on Mondays
and on Fridays? Assume that police give no tickets on weekends.
Solution: Suppose that the answer is negative and the campus police patrol the parking lot
randomly; that is, the parking lot is patrolled every day with the same probability. Let A be the
event that out of 10 tickets given on random days, none is issued on a Monday or on a Friday.
If P (A) is very small, we can conclude that the campus police do not patrol the parking lot
on these two days. Otherwise, we conclude that what happened is accidental and police patrol
the parking lot randomly. To find P (A), note that since each ticket has five possible days of
being issued, there are 510 possible ways for all tickets to have been issued. Of these, in only
310 ways no ticket is issued on a Monday or on a Friday. Thus P (A) = 310 /510 ≈ 0.006,
a rather small probability. Therefore, it is reasonable to assume that the campus police do not
patrol the parking lot on these two days. Example 2.8 (Standard Birthday Problem) What is the probability that at least two
students of a class of size n have the same birthday? Compute the numerical values of such
probabilities for n = 23, 30, 50, and 60. Assume that the birth rates are constant throughout
the year and that each year has 365 days.
Solution: There are 365 possibilities for the birthdays of each
Therefore, the
of the n students.
sample space has 365n points. In 365 × 364 × 363 × · · · × 365 − (n − 1) ways the birthdays
of no two of the n students coincide. Hence P (n), the probability that no two students have
the same birthday, is
365 × 364 × 363 × · · · × 365 − (n − 1)
,
P (n) =
365n
and therefore the desired probability is 1 − P (n). For n = 23, 30, 50, and 60 the answers are
0.507, 0.706, 0.970, and 0.995, respectively. Example 2.9 Terra was born on July 4. Assuming that the birth rates are constant throughout
the year and each year has 365 days, for Terra to have a 50% chance of meeting at least one
person with her birthday, how many random people does she need to meet?
Solution: Let n be the minimum number of people that Terra needs to meet to have a 50%
chance of meeting at least one person with her birthday. There are 365 possibilities for the
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 49 — #65
✐
✐
Section 2.2
Counting Principle
49
birthdays of each of the n random people she needs to meet. Therefore, the sample space
has 365n points. Of these, in 364n cases none of those people was born on July 4. So the
probability that none
shares a birthday with Terra is 364n /365n . Thus the probability is
n
n
1 − 364 /365 that at least one of the n people was born on July 4. We must have
1−
or
364n
≥ 0.5,
365n
364 n
≤ 0.5.
365
This gives n ln(364/365) ≤ ln(0.5), which implies that n ≥ 252.65. So Terra needs to meet
253 random people to have a 50% chance of meeting at least one person with her birthday,
July 4. Remark 2.2 In probability and statistics studies, birthday problems similar to Examples 2.8
and 2.9 have been very popular since 1939, when introduced by von Mises. This is probably
because when solving such problems, numerical values obtained are often surprising. Persi Diaconis and Frederick Mosteller, two Harvard professors, have mentioned that they “find the
utility of birthday problems impressive as a tool for thinking about coincidences.” Diaconis
and Mosteller have illustrated basic statistical techniques for studying the fascinating, curious,
and complicated “Theory of Coincidences” in the December 1989 issue of the Journal of the
American Statistical Association. In their study, they have used birthday problems “as examples which make the point that in many problems our intuitive grasp of the odds is far off.”
Throughout this book, where appropriate, we will present some interesting versions of these
problems, but now that we have cited “coincidence,” let us examine this concept, just a little,
in conjunction with the birthday problems.
To be certain that, among a group, two individuals share the same birthday, by the pigeonhole principle, we need 367 people in a room since, including February 29, there are exactly
366 birthdays. To have a 50% chance of meeting someone with your own birthday, you need
to meet with 253 random people (see Example 2.9). These and similar facts might mislead our
intuition and astonish us with birthday coincidences such as the ones that require only 23 random people in a room for having a chance exceeding 50%, and 50 random people for a chance
exceeding 97% that at least two individuals share a birthday.
Note that the probability that among, say, 23 or 50 random people at least one has your
birthday is very low, 0.061 and 0.128, respectively. However, the probability that among a
group of 23 random people or among a group of 50 random people, at least two individuals
share a birthday exceeds 0.50 and 0.97, respectively. Similarly, if you purchase a lottery ticket,
the probability that you win the jackpot is negligible (see Example 2.22). But the probability
that, in a lottery, someone wins the jackpot is high. Such observations show that the probability
of coincidences when events are defined more broadly or more generically can be quite high.
Clearly, billions of events occur everyday all over the world, and there are billions of ways for
possible coincidences. No one can ever list all the possibilities, but inevitably some of them
will occur, and when they do, we recognize them, and they confound our intuition. We find
them astonishing because before they occur, their chances of occurrence are extremely small,
one in billions. To elaborate this further, let us examine a simple case. Consider a boy who
was born in Khulna, Bangladesh, on February 2 of a given year, and a girl who was born on
July 9 of the same year in Fort Wayne, Indiana. What are the chances for these infants to grow
up both becoming interested in mathematics, both continuing their education in math to the
highest level and each earning a Ph.D., and then both ending up teaching at a specific university
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 50 — #66
✐
✐
50
Chapter 2
Combinatorial Methods
in Massachusetts in the same department? Obviously, the probability of such a sequence of
coincidences is negligible. However, two professors who were born in those towns on February
2 and July 9 of 1969, in fact, both ended up as professors of mathematics at the same university
in Massachusetts. At American universities, such events, which have negligible probabilities,
occur all the time. No surprises there since we actually know that each person’s life is formed
by an extremely long series of events, each with an infinitesimal probability. We end this remark
with a few words by Diaconis and Mosteller in the abstract of their aforementioned paper.
... when enormous numbers of events and people and their interactions cumulate over time, almost any outrageous event is bound to occur. These sources
account for much of the force of synchronicity. Number of Subsets of a Set
Let A be a set. The set of all subsets of A is called the power set of A. As an important
application of the generalized counting principle, we now prove that the power set of a set with
n elements has 2n elements. This important fact has lots of good applications.
Theorem 2.3
A set with n elements has 2n subsets.
Proof: Let A = {a1 , a2 , a3 , . . . , an } be a set with n elements. Then there is a one-to-one
correspondence between the subsets of A and the sequences of 0’s and 1’s of length n: To a
subset B of A we associate a sequence b1 b2 b3 · · · bn , where bi = 0 if ai 6∈ B, and bi = 1 if
ai ∈ B. For example, if n = 3, we associate to the empty subset of A the sequence 000, to
{a2 , a3 } the sequence 011, and to {a1 } the sequence 100. Now, by the generalized counting
principle, the number of sequences of 0’s and 1’s of length n is 2 × 2 × 2 × · · · × 2 = 2n .
Thus the number of subsets of A is also 2n . Example 2.10 A restaurant advertises that it offers over 1000 varieties of pizza. If, at the
restaurant, it is possible to have on a pizza any combination of pepperoni, mushrooms, sausage,
green peppers, onions, anchovies, salami, bacon, olives, and ground beef, is the restaurant’s
advertisement true?
Solution: Any combination of the 10 ingredients that the restaurant offers can be put on a
pizza. Thus the number of different types of pizza that it is possible to make is equal to the number of subsets of the set {pepperoni, mushrooms, sausage, green peppers, onions, anchovies,
salami, bacon, olives, ground beef}, which is 210 = 1024. Therefore, the restaurant’s advertisement is true. Note that the empty subset of the set of ingredients corresponds to a plain
cheese pizza. Tree Diagrams
Tree diagrams are useful pictorial representations that break down a complex counting problem
into smaller, more tractable ones. They are used in situations where the number of possible
ways an experiment can be performed is finite. The following examples illustrate how tree
diagrams are constructed and why they are useful. A great advantage of tree diagrams is that
they systematically identify all possible cases.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 51 — #67
✐
✐
Section 2.2
Counting Principle
51
Example 2.11 Bill and John keep playing chess until one of them wins two games in a row
or three games altogether. In what percent of all possible cases does the game end because Bill
wins three games without winning two in a row?
Figure 2.1
Tree diagram of Example 2.11.
Solution: The tree diagram of Figure 2.1 illustrates all possible cases. The total number of
possible cases is equal to the number of the endpoints of the branches, which is 10. The number
of cases in which Bill wins three games without winning two in a row, as seen from the figure,
is one. So the answer is 10%. Note that the probability of this event is not 0.10 because not all
of the branches of the tree are equiprobable. Example 2.12 Cheyenne has $4. She decides to bet $1 on the flip of a fair coin four times.
What is the probability that (a) she breaks even; (b) she wins money?
Solution: The tree diagram of the Figure 2.2 illustrates various possible outcomes for
Cheyenne. The diagram has 16 endpoints, showing that the sample space has 16 elements.
In six of these 16 cases, Cheyenne breaks even and in five cases she wins money, so the desired
probabilities are 6/16 and 5/16, respectively. Figure 2.2
Tree diagram of Example 2.12.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 52 — #68
✐
✐
52
Chapter 2
Combinatorial Methods
EXERCISES
A
1.
How many six-digit numbers are there? How many of them contain the digit 5? Note
that the first digit of an n-digit number is nonzero.
2.
How many different five-letter codes can be made using a, b, c, d, and e? How many of
them start with ab?
3.
The population of a town is 20,000. If each resident has three initials, is it true that at
least two people have the same initials?
4.
In how many different ways can 15 offices be painted with four different colors?
5.
In flipping a fair coin 23 times, what is the probability of all heads or all tails?
6.
In how many ways can we draw five cards from an ordinary deck of 52 cards (a) with
replacement; (b) without replacement?
7.
In a word game, there are 6 tiles in a bag, each bearing one of the letters A, C, E, G, L,
N. Zanya draws 8 tiles, one at a time, randomly, writes down the letter each time, and
then returns the tile to the bag. What is the probability that, in the order drawn, the letters
form the word ELEGANCE?
8.
Two fair dice are thrown. What is the probability that the outcome is a 6 and an odd
number?
9.
From an ordinary deck of 52 cards, 4 cards are selected at random and placed on a table
face down. A boy, who would like to be a magician, guesses the denomination and suit
of all four cards randomly. What is the probability that he guesses at least one card
correctly?
10.
Mr. Smith has 12 shirts, eight pairs of slacks, eight ties, and four jackets. Suppose that
four shirts, three pairs of slacks, two ties, and two jackets are blue. (a) What is the
probability that an all-blue outfit is the result of a random selection? (b) What is the
probability that he wears at least one blue item tomorrow?
11.
A multiple-choice test has 15 questions, each having four possible answers, of which
only one is correct. If the questions are answered at random, what is the probability of
getting all of them right?
12.
Suppose that in a state, license plates have three letters followed by three numbers, in
a way that no letter or number is repeated in a single plate. Determine the number of
possible license plates for this state.
13.
A library has 800,000 books, and the librarian wants to encode each by using a code
word consisting of three letters followed by two numbers. Are there enough code words
to encode all of these books with different code words?
14.
How many n × m arrays (matrices) with entries 0 or 1 are there?
15.
How many divisors does 55,125 have?
Hint: 55,125 = 32 53 72 .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 53 — #69
✐
✐
Section 2.2
Counting Principle
53
16.
A delicatessen has advertised that it offers over 500 varieties of sandwiches. If at this deli
it is possible to have any combination of salami, turkey, bologna, corned beef, ham, and
cheese on French bread with the possible additions of lettuce, tomato, and mayonnaise,
is the deli’s advertisement true? Assume that a sandwich necessarily has bread and at
least one type of meat or cheese.
17.
For an international traveler, the only Automated Teller Machine (ATM) personal identification numbers (PINs) acceptable are those that do not begin with 0 and those that
do not consist of identical digits. If PINs allowed can have 4, 5, or 6 digits, what is the
probability that a 4-, 5-, or 6-digit number selected at random can be used as a PIN?
18.
How many four-digit numbers can be formed by using only the digits 2, 4, 6, 8, and 9?
How many of these have some digit repeated?
19.
In a mental health clinic there are 12 patients. A therapist invites all these patients to join
her for group therapy. How many possible groups could she get?
20.
Suppose that four cards are drawn successively from an ordinary deck of 52 cards, with
replacement and at random. What is the probability of drawing at least one king?
21.
A campus telephone extension has four digits. How many different extensions with no
repeated digits exist? Of these, (a) how many do not start with a 0; (b) how many do not
have 01 as the first two digits?
22.
There are N types of drugs sold to reduce acid indigestion. A random sample of n drugs
is taken with replacement. What is the probability that brand A is included?
Figure 2.3
Islands and connecting bridges of Exercise 23.
23.
A salesperson covers islands A, B, . . . , I . These islands are connected by the bridges
shown in the Figure 2.3. While on an island, the salesperson takes one of the possible
bridges at random and goes to another one. She does her business on this new island
and then takes a bridge at random to go to the next one. She continues this until she
reaches an island for the second time on the same day. She stays there overnight and
then continues her trips the next day. If she starts her trip from island I tomorrow, in
what percent of all possible trips will she end up staying overnight again at island I ?
24.
In North America, in the World Series, which is the most important baseball event each
year, the American League champion team plays with the National League champion
team a series of games. The first team to win four games will be the winner of the
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 54 — #70
✐
✐
54
Chapter 2
Combinatorial Methods
World Series championship and is awarded the Commissioner’s Trophy. Assuring that
it is equally likely for the teams to win, find the probability that the series (a) lasts 4
games; (b) lasts 5 games.
25.
To log into their computer accounts, Alec and Mildred must create passwords that begin
with a letter followed by 5 to 7 letters or numbers. If Mildred creates a password herself,
but Alec uses a random password creation tool, what is the probability that they end up
with the same password?
26.
From an ordinary deck of 52 cards, 15 cards are drawn randomly and without replacement. What is the probability that the first four cards are all clubs?
27.
A fair die is tossed eight times. What is the probability that the eighth outcome is not a
repetition?
B
28.
In a large town, Kennedy Avenue is a long north-south avenue with many intersections.
A drunken man is wandering along the avenue and does not really know which way he
is going. He is currently at an intersection O somewhere in the middle of the avenue.
Suppose that, at the end of each block, he either goes north with probability 1/2, or he
goes south with probability 1/2. Draw a tree diagram to find the probability that, after
walking four blocks, (a) he is back at intersection O; (b) he is only one block away from
intersection O .
29.
An integer is selected at random from the set {1, 2, . . . , 1, 000, 000}. What is the probability that it contains the digit 5?
30.
How many divisors does a natural number N have?
Hint: A natural number N can be written as pn1 1 pn2 2 · · · pnk k , where p1 , p2 , . . . , pk are
distinct primes.
31.
In tossing four fair dice, what is the probability of tossing, at most, one 3?
32.
A delicatessen advertises that it offers over 3000 varieties of sandwiches. If at this deli
it is possible to have any combination of salami, turkey, bologna, corned beef, and ham
with or without Swiss and/or American cheese on French, white, or whole wheat bread,
and possible additions of lettuce, tomato, and mayonnaise, is the deli’s advertisement
true? Assume that a sandwich necessarily has bread and at least one type of meat or
cheese.
33.
One of the five elevators in a building leaves the basement with eight passengers and
stops at all of the remaining 11 floors. If it is equally likely that a passenger gets off at
any of these 11 floors, what is the probability that no two of these eight passengers will
get off at the same floor?
34.
The elevator of a four-floor building leaves the first floor with six passengers and stops
at all of the remaining three floors. If it is equally likely that a passenger gets off at any
of these three floors, what is the probability that, at each stop of the elevator, at least one
passenger departs?
35.
A number is selected randomly from the set {0000, 0001, 0002, . . . , 9999}. What is
the probability that the sum of the first two digits of the number selected is equal to the
sum of its last two digits?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 55 — #71
✐
✐
Section 2.3
36.
Permutations
55
What is the probability that a random r -digit number (r ≥ 3) contains at least one 0, at
least one 1, and at least one 2?
Self-Quiz on Section 2.2
Time allotted: 20 Minutes
Each problem is worth 2.5 points.
1.
In 2010, three unrelated faculty members, Professors Rodriguez, Bucs, and Beineke
ended up sitting on three adjacent seats in a row at the spring commencement ceremony
of a university, where faculty occupied the chairs randomly. These professors somehow got into a discussion about birthday problems, and they were astonished when they
learned that all three of them had the same birthday. Calculate the probability of such a
coincidence. Assume that the birth rates are constant throughout the year and that each
year has 365 days.
2.
The chair of the industry-academic partnership of a town invites all 12 members of the
board and their spouses to his house for a Christmas party. If a board member may attend
without his spouse, but not vice versa, how many different groups can the chair get?
3.
In a pharmacy, there are 8 unrelated people standing in line to pick up their prescriptions.
Before the pharmacist gives the patients their prescription medications, following the
pharmacy’s policy, she asks them the month and day of their birthdays. What is the
probability that at least two of the last three of these 8 people have the same birthday?
Assume that the birth rates are constant throughout the year and that each year has 365
days.
4.
Cyrus and 27 other students are taking a course in probability this semester. If their professor chooses eight students at random and with replacement to ask them eight different
questions, what is the probability that one of them is Cyrus?
2.3
PERMUTATIONS
To count the number of outcomes of an experiment or the number of possible ways an event
can occur, it is often useful to look for special patterns. Sometimes patterns help us develop
techniques for counting. Two simple cases in which patterns enable us to count easily are
permutations and combinations. We study these two patterns in this section and the next.
Definition 2.1
An ordered arrangement of r objects from a set A containing n objects
(0 < r ≤ n) is called an r -element permutation of A, or a permutation of the elements of
A taken r at a time. The number of r -element permutations of a set containing n objects is
denoted by n Pr .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 56 — #72
✐
✐
56
Chapter 2
Combinatorial Methods
By this definition, if three people, Brown, Smith, and Jones, are to be scheduled for job
interviews, any possible order for the interviews is a three-element permutation of the set
{Brown, Smith, Jones}.
If, for example, A = {a, b, c, d}, then ab is a two-element permutation of A, acd is a
three-element permutation of A, and adcb is a four-element permutation of A. The order in
which objects are arranged is important. For example, ab and ba are considered different twoelement permutations, abc and cba are distinct three-element permutations, and abcd and cbad
are different four-element permutations.
To compute n Pr , the number of permutations of a set A containing n elements taken r at
a time (1 ≤ r ≤ n), we use the generalized counting principle: Since A has n elements, the
number of choices for the first object in the r -element permutation is n. For the second object,
the number of choices is the remaining n − 1 elements of A. For the third one, the number of
choices is the remaining n − 2, . . . , and, finally, for the r th object the number of choices is
n − (r − 1) = n − r + 1. Hence
n Pr = n(n − 1)(n − 2) · · · (n − r + 1).
(2.1)
An n-element permutation of a set with n objects is simply called a permutation. The
number of permutations of a set containing n elements, n Pn , is evaluated from (2.1) by putting
r = n.
(2.2)
n Pn = n(n − 1)(n − 2) · · · (n − n + 1) = n!.
The formula n! (the number of permutations of a set of n objects) has been well known
for a long time. Although it first appeared in the works of Persian and Arab mathematicians in
the twelfth century, there are indications that the mathematicians of India were aware of this
rule a few hundred years before Christ. However, the surprise notation ! used for “factorial”
was introduced by Christian Kramp in 1808. He chose this symbol perhaps because n! gets
surprisingly large even for small numbers. For example, we have that 18! ≈ 6.402373 × 1015 ,
a number that, according to Karl Smith,† is greater than six times “the number of all words ever
printed.” Comparing with the number of all words ever printed, we immediately realize that
the number of ways an ordinary deck of 52 cards can be arranged, 52! ≈ 8.07 × 1067 , is truly
astronomical. It is roughly 1000 times the estimated volume of the Milky Way in cubic inches,
which is approximately 9.76×1064 . Note that the volume of the Milky Way is approximated by
the volume of a disk with diameter 100,000 and thickness 1000 light years. Assuming that each
side of an ordinary die is half an inch, the Milky Way fits approximately 9.76 × 1064 /(1/8) ≈
7.8×1065 ordinary dice. Now imagine a fictional galaxy named Megantic, which, by volume, is
103.46 times larger than the Milky Way. This galaxy fits approximately 103.46 × 7.8 × 1065 ≈
8.07×1067 ≈ 52! ordinary dice. Furthermore, imagine that Megantic galaxy is filled with wellmixed identical ordinary dice all of which are white except one that is red. Then the probability
that a random shuffle of an ordinary deck of 52 cards results in a particular given order is
approximately equal to the drawing of the red dice in a random selection of a die from the
Megantic galaxy. These facts show that the probability that a random shuffle of an ordinary
deck of 52 cards results in an order that any other such deck has ever been, any time in history,
is almost 0.
There is a popular alternative for relation (2.1). It is obtained by multiplying both sides of
(2.1) by (n − r)! = (n − r)(n − r − 1) · · · 3 · 2 · 1. We get
n Pr · (n − r)!
†
= n(n − 1)(n − 2) · · · (n − r + 1) · (n − r)(n − r − 1) · · · 3 · 2 · 1 .
Karl J. Smith, The Nature of Mathematics, 12th ed., Brooks/Cole, Belmont, Calif., 2012, p. 39.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 57 — #73
✐
✐
Section 2.3
Permutations
57
This gives n Pr · (n − r)! = n!. Therefore,
The number of r -element permutations of a set containing n objects is
given by
n!
.
(2.3)
n Pr =
(n − r)!
Note that for r = n, this relation implies that n Pn = n!/0!. But by (2.2), n Pn = n!. Therefore,
for r = n, to make (2.3) consistent with (2.2), we define 0! = 1.
Example 2.13 Three people, Brown, Smith, and Jones, must be scheduled for job interviews. In how many different orders can this be done?
Solution: The number of different orders is equal to the number of permutations of the set
{Brown, Smith, Jones}. So there are 3! = 6 possible orders for the interviews. Example 2.14 Suppose that two anthropology, four computer science, three statistics, three
biology, and five music books are put on a bookshelf with a random arrangement. What is the
probability that the books of the same subject are together?
Solution: Let A be the event that all the books of the same subject are together. Then P (A) =
N (A)/N, where N (A) is the number of arrangements in which the books dealing with the
same subject are together and N is the total number of possible arrangements. Since there are
17 books and each of their arrangements is a permutation of the set of these books, N = 17!.
To calculate N (A), note that there are 2!×4!×3!×3!×5! arrangements in which anthropology
books are first, computer science books are next, then statistics books, after that biology, and
finally, music. Also, there are the same number of arrangements for each possible ordering of
the subjects. Since the subjects can be ordered in 5! ways, N (A) = 5! × 2! × 4! × 3! × 3! × 5!.
Hence
P (A) =
5! × 2! × 4! × 3! × 3! × 5!
≈ 6.996 × 10−8 . 17!
Example 2.15 If five boys and five girls sit in a row in a random order, what is the probability
that no two children of the same sex sit together?
Solution: There are 10! ways for 10 persons to sit in a row. In order that no two of the same
sex sit together, boys must occupy positions 1, 3, 5, 7, 9, and girls positions 2, 4, 6, 8, 10, or
vice versa. In each case there are 5!×5! possibilities. Therefore, the desired probability is equal
to
2 × 5! × 5!
≈ 0.008. 10!
We showed that the number of permutations of a set of n objects is n!. This formula is
valid only if all of the objects of the set are distinguishable from each other. Otherwise, the
number of permutations is different. For example, there are 8! permutations of the eight letters
ST AN F ORD because all of these letters are distinguishable from each other. But the number of permutations of the 8 letters BERKELEY is less than 8! since the second, the fifth,
and the seventh letters in BERKELEY are indistinguishable. If in any permutation of these
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 58 — #74
✐
✐
58
Chapter 2
Combinatorial Methods
letters we change the positions of these three indistinguishable E ’s with each other, no new
permutations would be generated. Let us calculate the number of permutations of the letters
BERKELEY . Suppose that there are x such permutations and we consider any particular
one of them, say BY ERELEK . If we label the E ’s: BY E1 RE2 LE3 K so that all of the letters are distinguishable, then by arranging E ’s among themselves we get 3! new permutations,
namely,
BY E1 RE2 LE3 K BY E2 RE3 LE1 K
BY E1 RE3 LE2 K BY E3 RE1 LE2 K
BY E2 RE1 LE3 K BY E3 RE2 LE1 K
which are otherwise all the same. Therefore, if all the letters were different, then for each one of
the x permutations we would have 3! times as many. That is, the total number of permutations
would have been x × 3!. But since eight different letters generate exactly 8! permutations, we
must have x×3! = 8!. This gives x = 8!/3!. We have shown that the number of distinguishable
permutations of the letters BERKELEY is 8!/3!. This sort of reasoning leads us to the
following general theorem.
Theorem 2.4 The number of distinguishable permutations of n objects of k different types,
where n1 are alike, n2 are alike, . . . , nk are alike and n = n1 + n2 + · · · + nk , is
n!
.
n1 ! × n2 ! × · · · × nk !
Example 2.16
and three c’s?
How many different 10-letter codes can be made using three a’s, four b’s,
Solution: By Theorem 2.4, the number of such codes is 10!/(3! × 4! × 3!) = 4200.
Example 2.17 In how many ways can we paint 11 offices so that four of them will be painted
green, three yellow, two white, and the remaining two pink?
Solution: Let “ggypgwpygwy ” represent the situation in which the first office is painted
green, the second office is painted green, the third one yellow, and so on, with similar representations for other cases. Then the answer is equal to the number of distinguishable permutations
of “ggggyyywwpp,” which by Theorem 2.4 is 11!/(4! × 3! × 2! × 2!) = 69, 300. Example 2.18
three heads?
A fair coin is flipped 10 times. What is the probability of obtaining exactly
Solution: The set of all sequences of H (heads) and T (tails) of length 10 forms the sample
space and contains 210 elements. Of all these, those with three H, and seven T, are desirable.
But the number of distinguishable sequences with three H’s and seven T’s is equal to
10
.
3! × 7!
Therefore, the probability of exactly three heads is
10! .
210 ≈ 0.12.
3! × 7!
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 59 — #75
✐
✐
Section 2.3
Permutations
59
EXERCISES
A
1.
In the popular TV show Who Wants to Be a Millionaire, contestants are asked to sort
four items in accordance with some norm: for example, landmarks in geographical order,
movies in the order of date of release, singers in the order of date of birth. What is the
probability that a contestant can get the correct answer solely by guessing?
2.
In New York City, a subway train arrives at the platform, and 12 people enter a subway
car that has only 6 empty seats. In how many ways can these passengers be seated?
3.
How many permutations of the set {a, b, c, d, e} begin with a and end with c?
4.
How many different messages can be sent by five dashes and three dots?
5.
Cynthia’s husband is away Monday through Friday, and she does not know how to cook.
If every evening she dines at one of her 8 favorite restaurants randomly, and during that
period she dines at a restaurant at most once, what is the probability that, one evening
she will eat at Wilbraham Pizzeria, one of her favorites?
6.
A three-year old child, Sheridan, has 8 square tiles, each bearing one of the letters in her
name. If she arranges the tiles randomly in a row all facing in the correct direction, what
is the probability that they show her name?
7.
Robert has eight guests, two of whom are Jim and John. If the guests will arrive in a
random order, what is the probability that John will not arrive right after Jim?
8.
Let A be the set of all sequences of 0’s, 1’s, and 2’s of length 12.
(a)
How many elements are there in A?
(b)
How many elements of A have exactly six 0’s and six 1’s?
(c)
How many elements of A have exactly three 0’s, four 1’s, and five 2’s?
9.
Professor Haste is somewhat familiar with six languages. To translate texts from one
language into another directly, how many one-way dictionaries does he need?
10.
In an exhibition, 20 cars of the same style that are distinguishable only by their colors,
are to be parked in a row, all facing a certain window. If four of the cars are blue, three
are black, five are yellow, and eight are white, how many choices are there?
11.
There are 25 hard metal bunk beds in an open bay barracks at a military base. In how
many ways can newly recruited soldiers be assigned to the 50 beds in the barracks? If the
sergeant divides the soldiers into two groups of 25 each and asks one group to occupy
the lower bunks and the other the upper bunks of the bunk beds, in how many ways can
this be done?
12.
At various yard sales, a woman has acquired five forks, of which no two are alike. The
same applies to her four knives and seven spoons. In how many different ways can
three place settings be chosen if each place setting consists of exactly one fork, one
knife, and one spoon? Assume that the arrangement of the place settings on the table is
unimportant.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 60 — #76
✐
✐
60
Chapter 2
13.
In a conference, Dr. Richman’s lecture is related to Dr. Chollet’s and should not precede
it. If there are six more speakers, how many schedules could be arranged?
Warning: Dr. Richman’s lecture is not necessarily scheduled right after Dr. Chollet’s
lecture.
14.
A dancing contest has 11 competitors, of whom three are Americans, two are Mexicans,
three are Russians, and three are Italians. If the contest result lists only the nationality of
the dancers, how many outcomes are possible?
15.
Six fair dice are tossed. What is the probability that at least two of them show the same
face?
16.
(a) Find the number of distinguishable permutations of the letters M ISSISSIP P I .
(b) In how many of these permutations P ’s are together? (c) In how many I ’s are together? (d) In how many P ’s are together, and I ’s are together? (e) In a random order of
the letters M ISSISSIP P I, what is the probability that all S ’s are together?
17.
A fair die is tossed eight times. What is the probability of exactly two 3’s, three 1’s, and
three 6’s?
18.
The letters in the word SUPERCALIFRAGILISTICEXPIALIDOCIOUS are arranged
randomly. (a) How many of the distinguishable arrangements begin with G and end with
X? (b) What is the probability that the outcome begins with G and ends with X?
19.
In drawing nine cards with replacement from an ordinary deck of 52 cards, what is the
probability of three aces of spades, three queens of hearts, and three kings of clubs?
20.
In a Napa Valley winery, guests are invited to a tasting room, and, for each of them, an
affable staff member pours 6 types of wine, which are sold at different prices, in 6 small
3-ounce stylish glasses, at random and one at a time. What is the probability that the first
three glasses of wine they serve a visitor (a) are, in order, the three most expensive of
the 6 brands he or she tastes; (b) are the most expensive brands, but not necessarily in
order of their prices?
21.
At a party, n men and m women put their drinks on a table and go out on the floor to
dance. When they return, none of them recognizes his or her drink, so everyone takes a
drink at random. What is the probability that each man selects his own drink?
22.
There are 20 chairs in a room numbered 1 through 20. If eight girls and 12 boys sit on
these chairs at random, what is the probability that the thirteenth chair is occupied by a
boy?
23.
Five fair dice are tossed. What is the probability of getting a sum of 27?
24.
Even with all the new advanced technological methods to send signals, some ships still
use flags for that purpose. Each flag has its own significance as a signal, and each arrangement of two or more flags conveys a special message. If a ship has 8 flags and 4
flagpoles, how many signals can it send?
25.
There are 12 students in a class. What is the probability that their birthdays fall in 12
different months? Assume that all months have the same probability of including the
birthday of a randomly selected person.
26.
If we put five math, six biology, eight history, and three literature books on a bookshelf
at random, what is the probability that all the math books are together?
Combinatorial Methods
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 61 — #77
✐
✐
Section 2.3
Permutations
61
27.
One of the five elevators in a building starts with seven passengers and stops at nine
floors. Assuming that it is equally likely that a passenger gets off at any of these nine
floors, find the probability that at least two of these passengers will get off at the same
floor.
28.
Five boys and five girls sit at random in a row. What is the probability that the boys are
together and the girls are together?
29.
If n balls are randomly placed into n cells, what is the probability that each cell will be
occupied?
30.
A town has six parks. On a Saturday, six classmates, who are unaware of each other’s
decision, choose a park at random and go there at the same time. What is the probability
that at least two of them go to the same park? Convince yourself that this exercise is the
same as Exercise 15, only expressed in a different context.
31.
A list of all permutations of abcdef is put in alphabetical order. What is the 601st entry
in the list?
Hint: The first 5! entries all begin with a.
32.
A club of 136 members is in the process of choosing a president, a vice president, a
secretary, and a treasurer. If two of the members are not on speaking terms and do not
serve together, in how many ways can these four people be chosen?
B
33.
Let S and T be finite sets with n and m elements, respectively.
(a)
(b)
(c)
How many functions f : S → T can be defined?
If m ≥ n, how many injective (one-to-one) functions f : S → T can be defined?
If m = n, how many surjective (onto) functions f : S → T can be defined?
34.
A fair die is tossed eight times. What is the probability of exactly two 3’s, exactly three
1’s, and exactly two 6’s?
35.
Suppose that 20 sticks are broken, each into one long and one short part. By pairing them
randomly, the 40 parts are then used to make 20 new sticks. (a) What is the probability
that long parts are all paired with short ones? (b) What is the probability that the new
sticks are exactly the same as the old ones?
36.
At a party, 15 married couples are seated at random at a round table. What is the probability that all men are sitting next to their wives? Suppose that of these married couples,
five husbands and their wives are older than 50 and the remaining husbands and wives
are all younger than 50. What is the probability that all men over 50 are sitting next
to their wives? Note that when people are sitting around a round table, only their seats
relative to each other matters. The exact position of a person is not important.
37.
A box contains five blue and eight red balls. Jim and Jack start drawing balls from the
box, respectively, one at a time, at random, and without replacement until a blue ball is
drawn. What is the probability that Jack draws the blue ball?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 62 — #78
✐
✐
62
Chapter 2
Combinatorial Methods
Self-Quiz on Section 2.3
Time allotted: 20 Minutes
Each problem is worth 2.5 points.
1.
At a university, all phone numbers begin with 782. If the remaining four digits are
equally likely to be 0, 1, . . . , 9, what is the probability that a randomly selected faculty member has a phone number consisting of 7 distinct digits?
2.
In an Extrasensory Perception (ESP) experiment, five people are each asked to think of
a card within the suit of hearts in an ordinary deck of 52 cards. What is the probability
that (a) no two think of the same card? (b) At least two of them think of the same card?
3.
One of Isabella’s iPhone playlists has 70 songs, 15 of which are by Lionel Richie. She
sets the iPhone to play the songs on that playlist. However, before doing that, Isabella
checks the shuffle button of her iPhone, and the iPhones begins to play the playlist’s
songs in a random order without repeating a song. What is the probability that the first
Lionel Richie song played is the 7th song played?
4.
A psychologist specializes in dissocial personality disorder and, for each patient, she
lists in order of strength, the 4 strongest traits of impulsivity, recklessness, deceitfulness,
exploitative behavior, low conscientiousness, irresponsible behavior, and high negative
emotionality. Suppose that, among patients, these traits are found in equal proportions.
2.4
(a)
How many different possibilities are there for the list of a random patient?
(b)
What is the probability that the list of a random patient includes impulsivity and
high negative emotionality?
COMBINATIONS
In many combinatorial problems, unlike permutations, the order in which elements are arranged
is immaterial. For example, suppose that in a contest there are 10 semifinalists and we want to
count the number of possible ways that three contestants enter the finals. If we argue that there
are 10 × 9 × 8 such possibilities, we are wrong since the contestants cannot be ordered. If A,
B, and C are three of the semifinalists, then ABC, BCA, ACB, BAC, CAB, and CBA are
all the same event and have the same meaning: “A, B, and C are the finalists.” The technique
known as combinations is used to deal with such problems.
Definition 2.2
An unordered arrangement of r objects from a set A containing n objects
(r ≤ n) is called an r -element combination of A, or a combination of the elements of A taken
r at a time.
Therefore, two combinations are different only if they differ in composition. Let x be the
number of r -element combinations of a set A of n objects. If all the permutations of each
r -element combination are found, then all the r -element permutations of A are found. Since for
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 63 — #79
✐
✐
Section 2.4
63
Combinations
each r -element combination of A there are r! permutations and the total number of r -element
permutations is n Pr , we have
x · r! = n Pr .
Hence x · r! = n!/(n − r)!, so x = n!/[(n − r)! r!]. Therefore, we have shown that
The number of r -element combinations of n objects is given by
n!
.
n Cr =
(n − r)! r!
Historically, a formula equivalent to n!/[(n − r)! r!] turned up in the works of the Indian
mathematician Bhaskara II (1114–1185) in the middle of the twelfth century. Bhaskara II used
his formula to calculate the number of possible medicinal preparations using six ingredients.
Therefore, the rule for calculation of the number of r -element combinations of n objects has
been known for a long time.
It is worthwhile to observe that n Cr is the number of subsets of size r that can be constructed from a set of size n. By Theorem 2.3, a set with n elements has 2n subsets. Therefore,
of these 2n subsets, the number of those that have exactly r elements is n Cr .
!
n
Notation: By the symbol
(read: n choose r ) we mean the number of all r -element
r
combinations of n objects. Therefore, for r ≤ n,
!
n
n!
.
=
r! (n − r)!
r
n
Observe that
0
and
!
=
n
n
!
!
!
n
n
= 1 and
=
= n. Also, for any 0 ≤ r ≤ n,
1
n−1
!
!
n
n
=
r
n−r
!
n+1
=
r
!
!
n
n
+
.
r
r−1
(2.4)
These relations can be proved algebraically or verified combinatorially. Let us prove
(2.4) by a combinatorial
argument. Consider a set of n + 1 objects, {a1 , a2 , . . . , an , an+1 }.
!
n+1
There are
r -element combinations of this set. Now we separate these r -element
r
combinations into two disjoint classes: one class consisting of all r -element combinations of {a1 , a2 , . . . , an } and another consisting of all (r − 1)-element
combinations of
!
n
{a1 , a2 , . . . , an } attached to an+1 . The latter class contains
elements and the forr−1
!
n
mer contains
elements, showing that (2.4) is valid.
r
Example 2.19 In how many ways can two mathematics and three biology books be selected
from eight mathematics and six biology books?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 64 — #80
✐
✐
64
Chapter 2
Combinatorial Methods
!
8
Solution: There are
possible ways to select two mathematics books and
2
ways to select three biology books. Therefore, by the counting principle,
!
!
6!
8
6
8!
×
= 560
×
=
6! 2! 3! 3!
2
3
6
3
!
possible
is the total number of ways in which two mathematics and three biology books can be selected.
Example 2.20 A random sample of 45 instructors from different state universities were
selected randomly and asked whether they are happy with their teaching loads. The responses
of 32 were negative. If Drs. Smith, Brown, and Jones were among those questioned, what is
the probability that all three of them gave negative responses?
!
45
Solution: There are
different possible groups with negative responses. If three of them
32
are Drs. Smith, Brown, and Jones, the other 29 are from the remaining 42 faculty members
questioned. Hence the desired probability is
!
42
29
! ≈ 0.35. 45
32
Example 2.21 In a small town, 11 of the 25 schoolteachers are pro-life, eight are pro-choice,
and the rest are indifferent. A random sample of five schoolteachers is selected for an interview.
What is the probability that (a) all of them are pro-choice; (b) all of them have the same opinion?
Solution:
(a)
There are
25
5
Of these, only
(b)
!
different ways to select random samples of size 5 out of 25 teachers.
!
8
are all pro-choice. Hence the desired probability is
5
!
8
5
! ≈ 0.0011.
25
5
By an argument similar to part (a), the desired probability equals
!
!
!
11
8
6
+
+
5
5
5
!
≈ 0.0099. 25
5
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 65 — #81
✐
✐
Section 2.4
Combinations
65
Example 2.22 In Maryland’s lottery, players pick six different integers between 1 and 49,
order of selection being irrelevant. The lottery commission then randomly selects six of these
as the winning numbers. A player wins the grand prize if all six numbers that he or she has
selected match the winning numbers. He or she wins the second prize if exactly five, and the
third prize if exactly four of the six numbers chosen match with the winning ones. Find the
probability that a certain choice of a bettor wins the grand, the second, and the third prizes,
respectively.
Solution: The probability of winning the grand prize is
1
49
6
! =
1
.
13, 983, 816
The probability of winning the second prize is
!
!
6
43
5
1
258
1
! =
≈
,
13, 983, 816
54, 200
49
6
and the probability of winning the third prize is
!
!
6
43
4
2
1
13, 545
! =
≈
. 13, 983, 816
1032
49
6
Example 2.23 From an ordinary deck of 52 cards, seven cards are drawn at random and
without replacement. What is the probability that at least one of the cards is a king?
!
52
Solution: In
ways seven cards can be selected from an ordinary deck of 52 cards. In
7
!
48
of these, none of the cards selected is a king. Therefore, the desired probability is
7
!
48
7
! = 0.4496.
P (at least one king) = 1 − P (no kings) = 1 −
52
7
Warning: A common mistake is to calculate this and similar probabilities as follows: To make
sure that there
! is at least one king among the seven cards drawn, we will first choose a king;
4
there are
possibilities. Then we choose the remaining six cards from the remaining 51
1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 66 — #82
✐
✐
66
Chapter 2
cards; there are
Combinatorial Methods
!
51
possibilities for this. Thus the answer is
6
!
!
4
51
1
6
! = 0.5385.
52
7
This solution is wrong because it counts some of the possible outcomes several times. For
example, the hand KH , 5C , 6D , 7H , KD , JC , and 9S is counted twice: once when KH is selected
as the first card from the kings and 5C , 6D , 7H , KD , JC , and 9S from the remaining 51, and
once when KD is selected as the first card from the kings and KH , 5C , 6D , 7H , JC , and 9S from
the remaining 51 cards. Example 2.24 What is the probability that a poker hand is a full house? A poker hand
consists of five randomly selected cards from an ordinary deck of 52 cards. It is a full house if
three cards are of one denomination and two cards are of another denomination: for example,
three queens and two 4’s.
!
52
Solution: The number of different poker hands is
. To count the number of full houses,
5
let us call a hand of type (Q,4) if it has three queens and two 4’s, with similar representations
for other types of full houses. Observe that (Q,4) and (4,Q) are different full houses, and types
such as (Q,Q) and (K,K) do not exist. Hence there are 13!× 12 different types of full houses.
!
4
4
Since for every particular type, say (4,Q), there are
ways to select three 4’s and
3
2
ways to select two Q’s, the desired probability equals
!
!
4
4
13 × 12 ×
×
3
2
!
≈ 0.0014. 52
5
Example 2.25 Show that the number of different ways n indistinguishable objects can be
placed into k distinguishable cells is
!
!
n+k−1
n+k−1
=
.
n
k−1
Solution: Let the n indistinguishable objects be represented by n identical oranges, and the k
distinguishable cells be represented by k people. We want to count the number of different ways
that n identical oranges can be divided among k people. To do this, add k − 1 identical apples
to the oranges. Then take the n + k − 1 apples and oranges and line them up in some random
order. Give all of the oranges preceding the first apple to the first person, all of the oranges
between the first and second apples to the second person, all of the oranges between the second
and third apples to the third person, and so on. Note that if, for example, an apple appears in
the beginning of the line, then the first person does not receive any oranges. Similarly, if two
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 67 — #83
✐
✐
Section 2.4
Combinations
67
apples appear next to each other, say, at the ith and (i + 1)st positions, then the (i + 1)st person
does not receive any oranges. This process establishes a one-to-one correspondence between
the ways n identical oranges can be divided among k people, and the number of distinguishable
permutations of n + k − 1 apples and oranges, of which the n oranges are identical and the
k − 1 apples are identical. By Theorem 2.4, the answer to this problem is
!
!
n+k−1
n+k−1
(n + k − 1)!
=
. =
k−1
n
n! (k − 1)!
Example 2.26 Let n be a positive integer, and let x1 +x2 +· · ·+xk = n be a given equation.
A vector (x1 , x2 , . . . , xk ) satisfying x1 + x2 + · · · + xk = n is said to be a nonnegative integer
solution of the equation if for each i, 1 ≤ i ≤ k, xi is a nonnegative integer. It is said to be a
positive integer solution of the equation if for each i, 1 ≤ i ≤ k, xi is a positive integer.
(a)
How many distinct nonnegative integer solutions does the equation x1 +x2 +· · ·+xk =
n have?
(b)
How many distinct positive integer solutions does the equation x1 + x2 + · · · + xk = n
have?
Solution:
(a)
If we think of x1 , x2 , . . . , xk as k cells, then the problem reduces to that of dividing n
identical objects (namely, n 1’s) into k cells. Hence, by Example 2.25, the answer is
!
!
n+k−1
n+k−1
=
.
n
k−1
(b)
For 1 ≤ i ≤ k, let yi = xi −1. Then, for each positive integer solution (x1 , x2 , . . . , xk )
of x1 + x2 + · · · + xk = n, there is exactly one nonnegative integer solution
(y1 , y2 , . . . , yk ) of y1 + y2 + · · · + yk = n − k, and conversely. Therefore, the number of positive integer solutions of x1 + x2 + · · · + xk = n is equal to the number of
nonnegative integer solutions of y1 + y2 + · · · + yk = n − k, which, by part (a), is
!
!
!
(n − k) + k − 1
n−1
n−1
=
=
. n−k
n−k
k−1
In Example 2.26, we have established the following:
For positive integers n and k, the number of distinct nonnegative integer
solutions of the equation x1 + x2 + · · · + xk = n is given by
!
n+k−1
.
k−1
The number
of distinct positive integer solutions of this equation is
!
n−1
.
k−1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 68 — #84
✐
✐
68
Chapter 2
Combinatorial Methods
Example 2.27 A stockholder is considering whether to invest $150,000 in five new stocks,
each in multiples of $10,000. In how many ways can he invest in these stocks if (a) he is
determined to buy at least $10,000 of each stock; (b) he does not necessarily select to invest in
all five stocks?
Solution:
i. We have
For 1 ≤ i ≤ 5, suppose that the stockholder decides to buy $(10, 000)xi of stock
10, 000x1 + 10, 000x2 + · · · + 10, 000x5 = 150, 000.
So the answer to (a) is the number of positive integer
solutions
!
! of the equation x1 + x2 + · · · +
15 − 1
14
x5 = 15, which, by Example 2.26, is
=
= 1001. The answer to part (b) is
5−1
4
the number of nonnegative integer!
solutions !
of the equation x1 + x2 + · · · + x5 = 15, which,
15 + 5 − 1
19
by Example 2.26, is
=
= 3876. 5−1
4
Example 2.28 An absentminded professor wrote n letters and sealed them in envelopes
before writing the addresses on the envelopes. Then he wrote the n addresses on the envelopes
at random. What is the probability that at least one letter was addressed correctly?
Solution: The total number of ways that one can write n addresses on n envelopes is n!; thus
the sample space contains n! points. Now we calculate the number of outcomes in which at least
one envelope is addressed correctly. To do this, let Ei be the event that the ith letter is addressed
correctly; then E1 ∪ E2 ∪ · · · ∪ En is the event that at least one letter is addressed correctly.
To calculate P (E1 ∪ E2 ∪ · · · ∪ En ), we use the inclusion-exclusion principle. To do so we
must calculate the probabilities of all possible intersections of the events from E1 , . . . , En , add
the probabilities that are obtained by intersecting an odd number of the events, and subtract all
the probabilities that are obtained by intersecting an even number of the events. Therefore, we
need to know the number of elements of Ei ’s, Ei Ej ’s, Ei Ej Ek ’s, and so on. Now Ei contains
(n − 1)! points, because when the ith letter is addressed correctly, there are (n − 1)! ways to
address the remaining n − 1 envelopes. So P (Ei ) = (n − 1)!/n!. Similarly, Ei Ej contains
(n − 2)! points, because if the ith and j th envelopes are addressed correctly, the remaining
n − 2 envelopes can be addressed in (n − 2)! ways. Thus P (Ei Ej ) = (n − 2)!/n!. Similarly,
P (Ei Ej Ek ) = (n − 3)!/n!, and so!on. Now in computing P (E1 ∪ E2!∪ · · · ∪ En ), there
n
n
are n terms of the form P (Ei ),
terms of the form P (Ei Ej ),
terms of the form
2
3
P (Ei Ej Ek ), and so on. Hence
!
n (n − 2)!
(n − 1)!
−
+ ···+
P (E1 ∪ E2 ∪ · · · ∪ En ) = n
n!
2
n!
!
!
n
−
(n
−
1)
!
n
n
1
(−1)n−2
+ (−1)n−1
.
n−1
n!
n n!
This expression simplifies to
P (E1 ∪ E2 ∪ · · · ∪ En ) = 1 −
1
1
1
(−1)n−1
+ − + ··· +
.
2! 3! 4!
n!
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 69 — #85
✐
✐
Section 2.4
Remark:
Combinations
69
P∞
Note that since ex = n=0 (xn /n!), if n → ∞, then
∞
[
1
1
(−1)n−1
1
+ ···
P
Ei = 1 − + − + · · · +
2! 3! 4!
n!
i=1
1
1
1
(−1)n
= 1 − 1 − 1 + − + − ··· +
+ ···
2! 3! 4!
n!
∞
X
(−1)n
1
=1−
= 1 − ≈ 0.632.
n!
e
n=0
Hence even if the number of the envelopes is very large, there is still a very good chance for at
least one envelope to be addressed correctly. One of the most important applications of combinatorics is that the formula for the
r -element combinations of n objects enables us to find an algebraic expansion for (x + y)n .
Theorem 2.5
For any integer n ≥ 0,
!
n
X
n
(x + y)n =
xn−i y i .
i
i=0
(Binomial Expansion)
Proof: By looking at some special cases, such as
(x + y)2 = (x + y)(x + y) = x2 + xy + yx + y 2
and
(x + y)3 = (x + y)(x + y)(x + y)
= x3 + x2 y + yx2 + xy 2 + yx2 + xy 2 + y 2 x + y 3 ,
it should become clear that when we carry out the multiplication
(x + y)n = (x + y)(x + y) · · · (x + y),
|
{z
}
n times
n−i i
y , 0 ≤ i ≤ n. Therefore, all we have to do!is to find
!
n
n
n−i i
out how many times the term x y appears, 0 ≤ i ≤ n. This is seen to be
=
n−i
i
n−i i
because x y emerges only whenever the x’s of n−i of the n factors of (x+y) are multiplied
by the y ’s of the remaining i factors of (x + y). Hence
!
!
!
n
n
n
(x + y)n =
xn +
xn−1 y +
xn−2 y 2 + · · ·
0
1
2
!
!
n
n n
n−1
+
xy
+
y . n−1
n
we obtain only terms of the form x
Remark 2.3
ith term is
Note that, by Theorem 2.5, the expansion of (x + y)n has n + 1 terms, and the
!
n
xn−i+1 y i−1 , 1 ≤ i ≤ n + 1. i−1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 70 — #86
✐
✐
70
Chapter 2
Combinatorial Methods
The binomial coefficients n Cr have a long history. Chinese mathematicians were using
them as early as the late eleventh century. The famous mathematician and poet of Persia, Omar
Khayyam, also rediscovered them (twelfth century), as did Nasir ad-Din Tusi in the following
century; they were rediscovered in Europe in the sixteenth century by the German mathematician Stifel as well as by the Italian mathematicians Tartaglia and Cardano. But perhaps the most
exhaustive study of these numbers was made by Pascal in the seventeenth century, and for that
reason they are usually associated with him.
What is the coefficient of x2 y 3 in the expansion of (2x + 3y)5 ?
Example 2.29
Solution: Let u = 2x and v =!3y; then (2x + 3y)5 = (u + v)5 . The coefficient of u2 v 3 in
5
the expansion of (u + v)5 is
and u2 v 3 = (22 · 33 )x2 y 3 ; therefore, the coefficient of x2 y 3
3
!
5
in the expansion of (2x + 3y)5 is
(22 · 33 ) = 1080. 3
!
!
!
!
!
n
n
n
n
n
Example 2.30 Evaluate the sum
+
+
+
+ ··· +
.
0
1
2
3
n
!
n
Solution: A set containing n elements has
, 0 ≤ i ≤ n, subsets with i elements. So the
i
given expression is the total number of the subsets of a set of n elements, and therefore it equals
2n . A second way to see this is to note that, by the binomial expansion, the given expression
equals
!
n
X
n n−i i
1 1 = (1 + 1)n = 2n . i
i=0
!
!
!
!
n
n
n
n
Example 2.31 Evaluate the sum
+2
+3
+ ··· + n
.
1
2
3
n
Solution:
n
i
i
!
!
n!
n · (n − 1)!
n−1
=
=n
.
=i·
i! (n − i)!
(i − 1)! (n − i)!
i−1
So
!
!
!
!
n
n
n
n
+2
+3
+ ··· + n
1
2
3
n
"
!
!
!
!#
n−1
n−1
n−1
n−1
= n
+
+
+ ··· +
= n · 2n−1 ,
0
1
2
n−1
by Example 2.30.
X
n 2
2n
n
=
.
n
i
i=0
Solution: We show this by using a combinatorial argument. For an analytic proof, see Exercise 65.
Let A = {a1 , a2 , . . . , an } and B = {b1 , b2 , . . . , bn } be two disjoint sets. The number of subsets of
Example 2.32
Prove that
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 71 — #87
✐
✐
Section 2.4
Combinations
71
2n
A ∪ B with n elements is
. On the other hand, any subset of A ∪ B with n elements is the
n
union of a subset of A withi elements
and a subset of B with n − i elements for some 0 ≤ i ≤ n.
n
n
Since for each i there are
such subsets, we have that the total number of subsets of
i
n−i
n X
n
n
n
n
A ∪ B with n elements is
. But since
=
, we have the identity. i
n−i
n−i
i
i=0
Example 2.33 In this example, we present an intuitive proof for the inclusion-exclusion principle
explained in Section 1.4:
P
n
[
i=1
n
n−1
n
n−2
X
X X
X n−1
X
Ai =
P (Ai ) −
P (Ai Aj ) +
i=1
i=1 j=i+1
n
X
P (Ai Aj Ak )
i=1 j=i+1 k=j+1
− · · · + (−1)n−1 P (A1 A2 · · · An ).
S
Let S be the sample space. If an outcome ω ∈
/ ni=1 Ai , then ω ∈
/ Ai , 1 ≤ i ≤ n. Therefore,
Sn the
probability of ω is not added to either side of the preceding equation. Suppose that ω ∈ i=1 Ai ;
then ω belongs to k of the Ai ’s for some k, 1 ≤ k ≤ n. Now, to the left side of the equation the
probability of ω is added exactly once. We will show that the same is trueon the right. Clearly, the
k
first term on the right side contains the probability of ω exactly k =
times. The second term
1 k
k
subtracts this probability
times, the third term adds it back, this time
times, and so on
2
3
until the kth term. Depending
on
whether k is even or odd, the kth term either subtracts or adds
k
the probability of ω exactly
times. The remaining terms do not contain the probability of ω.
k
Therefore, the total number of times the probability of ω is added to the right side is
k
k
k
k−1 k
−
+
− · · · (−1)
1
2
3
k
X
k k X
k k−i
k
k k−i
i
=−
1 (−1) =
−
1 (−1)i
i
0
i
i=1
i=0
= 1 − (1 − 1)k = 1,
where the next-to-the-last equation follows from the binomial expansion. This establishes the
inclusion-exclusion principle. Example 2.34 Suppose that we want to distribute n distinguishable balls into k distinguishable
cells so that n1 balls are distributed into the first cell, n2 balls into the second cell, . . . , nk balls into
the kth cell, wheren1 +
n2 + n3 + · · · + nk = n. To count the number of ways that this is possible,
n
note that we have
choices to distribute n1 balls into the first cell; for each choice of n1 balls
n1
n − n1
in the first cell, we then have
choices to distribute n2 balls into the second cell; for each
n2
n − n1 − n2
choice of n1 balls in the first cell and n2 balls in the second cell, we have
choices
n3
to distribute n3 balls into the third cell; and so on. Hence by the generalized counting principle the
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 72 — #88
✐
✐
72
Chapter 2
Combinatorial Methods
total number of such distributions is
n
n − n1
n − n1 − n2
n − n1 − n2 − · · · − nk−1
···
n1
n2
n3
nk
=
(n − n1 )!
(n − n1 − n2 )!
n!
×
×
× ···
(n − n1 )! n1 ! (n − n1 − n2 )! n2 ! (n − n1 − n2 − n3 )! n3 !
×
(n − n1 − n2 − n3 − · · · − nk−1 )!
n!
=
(n − n1 − n2 − n3 − · · · − nk−1 − nk )! nk !
n1 ! n2 ! n3 ! · · · nk !
because n1 + n2 + n3 + · · · + nk = n and (n − n1 − n2 − · · · − nk )! = 0! = 1.
As an application of Example 2.34, we state a generalization of the binomial expansion
(Theorem 2.5) and leave its proof as an exercise (see Exercise 44).
Theorem 2.6
(Multinomial Expansion)
In the expansion of
(x1 + x2 + · · · + xk )n ,
the coefficient of the term xn1 1 xn2 2 xn3 3 · · · xnk k , n1 + n2 + · · · + nk = n, is
n!
.
n1 ! n2 ! · · · nk !
Therefore,
(x1 + x2 + · · · + xk )n =
X
n!
xn1 1 xn2 2 xn3 3 · · · xnk k .
n ! n2 ! · · · nk !
n1 +n2 +···+n =n 1
k
Note that the sum is taken over all nonnegative integers n1 , n2 , . . . , nk such that
n1 + n2 + · · · + nk = n.
EXERCISES
A
1.
Jim has 20 friends. If he decides to invite six of them to his birthday party, how many
choices does he have?
2.
Each state of the 50 in the United States has two senators. In how many ways may a
majority be achieved in the U.S. Senate? Ignore the possibility of absence or abstention.
Assume that all senators are present and voting.
3.
A panel consists of 20 men and 25 women. How many choices do we have to obtain a
jury of six men and six women from this panel?
4.
There are 24 eggs in a rectangular egg carton arranged in 4 rows of 6 eggs each. Three
eggs are selected at random from the carton to make an omelet. What is the probability
that the eggs are selected from the corners of the carton?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 73 — #89
✐
✐
Section 2.4
Combinations
73
5.
From an ordinary deck of 52 cards, five are drawn randomly. What is the probability of
drawing exactly three face cards?
6.
As part of an English department’s requirements, students must take 5 courses out of a
list of 9, of which 3 are advanced courses in Shakespeare. If it is required that at least one
of these 5 courses be a Shakespeare course, how many choices does an English student
have for this particular requirement?
7.
A random sample of n elements is taken from a population of size N without replacement. What is the probability that a fixed element of the population is included? Simplify
your answer.
8.
In a box of 12 fuses, there are 3 that are defective. If a quality control engineer tests 4 of
these fuses at random, what is the probability that he does not find any of the defective
fuses?
9.
Judy puts one piece of fruit in her child’s lunch bag every day. If she has three oranges
and two apples for the next five days, in how many ways can she do this?
10.
Ann puts at most one piece of fruit in her child’s lunch bag every day. If she has only
three oranges and two apples for the next eight lunches of her child, in how many ways
can she do this?
11.
Lili has 20 friends. Among them are Kevin and Gerry, who are husband and wife. Lili
wants to invite six of her friends to her birthday party. If neither Kevin nor Gerry will go
to a party without the other, how many choices does Lili have?
12.
In a binary code, each code word has 6 bits, each of which is 0 or 1. What is the probability that a random code word (a) has three 0’s and three 1’s? (b) Begins with two 0’s?
(c) Ends in three 1’s?
13.
In a game, there are 12 tiles in a bag, each bearing one of the numbers 1 through 12.
Lida draws 6 tiles at random. What is the probability that the largest number she picks
out is 9?
14.
In front of Ilaria’s office there is a parking lot with 13 parking spots in a row. When cars
arrive at this lot, they park randomly at one of the empty spots. Ilaria parks her car in
the only empty spot that is left. Then she goes to her office. On her return she finds that
there are seven empty spots. If she has not parked her car at either end of the parking
area, what is the probability that both of the parking spaces that are next to Ilaria’s car
are empty?
15.
In addition to many scientific experiments that astronauts should operate while on the
space station, they must make sure that the station is in perfect condition. Suppose that
there are 13 astronauts on board, and on a specific day, three astronauts are needed to
replace broken equipment, five to check other equipment, and five to clean. In how many
ways can these tasks be divided?
16.
Find the coefficient of x9 in the expansion of (2 + x)12 .
17.
Find the coefficient of x3 y 4 in the expansion of (2x − 4y)7 .
√ 13
Find the 9th term of the binomial expansion of 1 + x .
9
5
Find the 4th term of the binomial expansion of x +
.
x
18.
19.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 74 — #90
✐
✐
74
Chapter 2
20.
A team consisting of three boys and four girls must be formed from a group of nine boys
and eight girls. If two of the girls are feuding and refuse to play on the same team, how
many possibilities do we have?
21.
There are 23 girls and 19 boys lined up randomly for a play in a kindergarten classroom.
What is the probability that the 12th child in the line is a girl?
22.
A fair coin is tossed 10 times. What is the probability of (a) five heads; (b) at least five
heads?
23.
If five numbers are selected at random from the set {1, 2, 3, . . . , 20}, what is the probability that their minimum is larger than 5?
24.
In North America, in the World Series, which is the most important baseball event each
year, the American League champion team plays a series of games with the National
League champion team. The first team to win four games will be the winner of the
World Series championship and is awarded the Commissioner’s Trophy. Assuring that
it is equally likely for the teams to win, find the probability that the series (a) lasts 6
games; (b) lasts 7 games.
25.
A soccer coach assigns his players randomly, 1 as a goalie, 4 as defenders, 3 as midfielders, and 3 as forwards. (a) How many choices does he have? (b) If only one player
is experienced enough to play as a goalie, what is the probability that in the random
division, he ends up to be the goalie?
26.
From a faculty of six professors, six associate professors, ten assistant professors, and
twelve instructors, a committee of size six is formed randomly. What is the probability
that (a) there are exactly two professors on the committee; (b) all committee members
are of the same rank?
27.
In a medical study, 25 patients with a duodenal ulcer have volunteered to be treated by
an experimental drug. The researcher has decided to treat 18 of the patients, randomly
selected, by the new drug and to give the remaining 7 a placebo as controls. If the duodenal ulcers of 17 of these patients are caused by the bacterium H. pylori, and the rest
by nonsteroidal anti-inflammatory drugs such as aspirin, what is the probability that the
control group (a) includes at least one patient whose duodenal ulcer is caused by the
bacterium H. pylori? (b) Includes at least two such patients?
28.
A lake contains 200 trout; 50 of them are caught randomly, tagged, and returned. If,
again, we catch 50 trout at random, what is the probability of getting exactly five tagged
trout?
29.
A round-robin tournament is a competition in which each participant plays against
every other participant exactly once. In a round-robin tournament with n contestants,
what is the total number of possible outcomes? An outcome lists out the winner and the
loser in each competition.
30.
A round-robin tournament, as defined in the previous exercise, is a competition in which
each participant plays against every other participant exactly once. Using a round-robin
tournament with n + 1 contestants, give a combinatorial argument for the following
identity:
!
n+1
1 + 2 + ··· + n =
.
2
Combinatorial Methods
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 75 — #91
✐
✐
Section 2.4
31.
Find the values of
n
X
i=0
2i
Combinations
75
!
!
n
X
n
n
and
xi
.
i
i
i=0
32.
A fair die is tossed six times. What is the probability of getting exactly two 6’s?
33.
Suppose that 12 married couples take part in a contest. If 12 persons each win a prize,
what is the !
probability that from every couple one of them is a winner? Assume that all
24
of the
possible sets of winners are equally probable.
12
34.
How many all capital, different words of any length, meaningful or meaningless, can we
create from the letters of “JUSTICE” without repeating a letter?
35.
Poker hands are classified into the following 10 nonoverlapping categories in increasing
order of likelihood. Calculate the probability of the occurrence of each class separately.
Recall that a poker hand consists of five cards selected randomly from an ordinary deck
of 52 cards.
Royal flush: The 10, jack, queen, king, and ace of the same suit.
Straight flush: All cards in the same suit, with consecutive denominations except for the royal flush.
Four of a kind: Four cards of one denomination and one card of a second denomination: for example, four 8’s and a jack.
Full house: Three cards of one denomination and two cards of a second
denomination: for example, three 4’s and two queens.
Flush: Five cards all in one suit but not a straight or royal flush.
Straight: Cards of distinct consecutive denominations, not all in one suit:
for example, 3 of hearts, 4 of hearts, 5 of spades, 6 of hearts, and 7 of
clubs.
Three of a kind: Three cards of one denomination, a fourth card of a
second denomination, and a fifth card of a third denomination.
Two pairs: Two cards from one denomination, another two from a second denomination, and the fifth card from a third denomination.
One pair: Two cards from one denomination, with the third, fourth, and
fifth cards from a second, third, and fourth denomination, respectively: for
example, two 8’s, a king, a 5, and a 4.
None of the above.
36.
There are 12 nuts and 12 bolts in a box. If the contents of the box are divided between
two handymen, what is the probability that each handyman will get six nuts and six
bolts?
37.
A young businessman is obsessed with buying life insurance policies to protect his family. He decides to allocate $7200 a year toward all or some of the Term, Universal,
Indexed Universal, Survivorship Universal, and Variable Universal life insurance premiums. If he decides to pay premiums in units of $600 for each policy, how many choices
does he have (a) if he insists in spending the entire $7200 toward life insurance premiums; (b) if he does not necessarily want to spend the whole amount?
38.
A history professor who teaches three sections of the same course every semester decides
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 76 — #92
✐
✐
76
Chapter 2
Combinatorial Methods
to make several tests and use them for the next 10 years (20 semesters) as final exams.
The professor has two policies: (1) not to give the same test to more than one class in a
semester, and (2) not to repeat the same combination of three tests for any two semesters.
Determine the minimum number of different tests that the professor should prepare.
39.
A four-digit number is selected at random. What is the probability that its ones place is
less than its tens place, its tens place is less than its hundreds place, and its hundreds
place is less than its thousands place? Note that the first digit of an n-digit number is
nonzero.
40.
Using Theorem 2.6, expand (x + y + z)2 .
41.
What is the coefficient of x2 y 3 z 2 in the expansion of (2x − y + 3z)7 ?
42.
What is the coefficient of x3 y 7 in the expansion of (2x − y + 3)13 ?
43.
An ordinary deck of 52 cards is dealt, 13 each, to four players at random. What is the
probability that each player receives 13 cards of the same suit?
44.
Using induction, binomial expansion, and the identity
!
n!
n
(n − n1 )!
=
,
n1 ! n2 ! · · · nk !
n1 n2 ! n3 ! · · · nk !
prove the formula of multinomial expansion (Theorem 2.6).
45.
A staircase is to be constructed between M and N (see Figure 2.4). The distances from
M to L, and from L to N, are 5 and 2 meters, respectively. If the height of a step is
25 centimeters and its width can be any integer multiple of 50 centimeters, how many
different choices do we have?
Figure 2.4
Staircase of Exercise 45.
46.
Each state of the 50 in the United States has two senators. What is the probability that
in a random committee of 50 senators (a) Maryland is represented; (b) all states are
represented?
47.
(Newton–Pepys problem) In 1693, Samuel Pepys, an English Naval administrator
and member of the Parliament, corresponded with Isaac Newton concerning the following wager he intended to make: Which of the following trials is more likely than the
other two?
(a)
At least one 6 when 6 dice are tossed.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 77 — #93
✐
✐
Section 2.4
(b)
At least two 6’s when 12 dice are tossed.
(c)
At least three 6’s when 18 dice are tossed.
Combinations
77
Isaac Newton calculated the probabilities of these events correctly and provided Samuel
Pepys with correct answers. He also justified his answers intuitively; however, his justifications were invalid. Calculate the probability of the events above posed by Samuel
Pepys.
B
48.
Prove the binomial expansion formula
by!induction. !
!
n
n
n+1
Hint: Use the identity
+
=
.
k−1
k
k
49.
A class contains 30 students. What is the probability that there are six months each
containing the birthdays of two students, and six months each containing the birthdays
of three students? Assume that all months have the same probability of including the
birthday of a randomly selected person.
50.
In a closet there are 10 pairs of shoes. If six shoes are selected at random, what is
the probability of (a) no complete pairs; (b) exactly one complete pair; (c) exactly two
complete pairs; (d) exactly three complete pairs?
51.
An ordinary deck of 52 cards is divided into two equal sets randomly. What is the probability that each set contains exactly 13 red cards?
52.
In a box, there are 12 balls, identical in every way, except that they are numbered 1
through 12. We draw 6 balls at random and without replacement one by one. What is the
probability that the numbers on the balls appear in increasing order but not necessarily
consecutive?
Hint: There is a one-to-one correspondence between all the possible ways that
the numbers on the balls drawn are in increasing order and the set of subsets of
{1, 2, . . . , 12} with exactly 6 elements.
53.
A train consists of n cars. Each of m passengers (m > n) will choose a car at random
to ride in. What is the probability that (a) there will be at least one passenger in each car;
(b) exactly r (r < n) cars remain unoccupied?
54.
Suppose that n indistinguishable balls are placed at random into n distinguishable cells.
What is the probability that exactly one cell remains empty?
55.
At a departmental party, Dr. James P. Coughlin distributes 7 baby photos of 7 of his
colleagues among the guests with a list of the names of those colleagues. Each guest is
invited to identify the babies in the photos. What is the probability that a participant can
identify exactly 3 of the 7 photos correctly by purely random guessing?
56.
Show that
Hint:
!
!
!
n
n+1
n+r
+
+ ··· +
=
0
1
r
!
!
!
n
n+1
n
=
−
.
r
r
r−1
!
n+r+1
.
r
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 78 — #94
✐
✐
78
Chapter 2
57.
Prove that
Combinatorial Methods
!
!
!
!
!
n
n
n
k n
n n
−
+
− · · · + (−1)
+ · · · + (−1)
= 0.
0
1
2
k
n
58.
By a combinatorial argument, prove that for r ≤ n and r ≤ m,
!
! !
!
!
! !
n+m
m
n
m
n
m
n
=
+
+ ··· +
.
r
0
r
1
r−1
r
0
59.
Evaluate the following sum:
!
!
!
!
1 n
1
n
1 n
n
+
+ ··· +
.
+
2 1
3 2
n+1 n
0
60.
Suppose that five points are selected at random from the interval (0, 1). What is the
probability that exactly two of them are between 0 and 1/4?
Hint: For any point there are four equally likely possibilities: to fall into (0, 1/4),
[1/4, 1/2), [1/2, 3/4), and [3/4, 1).
61.
A lake has N trout, and t of them are caught at random, tagged, and returned. We catch
n trout at a later time randomly and observe that m of them are tagged.
(a)
Find PN , the probability of what we observed actually happening.
(b)
To estimate the number of trout in the lake, statisticians find the value of N that
maximizes PN . Such a value is called the maximum likelihood estimator of N .
Show that the maximum of PN is [nt/m], where by [nt/m] we mean the greatest
integer less than or equal to nt/m. That is, prove that the maximum likelihood
estimator of the number of trout in the lake is [nt/m].
Hint: Investigate for what values of N the probability PN is increasing and for what
values it is decreasing.
62.
63.
64.
Let n be a positive integer. A random sample of four elements is taken from the set
0, 1, 2, . . . , n , one at a time and with replacement. What is the probability that the
sum of the first two elements is equal to the sum of the last two elements?
For a given position with n applicants, m applicants are equally qualified and n − m
applicants are not qualified at all. Assume that a recruitment process is considered to
be fair if the probability that a qualified applicant is hired is 1/m, and the probability
that an unqualified applicant is hired is 0. One fair recruitment process is to interview all
applicants, identify the qualified ones, and then hire one of the m qualified applicants
randomly. A second fair process, which is more efficient, is to interview applicants in a
random order and employ the first qualified applicant encountered. For the fairness of
the second recruitment process,
(a)
present an intuitive argument;
(b)
give a rigorous combinatorial proof.
For k ≥ 1, n ≥ k, how many distinct positive integer vectors (x1 , x2 , . . . , xk ) satisfy
the inequality x1 + x2 + · · · + xk ≥ n?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 79 — #95
✐
✐
Section 2.4
65.
Combinations
79
Using the binomial theorem, calculate the coefficient of xn in the expansion of
(1 + x)2n = (1 + x)n (1 + x)n to prove that
!
!2
n
X
2n
n
=
.
n
i
i=0
For a combinatorial proof of this relation, see Example 2.32.
66.
67.
68.
An absentminded professor wrote n letters and sealed them in envelopes without writing
the addresses on them. Then he wrote the n addresses on the envelopes at random. What
is the probability that exactly k of the envelopes were addressed correctly?
Hint: Consider a particular set of k letters. Let M be the total number of ways that
only!these k letters can be addressed correctly. The desired probability is the quantity
P
n
i
M/n!; using Example 2.28, argue that M satisfies n−k
i=2 (−1) /i! = M/(n−k)!.
k
A fair coin is tossed n times. Calculate the probability of getting no successive heads.
Hint: Let xi be the number of sequences of H’s and T’s of length i with no successive
H’s. Show that xi satisfies xi = xi−1 + xi−2 , i ≥ 2, where x0 = 1 and x1 = 2. The
answer is xn /2n . Note that {xi }∞
i=1 is a Fibonacci-type sequence.
What is the probability that the birthdays of at least two students of a class of size n are
at most k days apart? Assume that the birthrates are constants throughout the year and
that each year has 365 days.
Self-Quiz on Section 2.4
Time allotted: 20 Minutes
Each problem is worth 2.5 points.
1.
For an experiment, Sheri, a neuroscientist, needs to select 5 rats from each of the 5
groups of rats being studied. If each group has 12 rats, how many choices does Sheri
have?
2.
A broadcasting company recruits college students to help with their Olympic coverage.
If they choose 4 athletes randomly from their 7 finalists, what is the probability that
(a) the 4 tallest of the finalists are selected? (b) Exactly 3 of the 4 athletes selected are
the tallest of the finalists? Assume that no two of the finalists are of the same exact
height.
3.
There are 8 chairs in a row, on the upper deck of a boat, attached to the floor by screws
and nails. A sociologist observed that from the first 4 passengers who sat in that row
no two sat next to each other. Based on this observation, can she conclude that, when
possible, passengers avoid taking seats next to occupied ones?
Hint: Suppose that passengers occupied the seats at random and calculate the probability of what the sociologist observed.
4.
For a chess tournament, 16 individuals are to be divided into 8 groups of 2 each. In how
many ways can this be done?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 80 — #96
✐
✐
80
Chapter 2
2.5
STIRLING’s FORMULA
Combinatorial Methods
To estimate n! for large values of n, the following formula is used. Discovered in 1730 by
James Stirling (1692–1770), it appeared in his book Methodus Differentialis (Bowyer, London,
1730). Also, independently, a slightly less general version of the formula was discovered by
De Moivre, who used it to prove the central limit theorem.
Theorem 2.7
(Stirling’s Formula)
n! ∼
√
2πn nn e−n ,
n!
where the sign ∼ means lim √
= 1.
n→∞
2πn nn e−n
Stirling’s formula, which usually gives excellent approximations in numerical computations,
is also
√ often nused
to prove theoretical problems. Note that although the ratio R(n) =
n!
2πn n e−n becomes 1 at ∞, it is still close to 1, even for very small values of n.
The following table shows this fact.
n
n!
√
2πn nn e−n
R(n)
1
2
5
8
10
12
1
2
120
40,320
3,628,800
479,001,600
0.922
1.919
118.019
39,902.396
3,598,695.618
475,687,486.474
1.084
1.042
1.017
1.010
1.008
1.007
Be aware that, even though limn→∞ R(n) is 1, the difference between n! and
increases as n gets larger. In fact, it goes to ∞ as n → ∞.
Example 2.35
√
2πn nn e−n
Approximate the value of 2n (n!)2 (2n)! for large n.
Solution: By Stirling’s formula, we have
√
2n (n!)2
2n 2πn(nn )2 e−2n
πn
= n . ∼√
2n
−2n
(2n)!
2
4πn (2n) e
EXERCISE
1.
2n
Use Stirling’s formula to approximate
n
!
3
(2n)!
1
for large n.
and 22n
(4n)! (n!)2
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 81 — #97
✐
✐
Chapter 2
Summary
81
CHAPTER 2 SUMMARY
◮ Generalized Counting Principle: Let E1 , E2 , . . . , Ek be sets with n1 , n2 , . . . , nk
elements, respectively. Then there are n1 × n2 × n3 × · · · × nk ways in which we can, first,
choose an element of E1 , then an element of E2 , then an element of E3 , . . . , and finally an
element of Ek .
◮ Number of Subsets of a Set: A set with n elements has 2n subsets.
◮ r -Element Permutations: An ordered arrangement of r objects from a set A containing n objects (0 < r ≤ n) is called an r -element permutation of A, or a permutation of the
elements of A taken r at a time. The number of r -element permutations of a set containing n
n!
objects is denoted by n Pr . We have that n Pr =
. An n-element permutation of a set
(n − r)!
with n objects is simply called a permutation. The number of permutations of a set containing
n elements, n Pn , is n!.
◮ The number of distinguishable permutations of n objects of k different types, where n1
are alike, n2 are alike, . . . , nk are alike and n = n1 + n2 + · · · + nk , is
n!
.
n1 ! × n2 ! × · · · × nk !
◮ r -Element Combinations: An unordered arrangement of r objects from a set A containing n objects (r ≤ n) is called an r -element combination of A, or a combination of the
elements of A taken r at a time. !
The number of r -element permutations
of a set containing n
!
n
n
n!
objects is denoted by n Cr or
. We have that n Cr ≡
.
=
r
(n − r)! r!
r
◮ Binomial Expansion For any integer n ≥ 0,
!
n
X
n n−i i
(x + y) =
x y.
i
i=0
n
◮ Multinomial Expansion In the expansion of
(x1 + x2 + · · · + xk )n ,
the coefficient of the term xn1 1 xn2 2 xn3 3 · · · xnk k , n1 + n2 + · · · + nk = n, is
n!
.
n1 ! n2 ! · · · nk !
Therefore,
(x1 + x2 + · · · + xk )n =
X
n!
xn1 1 xn2 2 xn3 3 · · · xnk k .
n
!
n
!
·
·
·
n
!
2
k
n1 +n2 +···+n =n 1
k
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 82 — #98
✐
✐
82
Chapter 2
Combinatorial Methods
Note that the sum is taken over all nonnegative integers n1 , n2 , . . . , nk such that
n1 + n2 + · · · + nk = n.
REVIEW PROBLEMS
1.
Albert goes to the grocery store to buy fruit. There are seven different varieties of fruit,
and Albert is determined to buy no more than one of any variety. How many different
orders can he place?
2.
In how many ways can we divide 9 toys among 3 children evenly?
3.
Virginia has 1 one-dollar bill, 1 two-dollar bill, 1 five-dollar bill, 1 ten-dollar bill, and 1
twenty-dollar bill. She decides to give some money to her son Brian without asking for
change. How many choices does she have?
4.
To enhance children’s STEM skills, a teacher gives each student 12 Lego bricks but asks
them to use only 8 of them to construct a toy. If no two of the 12 bricks are identical,
how many choices does each child have for his or her 8 Lego bricks?
5.
If four fair dice are tossed, what is the probability that they will show four different
faces?
6.
From the 10 points that are placed on a circumference, two are selected randomly. What
is the probability that they are adjacent?
7.
A father buys nine different toys for his four children. In how many ways can he give
one child three toys and the remaining three children two toys each?
8.
A window dresser has decided to display five different dresses in a circular arrangement.
How many choices does she have?
9.
A student club has 50 members, of which 9 have volunteered themselves to serve on the
executive committee. One of the volunteers, Ruth, is known to be a bully. If the faculty
members in charge of the club choose the executive committee members at random from
the volunteers, what is the probability that Ruth is included?
10.
In a Napa Valley winery, guests are invited to a tasting room, and, for each of them, an
affable staff member pours 6 types of wine, which are sold at different prices, in 6 small
3-ounce stylish glasses, at random and one at a time. What is the probability that the first
glass of wine they serve a visitor is the most expensive brand and the last glass is the
least expensive one?
11.
A list of all permutations of 13579 is put in increasing order. What is the 100th number
in the list?
Hint: The first 4! = 24 numbers all begin with 1.
12.
At a departmental party, Dr. James P. Coughlin distributes 4 baby photos of 4 of his
colleagues among the guests with a list of the names of those colleagues. Each guest is
invited to identify the babies in the photos. What is the probability that a participant can
identify exactly 2 of the 4 photos correctly by purely random guessing?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 83 — #99
✐
✐
Chapter 2
Review Problems
83
13.
In a soccer tryout, 25 players compete. The coach has decided to cut 8 of the players,
assign 1 as a goalie, 4 as defenders, 3 as midfielders, 3 as forwards, and the remaining 6
as substitutes. How many choices does he have?
14.
After each lecture, a professor assigns 10 problems. However, he grades only 4 of them
randomly. If Salina’s solution to only one of the recent assignment problems is wrong,
what is the probability that she still gets a perfect score?
15.
Judy has three sets of classics in literature, each set having four volumes. In how many
ways can she put them on a bookshelf so that books of each set are not separated?
16.
Korina is returning from a trip to the United Kingdom and is packing 12 wrapped chocolate bars of which 3 are Aero Caramels, 4 are Chomps, 3 are Duncans, and 2 are Freddos.
The bars of the same brand are indistinguishable. In how many distinguishable ways can
Korina pack these chocolate bars?
17.
Suppose that 30 lawn mowers, of which seven have defects, are sold to a hardware store.
If the store manager inspects six of the lawn mowers randomly, what is the probability
that he finds at least one defective lawn mower?
18.
One of Zoey’s iPhone playlists has 12 songs, 5 by Beyoncé, 3 by Adele, and 4 by Céline
Dion. She sets the iPhone to play the songs of that playlist. However, before doing that
Zoey checks the shuffle button of her iPhone, and the iPhone plays the playlist’s songs
in a random order, without repeating a song. What is the probability that all songs by
Beyoncé are played back to back?
19.
In how many ways can 23 identical refrigerators be allocated among four stores so that
one store gets eight refrigerators, another four, a third store five, and the last one six
refrigerators?
20.
From a group of 8 male and 16 female potential jurors, a jury of 12 is selected at random.
What is the probability that all males are included?
21.
In how many arrangements of the letters BERKELEY are all three E ’s adjacent?
22.
Bill and John play in a backgammon tournament. A player is the winner if he wins three
games in a row or four games altogether. In what percent of all possible cases does the
tournament end because John wins four games without winning three in a row?
23.
For his calculus final exam, Dr. Channing, whose classes have a reputation for being
very easy, gives his students 20 problems with their complete solutions, and tells them
that he will choose 10 of them randomly for the final exam. Students who solve all 10
problems correctly will get an A, those who solve 9 correctly will get a B, and those
who solve only 8 correctly will get a C. Dr. Channing told them that he does not give
partial credit, and whoever solves fewer than 8 problems correctly, will get an F. Dolly,
who has not studied calculus all semester, has no idea how to do any of the problems.
However, she finds enough time to memorize only the solutions to 12 of the problems.
What is the probability that she gets A, B, C and F, respectively?
24.
For a card game tournament, 16 individuals must be divided into 4 groups of 4 each. In
how many ways can this be done?
25.
In a small town, both of the accidents that occurred during the week of June 8, 1988,
were on Friday the 13th. Is this a good excuse for a superstitious person to argue that
Friday the 13th’s are inauspicious?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 84 — #100
✐
✐
84
Chapter 2
26.
How many eight-digit numbers without two identical successive digits are there?
27.
A broadcasting company recruits college students to help with their Olympic coverage.
Suppose that three of their finalists are female athletes, and the remaining four are male
athletes. If they choose four of them randomly, what is the probability that (a) two of
them are female and two are male? (b) Two of them are the tallest females and two of
them are the tallest males of the finalists? Assume that no two of the finalists are of the
same height.
28.
A palindrome is a sequence of characters that reads the same forward or backward.
For example, rotator, madam, Hannah, the German name Otto, and an Indian language,
Malayalam, are palindromes. So are the following expressions: “Put up,” “Madam I’m
Adam,” “Was it a cat I saw?” and these two sentences in Latin concerning St. Martin,
Bishop of Tours: “Signa te, signa; temere me tangis et angis. Roma tibi subito motibus
ibit amor.” (Respectively: Cross, cross yourself; you annoy and vex me needlessly.
Through my exertions, Rome, your desire, will soon be near.) Determine the number
of palindromes containing 11 characters that can be made with (a) no letters repeating
more than twice; (b) one letter repeating three times and no other letter more than twice
(the chemical term detartrated is one such palindrome); (c) one letter repeating three
times, two each repeating two times, and one repeating four times.
Historical Remark: It is said that the palindrome was invented by a Greek poet named
Sotades in the third century B.C. It is also said that Sotades was drowned by order of a
king of the Macedonian dynasty, the reigning Ptolemy, who found him a real bore.
29.
In a lottery, players pick 5 different integers between 1 and 47, and the order of selection
is irrelevant. The lottery commission then randomly selects 5 of these as the winning
numbers. A player wins the grand prize if all 5 numbers that he or she has selected match
the winning numbers. Nicole was born on April 12, 1946. Based on her birthday, she has
purchased all tickets that include the numbers 4, 12, and 46. What is the probability that
Nicole wins the grand prize?
30.
Suppose that four women and two men enter a restaurant and sit at random around a
table that has four chairs on one side and another four on the other side. What is the
probability that the men are not all sitting on one side?
31.
If five Americans, five Italians, and five Mexicans sit randomly at a round table, what is
the probability that the persons of the same nationality sit together?
32.
In a bridge game, each of the four players gets 13 random cards. What is the probability
that every player has an ace?
33.
An urn contains 15 white and 15 black balls. Suppose that 15 persons each draw two
balls blindfolded from the urn without replacement. What is the probability that each of
them draws one white ball and one black ball?
34.
In a lottery the tickets are numbered 1 through N . A person purchases n (1 ≤ n ≤ N )
tickets at random. What is the probability that the ticket numbers are consecutive? (This
is a special case of a problem posed by Euler in 1763.)
35.
The chair of an academic department needs to form 3 search committees each consisting
of 5 tenured full professors. If the department has 18 such faculty, how many choices
does the chair have if he decides that no one can serve on more than one search committee? Note that a set of 5 professors serving on one search committee is distinguishable
from the same professors serving on another search committee.
Combinatorial Methods
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 85 — #101
✐
✐
Chapter 2
Self-Test Problems
85
36.
An ordinary deck of 52 cards is dealt, 13 each, at random among A, B, C, and D. What
is the probability that (a) A and B together get two aces; (b) A gets all the face cards; (c)
A gets five hearts and B gets the remaining eight hearts?
37.
To test if a computer program works properly, we run it with 12 different datasets, using
four computers, each running three datasets. If the datasets are distributed randomly
among different computers, how many possibilities are there?
38.
A four-digit number is selected at random. What is the probability that its ones place
is greater than its tens place, its tens place is greater than its hundreds place, and its
hundreds place is greater than its thousands place? Note that the first digit of an n-digit
number is nonzero.
From the set of integers 1, 2, 3, . . . , 100000 a number is selected at random. What is
the probability that the sum of its digits is 8?
Hint:
Establish a one-to-one correspondence between the set of integers from
{1, 2, . . . , 100000} the sum of whose digits is 8, and the set of possible ways 8 identical
objects can be placed into 5 distinguishable cells. Then use Example 2.25.
!
2n
Show that for n = 1, 2, 3, . . . ,
< 4n .
n
39.
40.
Self-Test on Chapter 2
Time allotted: 120 Minutes
Each problem is worth 10 points.
1.
What is the probability that exactly two students of a class of size 23 have the same
birthday? Assume that the birth rates are constant throughout the year and each year has
365 days.
2.
In a ternary code, each code has 7 bits, each of which is 0, 1, or 2. What is the probability
that a random code ends in 2 and has exactly three 0’s?
3.
The card game, bridge, featuring two teams of two players each, is played with an ordinary deck of 52 cards. The cards are dealt among the players, 13 each, randomly. What
is the probability that exactly 3 of the aces are in the hands of one team?
4.
Possible passwords for a computer account are strings of length 6 of numbers, capital
letters, and small letters, with no repetition allowed (all capital letters are distinguishable from all small letters. So a string such as AaBz8b has no repetition). What is the
probability that a password generated for this computer, by a random password creation
tool, has 2 digits, 2 capital letters, and 2 small letters?
5.
Each state of the 50 in the United States has two senators. For each state, the senator
with greater longevity in the chamber is the senior senator, and the other one is the
junior senator of that state. Since no state has both senators up for election in the same
year, every state always has a senior senator and a junior senator. If a committee of 12
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 86 — #102
✐
✐
86
Chapter 2
Combinatorial Methods
U.S. senators is selected at random, find the probability that (a) no two senators are from
the same state; (b) there are 6 senior senators and 6 junior senators, no two from the
same state, on the committee.
6.
There are 15 trees in a row of which 5 adjacent ones are infected by a fungal disease. Is
this logical evidence for an arborist to conclude that the fungus spread from one tree to
next through root-to-root contact?
Hint: Suppose that the disease strikes randomly, say, through airborne spores, and
calculate the probability that five adjacent trees become infected.
7.
Give a combinatorial proof for the relation
! !
!
!
n
k
n
n−1
=
.
k
1
1
k−1
Using this relation, show that
n
k
!
!
n n−1
.
=
k k−1
8.
There are 48 students to be randomly distributed among 3 different jewelry classes, 16
students per class. If 3 of the students are visually impaired, find the probability that
(a) each class gets one of the them; (b) all 3 end up in the same class.
9.
Let ℓ < k < m < n be positive integers. There will be n new movies released next
month. Suppose that a prominent movie critic, Mr. Wilbert, whose reviews are syndicated to more than 200 newspapers, is required to review exactly m movies next month
for publication.
10.
(a)
How many choices does he have?
(b)
If k of the new releases are specified, and Mr. Wilbert is required to review them
as part of his m reviews, how many choices does he have?
(c)
If k of the new releases are specified, and Mr. Wilbert is required to review at
least ℓ of them as part of his m reviews, how many choices does he have?
From a faculty of six professors, six associate professors, ten assistant professors, and
twelve instructors, a committee of size six is formed randomly. What is the probability
that there is at least one person from each rank on the committee?
Hint: Be careful, the answer is not
! !
!
!
!
6
6
10
12
30
1
1
1
1
2
!
= 1.397.
34
6
To find the correct answer, use the inclusion-exclusion principle explained in Section 1.4: Let A1 , A2 , A3 , and A4 be the events that there is no professor, no associate
professor, no assistant professor, and no instructor in the committee, respectively. The
desired probability is
P (Ac1 Ac2 Ac3 Ac4 ) = 1 − P (A1 ∪ A2 ∪ A3 ∪ A4 ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 87 — #103
✐
✐
Chapter 3
C onditional Probability
and I ndependence
3.1
CONDITIONAL PROBABILITY
To introduce the notion of conditional probability, let us first examine the following question:
Suppose that all of the freshmen of an engineering college took calculus and discrete math
last semester. Suppose that 70% of the students passed calculus, 55% passed discrete math,
and 45% passed both. If a randomly selected freshman is found to have passed calculus last
semester, what is the probability that he or she also passed discrete math last semester? To
answer this question, let A and B be the events that the randomly selected freshman passed
discrete math and calculus last semester, respectively. Note that the quantity we are asked to
find is not P (A), which is 0.55; it would have been if we were not aware that B has occurred.
Knowing that B has occurred changes the chances of the occurrence of A. To find the desired
probability, denoted by the symbol P (A | B) [read: probability of A given B ] and called
the conditional probability of A given B, let n be the number of all the freshmen in the
engineering college. Then (0.7)n is the number of freshmen who passed calculus, and (0.45)n
is the number of those who passed both calculus and discrete math. Therefore, of the (0.7)n
freshmen who passed calculus, (0.45)n of them passed discrete math as well. It is given that
the randomly selected student is one of the (0.7)n who passed calculus; we also want to find
the probability that he or she is one of the (0.45)n who passed discrete math. This is obviously
equal to (0.45)n/(0.7)n = 0.45/0.7. Hence
P (A | B) =
0.45
.
0.7
But 0.45 is P (AB) and 0.7 is P (B). Therefore, this example suggests that
P (A | B) =
P (AB)
.
P (B)
(3.1)
As another example, suppose that a bus arrives at a station every day at a random time
between 1:00 P.M. and 1:30 P.M. Suppose that at 1:10 the bus has not arrived and we want to
calculate the probability that it will arrive in not more than 10 minutes from now. Let A be
the event that the bus will arrive between 1:10 and 1:20 and B be the event that it will arrive
87
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 88 — #104
✐
✐
88
Chapter 3
Conditional Probability and Independence
between 1:10 and 1:30. Then the desired probability is
20 − 10
P (AB)
20 − 10
30 − 0
=
P (A | B) =
=
.
30 − 10
P (B)
30 − 10
30 − 0
(3.2)
The relation P (A | B) = P (AB)/P (A) discovered in (3.1) and (3.2) can also be verified
for other types of conditional probability problems. It is compatible with our intuition and is a
logical generalization of the existing relation between the conditional relative frequency of the
event A with respect to the condition B and the ratio of the relative frequencies of AB and the
relative frequency of B . For these reasons it is taken as a definition.
Definition 3.1
P (A | B), is
If P (B) > 0, the conditional probability of A given B, denoted by
P (A | B) =
P (AB)
.
P (B)
(3.3)
If P (B) = 0, formula (3.3) makes no sense, so the conditional probability is defined
only for P (B) > 0. This formula may be expressed in words by saying that the conditional
probability of A given that B has occurred is the ratio of the probability of joint occurrence
of A and B and the probability of B . It is important to note that (3.3) is neither an axiom
nor a theorem. It is a definition. As explained previously, this definition is fully justified and
is not made arbitrarily. Indeed, historically, it was used implicitly even before it was formally
introduced by De Moivre in the book The Doctrine of Chance.
Example 3.1 In a certain region of Russia, the probability that a person lives at least 80 years
is 0.75, and the probability that he or she lives at least 90 years is 0.63. What is the probability
that a randomly selected 80-year-old person from this region will survive to become 90?
Solution: Let A and B be the events that the person selected survives to become 90 and 80
years old, respectively. We are interested in P (A | B). By definition,
P (A | B) =
P (AB)
.
P (B)
Now, P (AB) = P (A) because “at least 80” and “at least 90” overlap with “at least 90,” which
is A. Hence
P (AB)
P (A)
0.63
P (A | B) =
=
=
= 0.84. P (B)
P (B)
0.75
Example 3.2 From the set of all families with two children, a family is selected at random
and is found to have a girl. What is the probability that the other child of the family is a girl?
Assume that in a two-child family all sex distributions are equally probable.
Solution: Let B and A be the events that the family has a girl and the family has two girls,
respectively. We are interested in P (A | B). Now, in a family with two children there are
four equally likely possibilities: (boy, boy), (girl, girl), (boy, girl), (girl, boy), where by, say,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 89 — #105
✐
✐
Section 3.1
Conditional Probability
89
(girl, boy), we mean that the older child is a girl and the younger is a boy. Thus P (B) =
3/4, P (AB) = 1/4. Hence
P (A | B) =
1/4
1
P (AB)
=
= . P (B)
3/4
3
Example 3.3 From the set of all families with two children, a child is selected at random
and is found to be a girl. What is the probability that the second child of this girl’s family is
also a girl? Assume that in a two-child family all sex distributions are equally probable.
Solution: Let B be the event that the child selected at random is a girl, and let A be the event
that the second child of her family is also a girl. We want to calculate P (A | B). Now the set
of all possibilities is as follows: The child is a girl with a sister, the child is a girl with a brother,
the child is a boy with a sister, and the child is a boy with a brother. Thus P (B) = 2/4 and
P (AB) = 1/4. Hence
P (A | B) =
P (AB)
1/4
1
=
= . P (B)
2/4
2
Remark 3.1 There is a major difference between Examples 3.2 and 3.3. In Example 3.2,
a family is selected and found to have a girl. Therefore, we have three equally likely possibilities: (girl, girl), (boy, girl), and (girl, boy), of which only one is desirable: (girl, girl). So
the required probability is 1/3. In Example 3.3, since a child rather than a family is selected,
families with (girl, girl), (boy, girl), and (girl, boy) are not equally likely to be chosen. In fact,
the probability that a family with two girls is selected equals twice the probability that a family
with one girl is selected. That is,
P (girl, girl) = 2P (girl, boy) = 2P (boy, girl).
This and the fact that the sum of the probabilities of (girl, girl), (girl, boy), (boy, girl) is 1 give
1
1
P (girl, girl) + P (girl, girl) + P (girl, girl) = 1,
2
2
which implies that
1
P (girl, girl) = . 2
Example 3.4 An English class consists of 10 Koreans, 5 Italians, and 15 Hispanics. A paper
is found belonging to one of the students of this class. If the name on the paper is not Korean,
what is the probability that it is Italian? Assume that names completely identify ethnic groups.
Solution: Let A be the event that the name on the paper is Italian, and let B be the event that
it is not Korean. To calculate P (A | B), note that AB = A because “the name being Italian,”
and “the name not being Korean” overlap with “the name being Italian,” which is A. Hence
P (AB) = P (A) = 5/30. Therefore,
P (A | B) =
P (AB)
5/30
1
=
= . P (B)
20/30
4
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 90 — #106
✐
✐
90
Chapter 3
Conditional Probability and Independence
Example 3.5 (Cromwell’s Rule) The British leading Bayesian statistician and decision
theorist, Dennis Lindley (1923-2013), stated a rule in subjective probability that we should
refrain from using probabilities 0 and 1 except when assigning probability to logically true or
false statements. He called this Cromwell’s Rule. Oliver Cromwell (1599-1658), an English
military and political leader, in 1650, in a letter to the synod of the Church of Scotland wrote†
“I beseech you, in the bowels of Christ, think it possible that you may be mistaken.” This wellknown phrase was the reason for Lindley naming his precept, Cromwell’s rule. The reason for
Cromwell’s rule is that if, for an event A, P (A) = 1, then for any event B with P (B) > 0,
P (A | B) = 1. That is, if you are certain that A is absolutely true, then no amount of evidence
to the contrary will change your mind. To prove this fact, note that, by Theorem 1.7,
P (B) = P (BA) + P (BAc ).
However, BAc ⊆ Ac . So 0 ≤ P (BAc ) ≤ P (Ac ) = 0, and hence P (BAc ) = 0. This implies
P (AB)
= 1. that P (B) = P (BA) and P (A | B) =
P (B)
Example 3.6 We draw eight cards at random from an ordinary deck of 52 cards. Given that
three of them are spades, what is the probability that the remaining five are also spades?
Solution: Let B be the event that at least three of them are spades, and let A denote the event
that all of the eight cards selected are spades. Then P (A | B) is the desired quantity and is
calculated from P (A | B) = P (AB)/P (B) as follows:
!
13
8
!
P (AB) =
52
8
since AB is the event that all eight cards selected are spades.
!
!
13
39
8
X
x
8−x
!
P (B) =
52
x=3
8
because at least three of the cards selected are spades if exactly three of them are spades, or
exactly four of them are spades, . . . , or all eight of them are spades. Hence
!
13
8
!
52
8
P (AB)
−6
P (A | B) =
=
!
! ≈ 5.44 × 10 . P (B)
13
39
8
X x
8−x
!
52
x=3
8
†
Carlyle, Thomas, ed. (1855). Oliver Cromwell’s letters and speeches. 1. New York: Harper. p. 448.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 91 — #107
✐
✐
Section 3.1
Conditional Probability
91
An important feature of conditional probabilities (stated in the following theorem) is that
they satisfy the same axioms that ordinary probabilities satisfy. This enables us to use the
theorems that are true for probabilities for conditional probabilities as well.
Theorem 3.1 Let S be the sample space of an experiment, and let B be an event of S with
P (B) > 0. Then
(a)
(b)
(c)
P (A | B) ≥ 0 for any event A of S .
P (S | B) = 1.
If A1 , A2 , . . . is a sequence of mutually exclusive events, then
P
∞
[
i=1
Ai | B =
∞
X
i=1
P (Ai | B).
The proof of this theorem is left as an exercise.
Reduction of Sample Space
Let B be an event of a sample space S with P (B) > 0. For a subset A of B, define Q(A) =
P (A | B). Then Q is a function from the set of subsets of B to [0, 1]. Clearly, Q(A) ≥ 0,
Q(B) = P (B | B) = 1 and, by Theorem 3.1, if A1 , A2 , . . . is a sequence of mutually
exclusive subsets of B, then
Q
∞
[
i=1
Ai = P
∞
[
i=1
Ai | B =
∞
X
i=1
P (Ai | B) =
∞
X
Q(Ai ).
i=1
Thus Q satisfies the axioms that probabilities do, and hence it is a probability function. Note
that while P is defined for all subsets of S, the probability function Q is defined only for
subsets of B . Therefore, for Q, the sample space is reduced from S to B . This reduction of
sample space is sometimes very helpful in calculating conditional probabilities. Suppose that
we are interested in P (E | B), where E ⊆ S . One way to calculate this quantity is to reduce
S to B and find Q(EB). It is usually easier to compute the unconditional probability Q(EB)
rather than the conditional probability P (E | B). Examples follow.
Example 3.7 A child mixes 10 good and three dead batteries. To find the dead batteries, his
father tests them one-by-one and without replacement. If the first four batteries tested are all
good, what is the probability that the fifth one is dead?
Solution: Using the information that the first four batteries tested are all good, we rephrase the
problem in the reduced sample space: Six good and three dead batteries are mixed. A battery is
selected at random: What is the probability that it is dead? The solution to this trivial question
is 3/9 = 1/3. Note that without reducing the sample space, the solution of this problem is not so
easy. Example 3.8 A farmer decides to test four fertilizers for his soybean fields. He buys 32 bags
of fertilizers, eight bags from each kind, and tries them randomly on 32 plots, eight plots from
each of fields A, B, C, and D, one bag per plot. If from type I fertilizer one bag is tried on field
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 92 — #108
✐
✐
92
Chapter 3
Conditional Probability and Independence
A and three on field B, what is the probability that two bags of type I fertilizer are tried on field
D?
Solution: Using the information that one and three bags of type I fertilizer are tried on fields A
and B, respectively, we rephrase the problem in the reduced sample space: 16 bags of fertilizers
(of which four are type I) are tried randomly, eight on field C and eight on field D. What is the
probability that two bags of type I fertilizer are tried on field D? The solution to this easier
unconditional version of the problem is
!
!
4
12
2
6
! ≈ 0.43. 16
8
Example 3.9 On a TV game show, there are three curtains. Behind two of the curtains there
is nothing, but behind the third there is a prize that the player might win. The probability that the
prize is behind a given curtain is 1/3. The game begins with the contestant randomly guessing
a curtain. The host of the show (master of ceremonies), who knows behind which curtain the
prize is, will then pull back a curtain other than the one chosen by the player and reveal that the
prize is not behind that curtain. The host will not pull back the curtain selected by the player,
nor will he pull back the one with the prize, if different from the player’s choice. At this point,
the host gives the player the opportunity to change his choice of curtain and select the other
one. The question is whether the player should change his choice. That is, has the probability
of the prize being behind the curtain chosen changed from 1/3 to 1/2 or it is still 1/3? If it is still
1/3, the contestant should definitely change his choice. Otherwise, there is no point in doing so.
Solution: We will show that the conditional probability that the contestant wins, given that
he always changes his original choice is 2/3. Therefore, the contestant should definitely change
his choice. To show this, suppose that the prize is behind curtain 1, and the contestant always
changes his choice. These two assumptions reduce the sample space. The elements of the reduced sample space can be described by 3-tuples (x, y, z), where x is the curtain the contestant
guesses first, y is the curtain the master of ceremonies pulls back, and z is the curtain the contestant switches to. For example, (3, 2, 1) represents a game in which the contestant guesses
curtain 3, the master of ceremonies pulls back curtain 2, and the contestant switches to curtain 1. Therefore, given that the prize is behind curtain 1 and the contestant changes his choice,
the reduced sample space S is
S = (1, 2, 3), (1, 3, 2), (2, 3, 1), (3, 2, 1) .
Note that the event that the contestant guesses curtain 2 is (2, 3, 1) , the event that
he guesses curtain 3 is (3, 2, 1) , whereas the event that he guesses curtain 1 is
(1, 2, 3), (1, 3, 2) . Given that the prize is behind curtain
1 and the contestant always switches
his original choice, the event that the contestant wins is (2, 3, 1), (3, 2, 1) . Since the contestant guesses a curtain with probability 1/3,
P
(2, 3, 1)
=P
(3, 2, 1)
=P
1
(1, 2, 3), (1, 3, 2) = .
3
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 93 — #109
✐
✐
Section 3.1
Conditional Probability
93
This shows that no matter what values are assigned to P (1, 2, 3) and P (1, 3, 2) ,
as long as their sum is 1/3, the conditional probability that the contestant wins, given that he
always changes his original choice, is
P
(2, 3, 1)
+P
(3, 2, 1)
=
2
. 3
By Theorem 3.1, the function Q defined above by Q(A) = P (A | B) is a probability
function. Therefore, Q(A) = P (A | B) satisfies the theorems stated for P . In particular, for
all choices of B, P (B) > 0,
1.
2.
3.
4.
5.
6.
P (∅ | B) = 0.
P (Ac | B) = 1 − P (A | B).
If C ⊆ A, then P (AC c | B) = P (A − C | B) = P (A | B) − P (C | B).
If C ⊆ A, then P (C | B) ≤ P (A | B).
P (A ∪ C | B) = P (A | B) + P (C | B) − P (AC | B).
P (A | B) = P (AC | B) + P (AC c | B).
7.
To calculate P (A1 ∪A2 ∪A3 ∪· · ·∪An | B), we calculate conditional probabilities of all
possible intersections of events from {A1 , A2 , . . . , An }, given B, add the conditional
probabilities obtained by intersecting an odd number of the events, and subtract the
conditional probabilities obtained by intersecting an even number of events.
8.
For any increasing or decreasing sequences of events {An , n ≥ 1},
lim P (An | B) = P ( lim An | B).
n→∞
n→∞
EXERCISES
A
1.
Suppose that 15% of the population of a country are unemployed women, and a total of
25% are unemployed. What percent of the unemployed are women?
2.
Suppose that 41% of Americans have blood type A, and 4% have blood type AB. If in
the blood of a randomly selected American soldier the A antigen is found, what is the
probability that his blood type is A? The A antigen is found only in blood types A and
AB.
3.
In a technical college all students are required to take calculus and physics. Statistics
show that 32% of the students of this college get A’s in calculus, and 20% of them get
A’s in both calculus and physics. Gino, a randomly selected student of this college, has
passed calculus with an A. What is the probability that he got an A in physics?
4.
Suppose that two fair dice have been tossed and the total of their top faces is found to be
divisible by 5. What is the probability that both of them have landed 5?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 94 — #110
✐
✐
94
Chapter 3
Conditional Probability and Independence
5.
A bus arrives at a station every day at a random time between 1:00 P.M. and 1:30 P.M. A
person arrives at this station at 1:00 and waits for the bus. If at 1:15 the bus has not yet
arrived, what is the probability that the person will have to wait at least an additional 5
minutes?
6.
Prove that P (A | B) > P (A) if and only if P (B | A) > P (B). In probability, if
for two events A and B, P (A | B) > P (A), we say that A and B are positively
correlated. If P (A | B) < P (A), A and B are said to be negatively correlated.
7.
A spinner is mounted on a wheel of unit circumference (radius 1/2π ). Arcs A, B, and
C of lengths 1/3, 1/2, and 1/6, respectively, are marked off on the wheel’s perimeter
(see Figure 3.1). The spinner is flicked and we know that it is not pointing toward C .
What is the probability that it points toward A?
Figure 3.1
Spinner of Exercise 7.
8.
In throwing two fair dice, what is the probability of a sum of 5 if they land on different
numbers?
9.
In a small lake, it is estimated that there are approximately 105 fish, of which 40 are
trout and 65 are carp. A fisherman caught eight fish; what is the probability that exactly
two of them are trout if we know that at least three of them are not?
10.
From 100 cards numbered 00, 01, . . . , 99, one card is drawn. Suppose that α and β
are the sum and the product, respectively, of the digits of the card selected. Calculate
P {α = i | β = 0} , i = 0, 1, 2, 3, . . . , 18.
11.
From families with three children, a family is selected at random and found to have a
boy. What is the probability that the boy has (a) an older brother and a younger sister;
(b) an older brother; (c) a brother and a sister? Assume that in a three-child family all
gender distributions have equal probabilities.
12.
In a study of the records of 831 women who died in 2016, an actuary observed that 185
of them died of cancer. Furthermore, she discovered that the mothers of 257 of the 831
women suffered from cancer, and from the mothers of all who died of cancer 93 had
died due to cancer. Find the probability that a randomly selected woman from this group
died of cancer if her mother did not have cancer.
13.
Show that if P (A) = 1, then P (B | A) = P (B).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 95 — #111
✐
✐
Section 3.1
Conditional Probability
95
14.
Prove that if P (A) = a and P (B) = b, then P (A | B) ≥ (a + b − 1)/b.
15.
Prove Theorem 3.1.
16.
Prove that if P (E | F ) ≥ P (G | F ) and P (E | F c ) ≥ P (G | F c ), then P (E) ≥
P (G).
17.
Suppose that 28 crayons, of which four are red, are divided randomly among Jack,
Marty, Sharon, and Martha (seven each). If Sharon has exactly one red crayon, what
is the probability that Marty has the remaining three?
18.
Adam and three of his friends are playing bridge. (a) If, holding a certain hand, Adam
announces that he has a king, what is the probability that he has at least one more king?
(b) If, for some other hand, Adam announces that he has the king of diamonds, what is
the probability that he has at least one more king? Compare parts (a) and (b) and explain
why the answers are not the same.
19.
An actuary studying the insurance preferences of homeowners in a region that is prone
to earthquakes, hurricanes, and floods has discovered that, for each of these perils, the
probability is 0.2 that a homeowner, selected randomly, has purchased coverage only
for that peril. Moreover, she has observed that, for any two of the perils, the probability is 0.14 that the homeowner has purchased coverage for exactly those two perils.
If one-third of the homeowners who have purchased a policy that covers earthquakes
and hurricanes also purchased flood insurance, what percentage of homeowners did not
purchase a policy that covered any of these perils?
B
20.
A number is selected at random from the set {1, 2, . . . , 10,000} and is observed to be
odd. What is the probability that it is (a) divisible by 3; (b) divisible by neither 3 nor 5?
21.
A retired person chooses randomly one of the six parks of his town everyday and goes
there for hiking. We are told that he was seen in one of these parks, Oregon Ridge, once
during the last 10 days. What is the probability that during this period he has hiked in
this park two or more times?
22.
A big urn contains 1000 red chips, numbered 1 through 1000, and 1750 blue chips,
numbered 1 through 1750. A chip is removed at random, and its number is found to be
divisible by 3. What is the probability that its number is also divisible by 5?
23.
There are three types of animals in a laboratory: 15 type I, 13 type II, and 12 type III.
Animals of type I react to a particular stimulus in 5 seconds, animals of types II and III
react to the same stimulus in 4.5 and 6.2 seconds, respectively. A psychologist selects 10
of these animals at random and finds that exactly four of them react to this stimulus in
6.2 seconds. What is the probability that at least two of them react to the same stimulus
in 4.5 seconds?
24.
In an international school, 60 students, of whom 15 are Korean, 20 are French, eight are
Greek, and the rest are Chinese, are divided randomly into four classes of 15 each. If
there is a total of eight French and six Korean students in classes A and B, what is the
probability that class C has four of the remaining 12 French and three of the remaining
nine Korean students?
25.
From the set of all families with two children, a family is selected at random and is
found to have a girl called Mary. We want to know the probability that both children of
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 96 — #112
✐
✐
96
Chapter 3
Conditional Probability and Independence
the family are girls. By Example 3.2, the probability should apparently be 1/3 because
presumably knowing the name of the girl should not make a difference. However, if we
ask a friend of that family whether Mary is the older or the younger child of the family,
upon receiving the answer, we can conclude that the probability of the other child being
a girl is 1/2. Therefore, there is no need to ask the friend, and we conclude already that
the probability is 1/2! What is the flaw in this argument? Explain.
Self-Quiz on Section 3.1
Time allotted: 20 Minutes
Each problem is worth 2.5 points.
1.
In a bridge game, played with a normal deck of 52 cards, hearts is the designated trump
suit. If the first 4 cards a player is dealt are ace, 5, 9, and the king, all of hearts, what is
the probability that none of the remaining 9 cards the player is dealt are hearts?
Hint: Reduce the sample space.
2.
For a certain loaded die, the probabilities of the possible outcomes are given by the
following table.
Outcome
1
2
3
4
5
6
Probability
0.27
0.15
0.17
0.2
0.05
0.16
If the die is tossed and the outcome is an odd number, what is the probability that it is 1?
3.
4.
The theaters of a town are showing seven comedies and nine dramas. Marlon has seen
five of the movies. If the first three movies he has seen are dramas, what is the probability
that the last two are comedies? Assume that Marlon chooses the shows at random and
sees each movie at most once.
Hint: Reduce the sample space.
For events E and F, suppose that P (EF ) = 0.23, P E ∪ F = 0.67, and
P (E | F ) = 0.46. Find P (F | E).
Often, when we think of a collection of events, we have a tendency to think about them in
either temporal or logical sequence. So, if, for example, a sequence of events A1 , A2 , . . . , An
occur in time or in some logical order, we can usually immediately write down the probabilities
P (A1 ), P (A2 | A1 ), . . . , P (An | A1 A2 · · · An−1 ) without much computation. However, we
may be interested in probabilities of intersection of events, or probabilities of events unconditional on the rest, or probabilities of earlier events, given later events. In the next three sections,
we will develop techniques for calculating such probabilities.
Suppose that, in a random phenomenon, events occur in either temporal or logical sequence. In Section 3.2 (The Multiplication Rule), we will discuss the probabilities of the intersection of events in terms of conditional probabilities of later events given earlier events. In
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 97 — #113
✐
✐
Section 3.2
The Multiplication Rule
97
Section 3.3 (Law of Total Probability), we will find the probabilities of unconditional events in
terms of conditional probabilities of later events given the earlier ones. In Section 3.4 (Bayes’
Formula), we will discuss a formula relating probabilities of earlier events, given later events,
to the conditional probabilities of later events given the earlier ones.
3.2
THE MULTIPLICATION RULE
The relation
P (A | B) =
P (AB)
P (B)
is also useful for calculating P (AB). If we multiply both sides of this relation by P (B) note
that P (B) > 0 , we get
P (AB) = P (B)P (A | B),
(3.4)
which means that the probability of the joint occurrence of A and B is the product of the
probability of B and the conditional probability of A given that B has occurred. If P (A) > 0,
then by letting A = B and B = A in (3.4), we obtain
P (BA) = P (A)P (B | A).
Since P (BA) = P (AB), this relation gives
P (AB) = P (A)P (B | A).
(3.5)
Thus, to calculate P (AB), depending on which of the quantities P (A | B) and P (B | A) is
known, we may use (3.4) or (3.5), respectively. The following example clarifies the usefulness
of these relations.
Example 3.10 Suppose that five good fuses and two defective ones have been mixed up. To
find the defective fuses, we test them one-by-one, at random and without replacement. What is
the probability that we are lucky and find both of the defective fuses in the first two tests?
Solution: Let D1 and D2 be the events of finding a defective fuse in the first and second tests,
respectively. We are interested in P (D1 D2 ). Using (3.5), we get
P (D1 D2 ) = P (D1 )P (D2 | D1 ) =
2 1
1
× = . 7 6
21
Relation (3.5) can be generalized for calculating the probability of the joint occurrence of
several events. For example, if P (AB) > 0, then
P (ABC) = P (A)P (B | A)P (C | AB).
(3.6)
To see this, note that P (AB) > 0 implies that P (A) > 0; therefore,
P (A)P (B | A)P (C | AB) = P (A)
P (AB) P (ABC)
= P (ABC).
P (A) P (AB)
The following theorem will generalize (3.5) and (3.6) to n events and can be shown in the same
way as (3.6) is shown.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 98 — #114
✐
✐
98
Chapter 3
Theorem 3.2
Conditional Probability and Independence
(The Multiplication Rule)
If P (A1 A2 A3 · · · An−1 ) > 0, then
P (A1 A2 A3 · · · An−1 An )
= P (A1 )P (A2 | A1 )P (A3 | A1 A2 ) · · · P (An | A1 A2 A3 · · · An−1 ).
Example 3.11 A consulting firm is awarded 43% of the contracts it bids on. Suppose that
Nordulf works for a division of the firm that gets to do 15% of the projects contracted for.
If Nordulf directs 35% of the projects submitted to his division, what percentage of all bids
submitted by the firm will result in contracts for projects directed by Nordulf?
Solution: Let A1 be the event that the firm is awarded a contract for some randomly selected
bid. Let A2 be the event that the contract will be sent to Nordulf’s division. Let A3 be the event
that Nordulf will direct the project. The desired probability is calculated as follows:
P (A1 A2 A3 ) = P (A1 )P (A2 | A1 )P (A3 | A1 A2 ) = (0.43)(0.15)(0.35) = 0.0226.
Hence 2.26% of all bids submitted by the firm will be directed by Nordulf.
Example 3.12 Suppose that five good and two defective fuses have been mixed up. To find
the defective ones, we test them one by one, at random and without replacement. What is the
probability that we find both of the defective fuses in exactly three tests?
Solution: Let D1 , D2 , and D3 be the events that the first, second, and third fuses tested are
defective, respectively. Let G1 , G2 , and G3 be the events that the first, second, and third fuses
tested are good, respectively. We are interested in P (G1 D2 D3 ∪D1 G2 D3 ), which we calculate
using the multiplication rule:
P (G1 D2 D3 ∪ D1 G2 D3 )
= P (G1 D2 D3 ) + P (D1 G2 D3 )
= P (G1 )P (D2 | G1 )P (D3 | G1 D2 ) + P (D1 )P (G2 | D1 )P (D3 | D1 G2 )
5 2 1 2 5 1
= × × + × × ≈ 0.095. 7 6 5 7 6 5
EXERCISES
A
1.
In a trial, the judge is 65% sure that Susan has committed a crime. Robert is a witness
who knows whether Susan is innocent or guilty. However, Robert is Susan’s friend and
will lie with probability 0.25 if Susan is guilty. He will tell the truth if she is innocent.
What is the probability that Robert will commit perjury?
2.
There are 14 marketing firms hiring new graduates. Kate randomly found the recruitment
ads of six of these firms and sent them her resume. If three of these marketing firms are
in Maryland, what is the probability that Kate did not apply to a marketing firm in
Maryland?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 99 — #115
✐
✐
Section 3.2
The Multiplication Rule
99
3.
In a game of cards, two cards of the same color and denomination form a pair. For
example, 8 of hearts and 8 of diamonds is one pair, king of spades and king of clubs
is another. If six cards are selected at random and without replacement, what is the
probability that there will be no pairs?
4.
If eight defective and 12 nondefective items are inspected one by one, at random and
without replacement, what is the probability that (a) the first four items inspected are
defective; (b) from the first three items at least two are defective?
5.
An actuary works for an auto insurance company whose customers all have insured at
least one car. She discovers that of all the customers who insure more than one car,
70% of them insure at least one SUV. Furthermore, she observes that 47% of all the
customers insure at least one SUV. If 65% of all customers insure more than one car,
what percentage of the customers insure exactly one non-SUV?
6.
There are five boys and six girls in a class. For an oral exam, their teacher calls them one
by one and randomly. (a) What is the probability that the boys and the girls alternate?
(b) What is the probability that the boys are called first? Compare the answers to parts
(a) and (b).
7.
In a lottery scratch-off game, each ticket has 10 coated circles in the middle and one
coated rectangle in the lower left corner. Underneath the coats of 4 of the circles, there is
a dollar sign, “$,” and underneath the remaining 6 circles is blank. A winning ticket is the
one with only those circles that contain the dollar sign scratched. The remaining circles
and the rectangle must be left untouched. The winning amount is written underneath
the coat of the rectangle, which will be scratched off by the lottery agent to reveal the
amount won. Harlan buys one such ticket and scratches off four of the circles at random.
Find the probability that (a) he wins; (b) he wins if the first two circles he scratches off
have the dollar sign underneath.
8.
An urn contains five white and three red chips. Each time we draw a chip, we look at
its color. If it is red, we replace it along with two new red chips, and if it is white, we
replace it along with three new white chips. What is the probability that, in successive
drawing of chips, the colors of the first four chips alternate?
9.
The law school of a university admits all applicants who have at least a 3.5 undergraduate
GPA and a Law School Admissions Test (LSAT) score of 154 or higher. Suppose that
of all students with a 3.5 or higher GPA, only 45% score at least 154 the first time they
take the LSAT test; of all those students who scored less than 154 the first time, 58%
will score 154 or higher the second time they take the test; and of all who scored lower
than 154 both the first and second times, 75% will score 154 or higher the third time
they take the test. Ladona has achieved the minimum GPA required by this law school.
What is the probability that she does not need to take the test more than three times to
earn at least the minimum LSAT score required by the law school?
10.
In the card game, bridge, played with an ordinary deck of 52 cards, all cards are dealt
among four players, 13 each, randomly. What is the probability that each player gets one
ace?
Hint: Let A1 be the event that the ace of hearts is dealt to one of the four players.
Let A2 be the event that the ace of hearts and ace of diamonds are dealt to two different
players. Let A3 be the event that the ace of hearts, ace of diamonds, and ace of spades are
dealt to three different players, and let A4 be the event that each player gets exactly one
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 100 — #116
✐
✐
100
Chapter 3
Conditional Probability and Independence
of the aces. Note that A4 ⊆ A3 ⊆ A2 ⊆ A1 , and the desired probability is P (A4 ) =
P (A1 A2 A3 A4 ).
11.
Suppose that 75% of all people with credit records improve their credit ratings within
three years. Suppose that 18% of the population at large have poor credit records, and of
those only 30% will improve their credit ratings within three years. What percentage of
the people who will improve their credit records within the next three years are the ones
who currently have good credit ratings?
B
12.
Cards are drawn at random from an ordinary deck of 52, one by one and without replacement. What is the probability that no heart is drawn before the ace of spades is
drawn?
13.
From an ordinary deck of 52 cards, cards are drawn one by one, at random and without
replacement. What is the probability that the fourth heart is drawn on the tenth draw?
Hint: Let F denote the event that in the first nine draws there are exactly three hearts,
and E be the event that the tenth draw is a heart. Use the multiplication rule: P (F E) =
P (F )P (E | F ).
14.
In a series of games, the winning number of the nth game, n = 1, 2, 3, . . . , is a number
selected at random from the set of integers {1, 2, . . . , n + 2}. Don bets on 1 in each
game and says that he will quit as soon as he wins. What is the probability that he has to
play indefinitely?
Hint: Let AnTbe the event
that Don loses the first n games. To calculate the desired
∞
probability, P
A
,
use
Theorem 1.8.
i
i=1
Self-Quiz on Section 3.2
Time allotted: 15 Minutes
Each problem is worth 5 points.
1.
An urn contains 6 blue and 4 red balls. Three balls are drawn at random and without
replacement. What is the probability that the balls drawn are alternatively of different
colors?
2.
On a given day, the first item produced by a manufacturer is defective with probability
p and non-defective with probability 1 − p. However, whether or not an item produced
afterward is defective depends only on the item that was produced right before it. It is
defective with probability p1 if the previous item produced is non-defective, and it is
defective with probability p2 if the previous item produced is defective. Find the probability that, on a randomly selected working day, the first four items produced by this
manufacturer are all non-defective.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 101 — #117
✐
✐
Section 3.3
3.3
Law of Total Probability
101
LAW OF TOTAL PROBABILITY
Sometimes it is not possible to calculate directly the probability of the occurrence of an event
A, but it is possible to find P (A | B) and P (A | B c ) for some event B . In such cases,
the following theorem, which is conceptually rich and has widespread applications, is used. It
states that P (A) is the weighted average of the probability of A given that B has occurred and
probability of A given that it has not occurred.
Theorem 3.3 (Law of Total Probability)
P (B c ) > 0. Then for any event A,
Let B be an event with P (B) > 0 and
P (A) = P (A | B)P (B) + P (A | B c )P (B c ).
Proof: By Theorem 1.7,
P (A) = P (AB) + P (AB c ).
(3.7)
Now P (B) > 0 and P (B c ) > 0. These imply that P (AB) = P (A | B)P (B) and
P (AB c ) = P (A | B c )P (B c ). Putting these in (3.7), we have proved the theorem. c
Figure 3.2
c
Tree diagram of Example 3.13.
Example 3.13 An insurance company rents 35% of the cars for its customers from agency
I and 65% from agency II. If 8% of the cars of agency I and 5% of the cars of agency II
break down during the rental periods, what is the probability that a car rented by this insurance
company breaks down?
Solution: Let A be the event that a car rented by this insurance company breaks down. Let
I and II be the events that it is rented from agencies I and II, respectively. Then by the law of
total probability,
P (A) = P (A | I)P (I) + P (A | II)P (II)
= (0.08)(0.35) + (0.05)(0.65) = 0.0605.
Tree diagrams facilitate solutions to this kind of problem. Let B and B c stand for breakdown
and not breakdown during the rental period, respectively. Then, as Figure 3.2 shows, to find
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 102 — #118
✐
✐
102
Chapter 3
Conditional Probability and Independence
the probability that a car breaks down, all we need to do is to compute, by multiplication,
the probability of each path that leads to a point B and then add them up. So, as seen from
the tree, the probability that the car breaks down is (0.35)(0.08) + (0.65)(0.05) = 0.06,
and the probability that it does not break down is (0.35)(0.92) + (0.65)(0.95) = 0.94.
Example 3.14 In a trial, the judge is 65% sure that Susan has committed a crime. Julie and
Robert are two witnesses who know whether Susan is innocent or guilty. However, Robert is
Susan’s friend and will lie with probability 0.25 if Susan is guilty. He will tell the truth if she
is innocent. Julie is Susan’s enemy and will lie with probability 0.30 if Susan is innocent. She
will tell the truth if Susan is guilty. What is the probability that, in the course of the trial, Robert
and Julie will give conflicting testimony?
Solution: Let I be the event that Susan is innocent. Let C be the event that Robert and Julie
will give conflicting testimony. By the law of total probability,
P (C) = P (C | I)P (I) + P (C | I c )P (I c )
= (0.30)(.35) + (0.25)(.65) = 0.2675. Example 3.15 (Gambler’s Ruin Problem) Two gamblers play the game of “heads or
tails,” in which each time a fair coin lands heads up player A wins $1 from B, and each time it
lands tails up, player B wins $1 from A. Suppose that player A initially has a dollars and player
B has b dollars. If they continue to play this game successively, what is the probability that (a)
A will be ruined; (b) the game goes forever with nobody winning?
Solution:
(a)
Let E be the event that A will be ruined if he or she starts with i dollars, and let pi =
P (E). Our aim is to calculate pa . To do so, we define F to be the event that A wins the
first game. Then
P (E) = P (E | F )P (F ) + P (E | F c )P (F c ).
In this formula, P (E | F ) is the probability that A will be ruined, given that he wins the
first game; so P (E | F ) is the probability that A will be ruined if his capital is i + 1;
that is, P (E | F ) = pi+1 . Similarly, P (E | F c ) = pi−1 . Hence
pi = pi+1 ·
1
1
+ pi−1 · .
2
2
(3.8)
Now p0 = 1 because if A starts with 0 dollars, he or she is already ruined. Also, if the
capital of A reaches a + b, then B is ruined; thus pa+b = 0. Therefore, we have to solve
the system of recursive equations (3.8), subject to the boundary conditions p0 = 1 and
pa+b = 0. To do so, note that (3.8) implies that
pi+1 − pi = pi − pi−1 .
Hence, letting p1 − p0 = α, we get
pi − pi−1 = pi−1 − pi−2 = pi−2 − pi−3 = · · · = p2 − p1 = p1 − p0 = α.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 103 — #119
✐
✐
Section 3.3
Law of Total Probability
103
Thus
p1 = p0 + α
p2 = p1 + α = p0 + α + α = p0 + 2α
p3 = p2 + α = p0 + 2α + α = p0 + 3α
..
.
pi = p0 + iα
..
.
Now p0 = 1 gives pi = 1 + iα. But pa+b = 0; thus 0 = 1 + (a + b)α. This gives
α = −1/(a + b); therefore,
pi = 1 −
a+b−i
i
=
.
a+b
a+b
In particular, pa = b/(a + b). That is, the probability that A will be ruined is b/(a + b).
(b)
The same method can be used with obvious modifications to calculate qi , the probability
that B is ruined if he or she starts with i dollars. The result is
qi =
a+b−i
.
a+b
Since B starts with b dollars, he or she will be ruined with probability qb = a/(a +
b). Thus the probability that the game goes on forever with nobody winning is
1 − (qb + pa ). But 1 − (qb + pa ) = 1 − a/(a + b) − b/(a + b) = 0. Therefore,
if this game is played successively, eventually either A is ruined or B is ruined. Remark 3.2 If the game is not fair and on each play gambler A wins $1 from B with
probability p, 0 < p < 1, p =
6 1/2, and loses $1 to B with probability q = 1 − p, relation
(3.8) becomes
pi = ppi+1 + qpi−1 ,
but pi = (p + q)pi , so (p + q)pi = ppi+1 + qpi−1 . This gives
q(pi − pi−1 ) = p(pi+1 − pi ),
which reduces to
pi+1 − pi =
q
(pi − pi−1 ).
p
Using this relation and following the same line of argument lead us to
pi =
h q i−1
p
+
q i−2
p
+ ··· +
q p
i
+ 1 α + p0 ,
where α = p1 − p0 . Using p0 = 1, we obtain
pi =
(q/p)i − 1
α + 1.
(q/p) − 1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 104 — #120
✐
✐
104
Chapter 3
Conditional Probability and Independence
After finding α from pa+b = 0 and substituting in this equation, we get
pi =
1 − (p/q)a+b−i
(q/p)i − (q/p)a+b
=
,
1 − (q/p)a+b
1 − (p/q)a+b
where the last equality is obtained by multiplying the numerator and denominator of the function by (p/q)a+b . In particular,
1 − (p/q)b
pa =
,
1 − (p/q)a+b
and, similarly,
qb =
1 − (q/p)a
.
1 − (q/p)a+b
In this case also, pa + qb = 1, meaning that eventually either A or B will be ruined and the
game does not go on forever. For comparison, suppose that A and B both start with $10. If they
play a fair game, the probability that A will be ruined is 1/2, and the probability that B will be
ruined is also 1/2. If they play an unfair game with p = 3/4, q = 1/4, then p10 , the probability
that A will be ruined, is almost 0.00002. To generalize Theorem 3.3, we will now state a definition.
Definition 3.2
Let {B1 , B2 , . . . , Bn } be a set of nonempty subsets of the
Snsample space S
of an experiment. If the events B1 , B2 , . . . , Bn are mutually exclusive and i=1 Bi = S, the
set {B1 , B2 , . . . , Bn } is called a partition of S .
Figure 3.3
Partition of the given sample space S.
For example, if S of Figure 3.3 is the sample space of an experiment, then B1 , B2 , B3 , B4 ,
and B5 of the same figure form a partition of S . As another example, consider the experiment
of drawing a card from an ordinary deck of 52 cards. Let B1 , B2 , B3 , and B4 denote the events
that the card is a spade, a club, a diamond, and a heart, respectively. Then {B1 , B2 , B3 , B4 } is
a partition of the sample space of this experiment. If Ai , 1 ≤ i ≤ 10, denotes the event that
the value of the card drawn is i, and A11 , A12 , and A13 are the events of jack, queen, and king,
respectively, then {A1 , A2 , . . . , A13 } is another partition of the same sample space. Note that
For an experiment with sample space S, for any event A, A and Ac both
nonempty, the set {A, Ac } is a partition.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 105 — #121
✐
✐
Section 3.3
Law of Total Probability
105
As observed, Theorem 3.3 is used whenever it is not possible to calculate P (A) directly,
but it is possible to find P (A | B) and P (A | B c ) for some event B . In many situations, it is
neither possible to find P (A) directly, nor possible to find a single event B that enables us to
use
P (A) = P (A | B)P (B) + P (A | B c )P (B c ).
In such situations Theorem 3.4, which is a generalization of Theorem 3.3 and is also called the
law of total probability, might be applicable.
Theorem 3.4 (Law of Total Probability) If {B1 , B2 , . . . , Bn } is a partition of the
sample space of an experiment and P (Bi ) > 0 for i = 1, 2, . . . , n, then for any event A of S,
P (A) = P (A | B1 )P (B1 ) + P (A | B2 )P (B2 ) + · · · + P (A | Bn )P (Bn )
n
X
=
P (A | Bi )P (Bi ).
i=1
More generally, let {B1 , B2 , . . .} be a sequence of mutually exclusive events of S such that
S
∞
i=1 Bi = S. Suppose that, for all i ≥ 1, P (Bi ) > 0. Then for any event A of S,
P (A) =
∞
X
i=1
P (A | Bi )P (Bi ).
Proof: Since B1 , B2 , . . . , Bn are mutually exclusive, Bi Bj = ∅ for i 6= j, it follows that
(ABi )(ABj ) = ∅ for i 6= j . Hence {AB1 , AB2 , . . . , ABn } is a set of mutually exclusive
events. Now
S = B1 ∪ B2 ∪ · · · ∪ Bn
gives
A = AS = AB1 ∪ AB2 ∪ · · · ∪ ABn ;
therefore,
P (A) = P (AB1 ) + P (AB2 ) + · · · + P (ABn ).
But P (ABi ) = P (A | Bi )P (Bi ) for i = 1, 2, . . . , n, so
P (A) = P (A | B1 )P (B1 ) + P (A | B2 )P (B2 ) + · · · + P (A | Bn )P (Bn ).
The proof of the more general case is similar.
When using this theorem, one should be very careful to choose B1 , B2 , B3 , . . . , so that
they form a partition of the sample space.
Example 3.16 Suppose that 80% of the seniors, 70% of the juniors, 50% of the sophomores,
and 30% of the freshmen of a college use the library of their campus frequently. If 30% of
all students are freshmen, 25% are sophomores, 25% are juniors, and 20% are seniors, what
percent of all students use the library frequently?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 106 — #122
✐
✐
106
Chapter 3
Conditional Probability and Independence
Figure 3.4
Tree diagram of Example 3.16.
Solution: Let U be the event that a randomly selected student is using the library frequently.
Let F, O, J, and E be the events that he or she is a freshman, sophomore, junior, or senior,
respectively. Then {F, O, J, E} is a partition of the sample space. Thus
P (U ) = P (U | F )P (F ) + P (U | O)P (O) + P (U | J)P (J) + P (U | E)P (E)
= (0.30)(0.30) + (0.50)(0.25) + (0.70)(0.25) + (0.80)(0.20) = 0.55.
Therefore, 55% of these students use the campus library frequently. The same calculation can
be carried out readily from the tree diagram of Figure 3.4, where U means they use the library
frequently and N means that they do not. Example 3.17 Suppose that the only parasite living in an aquatic habitat is a single-celled
organism, which after a second, with equal probabilities, either splits into two organisms, remains as is, or dies. Suppose that in subsequent seconds, all the living parasites in the habitat
follow the same behavior as the original parasite, independently of each other. What is the
probability that eventually the aquatic habitat will be clean with no parasites?
Solution: The aquatic habitat will eventually be clean if ultimately the parasite population
dies out. Let A be the event that this happens. Let T, R, and D be the events that, after a
second, the parasite splits into two organisms, remains the same, and dies, respectively. By the
law of total probability,
P (A) = P (A | T )P (T ) + P (A | R)P (R) + P (A | D)P (D)
2 1
1
1
= P (A) · + P (A) · + 1 · .
3
3
3
2
2
This gives P (A) − 2P (A) + 1 = 0, or equivalently P (A) − 1 = 0. So P (A) = 1, and
thus eventually the aquatic habitat will be clean with no parasite. Example 3.18 An urn contains 10 white and 12 red chips. Two chips are drawn at random
and, without looking at their colors, are discarded. What is the probability that a third chip
drawn is red?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 107 — #123
✐
✐
Section 3.3
Law of Total Probability
107
Solution: For i ≥ 1, let Ri be the event that the ith chip drawn is red and Wi be the event that
it is white. Intuitively, it should be clear that the two discarded chips provide no information,
so P (R3 ) = 12/22, the same as if it were the first chip drawn from the urn. To prove this
mathematically, note that {R2 W1 , W2 R1 , R2 R1 , W2 W1 } is a partition of the sample space;
therefore,
P (R3 ) = P (R3 | R2 W1 )P (R2 W1 ) + P (R3 | W2 R1 )P (W2 R1 )
+ P (R3 | R2 R1 )P (R2 R1 ) + P (R3 | W2 W1 )P (W2 W1 ).
(3.9)
Now
P (R2 W1 ) = P (R2 | W1 )P (W1 ) =
20
12 10
×
=
,
21 22
77
P (W2 R1 ) = P (W2 | R1 )P (R1 ) =
10 12
20
×
=
,
21 22
77
P (R2 R1 ) = P (R2 | R1 )P (R1 ) =
11 12
22
×
= ,
21 22
77
P (W2 W1 ) = P (W2 | W1 )P (W1 ) =
9
10
15
×
=
.
21 22
77
and
Substituting these values in (3.9), we get
P (R3 ) =
11 20 11 20 10 22 12 15
12
×
+
×
+
×
+
×
= . 20 77 20 77 20 77 20 77
22
EXERCISES
A
1.
If 5% of men and 0.25% of women are color blind, what is the probability that a randomly selected person is color blind?
2.
Suppose that 40% of the students of a campus are women. If 20% of the women and 16%
of the men of this campus are A students, what percent of all of them are A students?
3.
A random number is selected from the interval (0, 1]. Which one of the following is a
partition of the sample space of this experiment? Which one is not?
(a) (0, 1/2), (1/2, 1] ,
(b) (0, 1/3], (1/3, 1] ,
n i−1 i
o
(c) (0, 2/5], [2/5, 1] ,
(d)
, : n ≥ 1 is given and 1 ≤ i ≤ n .
n n
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 108 — #124
✐
✐
108
Chapter 3
Conditional Probability and Independence
4.
Jim has three cars of different models: A, B, and C. The probabilities that models A, B,
and C use over 3 gallons of gasoline from Jim’s house to his work are 0.25, 0.32, and
0.53, respectively. On a certain day, all three of Jim’s cars have 3 gallons of gasoline
each. Jim chooses one of his cars at random, and without paying attention to the amount
of gasoline in the car drives it toward his office. What is the probability that he makes it
to the office?
5.
One of the cards of an ordinary deck of 52 cards is lost. What is the probability that a
random card drawn from this deck is a spade?
6.
Two cards from an ordinary deck of 52 cards are missing. What is the probability that a
random card drawn from this deck is a spade?
7.
Of the patients in a hospital, 20% of those with, and 35% of those without myocardial
infarction have had strokes. If 40% of the patients have had myocardial infarction, what
percent of the patients have had strokes?
8.
Suppose that 37% of a community are at least 45 years old. If 80% of the time a person
who is 45 or older tells the truth, and 65% of the time a person below 45 tells the truth,
what is the probability that a randomly selected person answers a question truthfully?
9.
A person has six guns. The probability of hitting a target when these guns are properly
aimed and fired is 0.6, 0.5, 0.7, 0.9, 0.7, and 0.8, respectively. What is the probability of
hitting a target if a gun is selected at random, properly aimed, and fired?
10.
A factory produces its entire output with three machines. Machines I, II, and III produce
50%, 30%, and 20% of the output, but 4%, 2%, and 4% of their outputs are defective,
respectively. What fraction of the total output is defective?
11.
When traveling from Springfield, Massachusetts, to Baltimore, Maryland, 30% of the
time Charles takes I-91 south and then I-95 south, and 70% of the time he takes I-91
south, then I-84 west followed by I-81 south and I-83 south. Choosing the first option
takes Charles a random time between 5 hours and 30 minutes and 10 hours depending
upon the traffic. However, choosing the second option takes a random time between 6
hours and 30 minutes and 8 hours. The second option is a longer drive but has much
less traffic. What is the probability that on his next trip from Springfield to Baltimore,
Charles travels no more than 7 hours and 15 minutes?
12.
Solve the following problem, from the “Ask Marilyn” column of Parade Magazine, October 29, 2000.
I recently returned from a trip to China, where the government is so concerned about population growth that it has instituted strict laws about
family size. In the cities, a couple is permitted to have only one child.
In the countryside, where sons traditionally have been valued, if the first
child is a son, the couple may have no more children. But if the first child
is a daughter, the couple may have another child. Regardless of the sex
of the second child, no more are permitted. How will this policy affect the
mix of males and females?
To pose the question mathematically, what is the probability that a randomly selected
child from the countryside is a boy?
13.
At a gas station, 85% of the customers use 87 octane gasoline, 5% use 91 octane, and
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 109 — #125
✐
✐
Section 3.3
Law of Total Probability
109
10% use 93 octane. Suppose that, 70%, 85%, and 95% of 87 octane users, 91 octane
users, and 93 octane users fill their tanks, respectively. What is the probability that the
next customer entering this station to put gas in his or her vehicle will fill the tank?
14.
Suppose that five coins, of which exactly three are gold, are distributed among five persons, one each, at random, and one by one. Are the chances of getting a gold coin equal
for all participants? Why or why not?
15.
In a town, 7/9th of the men and 3/5th of the women are married. In that town, what
fraction of the adults are married? Assume that all married adults are the residents of the
town.
16.
A child gets lost in the Disneyland at the Epcot Center in Florida. The father of the child
believes that the probability of his being lost in the east wing of the center is 0.75 and in
the west wing is 0.25. The security department sends an officer to the east and an officer
to the west to look for the child. If the probability that a security officer who is looking
in the correct wing finds the child is 0.4, find the probability that the child is found.
17.
An actuary has discovered that, in her company, 65% of those who have only income
protection insurance and 80% of those who have only legal expense insurance will renew
their policies next year. Furthermore, she has observed that 87% of those who have both
of these policies will renew at least one of them next year. If 58% of policyholders of
this company have income protection insurance, 43% have legal expense insurance, and
23% have both policies, what percentage of the policy holders will renew at least one
policy next year?
18.
A number is selected at random from the set {1, 2, . . . , 20}. Then a second number is
selected randomly between 1 and the first number selected. What is the probability that
the second number is 5?
19.
Suppose that there exist N families on the earth
Pcand that the maximum number of children a family has is c. Let αj 0 ≤ j ≤ c,
j=0 αj = 1 be the fraction of families
with j children. Find the fraction of all children in the world who are the k th born of
their families (k = 1, 2, . . . , c).
20.
Let B be an event of a sample space S with P (B) > 0. For a subset A of S, define
Q(A) = P (A | B). By Theorem 3.1 we know that Q is a probability function. For E
and F, events of S P (F B) > 0 , show that Q(E | F ) = P (E | F B).
B
21.
Suppose that 40% of the students on a campus, who are married to students on the same
campus, are female. Moreover, suppose that 30% of those who are married, but not to
students at this campus, are also female. If one-third of the married students on this campus are married to other students on this campus, what is the probability that a randomly
selected married student from this campus is a woman?
Hint: Let M, C, and F denote the events that the random student is married, is married to a student on the same campus, and is female. For any event A, let Q(A) = P (A |
M ). Then, by Theorem 3.1, Q satisfies the same axioms that probabilities satisfy. Applying Theorem 3.3 to Q and, using the result of Exercise 20, we obtain
P (F | M ) = P (F | M C)P (C | M ) + P (F | M C c )P (C c | M ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 110 — #126
✐
✐
110
Chapter 3
22.
Suppose that the probability that a new seed planted in a specific farm germinates is
equal to the proportion of all planted seeds that germinated in that farm previously.
Suppose that the first seed planted in the farm germinated, but the second seed planted
did not germinate. For positive integers n and k (k < n), what is the probability that of
the first n seeds planted in the farm exactly k germinated?
23.
For n ≥S1, let E1 , E2 , . . . , En be events of a sample space. Find an n-element partin
tion of Si=1 Ei . That
Snis, find a set of mutually exclusive events {F1 , F2 , . . . , Fn } that
n
satisfies i=1 Fi = i=1 Ei .
24.
Conditional Probability and Independence
Suppose that 10 good and three dead batteries are mixed up. Jack tests them one by one,
at random and without replacement. But before testing the fifth battery he realizes that
he does not remember whether the first one tested is good or is dead. All he remembers
is that the last three that were tested were all good. What is the probability that the first
one is also good?
25.
A box contains 18 tennis balls, of which eight are new. Suppose that three balls are
selected randomly, played with, and after play are returned to the box. If another three
balls are selected for play a second time, what is the probability that they are all new?
26.
From families with three children, a child is selected at random and found to be a girl.
What is the probability that she has an older sister? Assume that in a three-child family
all sex distributions are equally probable.
Hint: Let G be the event that the randomly selected child is a girl, A be the event that
she has an older sister, and O, M, and Y be the events that she is the oldest, the middle,
and the youngest child of the family, respectively. For any subset B of the sample space
let Q(B) = P (B | G); then apply Theorem 3.3 to Q. (See also Exercises 20 and 21.)
27.
Suppose that three numbers are selected one by one, at random and without replacement
from the set of numbers {1, 2, 3, . . . , n}. What is the probability that the third number
falls between the first two if the first number is smaller than the second?
Self-Quiz on Section 3.3
Time allotted: 20 Minutes
1.
At a kindergarten, there is a box full of balls. Each child draws a ball at random to play
with and is not allowed to return it to the box to draw another one. When it is Natalie’s
turn to draw a ball, there are 3 red and 7 blue balls left in the box. Natalie loves to play
with a red ball and so does Nicole, who will draw right after Natalie. Does Natalie have
an advantage over Nicole for drawing a red ball? (3 points)
2.
There are three dice in a small box. The first die is unbiased; the second one is loaded,
and, when tossed, the probability of obtaining 6 is 3/8, and the probability of obtaining
each of the other faces is 1/8. The third die is also loaded, but the probability of obtaining
6 when tossed is 2/7, and the probability of each of the other five faces is 1/7. A die is
selected at random and tossed. What is the probability that the outcome is a 6?
(3 points)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 111 — #127
✐
✐
Section 3.4
3.
3.4
Bayes’ Formula
111
Seventy percent of all students of a college participate in a study abroad program. If
60% of the students of this college are male and 55% of the male students attend a
study abroad program, what percentage of female students participate in a study abroad
program? (4 points)
BAYES’ FORMULA
To introduce Bayes’ formula, let us first examine the following problem. In a bolt factory, 30,
50, and 20% of production is manufactured by machines I, II, and III, respectively. If 4, 5,
and 3% of the output of these respective machines is defective, what is the probability that a
randomly selected bolt that is found to be defective is manufactured by machine III? To solve
this problem, let A be the event that a random bolt is defective and B3 be the event that it is
manufactured by machine III. We are asked to find P (B3 | A). Now
P (B3 | A) =
P (B3 A)
.
P (A)
(3.10)
So, to calculate P (B3 | A), we need to know the quantities P (B3 A) and P (A). But neither
of these is given. However, since P (A | B3 ) and P (B3 ) are known we use the relation
P (B3 A) = P (A | B3 )P (B3 )
(3.11)
to find P (B3 A). To calculate P (A), we use the law of total probability. Let B1 and B2 be the
events that the bolt is manufactured by machines I and II, respectively. Then {B1 , B2 , B3 } is a
partition of the sample space; hence
P (A) = P (A | B1 )P (B1 ) + P (A | B2 )P (B2 ) + P (A | B3 )P (B3 ).
(3.12)
Substituting (3.11) and (3.12) in (3.10), we arrive at a formula, called Bayes’ formula, which
enables us to calculate P (B3 | A) readily:
P (B3 | A) =
P (B3 A)
P (A)
=
P (A | B3 )P (B3 )
P (A | B1 )P (B1 ) + P (A | B2 )P (B2 ) + P (A | B3 )P (B3 )
(3.13)
=
(0.03)(0.20)
≈ 0.14.
(0.04)(0.30) + (0.05)(0.50) + (0.03)(0.20)
(3.14)
Relation (3.13) is a particular case of the general form of Bayes’ formula (Theorem 3.5).
We will now explain how a tree diagram is used to write relation (3.14). Figure 3.5, in
which D stands for “defective” and N for “not defective,” is a tree diagram for this problem.
To find the desired probability all we need do is find (by multiplication) the probability of the
required path, the path from III to D, and divide it by the sum of the probabilities of the paths
that lead to D ’s.
In general, modifying the argument from which (3.13) was deduced, we arrive at the following theorem.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 112 — #128
✐
✐
112
Chapter 3
Conditional Probability and Independence
0.012 + 0.025 + 0.006 = 0.043
0.006/0.043 = 0.14
Figure 3.5
Tree diagram for relation (3.14).
Theorem 3.5 (Bayes’ Theorem)
Let {B1 , B2 , . . . , Bn } be a partition of the sample
space S of an experiment. If for i = 1, 2, . . . , n, P (Bi ) > 0, then for any event A of S with
P (A) > 0,
P (Bk | A) =
P (A | Bk )P (Bk )
.
P (A | B1 )P (B1 ) + P (A | B2 )P (B2 ) + · · · + P (A | Bn )P (Bn )
In statistical applications of Bayes’ theorem, B1 , B2 , . . . , Bn are called hypotheses,
P (Bi ) is called the prior probability of Bi , and the conditional probability P (Bi | A) is
called the posterior probability of Bi after the occurrence of A.
Note that, for any event B of S, B and B c both nonempty, the set {B, B c } is a partition
of S . Thus, by Theorem 3.5,
If P (B) > 0 and P (B c ) > 0, then for any event A of S with P (A) > 0,
P (B | A) =
P (A | B)P (B)
.
P (A | B)P (B) + P (A | B c )P (B c )
Similarly,
P (B c | A) =
P (A | B c )P (B c )
.
P (A | B)P (B) + P (A | B c )P (B c )
These are the simplest forms of Bayes’ formula. They are used whenever the quantities
P (A | B), P (A | B c ), and P (B) are given or can be calculated. The typical situation is
that A happens logically or temporally after B, so that the probabilities P (B) and P (A | B)
can be readily computed. Bayes’ formula is applicable when we know the probability of the
more recent event, given that the earlier event has occurred, P (A | B), and we wish to calculate
the probability of the earlier event, given that the more recent event has occurred, P (B | A).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 113 — #129
✐
✐
Section 3.4
Bayes’ Formula
113
In practice, Bayes’ formula is used when we know the effect of a cause and we wish to make
some inference about the cause.
Theorem 3.5, in its present form, is due to Laplace, who named it after Thomas
Bayes (1701–1761). Bayes, a prominent English philosopher and an ordained minister, did
a comprehensive study of the calculation of P (B | A) in terms of P (A | B). His work was
continued by other mathematicians such as Laplace and Gauss.
We will now present several examples concerning Bayes’ theorem. To emphasize the convenience of tree diagrams, in Example 3.23, we will use a tree diagram as well.
Example 3.19 In a study conducted three years ago, 82% of the people in a randomly selected sample were found to have “good” financial credit ratings, while the remaining 18%
were found to have “bad” financial credit ratings. Current records of the people from that sample show that 30% of those with bad credit ratings have since improved their ratings to good,
while 15% of those with good credit ratings have since changed to having a bad credit rating.
What percentage of people with good credit ratings now had bad ratings three years ago?
Solution: Let G be the event that a randomly selected person from the sample has a good
credit rating now. Let B be the event that he or she had a bad credit rating three years ago. The
desired quantity is P (B | G). By Bayes’ formula,
P (B | G) =
=
P (G | B)P (B)
P (G | B)P (B) + P (G | B c )P (B c )
(0.30)(.18)
= 0.072,
(0.30)(.18) + (.85)(.82)
where P (G | B c ) = 0.85, because the probability is 1 − 0.15 = 0.85 that a person with good
credit rating three years ago has a good credit rating now. Therefore, 7.2% of people with good
credit ratings now had bad ratings three years ago. Example 3.20 During a double homicide murder trial, based on circumstantial evidence
alone, the jury becomes 15% certain that a suspect is guilty. DNA samples recovered from the
murder scene are then compared with DNA samples extracted from the suspect. Given the size
and conditions of the recovered samples, a forensic scientist estimates that the probability of the
sample having come from someone other than the suspect is 10−9 . With this new information,
how certain should the jury be of the suspect’s guilt?
Solution: Let G and I be the events that the suspect is guilty and innocent, respectively. Let
D be the event that the recovered DNA samples from the murder scene match with the DNA
samples extracted from the suspect. Since {G, I} is a partition of the sample space, we can use
Bayes’ formula to calculate P (G | D), the probability that the suspect is the murderer in view
of the new evidence.
P (G | D) =
=
P (D | G)P (G)
P (D | G)P (G) + P (D | I)P (I)
1(.15)
= 0.9999999943.
1(.15) + 10−9 (.85)
This shows that P (I | D) = 1 − P (G | D) is approximately 5.67 × 10−9 , leaving no
reasonable doubt for the innocence of the suspect.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 114 — #130
✐
✐
114
Chapter 3
Conditional Probability and Independence
In some trials, prosecutors have argued that if P (D | I) is “small enough,” then there is
no reasonable doubt for the guilt of the defendant. Such an argument, called the prosecutor’s
fallacy, probably stems from confusing P (D | I) for P (I | D). One should pay attention
to the fact that P (D | I) is infinitesimal regardless of the suspect’s guilt or innocence. To
elaborate this further, note that in this example, P (I | D) is approximately 5.67 times larger
than P (D | I), which is 10−9 . This is because even without the DNA evidence, there is a 15%
chance that the suspect is guilty.
We will now present a situation, which clearly demonstrates that P (D | I) should not
be viewed as the probability of guilt in evaluating reasonable doubt. Suppose that the double
homicide in this example occurred in California, and there is no suspect identified. Suppose that
a DNA data bank identifies a person in South Dakota whose DNA matches what was recovered
at the California crime scene. Furthermore, suppose that the forensic scientist estimates that the
probability of the sample recovered at the crime scene having come from someone other than
the person in South Dakota is still 10−9 . If there is no evidence that this person ever traveled to
California, or had any motive for committing the double homicide, it is doubtful that he or she
would be indicted. In such a case P (D | I) is still 10−9 , but this quantity hardly constitutes
the probability of guilt for the person in South Dakota. The argument that since P (D | I) is
infinitesimal there is no reasonable doubt for the guilt of the person from South Dakota does
not seem convincing at all. In such a case, the quantity P (I | D), which can be viewed as the
probability of guilt, cannot even be estimated if nothing beyond DNA evidence exists. Example 3.21 On the basis of reconnaissance reports, Colonel Smith decides that the probability of an enemy attack against the left is 0.20, against the center is 0.50, and against the right
is 0.30. A flurry of enemy radio traffic occurs in preparation for the attack. Since deception is
normal as a prelude to battle, Colonel Brown, having intercepted the radio traffic, tells General
Quick that if the enemy wanted to attack on the left, the probability is 0.20 that he would have
sent this particular radio traffic. He tells the general that the corresponding probabilities for an
attack on the center or the right are 0.70 and 0.10, respectively. How should General Quick
use these two equally reliable staff members’ views to get the best probability profile for the
forthcoming attack?
Solution: Let A be the event that the attack would be against the left, B be the event that
the attack would be against the center, and C be the event that the attack would be against the
right. Let ∆ be the event that this particular flurry of radio traffic occurs. Colonel Brown has
provided information of conditional probabilities of a particular flurry of radio traffic given that
the enemy is preparing to attack against the left, the center, and the right. However, Colonel
Smith has presented unconditional probabilities, on the basis of reconnaissance reports, for
the enemy attacking against the left, the center, and the right. Because of these, the general
should take the opinion of Colonel Smith as prior probabilities for A, B, and C . That is,
P (A) = 0.20, P (B) = 0.50, and P (C) = 0.30. Then he should calculate P (A | ∆),
P (B | ∆), and P (C | ∆) based on Colonel Brown’s view, using Bayes’ theorem, as follows.
P (A | ∆) =
=
P (∆ | A)P (A)
P (∆ | A)P (A) + P (∆ | B)P (B) + P (∆ | C)P (C)
(0.2)(0.2)
(0.2)(0.2)
=
≈ 0.095.
(0.2)(0.2) + (0.7)(0.5) + (0.1)(0.3)
0.42
Similarly,
P (B | ∆) =
(0.1)(0.3)
(0.7)(0.5)
≈ 0.83 and P (C | ∆) =
≈ 0.071. 0.42
0.42
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 115 — #131
✐
✐
Section 3.4
Bayes’ Formula
115
Example 3.22 In logic, the contrapositive of “A implies B ” is defined to be not B implies
not A. We know that a statement and its contrapositive are equivalent. Show the analogue of
this fact in probability. That is, show that if A and B are two events with probabilities that are
neither 0 nor 1, then P (B | A) = 1 implies that P (Ac | B c ) = 1.
Proof:
Note that P (B c | A) = 1 − P (B | A) = 0. So by Bayes’ theorem,
P (Ac | B c ) =
P (B c | Ac )P (Ac )
P (B c | Ac )P (Ac )
=
= 1. P (B c | Ac )P (Ac ) + P (B c | A)P (A)
P (B c | Ac )P (Ac )
Example 3.23 A box contains seven red and 13 blue balls. Two balls are selected at random
and are discarded without their colors being seen. If a third ball is drawn randomly and observed
to be red, what is the probability that both of the discarded balls were blue?
Solution: Let BB, BR, and RR be the events that the discarded balls are blue and blue,
blue and red, red and red, respectively. Also, let R be the event that the third ball drawn is red.
Since {BB, BR, RR} is a partition of sample space, Bayes’ formula can be used to calculate
P (BB | R).
P (BB | R) =
P (R | BB)P (BB)
.
P (R | BB)P (BB) + P (R | BR)P (BR) + P (R | RR)P (RR)
Now
P (BB) =
39
13 12
×
=
,
20 19
95
P (RR) =
7
6
21
×
=
,
20 19
190
and
P (BR) =
13
7
7
13
91
×
+
×
=
,
20 19 20 19
190
where the last equation follows since BR is the union of two disjoint events: namely, the first
ball discarded was blue, the second was red, and vice versa. Thus
7
39
×
18 95
P (BB | R) =
≈ 0.46.
7
39
6
91
5
21
×
+
×
+
×
18 95 18 190 18 190
This result can be found from the tree diagram of Figure 3.6 as well. Alternatively, reducing
sample space, given that the third ball was red, there are 13 blue and six red balls which could
have been discarded. Thus
P (BB | R) =
13 12
×
≈ 0.46. 19 18
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 116 — #132
✐
✐
116
Chapter 3
Conditional Probability and Independence
6 91
7 39
5 21
·
+
·
+
·
≈ 0.35,
18 190 18 190 18 95
7 39
0.16
·
≈ 0.16,
≈ 0.46.
18 95
0.35
Figure 3.6
Tree diagram for Example 3.23.
EXERCISES
A
1.
In transmitting dot and dash signals, a communication system changes 1/4 of the dots
to dashes and 1/3 of the dashes to dots. If 40% of the signals transmitted are dots and
60% are dashes, what is the probability that a dot received was actually a transmitted
dot?
2.
On a multiple-choice exam with four choices for each question, a student either knows
the answer to a question or marks it at random. If the probability that he or she knows
the answers is 2/3, what is the probability that an answer that was marked correctly was
not marked randomly?
3.
Suppose that 20% of the e-mails Derek receives are spam. Suppose that 12% of the spam
e-mails and 0.05% of the non-spam e-mails are concerning the e-mail storage in Derek’s
e-mail account. Derek has just received an e-mail concerning his e-mail storage. What
is the probability that it is spam?
4.
A judge is 65% sure that a suspect has committed a crime. During the course of the
trial, a witness convinces the judge that there is an 85% chance that the criminal is lefthanded. If 23% of the population is left-handed and the suspect is also left-handed, with
this new information, how certain should the judge be of the guilt of the suspect?
5.
When Professor Wagoner teaches calculus, he only grades 25% of the students’ exam
papers, randomly selected. The remaining papers are graded by his teaching assistants
(TA’s). Suppose that 78% of the papers graded by Professor Wagoner and 86% of the
papers graded by the TA’s get a passing grade. If Harris, a calculus student of Professor
Wagoner, got a passing grade on his last exam, what is the probability that his exam
paper was graded by a TA?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 117 — #133
✐
✐
Section 3.4
Bayes’ Formula
117
6.
When traveling from Springfield, Massachusetts, to Baltimore, Maryland, 30% of the
time Charles takes I-91 south and then I-95 south, and 70% of the time he takes I-91
south, then I-84 west followed by I-81 south and I-83 south. Choosing the first option
takes Charles a random time between 5 hours and 30 minutes and 10 hours depending
upon the traffic. However, choosing the second option takes a random time between 6
hours and 30 minutes and 8 hours. The second option is a longer drive but has much less
traffic. We know that, on his last trip, it took Charles less than 7 hours and 15 minutes
to travel from Springfield to Baltimore. What is the probability that he chose the second
option?
7.
In a trial, the judge is 65% sure that Susan has committed a crime. Julie and Robert
are two witnesses who know whether Susan is innocent or guilty. However, Robert is
Susan’s friend and will lie with probability 0.25 if Susan is guilty. He will tell the truth
if she is innocent. Julie is Susan’s enemy and will lie with probability 0.30 if Susan is
innocent. She will tell the truth if Susan is guilty. What is the probability that Susan is
guilty if Robert and Julie give conflicting testimony?
8.
Suppose that 5% of the men and 2% of the women working for a corporation make over
$120,000 a year. If 30% of the employees of the corporation are women, what percent
of those who make over $120,000 a year are women?
9.
There are three dice in a small box. The first die is unbiased; the second one is loaded,
and, when tossed, the probability of obtaining 6 is 3/8, and the probability of obtaining
each of the other faces is 1/8. The third die is also loaded, but the probability of obtaining
6 when tossed is 2/7, and the probability of each of the other five faces is 1/7. A die was
selected from the box at random and was tossed. If the outcome was a 6, for i = 1, 2, 3,
find the probability that the die selected and tossed was the ith die.
10.
A stack of cards consists of six red and five blue cards. A second stack of cards consists
of nine red cards. A stack is selected at random and three of its cards are drawn. If all of
them are red, what is the probability that the first stack was selected?
11.
Suppose that currently it is a bull market, and in the stock market, share prices are
rising. Furthermore, suppose that the Nasdaq Composite index closes at higher points
86% of the trading days. Iniko is a Nasdaq financial expert, and when she predicts that
the Nasdaq Composite index closes at a lower or equal point the next trading day, her
forecast will turn out to be true 90% of the time. However, her prediction is confirmed
only 75% of the time that the Nasdaq Composite index closes at a higher point the
next trading day. Tomorrow is a trading day, and Iniko has predicted that the Nasdaq
Composite index will close at a higher point than it did last time. What is the probability
that her prediction will turn out to be true?
12.
Based on an insurance company’s evaluation, with probability 0.85, Whitney is a safe
driver, and with probability 0.15 she is not. Suppose that a safe driver avoids at-fault
accidents in the next year with probability 0.87, and an unsafe driver avoids at-fault
accidents in that period with probability 0.54. If next year Whitney gets involved in a
car accident that is her fault, how should the insurance company revise the probability
of Whitney being a safe driver?
13.
An actuary has calculated that, for the age group 16-25, the probability is 0.08 that a
driver whose car is insured by her company gets involved in a car accident within a
year. For the age groups 26-35, 36-65, and 66-100, the corresponding probabilities are
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 118 — #134
✐
✐
118
Chapter 3
Conditional Probability and Independence
0.03, 0.02, and 0.04, respectively. Suppose that the percentages of the company’s insured
drivers belonging to various age groups are given by the following table:
Age Group
16-25
26-35
36-65
66-100
Percentage of Insured Drivers
10%
15%
45%
30%
What proportion of the company’s insured drivers who get involved in an accident within
a year belong to the age group 16-25?
14.
A certain cancer is found in one person in 5000. If a person does have the disease, in
92% of the cases the diagnostic procedure will show that he or she actually has it. If a
person does not have the disease, the diagnostic procedure in one out of 500 cases gives
a false positive result. Determine the probability that a person with a positive test result
has the cancer.
15.
At a grocery store, an absent-minded, honest professor, Dexter, hands the playful cashier,
Hans, a dollar bill. Hans puts the bill in the cash register drawer and gives Dexter $2.75
in change. Dexter, who really does not remember whether he handed him a $10 bill or a
$20 bill, claims that he gave Hans a $20 bill and demands an additional $10 back. Hans
is an honest man as well, and each time dollar bills are given to him, he examines them
and notes their amounts before putting them in the drawer. However, 8% of the time,
due to his playfulness, he does not remember correctly how much money was handed
to him. If the cash register has 13 $20 bills and 18 $10 bills, what is the probability that
Dexter actually handed Hans a $20 bill?
16.
Urns I, II, and III contain three pennies and four dimes, two pennies and five dimes,
three pennies and one dime, respectively. One coin is selected at random from each urn.
If two of the three coins are dimes, what is the probability that the coin selected from
urn I is a dime?
17.
In a study it was discovered that 25% of the paintings of a certain gallery are not original.
A collector in 15% of the cases makes a mistake in judging if a painting is authentic or a
copy. If she buys a piece thinking that it is original, what is the probability that it is not?
18.
There are three identical cards that differ only in color. Both sides of one are black, both
sides of the second one are red, and one side of the third card is black and its other side
is red. These cards are mixed up and one of them is selected at random. If the upper side
of this card is red, what is the probability that its other side is black?
19.
With probability of 1/6 there are i defective fuses among 1000 fuses (i = 0, 1, 2, 3, 4,
5). If among 100 fuses selected at random, none was defective, what is the probability
of no defective fuses at all?
20.
Solve the following problem, asked of Marilyn Vos Savant in the “Ask Marilyn” column
of Parade Magazine, February 18, 1996.
Say I have a wallet that contains either a $2 bill or a $20 bill (with equal
likelihood), but I don’t know which one. I add a $2 bill. Later, I reach into
my wallet (without looking) and remove a bill. It’s a $2 bill. There’s one bill
remaining in the wallet. What are the chances that it’s a $2 bill?
21.
Some studies have shown that 1 in 177 Americans have celiac disease. Nevertheless, in
general, for ordinary physicians, it is hard to diagnose it and, as a result, only 20% of
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 119 — #135
✐
✐
Section 3.4
Bayes’ Formula
119
the patients that have it are actually diagnosed with the disease. The good news is that of
those who do not have celiac disease, 98% test negative. Hoyt is tested positive for this
disease. However, to make sure, he is tested for a second time, and the result is positive
again. Considering the results of both tests, what is the probability that Hoyt has celiac
disease?
B
22.
There are two stables on a farm, one that houses 20 horses and 13 mules, the other with
25 horses and eight mules. Without any pattern, animals occasionally leave their stables
and then return to their stables. Suppose that during a period when all the animals are in
their stables, a horse comes out of a stable and then returns. What is the probability that
the next animal coming out of the same stable will also be a horse?
23.
An urn contains five red and three blue chips. Suppose that four of these chips are selected at random and transferred to a second urn, which was originally empty. If a random chip from this second urn is blue, what is the probability that two red and two blue
chips were transferred from the first urn to the second urn?
24.
There are three dice in a small box. The first die is unbiased; the second one is loaded,
and, when tossed, the probability of obtaining 6 is 3/8, and the probability of obtaining
each of the other faces is 1/8. The third die is also loaded, but the probability of obtaining
6 when tossed is 2/7, and the probability of each of the other five faces is 1/7. A die was
selected from the box at random and was tossed, and the outcome was 6. If the die is
tossed once more, what is the probability that the outcome is 6 again?
25.
The advantage of a certain blood test is that 90% of the time it is positive for patients
having a certain disease. Its disadvantage is that 25% of the time it is also positive in
healthy people. In a certain location 30% of the people have the disease, and anybody
with a positive blood test is given a drug that cures the disease. If 20% of the time
the drug produces a characteristic rash, what is the probability that a person from this
location who has the rash had the disease in the first place?
Self-Quiz on Section 3.4
Time allotted: 20 Minutes
1.
In a small town in Massachusetts, of the 60 students who took the SAT the last time that
it was offered, 17 had attended an SAT test prep course. If taking this course increases
the chances of a student to score 1200 or higher from 20% to 35%, what is the probability
that a student who scored higher than 1200 had attended the prep course? (3 points)
2.
A box has 10 coins of which 3 are gold. Eileen selects a coin at random first. Then
Bernice draws a coin randomly and finds that it is gold. What is the probability that
Eileen’s coin is also gold? (3 points)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 120 — #136
✐
✐
120
3.
3.5
Chapter 3
Conditional Probability and Independence
At a gas station, 85% of the customers use 87 octane gasoline, 5% use 91 octane, and
10% use 93 octane. Suppose that, 70%, 85%, and 95% of 87 octane users, 91 octane
users, and 93 octane users fill their tanks, respectively. At this station, a customer just
filled his tank. What is the probability that he used 93 octane gasoline? (4 points)
INDEPENDENCE
Let A and B be two events of a sample space S, and assume that P (A) > 0 and P (B) > 0.
We have seen that, in general, the conditional probability of A given B is not equal to the
probability of A. However, if it is, that is, if P (A | B) = P (A), we say that A is independent
of B . This means that if A is independent of B, knowledge regarding the occurrence of B does
not change the chance of the occurrence of A. The relation P (A | B) = P (A) is equivalent
to the relations P (AB)/P (B) = P (A), P (AB) = P (A)P (B), P (BA)/P (A) = P (B),
and P (B | A) = P (B). The equivalence of the first and last of these relations implies that
if A is independent of B, then B is independent of A. In other words, if knowledge regarding
the occurrence of B does not change the chance of occurrence of A, then knowledge regarding
the occurrence of A does not change the chance of occurrence of B . Hence independence is
a symmetric relation on the set of all events of a sample space. As a result of this property,
instead of making the definitions “A is independent of B ” and “B is independent of A,” we
simply define the concept of the “independence of A and B .” To do so, we take P (AB) =
P (A)P (B) as the definition. We do this because a symmetrical definition relating A and B
does not readily follow from either of the other relations given i.e., P (A | B) = P (A) or
P (B | A) = P (B) . Moreover, these relations require either that P (B) > 0 or P (A) > 0,
whereas our definition does not.
Definition 3.3
Two events A and B are called independent if
P (AB) = P (A)P (B).
If two events are not independent, they are called dependent. If A and B are independent, we
say that {A, B} is an independent set of events.
Note that in this definition we did not require P (A) or P (B) to be strictly positive. Hence
by this definition any event A with P (A) = 0 or 1 is independent of every event B (see
Exercise 16).
Example 3.24 In the experiment of tossing a fair coin twice, let A and B be the events of
getting heads on the first and second tosses, respectively. Intuitively, it is clear that A and B
are independent. To prove this mathematically, note that P (A) = 1/2 and P (B) = 1/2. But
since the sample space of this experiment consists of the four equally probable events: HH,
HT, TH, and TT, we have P (AB) = P (HH) = 1/4. Hence P (AB) = P (A)P (B) is valid,
implying the independence of A and B . It is interesting to know that Jean Le Rond d’Alembert
(1717–1783), a French mathematician, had argued that, since in the experiment of tossing a fair
coin twice the possible number of heads is 0, 1, and 2, the probability of no heads, one heads,
and two heads, each is 1/3. Example 3.25 In the experiment of drawing a card from an ordinary deck of 52 cards,
let A and B be the events of getting a heart and an ace, respectively. Whether A and B are
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 121 — #137
✐
✐
Section 3.5
Independence
121
independent cannot be answered easily on the basis of intuition alone. However, using the
defining formula, P (AB) = P (A)P (B), this can be answered at once since P (AB) = 1/52,
P (A) = 1/4, P (B) = 1/13, and 1/52 = 1/4 × 1/13. Hence A and B are independent
events. Example 3.26 An urn contains five red and seven blue balls. Suppose that two balls are
selected at random and with replacement. Let A and B be the events that the first and the
second balls are red, respectively. Then, using the counting principle, Theorem 2.1, we get
P (AB) = (5×5)/(12×12). Now P (AB) = P (A)P (B) since P (A) = 5/12 and P (B) =
5/12. Thus A and B are independent. If we do the same experiment without replacement, then
P (B | A) = 4/11 while
P (B) = P (B | A)P (A) + P (B | Ac )P (Ac )
5
5
7
5
4
×
+
×
=
,
=
11 12 11 12
12
which might be quite surprising to some. But it is true. If no information is given on the
outcome of the first draw, there is no reason for the probability of the second ball being red to
differ from 5/12. Thus P (B | A) 6= P (B), implying that A and B are dependent. Example 3.27 In the experiment of selecting a random number from the set of natural numbers {1, 2, 3, . . . , 100}, let A, B, and C denote the events that they are divisible by 2, 3, and
5, respectively. Clearly, P (A) = 1/2, P (B) = 33/100, P (C) = 1/5, P (AB) = 16/100,
and P (AC) = 1/10. Hence A and B are dependent while A and C are independent. Note
that if the random number is selected from {1, 2, 3, . . . , 300}, then each of {A, B}, {A, C},
and {B, C} is an independent set of events. This is because 300 is divisible by 2, 3, and 5, but
100 is not divisible by 3. Figure 3.7
Spinner of Example 3.28.
Example 3.28 A spinner is mounted on a wheel. Arcs A, B, and C, of equal length, are
marked off on the wheel’s perimeter (see Figure 3.7). In a game of chance, the spinner is flicked,
and depending on whether it stops on A, B, or C, the player wins 1, 2, or 3 points, respectively.
Suppose that a player plays this game twice. Let E denote the event that he wins 1 point in the
first game and any number of points in the second. Let F be the event that he wins a total of
3 points in both games, and G be the event that he wins a total of 4 points in both games. The
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 122 — #138
✐
✐
122
Chapter 3
Conditional Probability and Independence
sample space of this experiment has 3 × 3 = 9 elements, E = (1, 1), (1, 2), (1, 3) , F =
(1, 2), (2, 1) , and G = (1, 3), (2, 2), (3, 1) . Therefore, P (E) = 1/3, P (G) = 1/3,
P (F ) = 2/9, P (F E) = 1/9, and P (GE) = 1/9. These show that E and G are independent,
whereas E and F are not. To justify these intuitively, note that if we are interested in a total of
3 points, then getting 1 point in the first game is good luck, since obtaining 3 points in the first
game makes it impossible to win a sum of 3. However, if we are interested in a sum of 4 points,
it does not matter what we win in the first game. We have the same chance of obtaining a sum
of 4 if the first game results in any of the numbers 1, 2, or 3. We now prove that if A and B are independent events, so are A and B c .
Theorem 3.6
Proof:
If A and B are independent, then A and B c are independent as well.
By Theorem 1.7,
P (A) = P (AB) + P (AB c ).
Therefore,
P (AB c ) = P (A) − P (AB) = P (A) − P (A)P (B)
= P (A) 1 − P (B) = P (A)P (B c ). Corollary
If A and B are independent, then Ac and B c are independent as well.
Proof: A and B are independent, so by Theorem 3.6 the events A and B c are independent.
Now, using the same theorem again, we have that Ac and B c are independent. Thus, if A and B are independent, knowledge about the occurrence or nonoccurrence of A
does not change the chances of the occurrence or nonoccurrence of B, and vice versa.
Remark 3.3 If A and B are mutually exclusive events and P (A) > 0, P (B) > 0, then
they are dependent. This is because, if we are given that one has occurred, the chance of the
occurrence of the other one is zero. That is, the occurrence of one of them precludes the occurrence of the other. For example, let A be the event that the next president of the United States
is a Democrat and B be the event that he or she is a Republican. Then A and B are mutually
exclusive; hence they are dependent. If A occurs, that is, if the next president is a Democrat,
the probability that B occurs, that is, he or she is a Republican is zero, and vice versa. The following example shows that if A is independent of B and if A is independent of C,
then A is not necessarily independent of BC or of B ∪ C .
Example 3.29 Dennis arrives at his office every day at a random time between 8:00 A.M. and
9:00 A.M. Let A be the event that Dennis arrives at his office tomorrow between 8:15 A.M. and
8:45 A.M. Let B be the event that he arrives between 8:30 A.M. and 9:00 A.M., and let C be the
event that he arrives either between 8:15 A.M. and 8:30 A.M. or between 8:45 A.M. and 9:00 A.M.
Then AB, AC, BC, and B ∪C are the events that Dennis arrives at his office between 8:30 and
8:45, 8:15 and 8:30, 8:45 and 9:00, and 8:15 and 9:00, respectively. Thus P (A) = P (B) =
P (C) = 1/2 and P (AB) = P (AC) = 1/4. So P (AB) = P (A)P (B) and P (AC) =
P (A)P (C); that is, A is independent of B and it is independent of C . However, since BC
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 123 — #139
✐
✐
Section 3.5
Independence
123
and A are mutually exclusive, they are dependent. Also, P (A | B ∪ C) = 2/3 6= P (A). Thus
B ∪ C and A are dependent as well. Example 3.30 (Jailer’s Paradox) The jailer of a prison in which Alex, Ben, and Tim
are held is the only person, other than the judge, who knows which of these three prisoners is
condemned to death, and which two will be freed. The prisoners know that exactly two of them
will go free; they do not know which two. Alex has written a letter to his fiancée. Just in case
he is not one of the two who will be freed, he wants to give the letter to a prisoner who goes
free to deliver. So Alex asks the jailer to tell him the name of one of the two prisoners who will
go free. The jailer refuses to give that information to Alex. To begin with, he is not allowed
to tell Alex whether Alex goes free or not. Putting Alex aside, he argues that, if he reveals the
name of a prisoner who will go free, then the probability of Alex dying increases from 1/3 to
1/2. He does not want to do that.
As Zweifel notes in the June 1986 issue of Mathematics Magazine, page 156, “this seems
intuitively suspect, since the jailer is providing no new information to Alex, so why should
his probability of dying change?” Zweifel is correct. Just revealing to Alex that Ben goes free,
or just revealing to him that Tim goes free is not the type of information that changes the
probability of Alex dying. What changes the probability of Alex dying is telling him whether
both Ben and Tim are going free or exactly one of them is going free.
To explain this paradox, we will show that, under suitable conditions, if the jailer tells Alex
that Tim goes free, still the probability is 1/3 that Alex dies. Telling Alex that Tim is going free
reveals to Alex that the probability of Ben dying is 2/3. Similarly, telling Alex that Ben is going
free reveals to Alex that the probability of Tim dying is 2/3. To show these facts, let A, B, and
T be the events that “Alex dies,” “Ben dies,” and “Tim dies.” Let
ω1 = Tim dies, and the jailer tells Alex that Ben goes free
ω2 = Ben dies, and the jailer tells Alex that Tim goes free
ω3 = Alex dies, and the jailer tells Alex that Ben goes free
ω4 = Alex dies, and the jailer tells Alex that Tim goes free.
The sample space of all possible episodes is S = {ω1 , ω2 , ω3 , ω4 }. Now, if Tim dies, with
probability 1, the jailer will tell Alex that Ben goes free. Therefore, ω1 occurs if and only if
Tim dies. This implies that P (ω1 ) = 1/3. Similarly, if Ben dies, with probability 1, the jailer
will tell Alex that Tim goes free. Therefore, ω2 occurs if and only if Ben dies. This shows that
P (ω2 ) = 1/3. To assign probabilities to ω3 and ω4 , we will make two assumptions: (1) If Alex
is the one who is scheduled to die, then the event that the jailer will tell Alex that Ben goes free
is independent of the event that the jailer will tell Alex that Tim goes free, and (2) if Alex is the
one who is scheduled to die, then the probability that the jailer will tell Alex that Ben goes free
is 1/2; hence the probability that he will tell Alex that Tim goes free is 1/2 as well. Under these
conditions, P (ω3 ) = P (ω4 ) = 1/6. Let J be the event that “the jailer tells Alex that Tim goes
free;” then
1
P (ω4 )
1
P (AJ)
6
=
=
= ,
P (A | J) =
P (J)
P (ω2 ) + P (ω4 )
3
1 1
+
3 6
This shows that if the jailer reveals no information about the fate of Alex, telling Alex the name
of one prisoner who goes free does not change the probability of Alex dying; it remains to be
1/3.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 124 — #140
✐
✐
124
Chapter 3
Conditional Probability and Independence
Note that the decision as which of Ben, Tim, or Alex is condemned to death, and which
two will be freed, has been made by the judge. The jailer has no option to control Alex’s fate.
With probability 1, the jailer knows which one of the three dies and which two will go free. It
is Alex who doesn’t know any of these probabilities. Alex can only analyze these probabilities
based on the information he receives from the jailer. If Alex is dying, and the jailer disobeys the
conditions (1) and (2), then he could decide to tell Alex, with arbitrary probabilities, which of
Ben or Tim goes free. In such a case, P (A | J) will no longer equal 1/3. It will vary depending
on the probabilities of J and J c . If the jailer insists on not giving extra information to Alex, he
should obey conditions (1) and (2). Furthermore, the jailer should tell Alex what his rules are
when he reveals information. Alex can only analyze the probability of dying based on the full
information he receives from the jailer. If the jailer does not reveal his rules to Alex, there is no
way for Alex to know whether the probability that he is the unlucky prisoner has changed or
not.
Zweifel analyzes this paradox by using Bayes’ formula:
P (A | J) =
P (J | A)P (A)
P (J | A)P (A) + P (J | B)P (B) + P (J | T )P (T )
1 1
×
1
2 3
=
= .
3
1 1
1
1
× +1× +0×
2 3
3
3
Similarly, if D is the event that “the jailer tells Alex that Ben goes free,” then P (A | D) = 1/3.
In his explanation of this paradox Zweifel writes
Aside from the formal application of Bayes’ theorem, one would like to understand this “paradox” from an intuitive point of view. The crucial point which
can perhaps make the situation clear is the discrepancy between P (J | A)
and P (J | B) noted above. Bridge players call this the “Principle of Restricted
Choice.” The probability of a restricted choice is obviously greater than that of
a free choice, and a common error made by those who attempt to solve such
problems intuitively is to overlook this point. In the case of the jailer’s paradox,
if the jailer says “Tim will go free” this is twice as likely to occur when Ben
is scheduled to die (restricted choice; jailer must say “Tim”) as when Alex is
scheduled to die (free choice; jailer could say either “Tim” or “Ben”).
The jailer’s paradox and its solution have been around for a long time. Despite this, when
the same problem in a different context came up in the “Ask Marilyn” column of Parade Magazine on September 9, 1990 (see Example 3.9), and Marilyn Vos Savant† gave the correct answer
to the problem in the December 2, 1990, issue, she was taken to task by three mathematicians.
By the time the February 17, 1991 issue of Parade Magazine was published, Vos Savant had
received 2000 letters on the problem, of which 92% of general respondents and 65% of university respondents opposed her answer. This story reminds us of the statement of De Moivre in
his dedication of The Doctrine of Chance:
Some of the Problems about Chance having a great appearance of Simplicity,
the Mind is easily drawn into a belief, that their Solution may be attained by
†
Ms. Vos Savant is the writer of a reader-correspondence column and is listed in the Guinness Book of
Records Hall of Fame for “highest IQ.”
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 125 — #141
✐
✐
Section 3.5
Independence
125
the mere Strength of natural good Sense; which generally proving otherwise
and the Mistakes occasioned thereby being not unfrequent, ‘tis presumed that
a Book of this Kind, which teaches to distinguish Truth from what seems so
nearly to resemble it, will be looked upon as a help to good Reasoning.
We now extend the concept of independence to three events: A, B, and C are called independent if knowledge about the occurrence of any of them, or the joint occurrence of any two of
them, does not change the chances of the occurrence of the remaining events. That is, A, B,
and C are independent if {A, B}, {A, C}, {B, C}, {A, BC}, {B, AC}, and {C, AB} are
all independent sets of events. Hence A, B, and C are independent if
P (AB) = P (A)P (B),
P (AC) = P (A)P (C),
P (BC) = P (B)P (C),
P A(BC) = P (A)P (BC),
P B(AC) = P (B)P (AC),
P C(AB) = P (C)P (AB).
Now note that these relations can be reduced since the first three and the relation P (ABC) =
P (A)P (B)P (C) imply the last three relations. Hence the definition of the independence of
three events can be shortened as follows.
Definition 3.4
The events A, B, and C are called independent if
P (AB) = P (A)P (B),
P (AC) = P (A)P (C),
P (BC) = P (B)P (C),
P (ABC) = P (A)P (B)P (C).
If A, B, and C are independent events, we say that {A, B, C} is an independent set of events.
The following example demonstrates that P (ABC) = P (A)P (B)P (C), in general, does
not imply that {A, B, C} is a set of independent events.
Example 3.31 Let an experiment consist of throwing a die twice. Let A be the event that
in the second throw the die lands 1, 2, or 5; B the event that in the second throw it lands 4,
5, or 6; and C the event that the sum of the two outcomes is 9. Then P (A) = P (B) = 1/2,
P (C) = 1/9, and
1
1
6= = P (A)P (B),
6
4
1
1
P (AC) =
6=
= P (A)P (C),
36
18
1
1
P (BC) =
6=
= P (B)P (C),
12
18
P (AB) =
while
P (ABC) =
1
= P (A)P (B)P (C).
36
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 126 — #142
✐
✐
126
Chapter 3
Conditional Probability and Independence
Thus the validity of P (ABC) = P (A)P (B)P (C) is not sufficient for the independence of
A, B, and C . If A, B, and C are three events and the occurrence of any of them does not change the
chances of the occurrence of the remaining two, we say that A, B, and C are pairwise independent. Thus {A, B, C} forms a set of pairwise independent events if P (AB) = P (A)P (B),
P (AC) = P (A)P (C), and P (BC) = P (B)P (C). The difference between pairwise independent events and independent events is that, in the former, knowledge about the joint occurrence of any two of them may change the chances of the occurrence of the remaining one, but
in the latter it would not. The following example illuminates the difference between pairwise
independence and independence.
Example 3.32 A regular tetrahedron is a body that has four faces and, if it is tossed, the
probability that it lands on any face is 1/4. Suppose that one face of a regular tetrahedron has
three colors: red, green, and blue. The other three faces each have only one color: red, blue, and
green, respectively. We throw the tetrahedron once and let R, G, and B be the events that the
face on which it lands contains red, green, and blue, respectively. Then P (R | G) = 1/2 =
P (R), P (R | B) = 1/2 = P (R), and P (B | G) = 1/2 = P (B). Thus the events R,
B, and G are pairwise independent. However, R, B, and G are not independent events since
P (R | GB) = 1 =
6 P (R). The independence of more than three events may be defined in a similar manner. A set
of n events A1 , A2 , . . . , An is said to be independent if knowledge about the occurrence of
any of them or the joint occurrence of any number of them does not change the chances of
the occurrence of the remaining events. If we write this definition in terms of formulas, we get
many equations. Similar to the case of three events, if we reduce the number of these equations
to the minimum number that can be used to have all of the formulas satisfied, we reach the
following definition.
Definition 3.5 The set of events {A1 , A2 , . . . , An } is called independent if for every subset
{Ai1 , Ai2 , . . . , Aik }, k ≥ 2, of {A1 , A2 , . . . , An },
P (Ai1 Ai2 · · · Aik ) = P (Ai1 )P (Ai2 ) · · · P (Aik ).
(3.15)
This definition is not in fact limited to finite sets and is extended to infinite sets of
events
∞
(countable or uncountable) in the obvious way. For example, the sequence of events Ai i=1
is called independent if for any of its subsets {Ai1 , Ai2 , . . . , Aik }, k ≥ 2, (3.15) is valid.
By Definition 3.5, events A1 , A2 , . . . , An are independent if, for all combinations
1 ≤ i < j < k < · · · ≤ n, the relations
P (Ai Aj ) = P (Ai )P (Aj ),
P (Ai Aj Ak ) = P (Ai )P (Aj )P (Ak ),
..
.
P (A1 A2 · · · An ) = P (A1 )P (A2 ) · · · P (An )
!
n
are valid. Now we see that the first line stands for
equations, the second line stands for
2
!
!
n
n
equations, . . . , and the last line stands for
equations. Therefore, A1 , A2 , . . . , An
3
n
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 127 — #143
✐
✐
Section 3.5
Independence
127
!
!
!
n
n
n
+
+ ··· +
relations are satisfied. Note
are independent if all of the above
3
n
2
that, by the binomial expansion (see Theorem 2.5 and Example 2.30),
!
!
!
!
!
n
n
n
n
n
+
+ ··· +
= (1 + 1)n −
−
= 2n − n − 1.
3
n
0
1
2
Thus the number of these equations is 2n − n − 1. Although these equations seem to be
cumbersome to check, it usually turns out that they are obvious and checking is not necessary.
Figure 3.8
Electric circuit of Example 3.33.
Example 3.33 Figure 3.8 shows an electric circuit in which each of the switches located at
1, 2, 3, and 4 is independently closed or open with probabilities p and 1 − p, respectively. If a
signal is fed to the input, what is the probability that it is transmitted to the output?
Solution: Let Ei be the event that the switch at location i is closed, 1 ≤ i ≤ 4. A signal
fed to the input will be transmitted to the output if at least one of the events E1 E2 and E3 E4
occurs. Hence the desired probability is
P (E1 E2 ∪ E3 E4 ) = P (E1 E2 ) + P (E3 E4 ) − P (E1 E2 E3 E4 )
= p2 + p2 − p4 = p2 (2 − p2 ). Example 3.34 We draw cards, one at a time, at random and successively from an ordinary
deck of 52 cards with replacement. What is the probability that an ace appears before a face
card?
Solution: We will explain two different techniques that may be used to solve this type of
problems. For a third technique, see Exercise 31, Section 12.3.
Technique 1: Let E be the event of an ace appearing before a face card. Let A, F, and
B be the events of ace, face card, and neither in the first experiment, respectively. Then, by the
law of total probability,
P (E) = P (E | A)P (A) + P (E | F )P (F ) + P (E | B)P (B).
Thus
P (E) = 1 ×
4
12
36
+0×
+ P (E | B) × .
52
52
52
(3.16)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 128 — #144
✐
✐
128
Chapter 3
Conditional Probability and Independence
Now note that since the outcomes of successive experiments are all independent of each other,
when the second experiment begins, the whole probability process starts all over again. Therefore, if in the first experiment neither a face card nor an ace are drawn, the probability of E
before doing the first experiment and after it would be the same; that is, P (E | B) = P (E).
Thus Equation (3.16) gives
P (E) =
36
4
+ P (E) × .
52
52
Solving this equation for P (E), we obtain P (E) = 1/4, a quantity expected because the
number of face cards is three times the number of aces (see also Example 3.35).
Technique 2: Let An be the event that no face card or ace appears on the first
(nS
− 1) drawings, and the nth draw is an ace. Then the event of “an ace before a face card”
∞
is n=1 An . Now {An , n ≥ 1} forms a sequence of mutually exclusive events because, if
n 6= m, simultaneous occurrence of An and Am is the impossible event that an ace appears for
the first time in the nth and mth draws. Hence
P
∞
[
An =
n=1
∞
X
P (An ).
n=1
To compute P (An ), note that P (an ace on any draw) = 1/13 and P (no face card and no ace
in any trial) = 9/13. By the independence of trials we obtain
P (An ) =
Therefore,
P
∞
[
9 n−1 1 .
13
13
∞ ∞
X
9 n−1 1 1 X 9 n−1
An =
=
13
13
13 n=1 13
n=1
n=1
=
1
1
1
·
= .
13 1 − 9/13
4
P
Here P is calculated from the geometric series theorem: For a 6= 0, |r| < 1, the geometric
∞
series n=m ar n converges to ar m /(1 − r).
It is interesting to observe that, in this problem, if cards are drawn without replacement,
even though the trials are no longer independent, the answer would still be the same. For 1 ≤
n ≤ 37, let En be the event that no face card or ace appears on the first n − 1 drawings; let
Fn be the
that the nth draw is an ace. Then the event of “an ace appearing before a face
Sevent
37
card” is n=1 En Fn . Clearly, {En Fn , 1 ≤ n ≤ 37} forms a sequence of mutually exclusive
events. Hence
P
37
[
n=1
37
37
X
X
En Fn =
P (En Fn ) =
P (En )P (Fn | En )
n=1
=
37
X
n=1
!
n=1
36
n−1
4
1
!×
= . 52 − (n − 1)
4
52
n−1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 129 — #145
✐
✐
Section 3.5
Independence
129
Example 3.35 Let S be the sample space of a repeatable experiment. Let A and B be
mutually exclusive events of S . Prove that, in independent trials of this experiment, the event
A occurs before the event B with probability P (A)/[P (A) + P (B)].
Proof: Let P (A) = p and P (B) = q . Let En be the event that none of A and B occurs in
the first n − 1 trials and the outcome of the nth experiment is A. The desired probability is
P
∞
[
n=1
∞
∞
X
X
P (En ) =
(1 − p − q)n−1 p =
En =
n=1
n=1
p
p
=
. 1 − (1 − p − q)
p+q
Example 3.36 Show that in successive tosses of a fair die indefinitely, the probability of
obtaining no 6 is 0.
Solution:
Clearly,
For n ≥ 1, let En be the event of at least one 6 in the first n tosses of the die.
E1 ⊆ E2 ⊆ · · · ⊆ En ⊆ En+1 ⊆ · · · .
S∞
Therefore, En ’s form an increasing sequence of events. Note that limn→∞ En = i=1 Ei
is the event that in successive tosses of the die indefinitely, eventually a 6 will occur. By the
Continuity of the Probability Function (Theorem 1.8), we have
h
5 n i
5 n
P lim En = lim P (En ) = lim 1 −
= 1 − lim
= 1 − 0 = 1.
n→∞
n→∞
n→∞
n→∞ 6
6
This shows that, with probability 1, eventually a 6 will occur. Therefore, the probability of no
6 ever is 0. Example 3.37 Adam tosses a fair coin n + 1 times, Andrew tosses the same coin n times.
What is the probability that Adam gets more heads than Andrew?
Solution: Let H1 and H2 be the number of heads obtained by Adam and Andrew, respectively.
Also, let T1 and T2 be the number of tails obtained by Adam and Andrew, respectively. Since
the coin is fair,
P (H1 > H2 ) = P (T1 > T2 ).
But
P (T1 > T2 ) = P (n + 1 − H1 > n − H2 ) = P (H1 ≤ H2 ).
Therefore, P (H1 > H2 ) = P (H1 ≤ H2 ). So
P (H1 > H2 ) + P (H1 ≤ H2 ) = 1
implies that
P (H1 > H2 ) = P (H1 ≤ H2 ) =
1
.
2
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 130 — #146
✐
✐
130
Chapter 3
Conditional Probability and Independence
Note that a combinatorial solution to this problem is neither elegant nor easy to handle:
P (H1 > H2 ) =
n
X
i=0
=
P (H1 > H2 | H2 = i)P (H2 = i)
n
n+1
X
X
P (H1 = j)P (H2 = i)
i=0 j=i+1
=
(n + 1)!
n!
n
n+1
X
X j! (n + 1 − j)! i! (n − i)!
=
1
2n
! !
n+1
n
.
j
i
2n+1
i=0 j=i+1
n
n+1
X
X
22n+1 i=0 j=i+1
However, comparing these two solutions, we obtain the following interesting identity.
! !
n
n+1
X
X n+1
n
= 22n . j
i
i=0 j=i+1
EXERCISES
A
1.
Jean le Rond d’Alembert, a French mathematician, believed that in successive flips of a
fair coin, after a long run of heads, a tail is more likely. Do you agree with d’Alembert
on this? Explain.
2.
Clark and Anthony are two old friends. Let A be the event that Clark will attend Anthony’s funeral. Let B be the event that Anthony will attend Clark’s funeral. Are A and
B independent? Why or why not?
3.
In a certain country, the probability that a fighter plane returns from a mission without
mishap is 49/50, independent of other missions. In a conversation, Mia concluded that
any pilot who flew 49 consecutive missions without mishap should be returned home
before the fiftieth mission. But, on considering the matter, Jim concluded that the probability of a randomly selected pilot being able to fly 49 consecutive missions safely is
(49/50)49 = 0.3716017. In other words, the odds are almost two to one against an
ordinary pilot performing the feat that this pilot has already performed. Hence the pilot
would seem to be more skillful than most and thus has a better chance of surviving the
50th mission. Who is right, Mia, Jim, or neither of them? Explain.
4.
A fair die is rolled twice. Let A denote the event that the sum of the outcomes is odd,
and B denote the event that it lands 2 on the first toss. Are A and B independent? Why
or why not?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 131 — #147
✐
✐
Section 3.5
Independence
131
5.
The only information revealed concerning the three children of the new mayor of a large
town is that their names are Daly, Emmett, and Karina. Is the event that Daly is younger
than Emmett independent of the event that Karina is younger than Emmett?
6.
Cal Ripken is a distinguished American baseball player who played 21 seasons in Major
League Baseball for the Baltimore Orioles (1981-2001). Among his many achievements,
the most important one is that he holds the record for playing in 2,632 consecutive games
over more than 16 years. Suppose that, independently of other games, the probability is
0.005 (5 in 1000) that a baseball player cannot play a game due to illness, injury, family
emergency, or other emergencies. What is the probability that a baseball player breaks
Cal Ripken’s record?
7.
An urn has three red and five blue balls. Suppose that eight balls are selected at random
and with replacement. What is the probability that the first three are red and the rest are
blue balls?
8.
Suppose that two points are selected at random and independently from the interval
(0, 1). What is the probability that the first one is less than 3/4, and the second one is
greater than 1/4?
9.
According to a recent mortality table, the probability that a 35-year-old U.S. citizen will
live to age 65 is 0.725. (a) What is the probability that John and Jim, two 35-year-old
Americans who are not relatives, both live to age 65? (b) What is the probability that
neither John nor Jim lives to that age?
10.
The Italian mathematician Giorlamo Cardano once wrote that if the odds in favor of an
event are 3 to 1, then the odds in favor of the occurrence of that event in two consecutive
independent experiments are 9 to 1. (He squared 3 and 1 to obtain 9 to 1.) Was Cardano
correct?
11.
Consider the four “unfolded” dice in Figure 3.9 designed by Stanford professor Bradley
Effron. Clearly, none of these dice is an ordinary die with sides numbered 1 through 6.
A game consists of two players each choosing one of these four dice and rolling it. The
player rolling a larger number is the winner. Show that if all four dice are fair, it is twice
as likely for die A to beat die B, twice as likely for die B to beat die C, twice as likely
for die C to beat D, and surprisingly, it is twice as likely for die D to beat die A.
Figure 3.9
12.
†
Effron dice.
(Chevalier de Méré’s Paradox† ) In the seventeenth century in France there were
two popular games, one to obtain at least one 6 in four throws of a fair die and the other
Some scholars consider this problem to be the inception of modern probability theory.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 132 — #148
✐
✐
132
Chapter 3
Conditional Probability and Independence
to bet on at least one double 6 in 24 throws of two fair dice. French nobleman and mathematician Chevalier de Méré argued that the probabilities of a 6 in one throw of a fair
die and a double 6 in one throw of two fair dice are 1/6 and 1/36, respectively. Therefore,
the probability of at least one 6 in four throws of a fair die and at least one double 6 in
24 throws of two fair dice are 4 × 1/6 = 2/3 and 24 × 1/36 = 2/3, respectively.
However, even then experience had convinced gamblers that the probability of winning
the first game was higher than winning the second game, a contradiction. Explain why
de Méré’s argument is false, and calculate the correct probabilities of winning in these
two games.
13.
In data communications, a message transmitted from one end is subject to various
sources of distortion and may be received erroneously at the other end. Suppose that
a message of 64 bits (a bit is the smallest unit of information and is either 1 or 0) is
transmitted through a medium. If each bit is received incorrectly with probability 0.0001
independently of the other bits, what is the probability that the message received is free
of error?
14.
Find an example in which P (AB) < P (A)P (B).
15.
Show that if an event A is independent of itself, then P (A) = 0 or 1.
16.
(a)
Show that if P (A) = 1, then P (AB) = P (B).
(b)
Prove that any event A with P (A) = 0 or P (A) = 1 is independent of every
event B .
17.
Show that if A and B are independent and A ⊆ B, then either P (A) = 0 or P (B) = 1.
18.
Suppose that 55% of the customers of a shoestore buy black shoes. Find the probability
that at least one of the next six customers who purchase a pair of shoes from this store
will buy black shoes. Assume that these customers decide independently.
19.
An actuary studying the insurance preferences of homeowners in California discovers
that the event that a homeowner purchases earthquake coverage is independent of the
event that he or she purchases flood insurance. Furthermore, the actuary observes that a
homeowner is five times as likely to purchase earthquake coverage as flood insurance.
If the probability is 0.19 that a homeowner purchases both earthquake and flood coverage, what is the probability that a homeowner purchases neither earthquake nor flood
insurance?
20.
Three missiles are fired at a target and hit it independently, with probabilities 0.7, 0.8,
and 0.9, respectively. What is the probability that the target is hit?
21.
In the ball bearing manufacturing process, for each item being made, two types of defects
will occur, independently of each other, with probabilities 0.03 and 0.05, respectively.
What is the probability that a randomly selected ball bearing has at least one defect?
22.
In his book, Probability 1, published by Harcourt Brace and Company, 1998, Amir
Aczel estimates that the probability of life for any one given star in the known universe
is 0.000,000,000,000,05 independently of life for any other star. Assuming that there
are 100 billion galaxies in the universe and each galaxy has 300 billion stars, what is
the probability of life on at least one other star in the known universe? How does this
probability change if there were only a billion galaxies, each having 10 billion stars?
23.
In a tire factory, the quality control inspector examines a randomly chosen sample of 15
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 133 — #149
✐
✐
Section 3.5
Independence
133
tires. When more than one defective tire is found, production is halted, the existing tires
are recycled, and production is then resumed. The purpose of this process is to ensure
that the defect rate is no higher than 6%. Occasionally, by a stroke of bad luck, even if
the defect rate is no higher than 6%, more than one defective tire occurs among the 15
chosen. Determine the fraction of the time that this happens.
24.
In the experiment of rolling two fair dice successively, what is the probability that a sum
of 5 appears before a sum of 7?
Hint: See Example 3.35.
25.
An experiment consists of first tossing a fair coin and then drawing a card randomly
from an ordinary deck of 52 cards with replacement. If we perform this experiment
successively, what is the probability of obtaining heads on the coin before an ace from
the cards?
Hint: See Example 3.35.
26.
In a community of M men and W women, m men and w women smoke, where m ≤ M
and w ≤ W. If a person is selected at random and A and B are the events that the person
is a man and smokes, respectively, under what conditions are A and B independent?
27.
Prove that if A, B, and C are independent, then A and B ∪ C are independent. Also
show that A − B and C are independent.
28.
A fair die is rolled six times. If on the ith roll, 1 ≤ i ≤ 6, the outcome is i, we say that
a match has occurred. What is the probability that at least one match occurs?
29.
There are n cards in a box numbered 1 through n. We draw cards successively and at
random with replacement. If the ith draw is the card numbered i, we say that a match has
occurred. (a) What is the probability of at least one match in n trials? (b) What happens
if n increases without bound?
30.
In a certain county, 15% of patients suffering heart attacks are younger than 40, 20% are
between 40 and 50, 30% are between 50 and 60, and 35% are above 60. On a certain
day, 10 unrelated patients suffering heart attacks are transferred to a county hospital. If
among them there is at least one patient younger than 40, what is the probability that
there are two or more patients younger than 40?
31.
When a loaded die is rolled, the outcome is 6 with probability p. Find the probability
that in successive rolls of the die, all outcomes are 6’s.
32.
If the events A and B are independent and the events B and C are independent, is it true
that the events A and C are also independent? Why or why not?
33.
From the set of all families with three children a family is selected at random. Let A be
the event that “the family has children of both sexes” and B be the event that “there is at
most one girl in the family.” Are A and B independent? Answer the same question for
families with two children and families with four children. Assume that for any family
size all sex distributions have equal probabilities.
34.
An event occurs at least once in four independent trials with probability 0.59. What is
the probability of its occurrence in one trial?
35.
Let {A1 , A2 , . . . , An } be an independent set of events and P (Ai ) = pi , 1 ≤ i ≤ n.
(a) What is the probability that at least one of the events A1 , A2 , . . . , An occurs?
(b) What is the probability that none of the events A1 , A2 , . . . , An occurs?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 134 — #150
✐
✐
134
Chapter 3
36.
From the set of all families with two children, a family is selected at random and is found
to have a girl who was born in January. What is the probability that the other child of
the family is a girl? Assume that in a two-child family all sex distributions are equally
probable, gender of a child is independent of the month he or she is born, and all months
are equally likely to be a birth month of a randomly selected child.
37.
Figure 3.10 shows an electric circuit in which each of the switches located at 1, 2, 3, 4,
5, and 6 is independently closed or open with probabilities p and 1 − p, respectively. If
a signal is fed to the input, what is the probability that it is transmitted to the output?
Conditional Probability and Independence
Figure 3.10
38.
Electric circuit of Exercise 37.
Cards are drawn at random and with replacement from an ordinary deck of 52 cards,
successively and indefinitely. Using Theorem 1.8, show that the probability that the ace
of hearts never occurs is 0.
B
39.
An urn contains two red and four white balls. Balls are drawn from the urn successively,
at random and with replacement. What is the probability that exactly three whites occur
in the first five trials?
40.
A fair coin is tossed n times. Show that the events “at least two heads” and “one or two
tails” are independent if n = 3 but dependent if n = 4.
41.
A fair coin is flipped indefinitely. What is the probability of (a) at least one head in the
first n flips; (b) exactly k heads in the first n flips; (c) getting heads in all of the flips
indefinitely?
42.
If two fair dice are tossed six times, what is the probability that the sixth sum obtained
is not a repetition?
43.
Suppose that an airplane passenger whose itinerary requires a change of airplanes in
Ankara, Turkey, has a 4% chance of losing each piece of his or her luggage independently. Suppose that the probability of losing each piece of luggage in this way is 5% at
Da Vinci airport in Rome, 5% at Kennedy airport in New York, and 4% at O’Hare airport
in Chicago. Dr. Keating travels from Bombay to San Francisco with one piece of luggage in the baggage compartment. He changes airplanes in Ankara, Rome, New York,
and Chicago.
(a)
What is the probability that his luggage does not reach his destination with him?
(b)
If the luggage does not reach his destination with him, what is the probability that
it was lost at Da Vinci airport in Rome?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 135 — #151
✐
✐
Section 3.5
Independence
135
44.
(The Game of Craps) In the game of craps, the player rolls two unbiased dice. If
the sum of the outcomes is 2, 3, or 12, she loses. If it is 7 or 11, she wins. However, if the
sum is one of the numbers 4, 5, 6, 8, 9, or 10, she will continue rolling repeatedly until
either she rolls the sum that she initially obtained, in which case she wins, or she rolls a
7, in which case she loses. What is the probability that in a game of craps, a player will
win?
45.
An urn contains nine red and one blue balls. A second urn contains one red and five blue
balls. One ball is removed from each urn at random and without replacement, and all of
the remaining balls are put into a third urn. If we draw two balls randomly from the third
urn, what is the probability that one of them is red and the other one is blue?
46.
From a population of people with unrelated birthdays, 30 people are selected at random.
What is the probability that exactly four people of this group have the same birthday and
that all the others have different birthdays (exactly 27 birthdays altogether)? Assume
that the birthrates are constant throughout the first 365 days of a year but that on the
366th day it is one-fourth that of the other days.
47.
In a contest, contestants A, B, and C are each asked, in turn, a general scientific question. If a contestant gives a wrong answer to a question, he drops out of the game. The
remaining two will continue to compete until one of them drops out. The last person
remaining is the winner. Suppose that a contestant knows the answer to a question independently of the other contestants, with probability p. Let CAB represent the event
that C drops out first, A next, and B wins, with similar representations for other cases.
Calculate and compare the probabilities of ABC, BCA, and CAB .
48.
Figure 3.11 shows an electric circuit in which each of the switches located at 1, 2, 3, 4,
and 5 is independently closed or open with probabilities p and 1 − p, respectively. If a
Figure 3.11
Electric circuit of Exercise 48.
signal is fed to the input, what is the probability that it is transmitted to the output?
49.
Hemophilia is a hereditary disease. If a mother has it, then with probability 1/2, any of
her sons independently will inherit it. Otherwise, none of the sons becomes hemophilic.
Julie is the mother of two sons, and from her family’s medical history it is known that,
with the probability 1/4, she is hemophilic. What is the probability that (a) her first son
is hemophilic; (b) her second son is hemophilic; (c) none of her sons are hemophilic?
Hint: Let H, H1 , and H2 denote the events that the mother, the first son, and the
second son are hemophilic, respectively. It should be clear that the events H1 and H2
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 136 — #152
✐
✐
136
Chapter 3
Conditional Probability and Independence
are conditionally independent given H . That is, if we are given that the mother is
hemophilic, knowledge about one son being hemophilic does not change the chance
of the other son being hemophilic. However, H1 and H2 are not independent. This is
because if we know that one son is hemophilic, the mother is hemophilic and therefore
with probability 1/2 the other son is also hemophilic.
50.
(Laplace’s Law of Succession) Suppose that n + 1 urns are numbered 0 through
n, and the ith urn contains i red and n − i white balls, 0 ≤ i ≤ n. An urn is selected
at random, and then the balls it contains are removed one by one, at random, and with
replacement. If the first m balls are all red, what is the probability that the (m + 1)st
ball is also red?
Hint: Let Ui be the event that the ith urn is selected, Rm the event that the first m
balls drawn are all red, and R the event that the (m + 1)st ball drawn is red. Note that R
and Rm are conditionally independent given Ui ; that is, given that the ith urn is selected,
R and Rm are independent. Hence
P (R | Rm Ui ) = P (R | Ui ) =
i
.
n
To find P (R | Rm ), use
P (R | Rm ) =
n
X
i=0
P (R | Rm Ui )P (Ui | Rm ),
where P (Ui | Rm ) is obtained from Bayes’ theorem. The final answer is
n m+1
X
i
n
P (R | Rm ) = i=0n .
X k m
n
k=0
Note that if we consider the function f (x) = xm+1 on the interval [0, 1], the right
Riemann sum of this function on the partition 0 = x0 < x1 < x2 < · · · < xn =
1, xi = i/n is
n
X
i=1
f (xi )(xi − xi−1 ) =
n
X
i=1
xm+1
i
n
1
1 X i m+1
=
.
n
n i=0 n
R1
Pn
Therefore, for large n, (1/n) i=0 (i/n)m+1 is approximately equal to 0 xm+1 dx =
R
P
1
1/(m + 2). Similarly, (1/n) nk=0 (k/n)m is approximately 0 xm dx = 1/(m + 1).
Hence
n
1 X i m+1
n
n
m+1
1/(m + 2)
=
.
P (R | Rm ) = i=0n ≈
m
X
1/(m
+
1)
m+2
k
1
n k=0 n
Laplace designed and solved this problem for philosophical reasons. He used it to argue
that the sun will rise tomorrow with probability (m + 1)/(m + 2) if we know that
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 137 — #153
✐
✐
Chapter 3
Summary
137
it has risen in the m preceding days. Therefore, according to this, as time passes by,
the probability that the sun rises again becomes higher, and the event of the sun rising
becomes more and more certain day after day. It is clear that, to argue this way, we should
accept the problematic assumption that the phenomenon of the sun rising is “random,”
and the model of this problem is identical to its random behavior.
Self-Quiz on Section 3.5
Time allotted: 20 Minutes
Each problem is worth 2.5 points.
1.
Suppose that the Dow-Jones Industrial Average (DJIA) rises 52% of trading days and
falls 48% of those days. What is the probability that, in the next 5 trading days, the
DJIA does not rise on any two consecutive trading days and does not fall on any two
consecutive trading days?
2.
A specific type of missile fired at a target hits it with probability 0.6. Find the minimum
number of such missiles to be fired to have a probability of at least 0.95 of hitting the
target.
3.
In a ball bearing manufacturing process, after an item is manufactured, two inspectors
will inspect it for defects. Suppose that each inspector finds a defect, independently of
the other inspector, with probability 0.93. What is the probability that a defect is found
by only one inspector?
4.
Suppose that 48% of Dr. Darabi’s patients visit his office for a dental cleaning, 18% visit
his office for orthodontic work, 12% for an extraction, 10% for a root canal, and the rest
for other dental issues. If, for a person, the need for orthodontic work, an extraction, and
a root canal are independent of each other, what is the probability that a random patient
who just left Dr. Darabi’s office needed at least one of these three dental services?
CHAPTER 3 SUMMARY
◮ Conditional Probability If P (B) > 0, the conditional probability of A given B,
denoted by P (A | B), is defined to be P (A | B) =
P (AB)
.
P (B)
◮ Reduction of Sample Space Let B be an event of a sample space S with P (B) > 0.
For a subset A of B, define Q(A) = P (A | B). Then Q is a function from the set of subsets of
B to [0, 1] and it satisfies the axioms that probabilities do, and hence it is a probability function.
Note that while P is defined for all subsets of S, the probability function Q is defined only for
subsets of B . Therefore, for Q, the sample space is reduced from S to B . This reduction of
sample space is sometimes very helpful in calculating conditional probabilities. Suppose that
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 138 — #154
✐
✐
138
Chapter 3
Conditional Probability and Independence
we are interested in P (E | B), where E ⊆ S . One way to calculate this quantity is to reduce
S to B and find Q(EB). It is usually easier to compute the unconditional probability Q(EB)
rather than the conditional probability P (E | B).
◮ The Multiplication Rule If P (B) > 0, then P (AB) = P (B)P (A | B). In general, if
P (A1 A2 A3 · · · An−1 ) > 0, then
P (A1 A2 A3 · · · An−1 An )
= P (A1 )P (A2 | A1 )P (A3 | A1 A2 ) · · · P (An | A1 A2 A3 · · · An−1 ).
◮ Partition of a Sample Space
Let {B1 , B2 , . . . , Bn } be a set of nonempty subsets of
the
sample
space
S
of
an
experiment.
If
the events B1 , B2 , . . . , Bn are mutually exclusive and
Sn
i=1 Bi = S, the set {B1 , B2 , . . . , Bn } is called a partition of S .
Let B be an event with P (B) > 0 and P (B c ) > 0. Then for
◮ Law of Total Probability
any event A, we have that
P (A) = P (A | B)P (B) + P (A | B c )P (B c ).
In general, if {B1 , B2 , . . . , Bn } is a partition of the sample space of an experiment and
P (Bi ) > 0 for i = 1, 2, . . . , n, then for any event A of S,
P (A) = P (A | B1 )P (B1 ) + P (A | B2 )P (B2 ) + · · · + P (A | Bn )P (Bn )
n
X
=
P (A | Bi )P (Bi ).
i=1
EvenSmore generally, let {B1 , B2 , . . .} be a sequence of mutually exclusive events of S such
∞
that i=1 Bi = S. Suppose that, for all i ≥ 1, P (Bi ) > 0. Then for any event A of S,
P (A) =
∞
X
i=1
◮ Bayes’ Formula
P (A | Bi )P (Bi ).
If P (B) > 0 and P (B c ) > 0, then for any event A of S with
P (A) > 0,
P (B | A) =
P (A | B)P (B)
.
P (A | B)P (B) + P (A | B c )P (B c )
In general, Let {B1 , B2 , . . . , Bn } be a partition of the sample space S of an experiment. If for
i = 1, 2, . . . , n, P (Bi ) > 0, then for any event A of S with P (A) > 0,
P (Bk | A) =
P (A | Bk )P (Bk )
.
P (A | B1 )P (B1 ) + P (A | B2 )P (B2 ) + · · · + P (A | Bn )P (Bn )
◮ Independence of Events
Two events A and B are called independent if P (AB) =
P (A)P (B). If A and B are independent, and P (B) > 0, then P (A | B) = P (A); that
is, if A is independent of B, knowledge regarding the occurrence of B does not change the
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 139 — #155
✐
✐
Chapter 3
Review Problems
139
chance of the occurrence of A. If two events are not independent, they are called dependent. If
A and B are independent, then A and B c are independent as well. As a result, If A and B are
independent, then so are Ac and B c . In general, the set of events {A1 , A2 , . . . , An } is called
independent if for every subset {Ai1 , Ai2 , . . . , Aik }, k ≥ 2, of {A1 , A2 , . . . , An },
P (Ai1 Ai2 · · · Aik ) = P (Ai1 )P (Ai2 ) · · · P (Aik ).
REVIEW PROBLEMS
1.
Two persons arrive at a train station, independently of each other, at random times between 1:00 P.M. and 1:30 P.M. What is the probability that one will arrive between 1:00 P.M.
and 1:12 P.M., and the other between 1:17 P.M. and 1:30 P.M.?
2.
From the students of a college that does not offer graduate programs, a student is selected at random. Let E1 , E2 , E3 , and E4 denote the events that the student is a freshman, sophomore, junior, and senior, respectively. Let A be the event that the randomly
selected student’s grade point average (GPA) is A+ , A, or A− . Define the events B, C,
D similarly. Let F be the event that the student’s GPA is F. Which of the following sets
of events is a partition of the sample space of this experiment? Which of them is not?
(a) {E1 , E2 , E3 , E4 }, (b) {A, B, C, D}, (c) {E1 , E3 , B, C}, (d) {A, B, C, D, F }.
3.
A polygraph operator detects innocent suspects as being guilty 3% of the time. If during
a crime investigation six innocent suspects are examined by the operator, what is the
probability that at least one of them is detected as guilty?
4.
Suppose that 5% of men and 0.25% of women are color blind. In Belavia, only 42%
of the persons 65 and older are male. What is the probability that a randomly selected
person from this age group is color blind?
5.
In statistical surveys where individuals are selected randomly and are asked questions,
experience has shown that only 48% of those under 25 years of age, 67% between 25 and
50, and 89% above 50 will respond. A social scientist is about to send a questionnaire to
a group of randomly selected people. If 30% of the population are younger than 25 and
17% are older than 50, what percent will answer her questionnaire?
6.
Let n > 1 be an integer. Suppose that a fair coin is tossed independently and repeatedly.
Find the probability of tails on the first toss or heads on the nth toss.
7.
Every morning, Galya flips a fair coin to decide whether she wants to take the bus to
school or she wants to drive there. No matter what her means of transport, if she leaves
her house at or before 7:00 a.m., she gets to school on time. However, if she leaves 15
minutes later, at 7:15 a.m., then the probability is 0.80 to arrive at school on time if she
drives her own car, and the probability is 0.62 to get to school on time if she takes the
bus. On a certain day, Galya left her house at 7:15 and got to school on time. What is the
probability that she took the bus on that day?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 140 — #156
✐
✐
140
Chapter 3
Conditional Probability and Independence
8.
Diseases D1 , D2 , and D3 cause symptom A with probabilities 0.5, 0.7, and 0.8, respectively. If 5% of a population have disease D1 , 2% have disease D2 , and 3.5% have
disease D3 , what percent of the population have symptom A? Assume that the only possible causes of symptom A are D1 , D2 , and D3 and that no one carries more than one
of these three diseases.
9.
Professor Stern has three cars. The probability that on a given day car 1 is operative
is 0.95, that car 2 is operative is 0.97, and that car 3 is operative is 0.85. If Professor
Stern’s cars operate independently, find the probability that on next Thanksgiving day
(a) all three of his cars are operative; (b) at least one of his cars is operative; (c) at most
two of his cars are operative; (d) none of his cars is operative.
10.
A bus traveling from Baltimore to New York breaks down at a random location. If the
bus was seen running at Wilmington, what is the probability that the breakdown occurred
after passing through Philadelphia? The distances from New York, Philadelphia, and
Wilmington to Baltimore are, respectively, 199, 96, and 67 miles.
11.
Stacy and George are playing the heads or tails game with a fair coin. The coin is flipped
repeatedly until either the fifth heads or the fifth tails appears. If the fifth heads occurs
first, Stacy wins the game. Otherwise, George is the winner. Suppose that after the fifth
flip, three heads and two tails have occurred. What is the probability that Stacy wins this
game?
12.
Roads A, B, and C are the only escape routes from a state prison. Prison records show
that, of the prisoners who tried to escape, 30% used road A, 50% used road B, and 20%
used road C. These records also show that 80% of those who tried to escape via A, 75%
of those who tried to escape via B, and 92% of those who tried to escape via C were
captured. What is the probability that a prisoner who succeeded in escaping used road
C?
13.
From an ordinary deck of 52 cards, 10 cards are drawn at random. If exactly four of
them are hearts, what is the probability of at least one spade being among them?
14.
A fair die is thrown twice. If the second outcome is 6, what is the probability that the
first one is 6 as well?
15.
Suppose that 10 dice are thrown and we are told that among them at least one has landed
6. What is the probability that there are two or more sixes?
16.
Urns I and II contain three pennies and four dimes, and two pennies and five dimes,
respectively. One coin is selected at random from each urn. If exactly one of them is a
dime, what is the probability that the coin selected from urn I is the dime?
17.
There are 49 unbiased dice and one loaded die in a basket. For the loaded die, when
tossed the probability of obtaining 6 is 3/8, and the probability of obtaining each of the
other five faces is 1/8. A die is selected at random and tossed five times. If all five times
it lands on 6, what is the probability that the die selected at random is the loaded one?
18.
Six fair dice are tossed independently. Find the probability that the number of 1’s minus
the number of 2’s will be 3.
19.
An experiment consists of first tossing an unbiased coin and then rolling a fair die. If we
perform this experiment successively, what is the probability of obtaining a heads on the
coin before a 1 or 2 on the die?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 141 — #157
✐
✐
Chapter 3
Review Problems
141
20.
Suppose that the Dow-Jones Industrial Average (DJIA) rises 52% of trading days and
falls 48% of those days. Elmer is a stock market expert, and his forecasts will turn out
to be true 68% of the time when he predicts that the DJIA will rise the next trading day.
However, when he predicts that it will fall the next trading day, only 59% of the time his
forecasts will come true. Tomorrow is a trading day, and Elmer has predicted that the
DJIA will fall. What is the probability that Elmer’s prediction will turn out to be wrong,
and the DJIA will rise tomorrow?
21.
Urn I contains 25 white and 20 black balls. Urn II contains 15 white and 10 black balls.
An urn is selected at random and one of its balls is drawn randomly and observed to be
black and then returned to the same urn. If a second ball is drawn at random from this
urn, what is the probability that it is black?
22.
An urn contains nine red and one blue balls. A second urn contains one red and five blue
balls. One ball is removed from each urn at random and without replacement, and all
of the remaining balls are put into a third urn. What is the probability that a ball drawn
randomly from the third urn is blue?
23.
A fair coin is tossed. If the outcome is heads, a red hat is placed on Lorna’s head. If it
is tails, a blue hat is placed on her head. Lorna cannot see the hat. She is asked to guess
the color of her hat. Is there a strategy that maximizes Lorna’s chances of guessing correctly?
Hint: Suppose that Lorna chooses the color red with probability α and blue with probability 1 − α. Find the probability that she guesses correctly.
24.
A child is lost at Epcot Center in Florida. The father of the child believes that the probability of his being lost in the east wing of the center is 0.75, and in the west wing 0.25.
The security department sends three officers to the east wing and two to the west to
look for the child. Suppose that an officer who is looking in the correct wing (east or
west) finds the child, independently of the other officers, with probability 0.4. Find the
probability that the child is found.
25.
Solve the following problem, asked of Marilyn Vos Savant in the “Ask Marilyn” column
of Parade Magazine, August 9, 1992.
Three of us couples are going to Lava Hot Springs next weekend. We’re
staying two nights, and we’ve rented two studios, because each holds
a maximum of only four people. One couple will get their own studio on
Friday, a different couple on Saturday, and one couple will be out of luck.
We’ll draw straws to see which are the two lucky couples. I told my wife
we should just draw once, and the loser would be the couple out of luck
both nights. I figure we’ll have a two-out-of-three (66 32 %) chance of winning one of the nights to ourselves. But she contends that we should
draw straws twice—first on Friday and then, for the remaining two couples only, on Saturday—reasoning that a one-in-three (33 31 %) chance for
Friday and a one-in-two (50%) chance for Saturday will give us better
odds. Which way should we go?
26.
A student at a certain university will pass the oral Ph.D. qualifying examination if at
least two of the three examiners pass her or him. Past experience shows that (a) 15% of
the students who take the qualifying exam are not prepared, and (b) each examiner will
independently pass 85% of the prepared and 20% of the unprepared students. Kevin took
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 142 — #158
✐
✐
142
Chapter 3
Conditional Probability and Independence
his Ph.D. qualifying exam with Professors Smith, Brown, and Rose. What is the probability that Professor Rose has passed Kevin if we know that neither Professor Brown nor
Professor Smith has passed him? Let S, B, and R be the respective events that Professors Smith, Brown, and Rose have passed Kevin. Are these three events independent?
Are they conditionally independent given that Kevin is prepared? (See Exercises 49 and
50, Section 3.5.)
Self-Test on Chapter 3
Time allotted: 120 Minutes
Each problem is worth 10 points.
1.
Currently, negotiations on complicated border disputes are going on between
Chernarus and Carpathia. Suppose that a summit between the foreign ministers of the
two countries will resolve all the border issues with probability 1/4. Furthermore, suppose that, for i ≥ 2, if the previous i − 1 summits are unsuccessful, the ith summit
1
will be successful with probability 1 − . What is the probability that no more than 4
i
summits are necessary to successfully resolve the border disputes between Chernarus
and Carpathia?
2.
Suppose that a system of five components is functional if at least one of the components
A1 , A2 , and A3 and both components B1 and B2 are operable. Suppose that a component is operable, independently of other components, with probability 0.85. Find the
probability that the system functions.
3.
In Helsinki, Finland, Hunter is the fourth in line at a station to board a city tour minibus
having 16 passenger seats. A minibus with all empty passenger seats arrives at the station. Suppose that each passenger boarding the minibus selects one of the unoccupied
seats at random. If Hunter’s favorite seat is the window seat in the front row opposite
from the minibus driver, what is the probability that when he boards the minibus his
favorite seat is unoccupied?
4.
Tasha has two dice, one is unbiased, but the other one is loaded so that the probability
of 6 is twice the probability of any of the other five faces, which are all equally likely.
Tasha picks one of the two dice randomly and tosses it three times. If all three times it
lands on 6, what is the probability that the die Tasha picked is the unbiased one?
5.
Suppose that before death, an individual in a population of living organisms gives birth
to 0, 1, or 2 new individuals, independently of other organisms with probabilities 1/4,
1/2, and 1/4, respectively. What is the probability that the second generation offspring
of such a living organism consists of at least one individual?
6.
In a metropolis, consecutive traffic lights are coordinated so that a driver who finds the
first light green will find the second light also green with probability 0.75. Similarly, a
driver who finds the first traffic light yellow or red will also find the second traffic light
yellow or red with probability 0.75. Suppose that 50% of the time a traffic light is green.
Find the probability that a driver has to stop at least once for the red or yellow lights
when passing through two consecutive traffic lights.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 143 — #159
✐
✐
Chapter 3
Self-Test Problems
143
7.
In a lottery scratch-off game, each ticket has 10 coated circles in the middle and one
coated rectangle in the lower left corner. Underneath the coats of 4 of the circles, there is
a dollar sign, “$,” and underneath the remaining 6 circles is blank. A winning ticket is the
one with only those circles that contain the dollar sign scratched. The remaining circles
and the rectangle must be left untouched. The winning amount is written underneath the
coat of the rectangle, which will be scratched off by the lottery agent to reveal the amount
won. Harlan plays one such lottery scratch-off game and loses. What is the probability
that the first circle he scratched off which was blank underneath was the third circle?
8.
Suppose that of the individuals who have been exposed to the dust product of the mineral
asbestos, 3.3 per 1000 people develop mesothelioma. Suppose that Keith, a mine worker
exposed to such a dust product, was tested for mesothelioma, and the test result came
back positive. If 92% of the time the test is accurate, what is the probability that Keith
has mesothelioma?
9.
Vincent is a patient with the life threatening blood cancer leukemia, and he is in need of
a bone marrow transplant. He asks n people whether or not they are willing to donate
bone marrow to him if they are a close bone marrow match for him. Suppose that each
person’s response, independently of others, is positive with probability p1 and negative
with probability 1−p1 . Furthermore, suppose that, independently of the others, the probability is p2 that a person tested is a close match. (a) Find the probability that Vincent
finds at least one close bone marrow match among these n people. (b) Suppose that the
probability is 1 in 10,000 that two unrelated persons are compatible; that is, have a close
bone marrow match. If 70% of the people asked would respond positively to Vincent’s
request, and he asks 1000 people, what is the probability that he will find a match? What
if he asks 100,000 people?
10.
Starting with Hillary, two players, Hillary and Donald, take turns and roll a fair die. The
one who rolls a 6 first is the winner. What is the probability that Hillary will win?
Hint: For i ≥ 1, let Ai be the event that the first 6 is rolled on Hillary’s ith turn. Let
H be the event that
will win. Write P (H) in terms of P (Ai )’s. Note that for
P Hillary
i
|r| < 1, m ≥ 1, ∞
r
=
r m /(1 − r).
i=m
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 144 — #160
✐
✐
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 145 — #161
✐
✐
Chapter 4
D istribution F unctions and
Discrete R andom Variables
4.1
RANDOM VARIABLES
In real-world problems we are often faced with one or more quantities that do not have fixed
values. The values of such quantities depend on random actions, and they usually change from
one experiment to another. For example, the number of babies born in a certain hospital each
day is not a fixed quantity. It is a complicated function of many random factors that vary from
one day to another. So are the following quantities: the arrival time of a bus at a station, the
sum of the outcomes of two dice when thrown, the amount of rainfall in Seattle during a given
year, the number of earthquakes that occur in California per month, and the weight of grains
of wheat grown on a certain plot of land (it varies from one grain to another). In probability,
quantities introduced in these diverse examples are called random variables. The numerical
values of random variables are unknown. They depend on random elements occurring at the
time of the experiment and over which we have no control. For example, if in rolling two fair
dice, X is the sum, then X can only assume the values 2, 3, 4, . . . , 12 with the following
probabilities:
P (X = 2) = P (1, 1) = 1/36,
P (X = 3) = P (1, 2), (2, 1) = 2/36,
P (X = 4) = P (1, 3), (2, 2), (3, 1) = 3/36,
and, similarly,
Sum, i
5
6
7
8
9
10
11
12
P (X = i)
4/36
5/36
6/36
5/36
4/36
3/36
2/36
1/36
Clearly, {2, 3, 4, . . . , 12} is the set of possible values of X . Since X ∈ {2, 3, 4, . . . , 12},
P12
we should have i=2 P (X = i) = 1, which is readily verified. The numerical value of a
random variable depends on the outcome of the experiment. In this example, for instance, if
the outcome is (2, 3), then X is 5, and if it is (5, 6), then X is 11. X is not defined for points
that do not belong to S, the sample space of the experiment. Thus X is a real-valued function
145
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 146 — #162
✐
✐
146
Chapter 4
Distribution Functions and Discrete Random Variables
on S . However, not all real-valued functions on S are considered to be random variables. For
theoretical reasons, it is necessary that the inverse image of an interval in R be an event of S,
which motivates the following definition.
Definition 4.1
Let S be the sample space of an experiment. A real-valued function
X
:
S
→
R
is
called
a random variable of the experiment if, for each interval I ⊆ R ,
s : X(s) ∈ I is an event.
In probability, the set s : X(s) ∈ I
X ∈ I.
is often abbreviated as {X ∈ I}, or simply as
Example 4.1 Suppose that three cards are drawn from an ordinary deck of 52 cards, one
by one, at random and with replacement. Let X be the number of spades drawn; then X is a
random variable. If an outcome of spades is denoted by s, and other outcomes are represented
by t, then X is a real-valued function defined on the sample space
S = (s, s, s), (t, s, s), (s, t, s), (s, s, t), (s, t, t), (t, s, t), (t, t, s), (t, t, t) ,
by X(s, s, s) = 3, X(s, t, s) = 2, X(s, s, t) = 2, X(s, t, t) = 1, and so on. Now we
must determine the values that X assumes and the probabilities that are associated with them.
Clearly, X can take the values 0, 1, 2, and 3. The probabilities associated with these values are
calculated as follows:
27
3 3 3
× × =
,
4 4 4
64
P (X = 1) = P (s, t, t), (t, s, t), (t, t, s)
1 3
3 3 1 3 3 3 1 27
=
+
+
= ,
× ×
× ×
× ×
4 4 4
4 4 4
4 4 4
64
P (X = 2) = P (s, s, t), (s, t, s), (t, s, s)
1 1
9
3 1 3 1 3 1 1
=
+
+
= ,
× ×
× ×
× ×
4 4 4
4 4 4
4 4 4
64
1
P (X = 3) = P (s, s, s) =
.
64
P (X = 0) = P
(t, t, t)
=
If the cards are drawn without replacement, the probabilities associated with the values 0, 1, 2,
!
!
!
and 3 are
13
39
39
1
2
3
!,
! ,
P (X = 1) =
P (X = 0) =
52
52
3
3
P (X = 2) =
13
2
!
52
3
!
39
1
! ,
P (X = 3) =
13
3
52
3
!
!.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 147 — #163
✐
✐
Section 4.1
Random Variables
147
Therefore,
P (X = i) =
13
i
!
!
39
3−i
!
,
52
3
i = 0, 1, 2, 3.
Example 4.2 A bus stops at a station every day at some random time between 11:00 A.M.
and 11:30 A.M. If X is the actual arrival
is a random variable. It is a function
time of the bus, X
1
defined on the sample space S = t : 11 < t < 11 2 by X(t) = t. As we know from
β−α
= 2(β − α)
Section 1.7, P (X = t) = 0 for any t ∈ S, and P X ∈ (α, β) =
11 12 − 11
for any subinterval (α, β) of (11, 11 12 ). Example 4.3 In the United States, the number of twin births is approximately 1 in 90. Let
X be the number of births in a certain hospital until the first twins are born. X is a random
variable. Denote twin births by T and single births by N . Then X is a real-valued function
defined on the sample space
S = {T, N T, N N T, N N N T, . . .} by X( N
| NN
{z· · · N} T ) = i.
i−1
The set of all possible values of X is {1, 2, 3, . . .} and
P (X = i) = P ( N
| NN
{z· · · N} T ) =
i−1
89 i−1 1 .
90
90
Example 4.4 In a certain country, the draft-status priorities of eligible men are determined
according to their birthdays. Suppose that numbers 1 to 31 are assigned to men with birthdays
on January 1 to January 31, numbers 32 to 60 to men with birthdays on February 1 to February 29, numbers 61 to 91 to men with birthdays on March 1 to March 31, . . . , and finally,
numbers 336 to 366 to those with birthdays on December 1 to December 31. Then numbers are
selected at random, one by one, and without replacement, from 1 to 366 until all of them are
chosen. Those with birthdays corresponding to the first number drawn would have the highest
draft priority, those with birthdays corresponding to the second number drawn would have the
second-highest priority, and so on. Let X be the largest of the first 10 numbers selected. Then
X is a random variable that assumes the values 10, 11, 12, . . . , 366. The event X = i occurs
if the largest number among the first 10 is i, that is, if one of the first 10 numbers is i and the
other 9 are from 1 through i − 1. Thus
!
i−1
9
! ,
P (X = i) =
i = 10, 11, 12, . . . , 366.
366
10
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 148 — #164
✐
✐
148
Chapter 4
Distribution Functions and Discrete Random Variables
As an application, let us calculate P (X ≥ 336), the probability that, among individuals having
one of the 10 highest draft priorities, there are some with birthdays in December. We have that
!
i−1
366
X
9
! ≈ 0.592. P (X ≥ 336) =
366
i=336
10
Remark: On December 1, 1969, the Selective Service headquarters in Washington, D.C.
determined the draft priorities of 19-year-old males by using a method very similar to that
of Example 4.4. We will now explain how, from given random variables, new ones can be formed. Let
X and Y be two random variables over the same sample space S; then X : S → R and
Y : S → R are real-valued functions having the same domain. Therefore, we can form the
functions X + Y ; X − Y ; aX + bY, where a and b are constants; XY ; and X/Y, where
Y 6= 0. Since the domain of these real-valued functions is S, they are also random variables
defined on S . Hence the sum, the difference, linear combinations, the product, and the quotients
(if they exist) of random variables are themselves random variables. Similarly, if f : R → R is
an ordinary real-valued function, the composition of f and X, f ◦X : S → R, is also a random
variable. For example, let f : R → R be defined by f (x) = x2 ; then f ◦X : S → R is X 2 .
Hence X 2 is a random variable as well. Similarly, functions such as sin X, cos X 2 , eX , and
X 3 −2X are random variables.√
So are the following functions: X 2 +Y 2 ; (X 2 +Y 2 )/(2Y +1),
Y =
6 −1/2; sin X + cos Y ; X 2 + Y 2 ; and so on. Functions of random variables appear
naturally in probability problems. As an example, suppose that we choose a point at random
from the unit disk in the plane. If X and Y are the x and the y coordinates
of the point, X
√
and Y are random variables. The distance of (X, Y ) from the origin is X 2 + Y 2 , a random
variable that is a function of both X and Y .
Example 4.5 The diameter of a flat metal disk manufactured by a factory is a random number between 4 and 4.5. What is the probability that the area of such a flat disk chosen at random
is at least 4.41π ?
Solution: Let D be the diameter of the metal disk selected at random. D is a random variable,
and the area of the metal disk, π(D/2)2 , which is a function of D, is also a random variable.
We are interested in the probability of the event πD 2 /4 > 4.41π, which is
P
πD 2
4
> 4.41π = P (D 2 > 17.64) = P (D > 4.2).
Now to calculate P (D > 4.2), note that, since the length of D is a random number in the
interval (4, 4.5), the probability that it falls into the subinterval (4.2, 4.5) is
(4.5 − 4.2)/(4.5 − 4) = 3/5. Hence P (πD 2 /4 > 4.41π) = 3/5. Example 4.6 A random number is selected from the interval (0, π/2). What is the probability that its sine is greater than its cosine?
Solution: Let the number selected be X; then X is a random variable and therefore sin X and
cos X, which are functions of X, are also random variables. We are interested in the probability
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 149 — #165
✐
✐
Section 4.2
Distribution Functions
149
of the event sin X > cos X :
π π
−
1
π
2
4
= π
= ,
P (sin X > cos X) = P (tan X > 1) = P X >
4
2
−0
2
where the first equality holds since in the interval (0, π/2), cos X > 0, and the second equality
holds since, in this interval, tan X is strictly increasing. 4.2
DISTRIBUTION FUNCTIONS
Random variables are often used for the calculation of the probabilities of events. For example,
in the experiment of throwing two dice, if we are interested in a sum of at least 8, we define X
to be the sum and calculate P (X > 8). Other examples are the following:
1.
If a bus arrives at a random time between 10:00 A.M. and 10:30 A.M. at a station, and X is
the arrival time, then X < 10 16 is the event that the bus arrives before 10:10 A.M.
2.
If X is the price of gold per troy ounce on a random day, then X ≤ 400 is the event that
the price of gold remains at or below $400 per troy ounce.
3.
If X is the number of votes that the next Democratic presidential candidate will get, then
X ≥ 5 × 107 is the event that he or she will get at least 50 million votes.
4.
If X is the number of heads in 100 tosses of a coin, then 40 < X ≤ 60 is the event that
the number of heads is at least 41 and at most 60.
Usually, when dealing with a random variable X, for constants a and b (b < a), computation of one or several of the probabilities P (X = a), P (X < a), P (X ≤ a), P (X > b),
P (X ≥ b), P (b ≤ X ≤ a), P (b < X ≤ a), P (b ≤ X < a), and P (b < X < a) is
our ultimate goal. For this reason we calculate P (X ≤ t) for all t ∈ (−∞, +∞). As we will
show shortly, if P (X ≤ t) is known for all t ∈ R, then for any a and b, all of the probabilities
that are mentioned above can be calculated. In fact, since the real-valued function P (X ≤ t)
characterizes X, it tells us almost everything about X . This function is called the distribution
function of X .
Definition 4.2
If X is a random variable, then the function F defined on (−∞, +∞) by
F (t) = P (X ≤ t) is called the distribution function of X.
Since F “accumulates” all of the probabilities of the values of X up to and including t,
sometimes it is called the cumulative distribution function of X . The most important properties of the distribution functions are as follows:
1.
2.
F is nondecreasing; that is, if t < u, then F (t) ≤ F (u). To see this, note that the
occurrence of the event {X ≤ t} implies the occurrence of the event {X ≤ u}. Thus
{X ≤ t} ⊆ {X ≤ u} and hence P (X ≤ t) ≤ P (X ≤ u). That is, F (t) ≤ F (u).
limt→∞ F (t) = 1. To prove this, it suffices to show that for any increasing sequence
{tn } of real numbers that converges to ∞, limn→∞ F (tn ) = 1. This follows from the
continuity property of the probability function (see Theorem 1.8). The events {X ≤ tn }
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 150 — #166
✐
✐
150
Chapter 4
Distribution Functions and Discrete Random Variables
form an increasing sequence that converges to the event
that is, limn→∞ {X ≤ tn } = {X < ∞}. Hence
lim P (X ≤ tn ) = P
n→∞
which means that
∞
[
S∞
n=1 {X ≤ tn } = {X < ∞};
{X ≤ tn } = P (X < ∞) = 1,
n=1
lim F (t) = 1.
n→∞
3.
limt→−∞ F (t) = 0. The proof of this is similar to the proof that limt→∞ F (t) = 1.
4.
F is right continuous. That is, for every t ∈ R , F (t+) = F (t). This means that if tn is
a decreasing sequence of real numbers converging to t, then
lim F (tn ) = F (t).
n→∞
To prove this, note that since tn decreases
T∞ to t, the events {X ≤ tn } form a decreasing
sequence that converges to the event n=1 {X ≤ tn } = {X ≤ t}. Thus, by the continuity
property of the probability function,
lim P (X ≤ tn ) = P
n→∞
which means that
∞
\
{X ≤ tn } = P (X ≤ t),
n=1
lim F (tn ) = F (t).
n→∞
As mentioned previously, by means of F, the distribution function of a random variable
X, a wide range of probabilistic questions concerning X can be answered. Here are some
examples.
1.
To calculate P (X > a), note that P (X > a) = 1 − P (X ≤ a), thus
P (X > a) = 1 − F (a).
2.
To calculate P (a < X ≤ b), b > a, note that {a < X ≤ b} = {X ≤ b} − {X ≤ a}
and {X ≤ a} ⊆ {X ≤ b}. Hence, by Theorem 1.5,
P (a < X ≤ b) = P (X ≤ b) − P (X ≤ a) = F (b) − F (a).
3.
To calculate P (X < a), note that the
S∞sequence of the events {X ≤ a − 1/n} is an
increasing sequence that converges to n=1 {X ≤ a − 1/n} = {X < a}. Therefore, by
the continuity property of the probability function (Theorem 1.8),
∞ n
[
1
1 o
=P
X ≤a−
= P (X < a),
lim P X ≤ a −
n→∞
n
n
n=1
which means that
1
P (X < a) = lim F a −
.
n→∞
n
Hence P (X < a) is the left-hand limit of the function F as x → a; that is,
P (X < a) = F (a−).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 151 — #167
✐
✐
Section 4.2
Distribution Functions
151
To calculate P (X ≥ a), note that P (X ≥ a) = 1 − P (X < a). Thus
4.
P (X ≥ a) = 1 − F (a−).
Since {X = a} = {X ≤ a} − {X < a} and {X < a} ⊆ {X ≤ a}, we can write
5.
P (X = a) = P (X ≤ a) − P (X < a) = F (a) − F (a−).
Note that since F is right continuous, F (a) is the right-hand limit of F . This implies the
following important fact:
Let F be the distribution function of a random variable X; P (X = a) is the
difference between the right- and left-hand limits of F at a. If the function
F is continuous at a, these limits are the same and equal to F (a). Hence
P (X = a) = 0. Otherwise, F has a jump at a, and the magnitude of the
jump, F (a) − F (a−), is the probability that X = a.
As in cases 1 to 5, we can establish similar cases to obtain the following table.
Event
concerning X
X≤a
X>a
X<a
X≥a
X=a
Probability
of the event
in terms of F
F (a)
1 − F (a)
F (a−)
1 − F (a−)
F (a) − F (a−)
Event
concerning X
Probability
of the event
in terms of F
a<X ≤b
a<X <b
a≤X ≤b
a≤X <b
F (b) − F (a)
F (b−) − F (a)
F (b) − F (a−)
F (b−) − F (a−)
Example 4.7 The sales of a convenience store on a randomly selected day are X
thousand dollars, where X is a random variable with a distribution function of the following
form:
0
t<0
(1/2) t2
0≤t<1
F (t) =
2
k(4t − t ) 1 ≤ t < 2
1
t ≥ 2.
Suppose that this convenience store’s total sales on any given day are less than $2000.
(a)
Find the value of k .
(b)
Let A and B be the events that tomorrow the store’s total sales are between 500 and
1500 dollars, and over 1000 dollars, respectively. Find P (A) and P (B).
(c)
Are A and B independent events?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 152 — #168
✐
✐
152
Chapter 4
Distribution Functions and Discrete Random Variables
Solution: (a) Since X < 2, we have that P (X < 2) = 1, so F (2−) = 1. This gives
k(8 − 4) = 1, so k = 1/4.
3
1 1
3
=F
−F
(b)
P (A) = P
≤X≤
−
2
2
2
2
1
3
15 1
13
−F
=
− =
,
=F
2
2
16 8
16
3
1
P (B) = P (X > 1) = 1 − F (1) = 1 − = .
4
4
(c)
3
3
15
3
3
P (AB) = P 1 < X ≤
=F
− F (1) =
− =
. Since P (AB) 6=
2
2
16
4
16
P (A)P (B), A and B are not independent. Example 4.8
The distribution function of a random variable X is given by
0
x<0
x/4
0≤x<1
F (x) = 1/2
1≤x<2
1
x + 21 2 ≤ x < 3
12
1
x ≥ 3,
where the graph of F is shown in Figure 4.1. Compute the following quantities:
(a) P (X < 2);
(d) P (X > 3/2);
(b) P (X = 2);
(e) P (X = 5/2);
(c) P (1 ≤ X < 3);
(f) P (2 < X ≤ 7).
F(x)
1
3/4
2/3
1/2
1/4
0
Figure 4.1
x
1
2
3
Distribution function of Example 4.8.
Solution: (a) P (X < 2) = F (2−) = 1/2.
(b)
P (X = 2) = F (2) − F (2−) = (2/12 + 1/2) − 1/2 = 1/6.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 153 — #169
✐
✐
Section 4.2
(c)
(d)
Distribution Functions
153
P (1 ≤ X < 3) = P (X < 3) − P (X < 1) = F (3−) − F (1−) =
(3/12 + 1/2) − 1/4 = 1/2.
P (X > 3/2) = 1 − F (3/2) = 1 − 1/2 = 1/2.
(e)
P (X = 5/2) = 0 since F is continuous at 5/2 and has no jumps.
(f)
P (2 < X ≤ 7) = F (7) − F (2) = 1 − (2/12 + 1/2) = 1/3. Example 4.9 For the experiment of flipping a fair coin twice, let X be the number of tails
and calculate F (t), the distribution function of X, and then sketch its graph.
Solution: Since X assumes only the values 0, 1, and 2, we have F (t) = P (X ≤ t) = 0, if
t < 0. If 0 ≤ t < 1, then
F (t) = P (X ≤ t) = P (X = 0) = P {HH} = 1/4.
If 1 ≤ t < 2, then
F (t) = P (X ≤ t) = P (X = 0 or X = 1) = P {HH, HT, TH} = 3/4,
and if t ≥ 2, then P (X ≤ t) = 1. Hence
0
1/4
F (t) =
3/4
1
Figure 4.2 shows the graph of F .
t<0
0≤t<1
1≤t<2
t ≥ 2.
F(t)
1
3/4
1/4
0
Figure 4.2
t
1 2
Distribution function of Example 4.9.
Example 4.10 Suppose that a bus arrives at a station every day between 10:00 A.M. and
10:30 A.M., at random. Let X be the arrival time; find the distribution function of X and sketch
its graph.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 154 — #170
✐
✐
154
Chapter 4
Distribution Functions and Discrete Random Variables
Solution: The bus arrives at the station at random, between 10 and 10 12 , so if t ≤ 10, F (t) =
P (X ≤ t) = 0. Now if t ∈ 10, 10 12 , then
F (t) = P (X ≤ t) =
t − 10
= 2(t − 10),
10 12 − 10
and if t ≥ 10 12 , then F (t) = P (X ≤ t) = 1. Thus
0
F (t) = 2(t − 10)
1
The graph of F is shown in Figure 4.3.
t < 10
10 ≤ t < 10 12
t ≥ 10 12 .
F(t)
1
t
10
Figure 4.3
10.5
Distribution function of Example 4.10.
Remark 4.1 Suppose that F is a right-continuous, nondecreasing function on
(−∞, ∞) that satisfies limt→∞ F (t) = 1 and limt→−∞ F (t) = 0. It can be shown that
there exists a sample space S with a probability function and a random variable X over S such
that the distribution function of X is F . Therefore, a function is a distribution function if it
satisfies the conditions specified in this remark. EXERCISES
A
1.
Two fair dice are rolled and the absolute value of the difference of the outcomes is
denoted by X . What are the possible values of X, and the probabilities associated with
them?
2.
From an urn that contains five red, five white, and five blue chips, we draw two chips at
random. For each blue chip we win $1, for each white chip we win $2, but for each red
chip we lose $3. If X represents the amount that we either win or we lose, what are the
possible values of X and probabilities associated with them?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 155 — #171
✐
✐
Section 4.2
Distribution Functions
155
3.
In a society of population N, the probability is p that a person has a certain rare disease
independently of others. Let X be the number of people who should be tested until a
person with the disease is found, X = 0 if no one with the disease is found. What are
the possible values of X ? Determine the probabilities associated with these values.
4.
The side measurement of a plastic die, manufactured by factory A, is a random number
between 1 and 1 14 centimeters. What is the probability that the volume of a randomly
selected die manufactured by this company is greater than 1.424? Assume that the die
will always be a cube.
5.
F, the distribution function of a random variable X, is given by
0
t < −1
(1/4)t + 1/4
−1 ≤ t < 0
F (t) = 1/2
0≤t<1
(1/12)t + 7/12 1 ≤ t < 2
1
t ≥ 2.
(a)
Sketch the graph of F .
(b)
Calculate the following quantities: P (X < 1), P (X = 1), P (1 ≤ X < 2),
P (X > 1/2), P (X = 3/2), and P (1 < X ≤ 6).
6.
From families with three children a family is chosen at random. Let X be the number of
girls in the family. Calculate and sketch the distribution function of X . Assume that in a
three-child family all gender distributions are equally probable.
7.
A grocery store sells X hundred kilograms of rice every day, where the distribution of
the random variable X is of
the following form:
0
x<0
kx2
0≤x<3
F (x) =
2
k(−x + 12x − 3) 3 ≤ x < 6
1
x ≥ 6.
Suppose that this grocery store’s total sales of rice do not reach 600 kilograms on any
given day.
8.
(a)
Find the value of k .
(b)
What is the probability that the store sells between 200 and 400 kilograms of rice
next Thursday?
(c)
What is the probability that the store sells over 300 kilograms of rice next Thursday?
(d)
We are given that the store sold at least 300 kilograms of rice last Friday. What is
the probability that it did not sell more than 400 kilograms on that day?
Let X be a random variable with distribution function F . For p (0 < p < 1), Qp is said
to be a quantile of order p if
F (Qp −) ≤ p ≤ F (Qp ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 156 — #172
✐
✐
156
Chapter 4
Distribution Functions and Discrete Random Variables
In a certain country, the rate at which the price of oil per gallon changes from one year
to another has the following distribution function:
F (x) =
1
,
1 + e−x
−∞ < x < ∞.
Find Q0.50 , called the median of F ; Q0.25 , called the first quartile of F ; and Q0.75 ,
called the third quartile of F . Interpret these quantities.
9.
A random variable X is called symmetric about 0 if for all x ∈ R ,
P (X ≥ x) = P (X ≤ −x).
Prove that if X is symmetric about 0, then for all t > 0 its distribution function F
satisfies the following relations:
(a) P |X| ≤ t = 2F (t) − 1.
(b) P |X| > t = 2 1 − F (t) .
(c)
10.
11.
P (X = t) = F (t) + F (−t) − 1.
Determine if the following is a distribution function.
1
1 − e−t if t ≥ 0
π
F (t) =
0
if t < 0.
Determine if the following is a distribution function.
t
if t ≥ 0
F (t) = 1 + t
0
if t < 0.
12.
Determine if the following is a distribution function.
(
(1/2)et
t<0
F (t) =
1 − (3/4)e−t t ≥ 0.
13.
In the U.S., for fellowship, the Casualty Actuarial Society requires passing a series of
nine rigorous exams taken in order plus certain other learning objectives. Suppose that
the probability is p1 for a student to pass exam 1, and, for 2 ≤ i ≤ 9, if the student
has passed exam i − 1, then the probability is pi for him or her to pass exam i. Let
X = 0, if the student passes all of the exams, one after another, without failing any of
them. Otherwise, let X be the first exam that he or she fails. Find the probability mass
function of X .
Hint: Use Theorem 3.2, the multiplication rule.
14.
Airline A has commuter flights every 45 minutes from San Francisco airport to Fresno.
A passenger who wants to take one of these flights arrives at the airport at a random time.
Suppose that X is the waiting time for this passenger; find the distribution function of
X . Assume that seats are always available for these flights.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 157 — #173
✐
✐
Section 4.2
Distribution Functions
157
B
15.
A scientific calculator can generate two-digit random numbers. That is, it can choose a
number at random from the set {00, 01, 02, . . . , 99}. To obtain a random number from
the set {4, 5, . . . , 18}, show that we have to keep generating two-digit random numbers
until we obtain one between 4 and 18.
16.
In a small town there are 40 taxis, numbered 1 to 40. Three taxis arrive at random at a
station to pick up passengers. What is the probability that the number of at least one of
the taxis is less than 5?
17.
Let X be a randomly selected point from the interval (0, 3). What is the probability that
X 2 − 5X + 6 > 0?
18.
Let X be a random point selected from the interval (0, 1). Calculate F, the distribution
function of Y = X/(1 + X), and sketch its graph.
19.
In the United States, the number of twin births is approximately 1 in 90. At a certain
hospital let X be the number of births until the first twins are born. Find the first quartile, the median, and the third quartile of X . See Exercise 8 for the definitions of these
quantities.
20.
Let the time until a new car breaks down be denoted by X, and let
(
X if X ≤ 5
Y =
5 if X > 5.
Then Y is the life of the car, if it lasts less than 5 years, and is 5 if it lasts longer than 5
years. Calculate the distribution function of Y in terms of F, the distribution function of
X.
Self-Quiz on Section 4.2
Time allotted: 20 Minutes
1.
For what value of k, if any, is the function p(n) = k/n, n = 1, 2, 3, . . . , a probability
mass function? (3 points)
2.
Suppose that the distribution function of a random variable X is given by
F (t) =
0
t<1
1
1 − t2
Find P (X > 3) and P (X > 5 | X > 3).
t ≥ 1.
(3 points)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 158 — #174
✐
✐
158
3.
4.3
Chapter 4
Distribution Functions and Discrete Random Variables
There are 50 students enrolled in a class, and they arrive one at a time, independently
of each other. Suppose that the X th student is the first one who shares his birthday with
another student already present in the classroom. Find the probability mass function of
X . Assume that the birth rates are constant throughout the year and that each year has
365 days. (4 points)
DISCRETE RANDOM VARIABLES
In Section 4.1 we observed that the set of possible values of a random variable might be finite,
infinite but countable, or uncountable. For example, let X, Y, and Z be three random variables representing the respective number of tails in flipping a coin twice, the number of flips
until the first heads, and the amount of next year’s rainfall. Then the sets of possible values for
X, Y, and Z are the finite set {0, 1, 2}, the countable set {1, 2, 3, 4, . . .}, and the uncountable set {x : x ≥ 0}, respectively. Whenever the set of possible values that a random variable
X can assume is at most countable, X is called discrete. Therefore, X is discrete if either
the set of its possible values is finite or it is countably infinite. To each discrete random variable, a real-valued function p : R → R , defined by p(x) = P (X = x), is assigned and is
called the probability mass function of X . (It is also called the probability function of X
or the discrete probability function of X .) Since the set of values of X is countable, p(x)
is positive at most for a countable set. It is zero elsewhere; that is, if possible values of X are
x1 , x2 , x3 , . . . , then p(xi ) ≥ 0 (i = 1, 2, 3, . . .) and p(x) = 0 if x 6∈ {x1 , x2 , x3 , . . .}.
Now,
1 , x2 , x3 , . . .} is certain. Therefore, we have that
P∞ clearly, the occurrence of the eventP{x
∞
i=1 p(xi ) = 1.
i=1 P (X = xi ) = 1 or, equivalently,
Definition 4.3
The probability mass function p of a random variable X whose set of
possible values is {x1 , x2 , x3 , . . .} is a function from R to R that satisfies the following properties.
(a)
(b)
(c)
p(x) = 0 if x 6∈ {x1 , x2 , x3 , . . .}.
p(xi ) = P (X = xi ) and hence p(xi ) ≥ 0 (i = 1, 2, 3, . . .).
P∞
i=1 p(xi ) = 1.
Because of this definition, if, for a set {x1 , x2 , x3 , . . .}, there exists a function
p
:
R
P∞ → R such that p(xi ) ≥ 0 (i = 1, 2, 3, . . .), p(x) = 0, x 6∈ {x1 , x2 , x3 , . . .}, and
i=1 p(xi ) = 1, then p is called a probability mass function.
The probability mass function of a random variable is often demonstrated geometrically by
a set of vertical lines connecting the points (xi , 0) and (xi , p(xi )). For example, if X is the
number of heads in two flips of a fair coin, then X = 0, 1, 2 with p(0) = 1/4, p(1) = 1/2,
and p(2) = 1/4. Hence the graphical representation of p is as shown in Figure 4.4.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 159 — #175
✐
✐
Section 4.3
Discrete Random Variables
159
p(x)
1/2
1/4
0
Figure 4.4
1
x
2
Graph of the number of heads in two flips of a fair coin.
The distribution function F of a discrete random variable X, with the set of possible values
{x1 , x2 , x3 , . . .}, is a step function. Assuming that x1 < x2 < x3 < · · · , we have that if
t < x1 , then
F (t) = 0;
if x1 ≤ t < x2 , then
F (t) = P (X ≤ t) = P (X = x1 ) = p(x1 );
if x2 ≤ t < x3 , then
F (t) = P (X ≤ t) = P (X = x1 or X = x2 ) = p(x1 ) + p(x2 );
and in general, if xn−1 ≤ t < xn , then
F (t) =
n−1
X
p(xi ).
i=1
Thus F is constant in the intervals [xn−1 , xn ) with jumps at x1 , x2 , x3 , . . . . The magnitude of
the jump at xi is p(xi ).
Example 4.11
Can a function of the form
x
2
c
p(x) =
3
0
x = 1, 2, 3, . . .
elsewhere
be a probability mass function?
Solution: A probability mass function should have three properties: (1) p(x) must be zero
at all points except on a finite or a countable set. Clearly, this property
is satisfied. (2) p(x)
P
should be nonnegative. This
is
satisfied
if
and
only
if
c
≥
0
.
(3)
p(x
)
i = 1. This condition
P∞
is satisfied if and only if i=1 c(2/3)i = 1. This happens precisely when
c=
1
∞
X
i=1
(2/3)i
=
1
1
= ,
2
2/3
1 − 2/3
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 160 — #176
✐
✐
160
Chapter 4
Distribution Functions and Discrete Random Variables
where the second equality follows from the geometric series theorem. Thus only for c = 1/2,
a function of the given form is a probability mass function. Example 4.12 In the experiment of rolling a balanced die twice, let X be the maximum of
the two numbers obtained. Determine and sketch the probability mass function and the distribution function of X .
Solution: The possible values of X are 1, 2, 3, 4, 5, and 6. The sample space of this experiment consists of 36 equally likely outcomes. Hence the probability of any of them is 1/36.
Thus
p(1) = P (X = 1) = P (1, 1) = 1/36,
p(2) = P (X = 2) = P (1, 2), (2, 2), (2, 1) = 3/36,
p(3) = P (X = 3) = P (1, 3), (2, 3), (3, 3), (3, 2), (3, 1) = 5/36.
Similarly, p(4) = 7/36, p(5) = 9/36, and p(6) = 11/36; p(x) = 0 for x 6∈
{1, 2, 3, 4, 5, 6}. The graphical representation of p is shown in Figure 4.5. The distribution
function of X, F, is as follows (its graph is shown in Figure 4.6):
0
x<1
1/36 1 ≤ x < 2
4/36 2 ≤ x < 3
F (x) = 9/36 3 ≤ x < 4
16/36 4 ≤ x < 5
25/36 5 ≤ x < 6
1
x ≥ 6.
p(x)
11/36
9/36
7/36
5/36
3/36
1/36
1
Figure 4.5
2
3
4
5
6
x
Probability mass function of Example 4.12.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 161 — #177
✐
✐
Section 4.3
Discrete Random Variables
161
F(x)
1
25/36
16/36
9/36
4/36
1/36
x
1
Figure 4.6
2
3
4
5
6
Distribution function of Example 4.12.
Example 4.13 Let X be the number of births in a hospital until the first girl is born.
Determine the probability mass function and the distribution function of X . Assume that the
probability is 1/2 that a baby born is a girl.
Solution: Let p be the probability mass function of X and F be its distribution function. X
is a random variable that can assume any positive integer i. p(i) = P (X = i), and X = i
occurs if the first i − 1 births are all boys and the ith birth is a girl. Thus
i−1 i
1
1
1
=
i = 1, 2, 3, . . . ,
2
2
2
p(i) =
0
i 6= 1, 2, 3, . . . .
To determine F (t), note that for t < 1, F (t) = 0; for 1 ≤ t < 2, F (t) = 1/2; for 2 ≤ t < 3,
F (t) = 1/2 + 1/4 = 3/4; for 3 ≤ t < 4, F (t) = 1/2 + 1/4 + 1/8 = 7/8; and in general
for n − 1 ≤ t < n,
n−1
X 1 i
1
1
1
1
+ 2 + 3 + · · · + n−1 =
2 2
2
2
2
i=1
1 n−1
1 − (1/2)n
−1 = 1−
=
,
1 − 1/2
2
F (t) =
by the partial sum formula for geometric series. Thus
(
0
t<1
F (t) =
1 − (1/2)n−1 n − 1 ≤ t < n,
n = 2, 3, 4, . . . . ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 162 — #178
✐
✐
162
Chapter 4
Distribution Functions and Discrete Random Variables
EXERCISES
A
1.
Let p(x) = x/15, x = 1, 2, 3, 4, 5 be probability mass function of a random variable
X . Determine F, the distribution function of X, and sketch its graph.
2.
In the experiment of rolling a balanced die twice, let X be the minimum of the two
numbers obtained. Determine the probability mass function and the distribution function
of X and sketch their graphs.
3.
In the experiment of rolling a balanced die twice, let X be the sum of the two numbers
obtained. Determine the probability mass function of X .
4.
The distribution function of a random variable X is given by
0
if x < −2
1/2 if −2 ≤ x < 2
F (x) = 3/5 if 2 ≤ x < 4
8/9 if 4 ≤ x < 6
1
if x ≥ 6.
Determine the probability mass function of X and sketch its graph.
5.
Let X be the number of random numbers selected from {0, 1, 2, . . . , 9} independently
until 0 is chosen. Find the probability mass functions of X and Y = 2X + 1.
6.
A value i is said to be the mode of a discrete random variable X if it maximizes p(x),
the probability mass function of X . Find the modes of random variables X and Y with
probability mass functions
p(x) =
and
q(y) =
1 x
2
,
x = 1, 2, 3, . . . ,
1 y 3 4−y
4!
,
y! (4 − y)! 4
4
y = 0, 1, 2, 3, 4,
respectively.
7.
For each of the following, determine the value(s) of k for which p is a probability mass
function. Note that in parts (d) and (e), n is a positive integer.
(a)
p(x) = kx, x = 1, 2, 3, 4, 5.
(b)
p(x) = k(1 + x)2 , x = −2, 0, 1, 2.
(c)
p(x) = k(1/9)x , x = 1, 2, 3, . . . .
(d)
p(x) = kx, x = 1, 2, 3, . . . , n.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 163 — #179
✐
✐
Section 4.3
(e)
p(x) = kx2 , x = 1, 2, 3, . . . , n.
Hint:
Recall that
n
X
n(n + 1)
i=
,
2
i=1
n
X
i2 =
i=1
Discrete Random Variables
163
n(n + 1)(2n + 1)
.
6
8.
From 18 potential women jurors and 28 potential men jurors, a jury of 12 is chosen at
random. Let X be the number of women selected. Find the probability mass function of
X.
9.
Let p(x) = 3/4(1/4)x , x = 0, 1, 2, 3, . . . , be probability mass function of a random
variable X . Find F, the distribution function of X, and sketch its graph.
10.
In successive rolls of a fair die, let X be the number of rolls until the first 6 appears.
Determine the probability mass function and the distribution function of X .
11.
Let X be the number of claims filed by a randomly selected customer under a certain
homeowner’s insurance policy during a ten-year period. Suppose that an actuary has
estimated that p, the probability mass function of X, satisfies
p(n + 1) = 0.32p(n),
n ≥ 0.
What is the probability that a policy holder files at least two claims during the next ten
years?
12.
Suppose that the number of claims received by an insurance company in a given week
is independent of the number of claims received in any other week. An actuary has calculated that the probability mass function of the number of claims received in a random
week is
1 4 n
,
n ≥ 0.
p(n) =
5 5
Find the probability that the company will receive exactly 8 claims during the next two
weeks.
13.
A binary digit or bit is a zero or one. A computer assembly language can generate independent random bits. Let X be the number of independent random bits to be generated
until both 0 and 1 are obtained. Find the probability mass function of X
B
14.
Every Sunday, Bob calls Liz to see if she will play tennis with him on that day. If Liz has
not played tennis with Bob since i Sundays ago, the probability that she will say yes to
him is i/k, k ≥ 2, i = 1, 2, . . . , k . Therefore, if, for example, Liz does not play tennis
with Bob for k − 1 consecutive Sundays, then she will play with him next Sunday with
probability 1. Let Z be the number of weeks it takes Liz to play again with Bob since
they last played. Find the probability mass function of Z .
15.
Let X be the number of vowels (not necessarily distinct) among the first five letters of a
random arrangement of the following expression.
ELIZABETHTAYLOR
Find the probability mass function of X . Count the letter Y as a consonant.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 164 — #180
✐
✐
164
Chapter 4
16.
From a drawer that contains 10 pairs of gloves, six gloves are selected randomly. Let X
be the number of pairs of gloves obtained. Find the probability mass function of X .
17.
A fair die is tossed successively. Let X denote the number of tosses until each of the six
possible outcomes occurs at least once. Find the probability mass function of X .
Hint: For 1 ≤ i ≤ 6, let Ei be the event that the outcome i does not occur during the
first n tosses of the die. First calculate P (X > n) by writing the event X > n in terms
of E1 , E2 , . . . , E6 .
18.
To an engineering class containing 23 male and three female students, there are 13 work
stations available. To assign each work station to two students, the professor forms 13
teams one at a time, each consisting of two randomly selected students. In this process,
let X be the total number of students selected when the first team consisting of a male
and a female appears. Find the probability mass function of X .
Distribution Functions and Discrete Random Variables
Self-Quiz on Section 4.3
Time allotted: 20 Minutes
1.
The number of claims filed with a car insurance company, per week, is a random variable
with probability mass function
p(x) =
2.
4.4
Each problem is worth 5 points.
46
,
21(x + 2)(x + 3)
x = 0, 1, 2, . . . , 20.
(a)
Find the probability of at least one such claim next week.
(b)
If we are given that there were no more than 5 claims filed with the company two
weeks ago, find the probability that there were at least 4 claims filed.
We choose 13 numbers at random and without replacement from the set {1, 2, . . . , 100}.
Let X be the median of the numbers selected. Find the probability mass function of X .
Note that the median of the 13 numbers selected is the number in the middle when they
are put in order. For example, if the numbers chosen are 7, 12, 13, 25, 41, 48, 53, 59, 67,
77, 84, 90, and 96, then their median is 53.
EXPECTATIONS OF DISCRETE RANDOM VARIABLES
To clarify the concept of expectation, consider a casino game in which the probability of losing
$1 per game is 0.6, and the probabilities of winning $1, $2, and $3 per game are 0.3, 0.08, and
0.02, respectively. The gain or loss of a gambler who plays this game only a few times depends
on his luck more than anything else. For example, in one play of the game, a lucky gambler
might win $3, but he has a 60% chance of losing $1. However, if a gambler decides to play the
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 165 — #181
✐
✐
Section 4.4
Expectations of Discrete Random Variables
165
game a large number of times, his loss or gain depends more on the number of plays than on
his luck. A calculating player argues that if he plays the game n times, for a large n, then in
approximately (0.6)n games he will lose $1 per game, and in approximately (0.3)n, (0.08)n,
and (0.02)n games he will win $1, $2, and $3, respectively. Therefore, his total gain is
(0.6)n · (−1) + (0.3)n · 1 + (0.08)n · 2 + (0.02)n · 3 = (−0.08)n.
(4.1)
This gives an average of $ − 0.08, or about 8 cents of loss per game. The more the gambler
plays, the less luck interferes and the closer his loss comes to $0.08 per game. If X is the
random variable denoting the gain in one play, then the number −0.08 is called the expected
value of X . We write E(X) = −0.08. E(X) is the average value of X . That is, if we play
the game n times and find the average of the values of X, then as n → ∞, E(X) is obtained.
Since, for this game, E(X) < 0, we have that, on the average, the more we play, the more we
lose. If for some game E(X) = 0, then in the long run the player neither loses nor wins. Such
games are called fair. In this example, X is a discrete random variable with the set of possible
values {−1, 1, 2, 3}. The probability mass function of X, p(x), is given by
i
−1
1
2
3
p(i) = P (X = i)
0.6
0.3
0.08
0.02
and p(x) = 0 if x 6∈ {−1, 1, 2, 3}. Dividing both sides of (4.1) by n, we obtain
(0.6) · (−1) + (0.3) · 1 + (0.08) · 2 + (0.02) · 3 = −0.08.
Hence
−1 · p(−1) + 1 · p(1) + 2 · p(2) + 3 · p(3) = −0.08,
a relation showing that the expected value of X can be calculated directly by summing up the
product of possible values of X by their probabilities. This and similar examples motivate the
following general definition, which was first used casually by Pascal but introduced formally
by Huygens in the late seventeenth century.
Definition 4.4 The expected value of a discrete random variable X with the set of possible
values A and probability mass function p(x) is defined by
E(X) =
X
xp(x).
x∈A
We say that E(X) exists if this sum converges absolutely.
The expected value of a random variable X is also called the mean, or the mathematical
expectation, or simply the expectation of X . It is also occasionally denoted by E[X], EX,
µX , or µ.
P
Note that if each value x of X is weighted by p(x) = P (X = x), then x∈A xp(x) is
nothing but the weighted average of X . Similarly, if we think of a unit mass distributed along
the real line at the points of A so that the mass at x ∈ A is P (X = x), then E(X) is the
center of gravity.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 166 — #182
✐
✐
166
Chapter 4
Distribution Functions and Discrete Random Variables
Here are some examples to illuminate the notion of expected value, a fundamental concept
in probability and statistics.
Example 4.14 We flip a fair coin twice and let X be the number of heads obtained. What is
the expected value of X ?
Solution: The possible values of X are 0, 1, and 2, and the probability mass function of X is
given by p(0) = P (X = 0) = 1/4, p(1) = P (X = 1) = 1/2, p(2) = P (X = 2) = 1/4,
and p(x) = 0 if x 6∈ {0, 1, 2}. Thus
E(X) = 0 · p(0) + 1 · p(1) + 2 · p(2) = 0 ·
1
1
1
+ 1 · + 2 · = 1.
4
2
4
Therefore, we can expect an average of one head in every two flips.
Example 4.15 We write the numbers a1 , a2 , . . . , an on n identical balls and mix them in a
box. What is the expected value of a ball selected at random?
Solution: The set of possible values of X, the numbers written on the balls selected, is
{a1 , a2 , . . . , an }. The probability mass function of X is given by
p(a1 ) = p(a2 ) = · · · = p(an ) = 1/n,
and p(x) = 0 if x 6∈ {a1 , a2 , . . . , an }. Thus
E(X) =
X
x∈{a1 ,a2 ,...,an }
=
xp(x) = a1 ·
1
1
1
+ a2 · + · · · + an ·
n
n
n
a1 + a2 + · · · + an
.
n
Therefore, as expected, E(X) coincides with the average of the values a1 , a2 , . . . , an . That
is, if for a large number of times we draw balls at random and with replacement, record their
values, and then find their average, the result obtained is the arithmetic mean of the numbers
a1 , a2 , . . . , an . Example 4.16 A college mathematics department sends 8 to 12 professors to the annual
meeting of the American Mathematical Society, which lasts five days. The hotel at which the
conference is held offers a bargain rate of a dollars per day per person if reservations are
made 45 or more days in advance, but charges a cancellation fee of 2a dollars per person. The
department is not certain how many professors will go. However, from past experience it is
known that the probability of the attendance of i professors is 1/5 for i = 8, 9, 10, 11, and 12.
If the regular rate of the hotel is 2a dollars per day per person, should the department make any
reservations? If so, how many?
Solution: For i = 8, 9, 10, 11, and 12, let Xi be the total cost in dollars if the department
makes reservations for i professors. We should compute E(Xi ) for i = 8, 9, 10, 11, and
12. If E(Xj ) is the smallest of all, the department should make reservations for j professors.
To calculate E(X8 ), note that X8 only assumes the values 40a, 50a, 60a, 70a, and 80a,
which correspond to the cases that 8, 9, 10, 11, and 12 professors attend, respectively, while
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 167 — #183
✐
✐
Section 4.4
Expectations of Discrete Random Variables
167
reservations are made only for eight professors. Since the probability of any of these is 1/5, we
have
1
1
1
1
1
E(X8 ) = (40a) + (50a) + (60a) + (70a) + (80a) = 60a.
5
5
5
5
5
Similarly,
1
1
1
1
1
E(X9 ) = (42a) + (45a) + (55a) + (65a) + (75a) = 56.4a,
5
5
5
5
5
1
1
1
1
1
E(X10 ) = (44a) + (47a) + (50a) + (60a) + (70a) = 54.2a,
5
5
5
5
5
1
1
1
1
1
E(X11 ) = (46a) + (49a) + (52a) + (55a) + (65a) = 53.4a,
5
5
5
5
5
1
1
1
1
1
E(X12 ) = (48a) + (51a) + (54a) + (57a) + (60a) = 54a.
5
5
5
5
5
We see that X11 has the smallest expected value. Thus making 11 reservations is the most
reasonable policy. Example 4.17 In the lottery of a certain state, players pick six different integers between 1
and 49, the order of selection being irrelevant. The lottery commission then selects six of these
numbers at random as the winning numbers. A player wins the grand prize of $1,200,000 if
all six numbers that he has selected match the winning numbers. He wins the second and third
prizes of $800 and $35, respectively, if exactly five and four of his six selected numbers match
the winning numbers. What is the expected value of the amount a player wins in one game?
Solution: Let X be the amount that a player wins in one game. Then the possible values of
X are 1,200,000; 800; 35; and 0. The probabilities associated with these values are
P (X = 1, 200, 000) =
P (X = 800) =
6
5
!
49
6
43
1
!
!
≈ 0.000, 018,
1
! ≈ 0.000, 000, 072,
49
6
P (X = 35) =
!
!
6
43
4
2
! ≈ 0.000, 97,
49
6
and
P (X = 0) = 1 − 0.000, 000, 072 − 0.000, 018 − 0.000, 97 = 0.999, 011, 928.
Therefore,
E(X) ≈ 1, 200, 000(0.000, 000, 072) + 800(0.000, 018) + 35(0.000, 97)
+ 0(0.999, 011, 928) ≈ 0.13.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 168 — #184
✐
✐
168
Chapter 4
Distribution Functions and Discrete Random Variables
This shows that on the average players will win 13 cents per game. If the cost per game is 50
cents, then, on the average, a player will lose 37 cents per game. Therefore, a player who plays
10,000 games over several years will lose approximately $3700. Let X be a discrete random variable with a set of P
possible values A and probability
mass
function
p
.
We
say
that
E(X)
exists
if
the
sum
x∈A xp(x) converges, that is, if
P
x∈A xp(x) < ∞. We now present two examples of random variables whose mathematical expectations do not exist.
Example 4.18 (St. Petersburg Paradox) In a game, the player flips a fair coin successively until he gets a heads. If this occurs on the k th flip, the player wins 2k dollars. Therefore,
if the outcome of the first flip is heads, the player wins $2. If the outcome of the first flip is
tails but that of the second flip is heads, he wins $4. If the outcomes of the first two are tails but
the third one heads, he will win $8, and so on. The question is, to play this game, how much
should a person, who is willing to play a fair game, pay? To answer this question, let X be the
amount of money the player wins. Then X is a random variable with the set of possible values
{2, 4, 8, . . . , 2k , . . .} and
P (X = 2k ) =
Therefore,
E(X) =
∞
X
k=1
2k
1 k
2
1 k
2
=
,
∞
X
k=1
k = 1, 2, 3, . . . .
1 = 1 + 1 + 1 + · · · = ∞.
This result shows that the game remains unfair even if a person pays the largest possible amount
to play it. In other words, this is a game in which one always wins no matter how expensive it is
to play. To see what the flaw is, note that theoretically this game is not feasible to play because
it requires an enormous amount of money. In practice, however, the probability that a gambler
wins $2k for a large k is close to zero. Even for small values of k, winning is highly unlikely.
For example, to win $230 = 1, 073, 741, 824, you should get 29 tails in a row followed by a
head. The chance of this happening is 1 in 1,073,741,824, much less than 1 in a billion. The following example serves to illuminate further the concept of expected value. At the
same time, it shows the inadequacy of this concept as a central measure.
Example 4.19 Let X0 be the amount of rain that will fall in the United States on the next
Christmas day. For n > 0, let Xn be the amount of rain that will fall in the United States
on Christmas n years later. Let N be the smallest number of years that elapse before we get
a Christmas rainfall greater than X0 . Suppose that P (Xi = Xj ) = 0 if i 6= j, the events
concerning the amount of rain on Christmas days of different years are all independent, and the
Xn ’s are identically distributed. Find the expected value of N .
Solution: Since N is the first value for n for which Xn > X0 ,
P (N > n) = P (X0 > X1 , X0 > X2 , . . . , X0 > Xn )
= P max(X0 , X1 , X2 , . . . , Xn ) = X0 =
1
;
n+1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 169 — #185
✐
✐
Section 4.4
Expectations of Discrete Random Variables
169
where the last equality follows from the symmetry, there is no more reason for the maximum
to be at X0 than there is for it to be at Xi , 0 ≤ i ≤ n. Therefore,
P (N = n) = P (N > n − 1) − P (N > n) =
1
1
1
−
=
.
n n+1
n(n + 1)
From this it follows that
E(N ) =
∞
X
nP (N = n) =
n=1
∞
X
∞
X
n
1
=
= ∞.
n(n + 1) n=1 n + 1
n=1
Note that P (N > n − 1) = 1/n gives the probability that, in the United States, we will have
to wait more than, say, three years for a Christmas rainfall that is greater than X0 is only 1/4,
and the probability that we must wait more than nine years is only 1/10. Even with such low
probabilities, on average, it will still take infinitely many years before we will have more rain
on a Christmas day than we will have on next Christmas day. Example 4.20 The tanks of a country’s army are numbered 1 to N . In a war this country
loses n random tanks to the enemy, who discovers that the captured tanks are numbered. If
X1 , X2 , . . . , Xn are the numbers of the captured tanks, what is E(max Xi )? How can the
enemy use E(max Xi ) to find an estimate of N, the total number of this country’s tanks?
Solution: Let Y = max Xi ; then
P (Y = k) =
!
k−1
n−1
!
N
n
for k = n, n + 1, n + 2, . . . , N,
because if the maximum of Xi ’s is k, the numbers of the remaining n − 1 tanks are from 1 to
k − 1. Now
!
k−1
k
N
N
X
X
n−1
!
E(Y ) =
kP (Y = k) =
N
k=n
k=n
n
=
=
To calculate
N
X
k=n
1
N
n
n
N
n
!
!
N
X
k (k − 1)!
=
(n
−
1)!
(k
−
n)!
k=n
N
X
k=n
!
k
, note that
n
!
k
.
n
k
n
!
n
N
n
!
N
X
k!
n!
(k
− n)!
k=n
(4.2)
is the coefficient of xn in the polynomial (1 + x)k .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 170 — #186
✐
✐
170
Chapter 4
Therefore,
N
X
k=n
N
X
k
n
Distribution Functions and Discrete Random Variables
!
is the coefficient of xn in the polynomial
N
X
(1 + x)k . Since
k=n
k
n
(1 + x) = (1 + x)
k=n
N
−n
X
(1 + x)k = (1 + x)n
k=0
1
= (1 + x)N +1 − (1 + x)n ,
x
(1 + x)N −n+1 − 1
(1 + x) − 1
!
N
+
1
1
N
+1
n
, we have
and the coefficient of xn in the polynomial (1 + x)
− (1 + x) is
n+1
x
N
X
k=n
!
k
=
n
!
N +1
.
n+1
Substituting this result in (4.2), we obtain
!
N +1
(N + 1)!
n
n
n+1
(n + 1)! (N − n)!
n(N + 1)
!
E(Y ) =
=
=
.
n+1
N!
N
n! (N − n)!
n
n(N + 1)
for
n+1
n+1
N . We obtain N =
E(Y ) − 1. Therefore, if, for example, the enemy captures 12
n
tanks and the maximum number of tanks captured is 117, then assuming that E(Y ) is approximately equal to the value of Y observed, we get N ≈ (13/12) × 117 − 1 ≈ 126.
To estimate N, the total number of this country’s tanks, we solve E(Y ) =
The solutions to Examples 4.21 and 4.22 were given by Professor James Frykman from
Kent State University.
Example 4.21 An urn contains w white and b blue chips. A chip is drawn at random and
then is returned to the urn along with c > 0 chips of the same color. This experiment is then
repeated successively. Let Xn be the number of white chips drawn during the first n draws.
Show that E(Xn ) = nw/(w + b).
Solution: For n ≥ 1, let pn be the probability mass function of Xn . We will prove, by induction,
that E(Xn ) = nw/(w + b). For n = 1,
E(X1 ) = 0 ·
b
w
1·w
+1·
=
.
w+b
w+b
w+b
Suppose that E(Xn ) = nw/(w + b) for any integer n ≥ 1. To demonstrate that E(Xn+1 ) =
(n + 1)w/(w + b), note that
E(Xn+1 ) =
n+1
X
k=0
kpn+1 (k) = (n + 1)pn+1 (n + 1) +
n
X
kpn+1 (k).
(4.3)
k=1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 171 — #187
✐
✐
Section 4.4
Expectations of Discrete Random Variables
171
Now
pn+1 (n + 1) = P (Xn+1 = n + 1)
= P (Xn+1 = n + 1 | Xn = n)P (Xn = n)
w + nc
pn (n),
=
w + b + nc
(4.4)
and for 1 ≤ k ≤ n,
pn+1 (k) = P (Xn+1 = k) = P (Xn+1 = k | Xn = k)P (Xn = k)
+ P (Xn+1 = k | Xn = k − 1)P (Xn = k − 1)
=
b + (n − k)c
w + (k − 1)c
pn (k) +
pn (k − 1).
w + b + nc
w + b + nc
In relation (4.3), substituting (4.4) for pn+1 (n + 1) and (4.5) for pn+1 (k), we obtain
n
X
k b + (n − k)c
(n + 1)(w + nc)
E(Xn+1 ) =
pn (n) +
pn (k)
w + b + nc
w + b + nc
k=1
n
X
k w + (k − 1)c
+
pn (k − 1).
w + b + nc
k=1
(4.5)
(4.6)
A shift in index of the last sum in (4.6) gives
n
X
k w + (k − 1)c
pn (k − 1)
w + b + nc
k=1
=
=
n−1
X
1
(k + 1)(w + kc)pn (k)
w + b + nc k=0
n−1
n−1
X
X
1
k(w + kc)pn (k) +
(w + kc)pn (k)
w + b + nc k=0
k=0
n−1
n−1
n−1
X
X
X
1
k(w + kc)pn (k) +
wpn (k) +
kcpn (k)
=
w + b + nc k=1
k=0
k=1
n−1
n−1
X
X
1
=
k(w + kc + c)pn (k) + w
pn (k)
w + b + nc k=1
k=0
X
n
n
X
1
=
k(w + kc + c)pn (k) + w
pn (k)
w + b + nc k=1
k=1
− n(w + nc + c)pn (n) − wpn (n) .
(4.7)
Relations (4.6) and (4.7) combined yield what is needed to complete the proof:
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 172 — #188
✐
✐
172
Chapter 4
Distribution Functions and Discrete Random Variables
E(Xn+1 ) =
(n + 1)(w + nc) − n(w + nc + c) − w
pn (n)
w + b + nc
n
X
k b + (n − k)c + k(w + kc + c)
w
pn (k) +
+
w + b + nc
w + b + nc
k=1
n
w + b + nc + c X
w
=
kpn (k) +
w + b + nc k=1
w + b + nc
=
w
w + b + nc + c
E(Xn ) +
w + b + nc
w + b + nc
=
w + b + nc + c nw
w
(n + 1)w
·
+
=
. w + b + nc
w + b w + b + nc
w+b
Example 4.22 (Pólya’s Urn Model) An urn contains w white and b blue chips. A chip
is drawn at random and then is returned to the urn along with c > 0 chips of the same color.
Prove that if n = 2, 3, 4, . . . , such experiments are made, then at each draw the probability of
a white chip is still w/(w + b) and the probability of a blue chip is b/(w + b). This model was
first introduced in preliminary studies of “contagious diseases” and the spread of epidemics as
well as “accident proneness” in actuarial mathematics.
Solution: For all n ≥ 1, let Wn be the event that the nth draw is white and Bn be the event that
it is blue. We will show that P (Wn ) = w/(w + b). This implies that P (Bn ) = 1 − P (Wn ) =
b/(w + b). For all n ≥ 1, let Xn be the number of white chips drawn during the first n
draws. Let pn be the probability mass function of Xn . Clearly, P (W1 ) = w/(w + b). To show
that for n ≥ 2, P (Wn ) = w/(w + b), note that the events {Xn−1 = 0}, {Xn−1 = 1},
. . . , {Xn−1 = n − 1} form a partition of the sample space. Therefore, by the law of total
probability (Theorem 3.4),
P (Wn ) =
n−1
X
k=0
=
n−1
X
k=0
=
P (Wn | Xn−1 = k)P (Xn−1 = k)
P (Xn = k + 1 | Xn−1 = k)P (Xn−1 = k)
n−1
X
w + kc
pn−1 (k)
w + b + (n − 1)c
k=0
n−1
n−1
X
X
w
c
pn−1 (k) +
kpn−1 (k)
w + b + (n − 1)c k=0
w + b + (n − 1)c k=0
w
c
=
+
E(Xn−1 ).
w + b + (n − 1)c w + b + (n − 1)c
=
Now by Example 4.21, E(Xn−1 ) = (n − 1)w/(w + b). Hence, for n ≥ 2,
P (Wn ) =
w
c
(n − 1)w
w
+
·
=
. w + b + (n − 1)c w + b + (n − 1)c
w+b
w+b
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 173 — #189
✐
✐
Section 4.4
173
Expectations of Discrete Random Variables
We now discuss some elementary properties of the expected value of a discrete random
variable. Further properties of expected values are discussed in subsequent chapters.
Theorem 4.1 If X is a constant random variable, that is, if P (X = c) = 1 for a constant
c, then E(X) = c.
Proof:
There is only one possible value for X and that is c; hence E(X)
c · P (X = c) = c · 1 = c. =
Let g : R → R be a real-valued function and X be a discrete random P
variable with set of
possible values A and probability mass function p(x). Similar to E(X) = x∈A xp(x), there
P
is the important relation E g(X) = x∈A g(x)p(x), known as the law of the unconscious
statistician, which we now prove. This relation enables us to calculate the expected value of
the random variable g(X) without deriving its probability mass function. It implies that, for
example,
E(X 2 ) =
X
x2 p(x),
x∈A
E(X 2 − 2X + 4) =
E(X cos X) =
X
(x2 − 2x + 4)p(x),
x∈A
X
(x cos x)p(x),
x∈A
E(eX ) =
X
ex p(x).
x∈A
Theorem 4.2 Let X be a discrete random variable with set of possible values A and probability mass function p(x), and let g be a real-valued function. Then g(X) is a random variable
with
X
E g(X) =
g(x)p(x).
x∈A
Proof: Let S be the sample space. We are given that g : R → R is a real-valued function
and X, : S → A ⊆ R is a random variable with the set of possible valuesA. As we know,
g(X), the composition of g and X, is a function from S to the set g(A) = g(x) : x ∈ A .
Hence g(X) is a random variable with the possible set of values g(A). Now, by the definition
of expected value,
X
E g(X) =
zP g(X) = z .
z∈g(A)
Let g −1 {z} = x : g(x) = z , and notice that we are not claiming that g has an inverse
function. We are simply considering the set x : g(x) = z , which is called the inverse image
of z and is denoted by g −1 {z} . Now
P g(X) = z = P X ∈ g −1 {z} =
X
{x : x∈g−1 ({z})}
P (X = x) =
X
p(x).
{x : g(x)=z}
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 174 — #190
✐
✐
174
Chapter 4
Distribution Functions and Discrete Random Variables
Thus
X
X
z
zP g(X) = z =
E g(X) =
z∈g(A)
z∈g(A)
=
X
X
zp(x) =
z∈g(A) {x : g(x)=z}
=
X
X
X
p(x)
{x : g(x)=z}
X
g(x)p(x)
z∈g(A) {x : g(x)=z}
g(x)p(x),
x∈A
where the last equality follows from the fact that the sum over A can be performed in two
stages: We can first sum over all x with g(x) = z, and then over all z . Corollary Let X be a discrete random variable; g1 , g2 , . . . , gn be real-valued functions,
and let α1 , α2 , . . . , αn be real numbers. Then
E α1 g1 (X) + α2 g2 (X) + · · · + αn gn (X)
= α1 E g1 (X) + α2 E g2 (X) + · · · + αn E gn (X) .
Proof: Let the set of possible values of X be A, and its probability mass function be p(x).
Then, by Theorem 4.2,
E α1 g1 (X) + α2 g2 (X) + · · · + αn gn (X)
X
=
α1 g1 (x) + α2 g2 (x) + · · · + αn gn (x) p(x)
x∈A
= α1
X
x∈A
g1 (x)p(x) + α2
X
x∈A
g2 (x)p(x) + · · · + αn
X
gn (x)p(x)
x∈A
= α1 E g1 (X) + α2 E g2 (X) + · · · + αn E gn (X) . By this corollary, for example, we have relations such as the following:
E(2X 3 + 5X 2 + 7X + 4) = 2E(X 3 ) + 5E(X 2 ) + 7E(X) + 4
E(eX + 2 sin X + log X) = E(eX ) + 2E(sin X) + E(log X).
Moreover, this corollary implies that E(X) is linear. That is, if α, β ∈ R, then
E(αX + β) = αE(X) + β.
Example 4.23
The probability mass function of a discrete random variable X is given by
(
x/15 x = 1, 2, 3, 4, 5
p(x) =
0
otherwise.
What is the expected value of X(6 − X)?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 175 — #191
✐
✐
Section 4.4
Expectations of Discrete Random Variables
175
Solution: By Theorem 4.2,
2
3
4
5
1
+8·
+9·
+8·
+5·
= 7. E X(6 − X) = 5 ·
15
15
15
15
15
Example 4.24 A box contains 10 disks of radii 1, 2, . . . , and 10, respectively. What is the
expected value of the area of a disk selected at random from this box?
Solution: Let the radius of the disk be R; then R is a random variable with the probability
mass function p(x) = 1/10 if x = 1, 2, . . . , 10, and p(x) = 0 otherwise. E(πR2 ), the
desired quantity is calculated as follows:
E(πR2 ) = πE(R2 ) = π
10
X
i2
i=1
1
= 38.5π. 10
EXERCISES
A
1.
There is a story about Charles Dickens (1812–1870), the English novelist and one of the
most popular writers in the history of literature. It is known that Dickens was interested
in practical applications of mathematics. On the final day in March during a year in the
second half of the nineteenth century, he was scheduled to leave London by train and
travel about an hour to visit a very good friend. However, Mr. Dickens was aware of the
fact that in England there were, on the average, two serious train accidents each month.
Knowing that there had been only one serious accident so far during the month of March,
Dickens thought that the probability of a serious train accident on the last day of March
would be very high. Thus he called his friend and postponed his visit until the next day.
He boarded the train on April 1, feeling much safer and believing that he had used his
knowledge of mathematics correctly by leaving the next day. He did arrive safely! Is
there a fallacy in Dickens, argument? Explain.
2.
In a certain part of downtown Baltimore parking lots charge $7 per day. A car that is
illegally parked on the street will be fined $25 if caught, and the chance of being caught
is 60%. If money is the only concern of a commuter who must park in this location every
day, should he park at a lot or park illegally?
3.
Let X be a discrete random variable with probability mass function
4.
x
-2
0
2
4
p(x)
1/3
1/4
1/4
1/6
Find E 2(X − 1)(3 − X) .
In a lottery every week, 2,000,000 tickets are sold for $1 apiece. If 4000 of these tickets
pay off $30 each, 500 pay off $800 each, one ticket pays off $1,200,000, and no ticket
pays off more than one prize, what is the expected value of the winning amount for a
player with a single ticket?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 176 — #192
✐
✐
176
Chapter 4
Distribution Functions and Discrete Random Variables
5.
In a lottery, a player pays $1 and selects four distinct numbers from 0 to 9. Then, from an
urn containing 10 identical balls numbered from 0 to 9, four balls are drawn at random
and without replacement. If the numbers of three or all four of these balls matches the
player’s numbers, he wins $5 and $10, respectively. Otherwise, he loses. On the average,
how much money does the player gain per game? (Gain = win − loss.)
6.
An urn contains five balls, two of which are marked $1, two $5, and one $15. A game is
played by paying $10 for winning the sum of the amounts marked on two balls selected
randomly from the urn. Is this a fair game?
7.
Let X be a discrete random variable with the following probability mass function
x
π/6
π/4
π/3
π/2
p(x)
0.2
0.4
0.3
0.1
Find E(cos X).
8.
A box contains 20 fuses, of which five are defective. What is the expected number of
defective items among three fuses selected randomly?
9.
The demand for a certain weekly magazine at a newsstand is a random variable with
probability mass function p(i) = (10 − i)/18, i = 4, 5, 6, 7. If the magazine sells for
$a and costs $2a/3 to the owner, and the unsold magazines cannot be returned, how
many magazines should be ordered every week to maximize the profit in the long run?
P∞
It is well known that x=1 1/x2 = π 2 /6.
10.
11.
(a)
Show that p(x) = 6/(πx)2 , x = 1, 2, 3, . . . is the probability mass function of
a random variable X .
(b)
Prove that E(X) does not exist.
2
Show that p(x) = |x| + 1 /27, x = −2, −1, 0, 1, 2, is the probability mass
function of a random variable X .
Calculate E(X), E |X| , and E(2X 2 − 5X + 7).
(a)
(b)
12.
A box contains 10 disks of radii 1, 2, . . . , 10, respectively. What is the expected value
of the circumference of a disk selected at random from this box?
13.
The amount that an insurance policy pays for hospitalization is $a per day up to 4 days
and $(a/2) thereafter. Let X be the number of days a randomly selected policy holder
who needs hospitalization is hospitalized. If the probability mass function of X is given
by
8−n
p(n) =
,
1 ≤ n ≤ 7,
28
find the expected amount of money that the policy pays for each hospitalization.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 177 — #193
✐
✐
Section 4.4
14.
15.
Expectations of Discrete Random Variables
177
The distribution function of a random variable X is given by
0
if x < −3
3/8 if −3 ≤ x < 0
F (x) = 1/2 if 0 ≤ x < 3
3/4 if 3 ≤ x < 4
1
if x ≥ 4.
Calculate E(X), E X 2 − 2|X| , and E X|X| .
If X is a random number selected from the first 10 positive integers, what is the expected
value of X(11 − X)?
16.
Let X be the number of different birthdays among four persons selected randomly. Find
E(X).
17.
A newly married couple decides to continue having children until they have one of each
sex. If the events of having a boy and a girl are independent and equiprobable, how many
children should this
expect?
Pcouple
∞
Hint: Note that i=1 ir i = r/(1 − r)2 , |r| < 1.
B
18.
Suppose that there exist N families on the earth and that the maximum number of children a family
j = 0, 1, 2, . . . , c, let αj be the fraction of families with j
Pc has is c. For
children
j=0 αj = 1 . A child is selected at random from the set of all children in the
world. Let this child be the K th born of his or her family; then K is a random variable.
Find E(K).
19.
An ordinary deck of 52 cards is well-shuffled, and then the cards are turned face up one
by one until an ace appears. Find the expected number of cards that are face up.
20.
Suppose that n random integers are selected from {1, 2, . . . , N } with replacement.
What is the expected value of the largest number selected? Show that for large N the
answer is approximately nN/(n + 1).
21.
(a)
Show that
p(n) =
1
,
n(n + 1)
n ≥ 1,
is a probability mass function.
(b)
22.
Let X be a random variable with probability mass function p given in part (a);
find E(X).
To an engineering class containing 2n − 3 male and three female students, there are n
work stations available. To assign each work station to two students, the professor forms
n teams one at a time, each consisting of two randomly selected students. In this process,
let X be the number of students selected until a team of a male and a female is formed.
Find the expected value of X .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 178 — #194
✐
✐
178
Chapter 4
Distribution Functions and Discrete Random Variables
Self-Quiz on Section 4.4
Time allotted: 20 Minutes
1.
Each problem is worth 2.5 points.
Let X be a discrete random variable with probability mass function
x
0
1
3
7
p(x)
0.2
0.3
0.4
0.1
Find E(2X ).
2.
Two fair dice are tossed and the maximum of the outcomes is denoted by X . Find E(X).
3.
In a certain lottery, 15,000 tickets are sold. If there are 300 prizes of $5 each, 50 prizes
of $50 each, and one grand prize of $1,000, what is the fair price for each ticket?
4.
Let F be the distribution function of a random variable X . Find E(X).
0
t<2
0.2
2≤t<3
F (t) = 0.3
3≤t<5
0.8
5≤t<7
1
t≥7
4.5
VARIANCES AND MOMENTS OF DISCRETE RANDOM VARIABLES
Thus far, through many examples, we have explained the importance of mathematical expectation in detail. For instance, in Example 4.16, we have shown how expectation is applied in
decision making. Also, in Example 4.17, concerning the lottery, we showed that the expected
value of the winning amount per game gives an excellent estimation for the total amount a
player will win if he or she plays a large number of times. In these and many other situations,
mathematical expectation is the only quantity one needs to calculate. However, very frequently
we face situations in which the expected value by itself does not say much. In such cases more
information should be extracted from the probability mass function. As an example, suppose
that we are interested in measuring a certain quantity. Let X be the true value† of the quantity
minus the value obtained by measurement. Then X is the error of measurement. It is a random
variable with expected value zero, the reason being that in measuring a quantity a very large
number of times, positive and negative errors of the same magnitudes occur with equal probabilities. Now consider an experiment in which a quantity is measured several times, and the
†
True value is a nebulous concept. Here we shall use it to mean the average of a large number of measurements.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 179 — #195
✐
✐
Section 4.5
Variances and Moments of Discrete Random Variables
179
average of the errors is obtained to be a number close to zero. Can we conclude that the measurements are very close to the true value and thus are accurate? The answer is no because they
might differ from the true value by relatively large quantities but be scattered both in positive
and negative directions, resulting in zero expectation. Thus in this and similar cases, expectation by itself does not give adequate information, so additional measures for decision making
are needed. One such quantity is the variance of a random variable.
Variance measures the average magnitude of the fluctuations of a random variable from its
expected value. This is particularly important because random variables fluctuate from their
expected values. To mathematically define the variance of a random variable X, the first temptation
the expectation of the difference of X from its expected value, that is,
is to consider
E X − E(X) . But the difficulty with this quantity is that the positive and negative deviations
of X from E(X) cancel each other, and we always get 0. This can be seen mathematically
from the corollary of Theorem 4.2: Let E(X) = µ; then
E X − E(X) = E(X − µ) = E(X) − µ = E(X) − E(X) = 0.
Hence E X − E(X) is not an appropriate measure for the variance. However, if we consider
E |X − E(X)| instead, the problem of negative and positive deviations canceling each other
disappears. Since this quantity is the true average magnitude of the fluctuations of X from
E(X), it seems that it is the best candidate for an expression for the variance of X . But mathe
2 matically, E |X−E(X)| is difficult to handle; for this reason the quantity E X−E(X) ,
analogous to Euclidean distance in geometry, is used instead and is called the variance of X .
2 The square root of E X − E(X)
is called the standard deviation of X .
Definition 4.5 Let X be a discrete random variable with a set of possible values A, probability mass function p(x), and E(X) = µ. Then σX and Var(X), called the standard deviation
and the variance of X, respectively, are defined by
q σX = E (X − µ)2 and Var(X) = E (X − µ)2 .
Note that by this definition and Theorem 4.2,
X
Var(X) = E (X − µ)2 =
(x − µ)2 p(x).
x∈A
Let X be a discrete random variable with the set of possible values A and probability mass
function p(x). Suppose that the prediction of the value of X is in order, and if the value t is
predicted for X, then based on the error X −t, a penalty is charged. To minimize the penalty,
it seems reasonable to minimize E (X − t)2 . But
X
(x − t)2 p(x).
E (X − t)2 =
x∈A
Assuming that this series converges i.e., E(X 2 ) < ∞ , we differentiate it to find the mini
mum value of E (X − t)2 :
X
d d X
E (X − t)2 =
(x − t)2 p(x) =
−2(x − t)p(x) = 0.
dt
dt x∈A
x∈A
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 180 — #196
✐
✐
180
Chapter 4
This gives
Distribution Functions and Discrete Random Variables
X
xp(x) = t
x∈A
X
p(x) = t.
x∈A
P
Therefore, E (X − t)2 is a minimum for t = x∈A xp(x) = E(X) and the minimum value
2 is E X − E(X)
= Var(X). So the smaller that Var(X) is, the better E(X) predicts X .
We have that
Var(X) = min E (X − t)2 .
t
Earlier we mentioned that, if we think of a unit mass distributed along the real line at the
points of A so that the mass at x ∈ A is p(x) = P (X = x), then E(X) is the center of
gravity. As we know, since center of gravity does not provide any information about how the
mass is distributed around this center, the concept of moment of inertia is introduced. Moment
of inertia is a measure of dispersion (spread) of the mass distribution about the center of gravity.
E(X) is analogous to the center of gravity, and it too does not provide any information about
the distribution of X about this center of location. However, variance, the analog of the moment
of inertia, measures the dispersion, or spread, of a distribution about its expected value.
Related Historical Remark: In 1900, the Wright brothers were looking for a private location
with consistent wind to test their gliders. The data they got from the U.S. Weather Bureau
indicated that Kill Devil Hill, near Kitty Hawk, North Carolina, had, on average, suitable winds,
so they chose that location for their tests. However, the wind was not consistent. There were
many calm days and many days with strong winds that were not suitable for their tests. The
summary of the data was obtained by averaging undesirable extreme wind conditions. The
Weather Bureau statistics were misleading because they failed to utilize the standard deviation.
If the Wright brothers had been provided with the standard deviation of the wind speed at Kitty
Hawk, they would not have chosen that location for their tests. Example 4.25 Karen is interested in two games, Keno and Bolita. To play Bolita, she buys
a ticket for $1, draws a ball at random from a box of 100 balls numbered 1 to 100. If the ball
drawn matches the number on her ticket, she wins $75; otherwise, she loses. To play Keno,
Karen bets $1 on a single number that has a 25% chance to win. If she wins, they will return
her dollar plus two dollars more; otherwise, they keep the dollar. Let B and K be the amounts
that Karen gains in one play of Bolita and Keno, respectively. Then
E(B) = (74)(0.01) + (−1)(0.99) = −0.25,
E(K) = (2)(0.25) + (−1)(0.75) = −0.25.
Therefore, in the long run, it does not matter which of the two games Karen plays. Her gain
would be about the same. However, by virtue of
Var(B) = E (B − µ)2 = (74 + 0.25)2 (0.01) + (−1 + 0.25)2 (0.99) = 55.69
and
Var(K) = E (K − µ)2 = (2 + 0.25)2 (0.25) + (−1 + 0.25)2 (0.75) = 1.6875,
we can say that in Bolita, on average, the deviation of the gain from the expected value is much
higher than in Keno. In other words, the risk with Keno is far less than the risk with Bolita. In
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 181 — #197
✐
✐
Section 4.5
Variances and Moments of Discrete Random Variables
181
Bolita, the probability of winning is very small, but the amount one might win is high. In Keno,
players win more often but in smaller amounts. The following theorem states another useful formula for Var(X).
Theorem 4.3
2
Var(X) = E(X 2 ) − E(X) .
Proof: By the definition of variance,
Var(X) = E (X − µ)2 = E(X 2 − 2µX + µ2 )
= E(X 2 ) − 2µE(X) + µ2
= E(X 2 ) − 2µ2 + µ2 = E(X 2 ) − µ2
2
= E(X 2 ) − E(X) . One immediate application of this formula is that, since Var(X) ≥ 0, for any discrete
random variable X,
2
E(X) ≤ E(X 2 ).
2
The formula Var(X) = E(X 2 ) − E(X) is usually a better alternative for computing the
variance of X . Here is an example.
Example 4.26
die?
What is the variance of the random variable X, the outcome of rolling a fair
Solution: The probability mass function of X is given by p(x) = 1/6; x =1, 2, 3, 4, 5, 6,
and p(x) = 0, otherwise. Hence
E(X) =
6
X
x=1
2
E(X ) =
6
X
x=1
6
xp(x) =
7
1
1X
x = (1 + 2 + 3 + 4 + 5 + 6) = ,
6 x=1
6
2
6
x2 p(x) =
1X 2 1
91
x = (1 + 4 + 9 + 16 + 25 + 36) = .
6 x=1
6
6
Thus
2
35
91 49
−
= . Var(X) = E(X 2 ) − E(X) =
6
4
12
Suppose that a random variable X is constant; then E(X) = X and the deviations of X
from E(X) are 0. Therefore, the average deviation of X from E(X) is also 0. We have the
following theorem.
Theorem 4.4 Let X be a discrete random variable with the set of possible values A, and
mean µ. Then Var(X) = 0 if and only if X is a constant with probability 1.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 182 — #198
✐
✐
182
Chapter 4
Distribution Functions and Discrete Random Variables
Proof: We will show that Var(X) = 0 implies that X = µ with probability 1. Suppose not;
then there exists some k 6= µ such that p(k) = P (X = k) > 0. But then
Var(X) = (k − µ)2 p(k) +
X
x∈A−{k}
(x − µ)2 p(x) > 0,
which is a contradiction to Var(X) = 0. Conversely, if X is a constant c with
probability
1,
then X = E(X) = c = µ with probability 1. This implies that Var(X) = E (X − µ)2 = 0.
For constants a and b, a linear relation similar to E(aX + b) = aE(X) + b does not exist
for variance and for standard deviation. However, other important relations exist and are given
by the following theorem.
Theorem 4.5
that
Let X be a discrete random variable; then for constants a and b we have
Var(aX + b) = a2 Var(X),
σaX+b = |a|σX .
Proof: To see this, note that
2
Var(aX + b) = E (aX + b) − E(aX + b)
2
= E (aX + b) − aE(X) + b
2
= E a X − E(X)
2 = E a2 X − E(X)
2 = a2 E X − E(X)
= a2 Var(X).
Taking the square roots of both sides of this relation, we find that σaX+b = |a|σX .
Example
4.27
Suppose that, for a discrete random variable X, E(X) = 2 and
E X(X − 4) = 5. Find the variance and the standard deviation of −4X + 12.
Solution: By the Corollary of Theorem 4.2, E(X 2 − 4X) = 5 implies that
E(X 2 ) − 4E(X) = 5.
Substituting E(X) in this relation gives E(X 2 ) = 13. Hence, by Theorem 4.3,
2
Var(X) = E(X 2 ) − E(X) = 13 − 4 = 9,
√
σX = 9 = 3.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 183 — #199
✐
✐
Section 4.5
Variances and Moments of Discrete Random Variables
183
By Theorem 4.5,
Var(−4X + 12) = 16 Var(X) = 16 × 9 = 144,
σ−4X+12 = | − 4|σX = 4 × 3 = 12. Optional
As we know, variance measures the dispersion, or spread, of a distribution about its expected value. One way to find out which one of the two given random variables X and Y is
more dispersed, or spread, about an arbitrary point ω is to see which one is more concentrated
about ω . The following is a mathematical definition to this concept.
Definition 4.6
t > 0,
Let X and Y be two random variables and ω be a given point. If for all
P |Y − ω| ≤ t ≤ P |X − ω| ≤ t ,
then we say that X is more concentrated about ω than is Y .
A useful consequence of this definition is the following theorem, the proof of which we
leave as an exercise. This theorem should be intuitively clear.
Theorem 4.6 Suppose that X and Y are two random variables with E(X) = E(Y ) = µ.
If X is more concentrated about µ than is Y, then Var(X) ≤ Var(Y ).
MOMENTS
Let X be a random variable with expected value µ. Let c be a constant, n ≥ 0 be an integer,
and r > 0 be any real number, integral or not. The expected value of X, E(X), is also called
the first moment of X . In practice, expected values of some important functions of X have
n
n
also numerical and theoretical significance. Some of these functions
are
g(X)
= X , |X| ,
n
n
X − c , (X − c) , and (X − µ) . Provided that E g(X) < ∞, E g(X) in each of these
cases is defined as follows.
E g(X)
E(X n )
E |X|r
E(X − c)
E (X − c)n
E (X − µ)n
Definition
The nth moment of X
The r th absolute moment of X
The first moment of X about c
The nth moment of X about c
The nth central moment of X
⋆ Remark 4.2 Let X be a discrete random variable with probability mass function p(x)
and set of possible values A. Let n be a positive integer. It is important to know that if
E(X n+1 ) exists, then E(X n ) also exists. That is, the existence of higher moments implies
the existence of lower moments. In particular, this implies that if E(X 2 ) exists, then E(X)
and, hence, Var(X) exist. To prove this fact, note that, by definition, E(X n+1 ) exists if
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 184 — #200
✐
✐
184
P
Chapter 4
Distribution Functions and Discrete Random Variables
n+1
p(x)
x∈A |x|
< ∞. Let B = {x ∈ A : |x| < 1}; then B c = {x ∈ A : |x| ≥ 1}.
We have
X
x∈B
X
x∈B c
|x|n p(x) ≤
|x|n p(x) ≤
X
x∈B
p(x) ≤
X
x∈B c
X
p(x) = 1;
x∈A
|x|n+1 p(x) ≤
X
x∈A
|x|n+1 p(x) < ∞.
By these inequalities,
X
X
X
X
|x|n p(x) =
|x|n p(x) +
|x|n p(x) ≤ 1 +
|x|n+1 p(x) < ∞,
x∈A
x∈B c
x∈B
showing that E(X n ) also exists.
x∈A
EXERCISES
A
1.
Mr. Jones is about to purchase a business. There are two businesses available. The first
has a daily expected profit of $150 with standard deviation $30, and the second has a
daily expected profit of $150 with standard deviation $55. If Mr. Jones is interested in a
business with a steady income, which should he choose?
2.
The temperature of a material is measured by two devices. Using the first device, the
expected temperature is t with standard deviation 0.8; using the second device, the expected temperature is t with standard deviation 0.3. Which device measures the temperature more precisely?
3.
Find the variance of X, the random variable with probability mass function
(
|x − 3| + 1 /28
x = −3, −2, −1, 0, 1, 2, 3
p(x) =
0
otherwise.
4.
In the inventory of a multinational office supply retailing corporation, there are 7200
80-sheet smooth paper Renee pads made of acid free and ink-friendly paper in France.
The retailing corporation sells these pads only in packs of 12. Orders received by the
retailing corporation are from stationery stores, and the probability mass function of the
number of packs of pads per order is
x
3
4
5
10
p(x)
0.4
0.1
0.3
0.2
Find the expected value and the standard deviation of the number of pads left in the
inventory of the office supply corporation after the next order is shipped.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 185 — #201
✐
✐
Section 4.5
5.
6.
Variances and Moments of Discrete Random Variables
Find the variance and the standard deviation of a random variable X with distribution
function
0
x < −3
3/8
−3 ≤ x < 0
F (x) =
3/4
0≤x<6
1
x ≥ 6.
Let X be a discrete random variable with the set of possible values {x1 , x2 , . . . , xn };
X is called a discrete uniform random variable if
P (X = xi ) =
(a)
(b)
7.
185
1
,
n
1 ≤ i ≤ n.
Find
the special case, where xi = i, 1 ≤ i ≤ n. Note that
Pn E(X) and Var(X) forP
n
2
i
=
n(n
+
1)/2
and
i=1
i=1 i = n(n + 1)(2n + 1)/6.
Let X be a random integer from the set {1, 2, . . . N }. Find E(X), Var(X), and
σX .
Let X be the number of claims received, within a year, for an auto insurance policy
offered by an insurance company. Let p, the probability mass function of X , be given
by
n
25
30
35
40
45
50
55
p(n)
0.20
0.30
0.15
0.10
0.10
0.10
0.05
What percentage of the number of claims are within one standard deviation of the mean?
8.
9.
What are the expected number, the variance, and the standard deviation of the number
of spades in a poker hand? (A poker hand is a set of five cards that are randomly selected
from an ordinary deck of 52 cards.)
Suppose that X is a discrete random variable with E(X) = 1 and E X(X − 2) = 3.
Find Var(−3X + 5).
10.
In a game, Emily gives Harry three well-balanced quarters to flip. Harry will get to keep
all the ones that will land heads. He will return those landing tails. However, if all three
coins land tails, Harry must pay Emily two dollars. Find the expected value and the
variance of Harry’s net gain.
11.
Let X be a random variable defined by
P (X = −1) = P (X = 1) = 1/2.
Let Y be a random variable defined by
P (Y = −10) = P (Y = 10) = 1/2.
Which one of X and Y is more concentrated about 0 and why?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 186 — #202
✐
✐
186
Chapter 4
Distribution Functions and Discrete Random Variables
B
12.
13.
A drunken man has n keys, one of which opens the door to his office. He tries the keys
at random, one by one, and independently. Compute the mean and the variance of the
number of trials required to open the door if the wrong keys (a) are not eliminated; (b)
are eliminated.
√
For n = 1, 2, 3, . . . , let xn = (−1)n n. Let X be a discrete random variable with the
set of possible values A = {xn : n = 1, 2, 3, . . .} and probability mass function
p(xn ) = P (X = xn ) =
Show that even though
6
.
(πn)2
P∞
3
3
n=1 xn p(xn ) < ∞, E(X ) does not exist.
14.
Let X be a discrete random variable; let 0 < s < r. Show that if the r th absolute
moment of X exists, then the absolute moment of order s of X also exists.
15.
Let X and Y be two discrete random variables with the identical set of possible values
A = {a, b, c}, where a, b, and c are three different real numbers. Show that if E(X) =
E(Y ) and Var(X) =Var(Y ), then X and Y are identically distributed. That is,
P (X = t) = P (Y = t) for t = a, b, c.
16.
Let X and Y be two discrete random variables with the identical set of possible values
A = {a1 , a2 , . . . , an }, where a1 , a2 , . . . , an are n different real numbers. Show that if
E(X r ) = E(Y r ),
r = 1, 2, . . . , n − 1,
then X and Y are identically distributed. That is,
P (X = t) = P (Y = t) for t = a1 , a2 , . . . , an .
Self-Quiz on Section 4.5
Time allotted: 20 Minutes
1.
2.
Each problem is worth 5 points.
Two fair dice are tossed. Let X be the sum of the outcomes. Find Var(X) and σX .
Let X be a random variable with E(X) = 3 and E (X − 3)(4 − X) = −15. Find
Var(−3X + 8).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 187 — #203
✐
✐
Section 4.6
4.6
Standardized Random Variables
187
STANDARDIZED RANDOM VARIABLES
Let X be a random variable with mean µ and standard deviation σ . The random variable
X ∗ = (X − µ)/σ is called the standardized X . We have that
µ 1
µ
µ µ
= E(X) − = − = 0,
σ
σ
σ
σ
σ σ
1
µ
σ2
1
Var(X ∗ ) = Var X −
= 2 Var(X) = 2 = 1.
σ
σ
σ
σ
E(X ∗ ) = E
1
X−
When standardizing a random variable X, we change the origin to µ and the scale to the units
of standard deviation. The value that is obtained for X ∗ is independent of the units in which X
is measured. It is the number of standard deviation units by which X differs from E(X). For
example, let X be a random variable with mean 10 feet and standard deviation 2 feet. Suppose
that in a random observation we obtain X = 16; then X ∗ = (16 − 10)/2 = 3. This shows
that the distance of X from its mean is 3 standard deviation units regardless of the scale of
measurement. That is, if the same quantities are measured, say, in inches (12 inches = 1 foot),
then we will get the same standardized value:
X∗ =
16 × 12 − 10 × 12
= 3.
2 × 12
Standardization is particularly useful if two or more random variables with different distributions must be compared. Suppose that, for example, a student’s grade in a probability test is
72 and that her grade in a history test is 85. At first glance these grades suggest that the student
is doing much better in the history course than in the probability course. However, this might
not be true—the relative grade of the student in probability might be better than that in history.
To illustrate, suppose that the mean and standard deviation of all grades in the history test are
82 and 7, respectively, while these quantities in the probability test are 68 and 4. If we convert
the student’s grades to their standard deviation units, we find that her standard scores on the
probability and history tests are given by (72 − 68)/4 = 1 and (85 − 82)/7 = 0.43, respectively. These show that her grade in probability is 1 and in history is 0.43 standard deviation
unit higher than their respective averages. Therefore, she is doing relatively better in the probability course than in the history course. This comparison is most useful when only the means
and standard deviations of the random variables being studied are known. If the distribution
functions of these random variables are given, better comparisons might be possible.
We now prove that, for a random variable X, the standardized X, denoted by X ∗ , is independent of the units in which X is measured. To do so, let X1 be the observed value of X when
a different scale of measurement is used. Then for some α > 0, we have that X1 = αX + β,
and
(αX + β) − αE(X) + β
X1 − E(X1 )
∗
=
X1 =
σX1
σαX+β
α X − E(X)
X − E(X)
=
=
= X ∗.
ασX
σX
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 188 — #204
✐
✐
188
Chapter 4
Distribution Functions and Discrete Random Variables
EXERCISES
1.
Mr. Norton owns two appliance stores. In store 1 the number of TV sets sold by a salesperson is, on average, 13 per week with a standard deviation of five. In store 2 the number
of TV sets sold by a salesperson is, on average, seven with a standard deviation of four.
Mr. Norton has a position open for a person to sell TV sets. There are two applicants.
Mr. Norton asked one of them to work in store 1 and the other in store 2, each for one
week. The salesperson in store 1 sold 10 sets, and the salesperson in store 2 sold six sets.
Based on this information, which person should Mr. Norton hire?
2.
The mean and standard deviation in midterm tests of a probability course are 72 and 12,
respectively. These quantities for final tests are 68 and 15. What final grade is comparable to Velma’s 82 in the midterm.
CHAPTER 4 SUMMARY
◮ Random Variable
Let S be the sample space of an experiment. A real-valued function
X
:
S
→
R
is
called
a
random
variable of the experiment if, for each interval I ⊆ R ,
s : X(s) ∈ I is an event.
◮ Distribution Function
If X is a random variable, then the function F defined on
(−∞, +∞) by F (t) = P (X ≤ t) is called the distribution function of X . The most
important properties of a distribution function F are as follows: (i) F is nondecreasing;
(ii) limt→∞ F (t) = 1; (iii) limt→−∞ F (t) = 0; (iv) F is right continuous. That is, for every
t ∈ R , F (t+) = F (t). For a point a ∈ R, P (X = a) = F (a)−F (a−). Thus P (X = a) is
the difference between the right- and left-hand limits of F at a. If the function F is continuous
at a, these limits are the same and equal to F (a). Hence P (X = a) = 0. Otherwise, F has a
jump at a, and the magnitude of the jump, F (a) − F (a−), is the probability that X = a.
◮ Discrete Random Variable
Whenever the set of possible values that a random variable
X can assume is at most countable, X is called discrete. Therefore, X is discrete if either the
set of its possible values is finite or it is countably infinite. To each discrete random variable, a
real-valued function p : R → R , defined by p(x) = P (X = x), is assigned and is called the
probability mass function of X . Since the set of values of X is countable, p(x) is positive at
most for a countable set. It is zero elsewhere; that is, if possible values of X are x1 , x2 , x3 , . . . ,
then p(xi ) ≥ 0 (i = 1, 2, 3, . . .) and p(x) = 0 if x 6∈ {x1 , xP
2 , x3 , . . .}. Now, clearly, the
∞
occurrence ofPthe event {x1 , x2 , x3 , . . .} is certain. Therefore, i=1 P (X = xi ) = 1 or,
∞
equivalently, i=1 p(xi ) = 1.
◮ The distribution function F of a discrete random variable X, with the set of possible values
{x1 , x2 , x3 , . . .}, is a step function; that is, F is constant in the intervals [xn−1 , xn ) with jumps
at x1 , x2 , x3 , . . . . The magnitude of the jump at xi is p(xi ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 189 — #205
✐
✐
Chapter 4
Summary
189
◮ The Expected Value of a Discrete Random Variable
The expected value of a discrete random variable X with the set of possible values A and probability mass function p(x)
is defined by
E(X) =
X
xp(x).
x∈A
We say that E(X) exists if this sum converges
absolutely. Note that if each value x of X is
P
weighted by p(x) = P (X = x), then x∈A xp(x) is the weighted average of X .
If g is a real-valued function, then g(X) is a random variable with
X
E g(X) =
g(x)p(x).
x∈A
◮ If X is a constant random variable; that is, if P (X = c) = 1 for a constant c, then
E(X) = c.
◮ Let X be a discrete random variable; g1 , g2 , . . . , gn be real-valued functions, and let α1 ,
α2 , . . . , αn be real numbers. Then
E α1 g1 (X) + α2 g2 (X) + · · · + αn gn (X)
= α1 E g1 (X) + α2 E g2 (X) + · · · + αn E gn (X) .
This implies that E(X) is linear. That is, if α, β ∈ R, then
E(αX + β) = αE(X) + β.
◮ The Standard Deviation and Variance of a Discrete Random Variable Let X be
a discrete random variable with a set of possible values A, probability mass function p(x),
and E(X) = µ. Then σX and Var(X), called the standard deviation and the variance of X,
respectively, are defined by
q σX = E (X − µ)2 and Var(X) = E (X − µ)2 .
Variance measures the dispersion, or spread, of a distribution about its expected value. The
following formula is a better alternative for computing the variance of X .
2
Var(X) = E(X 2 ) − E(X) .
2
This relation implies that E(X) ≤ E(X 2 ).
◮ Var(X) = 0 if and only if X is a constant with probability 1.
◮ For constants a and b,
Var(aX + b) = a2 Var(X),
σaX+b = |a|σX .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 190 — #206
✐
✐
190
Chapter 4
Distribution Functions and Discrete Random Variables
◮ Standardized Random Variables Let X be a random variable with mean µ and standard deviation σ . The random variable X ∗ = (X − µ)/σ is called the standardized X . We
have that E(X ∗ ) = 0 and Var(X ∗ ) = 1. When standardizing a random variable X, we change
the origin to µ and the scale to the units of standard deviation. The value that is obtained for X ∗
is independent of the units in which X is measured. It is the number of standard deviation units
by which X differs from E(X). Standardization is particularly useful if two or more random
variables with different distributions must be compared.
REVIEW PROBLEMS
1.
An urn contains 10 chips numbered from 0 to 9. Two chips are drawn at random and
without replacement. What is the probability mass function of their total?
2.
A word is selected at random from the following poem of Persian poet and mathematician Omar Khayyām (1048–1131), translated by English poet Edward Fitzgerald (1808–
1883). Find the expected value of the length of the word.
The moving finger writes and, having writ,
Moves on; nor all your Piety nor Wit
Shall lure it back to cancel half a line,
Nor all your tears wash out a word of it.
3.
A statistical survey shows that only 2% of secretaries know how to use the highly sophisticated word processor language TEX. If a certain mathematics department prefers
to hire a secretary who knows TEX, what is the least number of applicants that should be
interviewed so as to have at least a 50% chance of finding one such secretary?
4.
An electronic system fails if both of its components fail. Let X be the time (in hours)
until the system fails. Experience has shown that
t −t/200
e
,
P (X > t) = 1 +
200
t ≥ 0.
What is the probability that the system lasts at least 200 but not more than 300 hours?
5.
A professor has prepared 30 exams of which 8 are difficult, 12 are reasonable, and 10
are easy. The exams are mixed up, and the professor selects four of them at random to
give to four sections of the course he is teaching. How many sections would be expected
to get a difficult test?
6.
The annual amount of rainfall (in centimeters) in a certain area is a random variable with
the distribution function
(
0
x<5
F (x) =
1 − (5/x2 )
x ≥ 5.
What is the probability that next year it will rain (a) at least 6 centimeters; (b) at most 9
centimeters; (c) at least 2 and at most 7 centimeters?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 191 — #207
✐
✐
Chapter 4
7.
Review Problems
191
Let X be the amount (in fluid ounces) of soft drink in a randomly chosen bottle from
company A, and Y be the amount of soft drink in a randomly chosen bottle from company B . A study has shown that the distributions of X and Y are as follows:
x
15.85
15.9
16
16.1
16.2
P (X = x)
0.15
0.21
0.35
0.15
0.14
P (Y = x)
0.14
0.05
0.64
0.08
0.09
Find E(X), E(Y ), Var(X), and Var(Y ) and interpret them.
8.
The fasting blood-glucose levels of 30 children are as follows.
58 62 80 58 64 76 80 80 80 58
62 64 76 76 58 64 62 80 58 58
80 64 58 62 76 62 64 80 62 76
Let X be the fasting blood-glucose level of a child chosen randomly from this group.
Find the distribution function of X .
9.
Experience shows that X, the number of customers entering a post office, during any
period of length t, is a random variable the probability mass function of which is of the
form
p(i) = k
(2t)i
,
i!
i = 0, 1, 2, . . . .
(a)
Determine the value of k .
(b)
Compute P (X < 4) and P (X > 1).
10.
From the set of families with three children a family is selected at random, and the
number of its boys is denoted by the random variable X . Find the probability mass
function and the distribution functions of X . Assume that in a three-child family all
gender distributions are equally probable.
11.
If the motor of a certain commercial grade dishwasher fails within the first year of purchase, the insurance company pays $2500 for its replacement or repair. Thereafter, each
year, the company’s payment for motor failure will be $500 less than the previous year
until the sixth year. If the motor fails during or after the sixth year, the company offers
no payments. Suppose that if such a motor has not failed by the first day of a year, the
probability that it fails during that year is 0.06. Find the expected value of the insurance
company’s payment for each such dishwasher motor.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 192 — #208
✐
✐
192
Chapter 4
Distribution Functions and Discrete Random Variables
Self-Test on Chapter 4
Time allotted: 120 Minutes
Each problem is worth 10 points.
1.
From a leap year calendar, a month is selected at random. Let X be the number of days
in that month. Find E(X) and σX . If the month is selected from a non-leap year, on
average, how many days does that month have?
2.
A large company has w women and m men in retirement age. If a random set of n,
n ≤ w + m, of these employees decides to retire next year, what is the probability mass
function of the number of women who will retire next year?
3.
In the front yard of a doctor’s clinic, there are exactly six parking spaces next to each
other in a row. Suppose that at a time when all six parking spaces are occupied, two of
the cars leave at random and their spots remain empty. Let X be the number of cars
between the empty spots. Find p, the probability mass function, and F, the distribution
function, of X .
4.
Let X be the number of bagels purchased by a random customer from the Longmeadow
Bagel Shop. Let F be the distribution function of X, and suppose that F (0) = 4/33,
F (1) = 16/33, F (2) = 26/33, and F (3) = 30/33. If P (X = 4) = P (X > 4), for
i = 0, 1, 2, 3, 4, find P (X = i).
5.
To help defray hospitalization expenses due to accidents, the insurance company, Joseph
Accident and Health, offers a hospital indemnity plan in which an injured customer is
paid a lump sum daily amount of a dollars per day up to three days of hospitalization and
a/2 dollars per day thereafter up to a maximum of 7 additional days. The probability
mass function of the number of days an injured customer is hospitalized is given by
1
(11 − x)
x = 0, 1, 2, . . . , 10
p(x) = 66
0
otherwise.
Find the expected payment of the Joseph Accident and Health to an injured customer
selected randomly?
6.
Let x be a nonnegative real number. By [x], we mean the greatest integer less than
or equal to x. For example, [1.2] = 1, [5.8] = 5, and [0.8] = 0. In a box, there
are n identical balls numbered from 1 to n. Let X be the number on a ball that is
drawn at random from the box. Find F, the distribution function of X by expressing
F (t) = P (X ≤ t) in terms of [t] for 1 ≤ t < n.
7.
The season finale of a Belavian reality television series, “Belavia’s Got Talent,” was held
in an oval auditorium which has 25 rows of seats. The first row has 10 seats, and each
succeeding row has 2 more seats than the previous row. All the seats in the auditorium
were occupied, and before doing some magic tricks, a magician, who was a finalist,
picked someone from the audience at random to help him with the performance. Let X
be the row number of the seat of the person picked from the audience by the magician.
Find the probability mass function of X .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 193 — #209
✐
✐
Chapter 4
8.
Self-Test Problems
193
At a university, professors are allowed to check out as many books as they wish from
the library. Let X be the number of books checked out by a random professor visiting
the library. Suppose that
P (X ≤ 1) = 7/13,
P (X = 1 or X = 2) = 6/13,
P (X > 2) = 3/13,
P (X = 3) = P (X > 3).
For i = 0, 1 , 2, 3, calculate P (X = i).
9.
At a department store, summer polo shirts are sold from April 1st until the end of
September. Suppose that for each polo shirt the department store sells the net profit
is $15.00, and for each shirt unsold by the end of September the net loss is $10.00. Also
suppose that, at the department store, for each i, 100 ≤ i ≤ 150, each year, the demand for polo shirts is i with probability 1/51. Determine the expected net profit of the
department store from selling polo shirts if it stocks 130 shirts each year.
10.
Under what conditions on α, β, γ, and θ is the following a distribution function?
1−α
t<0
β
0≤t<1
F (t) =
β−γ
1≤t<2
θ
t ≥ 2.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 194 — #210
✐
✐
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 195 — #211
✐
✐
Chapter 5
Special Discrete
Distributions
In this chapter we study some examples of discrete random variables. These random variables appear frequently in theory and applications of probability, statistics, and branches of
science and engineering.
5.1
BERNOULLI AND BINOMIAL RANDOM VARIABLES
Bernoulli trials, named after the Swiss mathematician James Bernoulli, are perhaps the simplest type of random variable. They have only two possible outcomes. One outcome is usually
called a success, denoted by s. The other outcome is called a failure, denoted by f . The experiment of flipping a coin is a Bernoulli trial. Its only outcomes are “heads” and “tails.” If we are
interested in heads, we may call it a success; tails is then a failure. The experiment of tossing
a die is a Bernoulli trial if, for example, we are interested in knowing whether the outcome is
odd or even. An even outcome may be called a success, and hence an odd outcome a failure,
or vice versa. If a fuse is inspected, it is either “defective” or it is “good.” So the experiment
of inspecting fuses is a Bernoulli trial. A good fuse may be called a success, a defective fuse a
failure.
The sample space of a Bernoulli trial contains two points, s and f . The random variable
defined by X(s) = 1 and X(f ) = 0 is called a Bernoulli random variable. Therefore, a
Bernoulli random variable takes on the value 1 when the outcome of the Bernoulli trial is a
success and 0 when it is a failure. If p is the probability of a success, then 1 − p (sometimes
denoted q ) is the probability of a failure. Hence the probability mass function of X is
1 − p ≡ q if x = 0
p(x) = p
(5.1)
if x = 1
0
otherwise.
Note that the same symbol p is used for the probability mass function and the Bernoulli parameter. This duplication should not be confusing since the p’s used for the probability mass
function often appear in the form p(x).
An accurate mathematical definition for Bernoulli random variables is as follows.
195
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 196 — #212
✐
✐
196
Chapter 5
Special Discrete Distributions
Definition 5.1
A random variable is called Bernoulli with parameter p if its probability
mass function is given by Equation (5.1).
From (5.1) it follows that the expected value of a Bernoulli random variable X, with
parameter p, is p, because
E(X) = 0 · P (X = 0) + 1 · P (X = 1) = P (X = 1) = p.
Also, since
E(X 2 ) = 0 · P (X = 0) + 1 · P (X = 1) = p,
we have
2
Var(X) = E(X 2 ) − E(X) = p − p2 = p(1 − p).
We will now summarize what we have shown.
For a Bernoulli random variable X with parameter p, 0 < p < 1,
q
E(X) = p, Var(X) = p(1 − p), σX = p(1 − p).
Example 5.1 If in a throw of a fair die the event of obtaining 4 or 6 is called a success, and
the event of obtaining 1, 2, 3, or 5 is called a failure, then
(
1 if 4 or 6 is obtained
X=
0 otherwise
is a Bernoulli random variable with the parameter p = 1/3. Therefore, its probability mass
function is
2/3 if x = 0
p(x) = 1/3 if x = 1
0
elsewhere.
The expected value of X is given by E(X) = p = 1/3, and its variance by Var(X) =
1/3(1 − 1/3) = 2/9. Let X1 , X2 , X3 , . . . be a sequence of Bernoulli random variables. If, for all ji ∈ {0, 1},
the sequence of events {X1 = j1 }, {X2 = j2 }, {X3 = j3 }, . . . are independent, we say that
{X1 , X2 , X3 , . . .} and the corresponding Bernoulli trials are independent.
Although Bernoulli trials are simple, if they are repeated independently, they may pose
interesting and even sometimes complicated questions. Consider an experiment in which n
Bernoulli trials are performed independently. The sample space of such an experiment, S, is
the set of different sequences of length n with x (x = 0, 1, . . . , n) successes (s’s) and (n − x)
failures (f ’s). For example, if, in an experiment, three Bernoulli trials are performed independently, then the sample space is
{f f f, sf f, f sf, f f s, f ss, sf s, ssf, sss}.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 197 — #213
✐
✐
Section 5.1
Bernoulli and Binomial Random Variables
197
If n Bernoulli trials all with probability of success p are performed independently, then X, the
number of successes, is one of the most important random variables. It is called a binomial
with parameters n and p. The set of possible values of X is {0, 1, 2, . . . , n}, it is defined
on the set S described previously, and its probability mass function is given by the following
theorem.
Theorem 5.1 Let X be a binomial random variable with parameters n and p. Then p(x),
the probability mass function of X, is
!
n x
p (1 − p)n−x if x = 0, 1, 2, . . . , n
x
p(x) = P (X = x) =
(5.2)
0
elsewhere.
Proof: Observe that the number of ways that, in n Bernoulli trials, x (x = 0, 1, 2, . . . , n)
successes can occur is equal to the number of different sequences of length
! n with x successes
n
(s’s) and (n − x) failures (f ’s). But the number of such sequences is
because the number
x
of distinguishable permutations of n objects
of two different types, where x are alike and
!
n
n!
=
(see Theorem 2.4). Since by the independence of
n − x are alike is
x! (n − x)!
x
the !
trials the probability of each of these sequences is px (1 − p)n−x , we have P (X = x) =
n x
p (1 − p)n−x . Hence (5.2) follows. x
Definition 5.2 The function p(x) given by Equation (5.2) is called the binomial probability
mass function with parameters (n, p).
The reason for this name is that the binomial expansion theorem (Theorem 2.5) guarantees
that p is a probability mass function:
!
n
n
X
X
n
n x
p(x) =
p (1 − p)n−x = p + (1 − p) = 1n = 1.
x
x=0
x=0
Example 5.2 A restaurant serves 8 entrées of fish, 12 of beef, and 10 of poultry. If customers
select from these entrées randomly, what is the probability that two of the next four customers
order fish entrées?
Solution: Let X denote the number of fish entrées (successes) ordered by the next four customers. Then X is binomial with the parameters (4, 8/30 = 4/15). Thus
!
4 4 2 11 2
P (X = 2) =
= 0.23. 2 15
15
Example 5.3 In a county hospital 10 babies, of whom six were boys, were born last Thursday. What is the probability that the first six births were all boys? Assume that the events that a
child born is a girl or is a boy are equiprobable.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 198 — #214
✐
✐
198
Chapter 5
Special Discrete Distributions
Solution: Let A be the event that the first six births were all boys and the last four all girls.
Let X be the number of boys; then X is binomial with parameters 10 and 1/2. The desired
probability is
P (A and X = 6)
P (A)
=
P (X = 6)
P (X = 6)
1 10
1
! 2
! ≈ 0.0048.
=
=
10 1 6 1 4
10
P (A | X = 6) =
6
2
2
6
Remark: We can also argue as follows: bbbbbbgggg, the!event that the first six were all boys,
10
and the remainder girls, is one sequence out of the
possible sequences of the sexes of
6
the births.
! Since each of these sequences has the same probability of occurrence, the answer is
. 10
1
. 6
Example 5.4 In a small town, out of 12 accidents that occurred in June 1986, four happened
on Friday the 13th. Is this a good reason for a superstitious person to argue that Friday the 13th
is inauspicious?
Solution: Suppose the probability that each accident occurs on Friday the 13th is 1/30, just as
on any other day. Then the probability of at least four accidents on Friday the 13th is
!
3
X
12 1 i 29 12−i
≈ 0.000, 493.
1−
30
30
i
i=0
Since the probability of four or more of these accidents occurring on Friday the 13th is
very small, this is a good reason for superstitious persons to argue that Friday the 13th is
inauspicious. Example 5.5 Suppose that jury members decide independently and that each with probability p (0 < p < 1) makes the correct decision. If the decision of the majority is final, which is
preferable: a three-person jury or a single juror?
Solution: Let X denote the number of persons who decide correctly among a three-person
jury. Then X is a binomial random variable with parameters (3, p). Hence the probability that
a three-person jury decides correctly is
!
!
3 2
3 3
P (X ≥ 2) = P (X = 2) + P (X = 3) =
p (1 − p) +
p (1 − p)0
2
3
= 3p2 (1 − p) + p3 = 3p2 − 2p3 .
Since the probability is p that a single juror decides correctly, a three-person jury is preferable
to a single juror if and only if
3p2 − 2p3 > p.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 199 — #215
✐
✐
Section 5.1
Bernoulli and Binomial Random Variables
199
This is equivalent to 3p − 2p2 > 1, so −2p2 + 3p − 1 > 0. But
−2p2 + 3p − 1 = 2(1 − p)(p − 1/2).
Since 1 − p > 0, 2(1 − p)(p − 1/2) > 0 if and only if p > 1/2. Hence a three-person jury
is preferable if p > 1/2. If p < 1/2, the decision of a single juror is preferable. For p = 1/2
there is no difference. Example 5.6 Machuca’s favorite bridge hands are those with at least two aces. Determine
the number of times he should play in order to have a chance of 90% or more to get at least
two favorite hands? Recall that a bridge hand consists of 13 randomly selected cards from an
ordinary deck of 52 cards.
Solution: We call a bridge hand a “success” if it is Machuca’s favorite. Let X be the number
of successes in n independent bridge hands. Clearly, X is a binomial random variable with
parameters (n, p), where
!
!
!
48
4
48
13
1
12
!−
! ≈ 0.257.
p=1−
52
52
13
13
We want to determine n so that P (X ≥ 2) ≥ 0.90. Since
P (X ≥ 2) = 1 − P (X = 0) − P (X = 1) ≥ 0.90,
we must have
P (X = 0) + P (X = 1) ≤ 0.10.
Therefore, n is the smallest integer satisfying
!
!
n
n
(0.257)0 (1 − 0.257)n +
(0.257)1 (1 − 0.257)n−1 ≤ 0.10.
0
1
This inequality reduces to
(0.743)n−1 (0.743 + 0.257n) ≤
1
.
10
By trial and error, we find that for n = 13, the left side is approximately 0.116 while for
n = 14, it is 0.091. Therefore n = 14 is the answer.
Remark: Note that n the number of times Machuca should play in order to get at least two
favorite hands with certainty (chance of 100%) satisfies the relation
(0.743)n−1 (0.743 + 0.257n) = 0.
This gives (0.743)n−1 = 0 or n = ∞. However, similar calculations show that, the number
of times he should play to get at least two favorite hands with probability of at least 0.99 is
only 23. ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 200 — #216
✐
✐
200
Chapter 5
Special Discrete Distributions
Example 5.7 A Canadian pharmaceutical company has developed three types of Ebola virus
vaccines. Suppose that vaccines I, II, and III are effective 92%, 88%, and 96% of the time,
respectively. Assume that a person reacts to the Ebola vaccine independently of other people.
The company sends to Guinea hundreds of qualified vaccine carrier containers containing 100
ml vaccine bottles each filled with one of the three types of the Ebola vaccines. Suppose that
the bottles of 40% of the containers are filled with only the type I vaccine, and the bottles of
35% and 25% of the containers are filled only with type II and type III vaccines, respectively.
In Guinea, a physician assistant takes one of the containers at random to a small village and
using its contents, vaccinates its entire 200-people population. It is certain that sooner or later
every member of the village will be exposed to the Ebola virus. What is the probability that at
most 6 of the villagers will become sick with Ebola?
Solution: Let X be the number of the villagers who will become sick with Ebola. Let A1 ,
A2 , and A3 , respectively, be the events that the physician assistant takes the container that has
type I, type II, and type III vaccines. The desired probability is
P (X ≤ 6) = P (X ≤ 6 | A1 )P (A1 ) + P (X ≤ 6 | A2 )P (A2 ) + P (X ≤ 6 | A3 )P (A3 )
!
6
X
200
=
(0.08)i (1 − 0.08)200−i · (0.40)+
i
i=0
!
6
X
200
(0.12)i (1 − 0.12)200−i · (0.35)+
i
i=0
!
6
X
200
(0.04)i (1 − 0.04)200−i · (0.25) ≈ 0.0783. i
i=0
Example 5.8 Let p be the probability that a randomly chosen person is pro-life, and let X
be the number of persons pro-life in a random sample of size n. Suppose that, in a particular
random sample of n persons, k are pro-life. Show that P (X = k) is maximum for p̂ = k/n.
That is, p̂ is the value of p that makes the outcome X = k most probable.
Solution: By definition of X,
P (X = k) =
!
n k
p (1 − p)n−k .
k
This gives
d
P (X = k) =
dp
=
!
n k−1
kp (1 − p)n−k − (n − k)pk (1 − p)n−k−1
k
!
n k−1
p (1 − p)n−k−1 k(1 − p) − (n − k)p .
k
d2
d
P (X = k) = 0, we obtain p = k/n. Now since 2 P (X = k) < 0, p̂ = k/n
dp
dp
is the maximum of P (X = k), and hence it is an estimate of p that makes the outcome x = k
most probable. Letting
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 201 — #217
✐
✐
Section 5.1
Bernoulli and Binomial Random Variables
201
Let X be a binomial random variable with parameters (n, p), 0 < p < 1, and probability
mass function p(x). We will now find the value of X at which p(x) is maximum. For any real
number t, let [t] denote
the largest
integer less than or equal to t. We will prove that p(x) is
maximum at x = (n + 1)p . To do so, we note that
n!
px (1 − p)n−x
(n − x)! x!
p(x)
=
p(x − 1)
=
n!
px−1 (1 − p)n−x+1
(n − x + 1)! (x − 1)!
(n − x + 1)p
x(1 − p)
(5.3)
(n + 1)p − x + x(1 − p)
(n + 1)p − xp + x − x
=
=
x(1 − p)
x(1 − p)
=
(n + 1)p − x
+ 1.
x(1 − p)
This equality shows that p(x) > p(x − 1) if and only if (n + 1)p − x > 0, or, equivalently,
if and only if x < (n + 1)p. Hence as x changes from 0 to (n + 1)p , p(x) increases. As x
changes from (n + 1)p to n, p(x) decreases. The maximum value of p(x) the peak of the
graphical representation of p(x) occurs at (n + 1)p .
What we have just shown is illustrated by the graphical representations of p(x) in
Figure 5.1 for binomial random variables with parameters (5, 1/2), (10, 1/2), and (20, 1/2).
p(x)
p(x)
0.3
0.3
n = 10
n=5
0.2
0.2
0.1
0.1
0
1
2
3 4
x
5
0
1 2
3
4
5 6 7
8
9 10
x
p(x)
0.3
0.2
n = 20
0.1
0
Figure 5.1
2
4
6
8 10 12 14 16 18 20
x
Examples of binomial probability mass functions.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 202 — #218
✐
✐
202
Chapter 5
Special Discrete Distributions
Expectations and Variances of Binomial Random Variables
Let X be a binomial random variable with parameters (n, p). Intuitively, we expect that the
expected value of X will be np. For example, if we toss a fair coin 100 times, we expect that
the average number of heads will be 50, which is 100 · 1/2 = np. Also, if we choose 10 fuses
from a set with 30% defective fuses, we expect the average number of defective fuses to be
np = 10(0.30) = 3. The formula E(X) = np can be verified directly from the definition of
mathematical expectation as follows:
!
n
n
X
X
n x
n!
px (1 − p)n−x
E(X) =
x
p (1 − p)n−x =
x
x!
(n
−
x)!
x
x=0
x=1
=
n
X
n!
px (1 − p)n−x
(x
−
1)!
(n
−
x)!
x=1
n
X
(n − 1)!
px−1 (1 − p)n−x
(x
−
1)!
(n
−
x)!
x=1
!
n
X
n − 1 x−1
= np
p (1 − p)n−x .
x
−
1
x=1
= np
Letting i = x − 1 (by reindexing this sum), we obtain
!
n−1
X n−1
n−1
E(X) = np
pi (1 − p)(n−1)−i = np p + (1 − p)
= np,
i
i=0
where the next-to-last equality follows from binomial expansion (Theorem 2.5). To calculate
the variance of X, from a procedure similar to the one we used to compute E(X), we obtain
(see Exercise 30)
!
n
X
2
2 n
E(X ) =
x
px (1 − p)n−x = n2 p2 − np2 + np.
x
x=1
Therefore,
2
Var(X) = E(X 2 ) − E(X) = −np2 + np = np(1 − p).
We have established that
If X is a binomial random variable with parameters n and p, then
q
E(X) = np, Var(X) = np(1 − p), σX = np(1 − p).
Example 5.9 (One-dimensional Random Walk) We will now examine an elementary
example of a random walk. In Chapter 12, we will revisit this concept and its applications.
Suppose that a particle is at 0 on the integer number line and suppose that at step 1, the particle
will move to 1 with probability p, 0 < p < 1, and will move to −1 with probability 1 − p.
Furthermore, if at step n, the particle is at i, then independently of the previous moves, it will
move 1 unit to the right to i + 1 with probability p and will move 1 unit to the left to i − 1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 203 — #219
✐
✐
Section 5.1
Bernoulli and Binomial Random Variables
203
with probability 1 − p. Let X be the position of the particle after n moves. Find the probability
mass function of X .
Solution: Let Y be the number of times the particle moves one unit to the right. Then Y
is a binomial random variable with parameters n and p. If Y = k, after the nth step, the
particle has moved k units to the right and n − k units to the left, k = 0, 1, . . . , n. So after n
steps, it will be at k − (n − k) = 2k − n. This shows that the set of possible values of X is
E = {−n, 2 − n, 4 − n, . . . 2n − n} and, for i ∈ E,
n + i
P (X = i) = P (2Y − n = i) = P Y =
2
!
n
=
p(n+i)/2 (1 − p)n−[(n+i)/2]
(n + i)/2
!
n
=
p(n+i)/2 (1 − p)(n−i)/2 . (n + i)/2
Example 5.10 A town of 100,000 inhabitants is exposed to a contagious disease. If the
probability that a person becomes infected is 0.04, what is the expected number of people who
become infected?
Solution: Assuming that people become infected independently of each other, the number of
those who will become infected is a binomial random variable with parameters 100,000 and
0.04. Thus the expected number of such people is 100, 000 × 0.04 = 4000. Note that here
“getting infected” is implicitly defined to be a success for mathematical simplicity! Therefore,
in the context of Bernoulli trials, what is called a success may be a failure in real life. Example 5.11 Two proofreaders, Ruby and Myra, read a book independently and found
r and m misprints, respectively. Suppose that the probability that a misprint is noticed by
Ruby is p and the probability that it is noticed by Myra is q, where these two probabilities
are independent. If the number of
misprints noticed by both Ruby and Myra is b, estimate the
number of unnoticed misprints. This problem was posed and solved by George Pólya (1888–
1985) in the January 1976 issue of the American Mathematical Monthly.
Solution: Let M be the total number of misprints in the book. The expected number of misprints that may be noticed by Ruby, Myra, and both of them are M p, M q, and M pq, respectively. Assuming that the number of misprints found are approximately equal to the expected
number, we have M p ≈ r, M q ≈ m, and M pq ≈ b. Therefore,
M=
rm
(M p)(M q)
≈
.
M pq
b
The number of unnoticed misprints is therefore estimated by
M − (r + m − b) ≈
(r − b)(m − b)
rm
− (r + m − b) =
. b
b
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 204 — #220
✐
✐
204
Chapter 5
Special Discrete Distributions
EXERCISES
A
1.
From an ordinary deck of 52 cards, cards are drawn at random and with replacement.
What is the probability that, of the first eight cards drawn, four are spades?
2.
The probability that a randomly selected person is female is 1/2. What is the expected
number of the girls in the first-grade classes of an elementary school that has 64 first
graders? What is the expected number of females in a family with six children?
3.
A graduate class consists of six students. What is the probability that exactly three of
them are born either in April or in October?
4.
Let X be a Bernoulli random variable with parameter p, 0 < p < 1. Find the probability
mass function of Y = 1 − X, E(Y ), and Var(Y ).
5.
In a state where license plates contain six digits, what is the probability that the license
number of a randomly selected car has two 9’s? Assume that each digit of the license
number is randomly selected from {0, 1, . . . , 9}.
6.
A box contains 30 balls numbered 1 through 30. Suppose that five balls are drawn at
random, one at a time, with replacement. What is the probability that the numbers of
two of them are prime?
7.
A manufacturer of nails claims that only 3% of its nails are defective. A random sample
of 24 nails is selected, and it is found that two of them are defective. Is it fair to reject
the manufacturer’s claim based on this observation?
8.
Let X be a discrete random variable with probability mass function p given by
x
p(x)
−1
2/9
0
4/9
1
3/9
other values
0
Find the probability mass functions of the random variables Y = |X| and Z = X 2 .
9.
Only 60% of certain kinds of seeds germinate when planted under normal conditions.
Suppose that four such seeds are planted, and X denotes the number of those that will
germinate. Find the probability mass functions of X and Y = 2X + 1.
10.
Suppose that the Internal Revenue Service will audit 20% of income tax returns reporting
an annual gross income of over $80,000. What is the probability that of 15 such returns,
at most four will be audited?
11.
If two fair dice are rolled 10 times, what is the probability of at least one 6 (on either
die) in exactly five of these 10 rolls?
12.
From the interval (0, 1), five points are selected at random and independently. What is
the probability that (a) at least two of them are less than 1/3; (b) the first decimal point
of exactly two of them is 3?
13.
Let X be a binomial random variable with parameters (n, p) and probability mass function p(x). Prove that if (n + 1)p is an integer, then p(x) is maximum at two different
points. Find both of these points.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 205 — #221
✐
✐
Section 5.1
Bernoulli and Binomial Random Variables
205
14.
On average, how many times should Ernie play poker in order to be dealt a straight flush
(royal flush included)? (See Exercise 35 of Section 2.4 for definitions of a royal and a
straight flush.)
15.
Suppose that each day the price of a stock moves up 1/8th of a point with probability 1/3
and moves down 1/8th of a point with probability 2/3. If the price fluctuations from one
day to another are independent, what is the probability that after six days the stock has
its original price?
16.
Even though there are some differences between the taste of Max Cola and the taste of
Golden Cola, most people cannot tell the difference. Pedro claims that in the absence
of brand information, he can say whether a cup of cola is Max or Golden by tasting the
drink. A statistician decided to test his claim. So she asked him to taste 6 cups of cola,
3 of which were Max and 3 of which were Golden, all in unmarked clear glasses, in a
random order. If Pedro has no ability to distinguish Max Cola from Golden Cola, what
is the probability that he could successfully identify the brand of four or more of the
samples based on pure guesses? That is, by guessing the brand of each sample randomly
and independently from other samples.
17.
From the set {x : 0 ≤ x ≤ 1}, 100 independent numbers are selected at random and
rounded to three decimal places. What is the probability that at least one of them is
0.345?
18.
Dr. Willis is teaching two sections of Calculus I, each with 25 students. Suppose that
each student will earn a passing grade, independently of other students, with probability
of 0.82. What is the probability that in one of Dr. Willis’s classes more than 22 students
will earn a passing grade and in the other one less than 23?
19.
A certain basketball player makes a foul shot with probability 0.45. Determine for what
value of k the probability of k baskets in 10 shots is maximum, and find this maximum
probability.
20.
What are the expected value and variance of the number of full house hands in n poker
hands? A poker hand consists of five randomly selected cards from an ordinary deck of
52 cards. It is a full house if three cards are of one denomination and two cards are of
another denomination: for example, three queens and two 4’s.
21.
Let X be the number of sixes obtained when a balanced die is tossed five times. Find
the probability mass function of Y = (X − 3)2 .
22.
A certain rare blood type can be found in only 0.05% of people. If the population of a
randomly selected group is 3000, what is the probability that at least two persons in the
group have this rare blood type?
23.
Edward’s experience shows that 7% of the parcels he mails will not reach their destination. He has bought two books for $20 each and wants to mail them to his brother. If he
sends them in one parcel, the postage is $5.20, but if he sends them in separate parcels,
the postage is $3.30 for each book. To minimize the expected value of his expenses (loss
+ postage), which way is preferable to send the books, as two separate parcels or as a
single parcel?
24.
Vincent is a patient with the life threatening blood cancer leukemia, and he is in need of
a bone marrow transplant. He asks n people whether or not they are willing to donate
bone marrow to him if they are a close bone marrow match for him. Suppose that each
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 206 — #222
✐
✐
206
Chapter 5
Special Discrete Distributions
person’s response, independently of others, is positive with probability p1 and negative
with probability 1 − p1 . Furthermore, suppose that, independently of the others, the
probability is p2 that a person tested is a close match. Find the probability that Vincent
finds at least one close bone marrow match among these n people.
Hint: For 0 ≤ i ≤ n, let Ai be the event that i of the n individuals respond positively
to Vincent’s request. Let X be the number of close bone marrow matches among those
who will be tested. To find P (X ≥ 1) = 1 − P (X = 0), calculate P (X = 0) by using
the Law of Total Probability, Theorem 3.4, applied to the partition {A0 , A1 , . . . An }.
25.
A woman and her husband want to have a 95% chance for at least one boy and at least
one girl. What is the minimum number of children that they should plan to have? Assume
that the events that a child is a girl and a boy are equiprobable and independent of the
gender of other children born in the family.
26.
A computer network consists of several stations connected by various media (usually
cables). There are certain instances when no message is being transmitted. At such “suitable instances,” each station will send a message with probability p independently of the
other stations. However, if two or more stations send messages, a collision will corrupt
the messages and they will be discarded. These messages will be retransmitted until they
reach their destination. Suppose that the network consists of N stations.
27.
(a)
What is the probability that at a “suitable instance” a message is initiated by one
of the stations and will go through without a collision?
(b)
Show that, to maximize the probability of a message going through with no collisions, exactly one message, on average, should be initiated at each “suitable
instance.”
(c)
Find the limit of the maximum probability obtained in (b) as the number of stations of the network grows to ∞.
Consider the following problem posed by Michael Khoury, U.S. Math Olympiad Team
Member, in “The Problem Solving Competition,” Oklahoma Publishing Company and
the American Society for the Communication of Mathematics, February 1999.
Bob is teaching a class with n students. There are n desks in the classroom numbered from 1 to n. Bob has prepared a seating chart, but the
students have already seated themselves randomly. Bob calls off the
name of the person who belongs in seat 1. This person vacates the seat
he or she is currently occupying and takes his or her rightful seat. If this
displaces a person already in the seat, that person stands at the front of
the room until he or she is assigned a seat. Bob does this for each seat
in turn.
Let X be the number of students standing at the front of the room after k, 1 ≤ k < n,
names have been called. Call a person “success” if he or she is standing. Is X a binomial
random variable? Why or why not?
B
28.
In a community, a persons are pro-choice, b (b < a) are pro-life, and n (n > a − b)
are undecided. Suppose that there will be a vote to determine the will of the majority
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 207 — #223
✐
✐
Section 5.1
Bernoulli and Binomial Random Variables
207
with regard to legalizing abortion. If by then all of the undecided persons make up their
minds, what is the probability that those who are pro-life will win? Assume that it is
equally likely that an undecided person eventually votes pro-choice or pro-life.
29.
A game often played in carnivals and gambling houses is called chuck-a-luck, where a
player bets on any number 1 through 6. Then three fair dice are tossed. If one, two, or
all three land the same number as the player’s, then he or she receives one, two, or three
times the original stake plus his or her original bet, respectively. Otherwise, the player
loses his or her stake. Let X be the net gain of the player per unit of stake. First find the
probability mass function of X; then determine the expected amount that the player will
lose per unit of stake.
30.
Let X be a binomial random variable with the parameters (n, p). Prove that
!
n
X
n
E(X 2 ) =
x2
px (1 − p)n−x = n2 p2 − np2 + np.
x
x=1
31.
Suppose that an aircraft engine will fail in flight with probability 1 − p independently of
the plane’s other engines. Also suppose that a plane can complete the journey successfully if at least half of its engines do not fail.
(a)
Is it true that a four-engine plane is always preferable to a two-engine plane?
Explain.
(b)
Is it true that a five-engine plane is always preferable to a three-engine plane?
Explain.
32.
The simplest error detection scheme used in data communication is parity checking.
Usually messages sent consist of characters, each character consisting of a number of
bits (a bit is the smallest unit of information and is either 1 or 0). In parity checking,
a 1 or 0 is appended to the end of each character at the transmitter to make the total
number of 1’s even. The receiver checks the number of 1’s in every character received,
and if the result is odd it signals an error. Suppose that each bit is received correctly
with probability 0.999, independently of other bits. What is the probability that a 7-bit
character is received in error, but the error is not detected by the parity check?
33.
In Exercise 32, suppose that a message consisting of six characters is transmitted. If each
character consists of seven bits, what is the probability that the message is erroneously
received, but none of the errors is detected by the parity check?
34.
How many games of poker occur until a preassigned player is dealt at least one straight
flush with probability of at least 3/4? (See Exercise 35 of Section 2.4 for a definition of
a straight flush.)
35.
The post office of a certain small town has only one clerk to wait on customers. The
probability that a customer will be served in any given minute is 0.6, regardless of the
time that the customer has already taken. The probability of a new customer arriving is
0.45, regardless of the number of customers already in line. The chances of two customers arriving during the same minute are negligible. Similarly, the chances of two
customers being served in the same minute are negligible. Suppose that we start with
exactly two customers: one at the postal window and one waiting on line. After 4 minutes, what is the probability that there will be exactly four customers: one at the window
and three waiting on line?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 208 — #224
✐
✐
208
Chapter 5
36.
(a)
Special Discrete Distributions
What is the probability of an even number of successes in n independent
Bernoulli trials?
Hint: Let rn be the probability of an even number of successes in n Bernoulli
trials. By conditioning on the first trial and using the law of total probability (Theorem 3.3), show that for n ≥ 1,
rn = p(1 − rn−1 ) + (1 − p)rn−1 .
Then prove that rn =
(b)
Prove that
[n/2]
X
k=0
37.
38.
1
1 + (1 − 2p)n .
2
!
n
1
p2k (1 − p)n−2k = 1 + (1 − 2p)n .
2k
2
An urn contains n balls whose colors, red or blue, are equally probable. For example,
the probability that all of the balls are red is (1/2)n . If in drawing k balls from the urn,
successively with replacement and randomly, no red balls appear, what is the probability
that the urn contains no red balls?
Hint: Use Bayes’ theorem.
While Rose always tells the truth, four of her friends, Albert, Brenda, Charles, and
Donna, tell the truth randomly only in one out of three instances, independent of each
other. Albert makes a statement. Brenda tells Charles that Albert’s statement is the truth.
Charles tells Donna that Brenda is right, and Donna says to Rose that Charles is telling
the truth and Rose agrees. What is the probability that Albert’s statement is the truth?
Self-Quiz on Section 5.1
Time allotted: 10 Minutes
Each problem is worth 5 points.
1.
What is the probability that at least two of the six members of a family are not born in
the fall? Assume that all seasons have the same probability of containing the birthday of
a person selected randomly.
2.
Suppose that a prescription drug causes 7 side effects each with probability of 0.05 independently of the others. What is the probability that a patient taking this drug experiences
at most 2 side effects.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 209 — #225
✐
✐
Section 5.2
5.2
Poisson Random Variables
209
POISSON RANDOM VARIABLES
Poisson as an Approximation to Binomial
From Section 5.1, it should have become clear that in many natural phenomena and real-world
problems, we are dealing with binomial random variables.
! Therefore, for possible values of
n x
parameters n, x, and p, we need to calculate p(x) =
p (1 − p)n−x . However, in many
x
cases, direct calculation of p(x) from this formula is not possible, because even for moderate
values of n, n! exceeds the largest integer that a computer can store. To overcome this difficulty,
indirect methods for calculation of p(x) are developed. For example, we may use the recursive
formula
(n − x + 1)p
p(x − 1),
p(x) =
x(1 − p)
obtained from (5.3), with the initial condition p(0) = (1 − p)n , to evaluate p(x). An important
study concerning calculation of p(x) is the one by the French mathematician Simeon-Denis
Poisson in 1837 in his book concerning the applications of probability to law. Poisson introduced the following procedure to obtain the formula that approximates the binomial probability
mass function when the number of trials is large (n → ∞), the probability of success is small
(p → 0), and the average number of successes remains a fixed quantity of moderate value
(np = λ for some constant λ).
Let X be a binomial random variable with parameters (n, p); then
!
λ i λ n−i
n i
n!
1−
P (X = i) =
p (1 − p)n−i =
(n − i)! i! n
n
i
λ n
i 1−
n(n − 1)(n − 2) · · · (n − i + 1) λ
n .
=
λ i
ni
i!
1−
n
i
n
Now, for large n and
is
appreciable λ, (1 − λ/n) is approximately 1, n(1 − λ/n)
−λ
x
approximately e
from calculus we know that limn→∞ (1 + x/n)
= e ; thus
limn→∞ (1 − λ/n)n = e−λ , and n(n − 1)(n − 2) · · · (n − i + 1) /ni is approximately 1,
because its numerator and denominator are both polynomials in n of degree i. Thus as n → ∞,
P (X = i) →
e−λ λi
.
i!
Three years after his book was published, Poisson died. The importance of this
approximation was left unknown until 1889, when the German-Russian mathematician
L. V. Bortkiewicz demonstrated its significance in probability theory and its applications.
Among other things, Bortkiewicz argued that since
∞
X
e−λ λi
i=0
i!
= e−λ
∞
X
λi
i!
i=0
= e−λ eλ = 1,
Poisson’s approximation by itself is a probability mass function. This and the introduction of
the Poisson processes in the twentieth century made the Poisson probability mass function one
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 210 — #226
✐
✐
210
Chapter 5
Special Discrete Distributions
of the three most important probability functions, the other two being the binomial and normal
probability functions.
Definition 5.3
A discrete random variable X with possible values 0, 1, 2, 3, . . . is called
Poisson with parameter λ, λ > 0, if
P (X = i) =
e−λ λi
,
i!
i = 0, 1, 2, 3, . . . .
Under the conditions specified in our discussion, binomial probabilities can be approximated by Poisson probabilities. Such approximations are generally good if p < 0.1 and
np ≤ 10. If np > 10, it would be more appropriate to use normal approximation, discussed in
Section 7.2.
Since a Poisson probability mass function is the limit of a binomial probability mass function, the expected value of a binomial random variable with parameters (n, p) is np, and
np = λ, it is reasonable to expect that the mean of a Poisson random variable with parameter λ is λ. To prove that this is the case, observe that
∞
X
∞
X
e−λ λi
E(X) =
iP (X = i) =
i
i!
i=0
i=1
= λe−λ
∞
∞
X
X
λi−1
λi
= λe−λ
= λe−λ eλ = λ.
(i
−
1)!
i!
i=1
i=0
The variance of a Poisson random variable X with parameter λ is also λ. To see this, note that
E(X 2 ) =
∞
X
i2 P (X = i) =
i=0
= λe−λ
= λe−λ
= λe−λ
∞
X
i=1
∞
X
i2
e−λ λi
i!
∞
X
iλi−1
1
d i
= λe−λ
(λ )
(i
−
1)!
(i
−
1)!
dλ
i=1
i=1
∞
∞
d h X λi i
d h X λi−1 i
= λe−λ
λ
dλ i=1 (i − 1)!
dλ i=1 (i − 1)!
d
(λeλ ) = λe−λ (eλ + λeλ ) = λ + λ2 ,
dλ
and hence
Var(X) = (λ + λ2 ) − λ2 = λ.
We have shown that
If X is a Poisson random variable with parameter λ, then
√
E(X) = Var(X) = λ,
σX = λ.
Some examples of binomial random variables that obey Poisson’s approximation are as
follows:
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 211 — #227
✐
✐
Section 5.2
Poisson Random Variables
211
1.
Let X be the number of babies in a community who grow up to at least 190 centimeters.
If a baby is called a success, provided that he or she grows up to the height of 190 or more
centimeters, then X is a binomial random variable. Since n, the total number of babies,
is large, p, the probability that a baby grows to the height of 190 centimeters or more, is
small, and np, the average number of such babies, is appreciable, X is approximately a
Poisson random variable.
2.
Let X be the number of winning tickets among the Maryland lottery tickets sold in Baltimore during one week. Then, calling winning tickets successes, we have that X is a
binomial random variable. Since n, the total number of tickets sold in Baltimore, is large,
p, the probability that a ticket wins, is small, and the average number of winning tickets is
appreciable, X is approximately a Poisson random variable.
3.
Let X be the number of misprints on a document page typed by a secretary. Then X is
a binomial random variable if a word is called a success, provided that it is misprinted!
Since misprints are rare events, the number of words is large, and np, the average number
of misprints, is of moderate value, X is approximately a Poisson random variable.
In the same way we can argue that random variables such as “the number of defective fuses
produced by a factory in a given year,” “the number of inhabitants of a town who live at least 85
years,” “the number of machine failures per day in a plant,” and “the number of drivers arriving
at a gas station per hour” are approximately Poisson random variables.
We will now give examples of problems that can be solved by Poisson’s approximations.
In the course of these examples, new Poisson random variables are also introduced.
Example 5.12 Every week the average number of wrong-number phone calls received by a
certain mail-order house is seven. What is the probability that they will receive (a) two wrong
calls tomorrow; (b) at least one wrong call tomorrow?
Solution: Assuming that the house receives a lot of calls, the number of wrong numbers
received tomorrow, X, is approximately a Poisson random variable with λ = E(X) = 1. Thus
1
e−1 · (1)n
=
,
n!
e · n!
and hence the answers to (a) and (b) are, respectively,
P (X = n) =
P (X = 2) =
1
≈ 0.18
2e
and
P (X ≥ 1) = 1 − P (X = 0) = 1 −
1
≈ 0.63. e
Example 5.13 Suppose that, on average, in every three pages of a book there is one typographical error. If the number of typographical errors on a single page of the book is a Poisson
random variable, what is the probability of at least one error on a specific page of the book?
Solution: Let X be the number of errors on the page we are interested in. Then X is a Poisson
random variable with E(X) = 1/3. Hence λ = E(X) = 1/3, and thus
P (X = n) =
(1/3)n e−1/3
.
n!
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 212 — #228
✐
✐
212
Chapter 5
Special Discrete Distributions
Therefore,
P (X ≥ 1) = 1 − P (X = 0) = 1 − e−1/3 ≈ 0.28. Example 5.14 The atoms of a radioactive element are randomly disintegrating. If every
gram of this element, on average, emits 3.9 alpha particles per second, what is the probability
that during the next second the number of alpha particles emitted from 1 gram is (a) at most 6;
(b) at least 2; (c) at least 3 and at most 6?
Solution: Every gram of radioactive material consists of a large number n of atoms. If for
these atoms the event of disintegrating and emitting an alpha particle during the next second
is called a success, then X, the number of alpha particles emitted during the next second, is a
binomial random variable. Now E(X) = 3.9, so np = 3.9 and p = 3.9/n. Since n is very
large, p is very small, and hence, to a very close approximation, X has a Poisson distribution
with the appreciable parameter λ = 3.9. Therefore,
P (X = n) =
(3.9)n e−3.9
,
n!
and hence (a), (b), and (c) are calculated as follows.
(a)
(b)
(c)
P (X ≤ 6) =
6
X
(3.9)n e−3.9
n!
n=0
≈ 0.899.
P (X ≥ 2) = 1 − P (X = 0) − P (X = 1) = 0.901.
P (3 ≤ X ≤ 6) =
6
X
(3.9)n e−3.9
n=3
n!
≈ 0.646. Example 5.15 Suppose that n raisins are thoroughly mixed in dough. If we bake k raisin
cookies of equal sizes from this mixture, what is the probability that a given cookie contains at
least one raisin?
Solution: Since the raisins are thoroughly mixed in the dough, the probability that a given
cookie contains any particular raisin is p = 1/k . If for the raisins the event of ending up in
the given cookie is called a success, then X, the number of raisins in the given cookie, is a
binomial random variable. For large values of k, p = 1/k is small. If n is also large but n/k
has a moderate value, it is reasonable to assume that X is approximately Poisson. Hence
P (X = i) =
λi e−λ
,
i!
where λ = np = n/k. Therefore, the probability that a given cookie contains at least one
raisin is
P (X 6= 0) = 1 − P (X = 0) = 1 − e−n/k . For real-world problems, it is important to know that in most cases numerical examples
show that, even for small values of n, the agreement between the binomial and the Poisson
probabilities is surprisingly good (see Exercise 11).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 213 — #229
✐
✐
Section 5.2
Poisson Random Variables
213
Poisson Processes
Earlier, the Poisson distribution was introduced to approximate binomial distribution for very
large n, very small p, and moderate np. On its own merit, the Poisson distribution appears in
connection with the study of sequences of random events occurring over time. Examples of
such sequences of random events are accidents occurring at an intersection, β -particles emitted from a radioactive substance, customers entering a post office, and California earthquakes.
Suppose that, starting from a time point labeled t = 0, we begin counting the number of events.
Then for each value of t, we obtain a number denoted by N (t), which is the number of events
that have occurred during [0, t]. For example, in studying the number of β -particles emitted
from a radioactive substance, N (t) is the number of β -particles emitted by time t. Similarly,
in studying the number of California earthquakes, N (t) is the number of earthquakes that have
occurred in the time interval [0, t]. Clearly, for each value of t, N (t) is a discrete random variable with the set of possible values {0, 1, 2, . . .}. To study the distribution of N (t), the number
of events occurring in [0, t], we make the following three simple and natural assumptions about
the way in which events occur.
Stationarity: For all n ≥ 0, and for any two equal time intervals ∆1 and ∆2 , the probability
of n events in ∆1 is equal to the probability of n events in ∆2 .
Independent Increments: For all n ≥ 0, and for any time interval (t, t + s), the probability of n events in (t, t + s) is independent of how many events have occurred earlier or
how they have occurred. In particular, suppose that the times 0 ≤ t1 < t2 < · · · < tk
are given. For 1 ≤ i < k − 1, let Ai be the event that ni events of the process occur in
[ti , ti+1 ). The independent increments mean that {A1 , A2 , . . . , Ak−1 } is an independent
set of events.
Orderliness: The occurrence of two or more events in a very small time interval is practically
impossible. This condition is mathematically expressed by limh→0 P N (h) > 1 /h =
0. This implies that as h → 0, the probability of two or more events, P N (h)
>
1
,
approaches 0 faster than h does. That is, if h is negligible, then P N (h) > 1 is even
more negligible.
Note that, by the stationarity property, the number of events in the interval (t1 , t2 ] has the
same distribution as the number of events in (t1 + s, t2 + s], s ≥ 0. This means that the random
variables N (t2 ) − N (t1 ) and N (t2 + s) − N (t1 + s) have the same probability mass function.
In other words, the probability of occurrence of n events during the interval of time from t1 to
t2 is a function of n and t2 − t1 and not of t1 and t2 independently.
The number of events in
(ti , ti+1 ], N (ti+1 ) − N (ti ), is called the increment in N (t) between ti and ti+1 . That is
why the second property in the preceding list is called independent-increments property. It is
also worthwhile to mention that stationarity and orderliness together imply the following fact,
proved in Section 12.2.
The simultaneous occurrence of two or more events is impossible. Therefore, under the aforementioned properties, events occur one at a time, not
in pairs or groups.
It is for this reason that the third property is called the orderliness property.
Suppose that random events occur in time in a way that the conditions discussed above—
stationarity, independent increments, and orderliness—are always satisfied. Then, if for some
interval of length t > 0, P N (t) = 0 = 0, we have that in any interval of length t at
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 214 — #230
✐
✐
214
Chapter 5
Special Discrete Distributions
least one event occurs. In such a case, it can be shown that in any interval of arbitrary length,
with probability
1, at least one event occurs. Similarly, if for some interval of length t > 0,
P N (t) = 0 = 1, then in any interval of length t no event will occur. In such a case, it can
be shown that in any interval of arbitrary length, with probability 1, no event occurs. To avoid
these uninteresting, trivial cases, throughout the book, we assume that, for all t > 0,
0 < P N (t) = 0 < 1.
We are now ready to state a celebrated theorem, for the validity of which we present a
motivating argument. The theorem is presented again, with a rigorous proof, in Chapter 12 as
Theorem 12.1.
Theorem 5.2 If random events occur in time in a way that the preceding conditions—
stationarity, independent increments,
and orderliness—are always satisfied, N (0) = 0 and,
for all t > 0, 0 < P N (t) = 0 < 1, then there exists a positive number λ such that
(λt)n e−λt
.
P N (t) = n =
n!
That
is, for all t > 0, N (t) is a Poisson
random variable with parameter λt. Hence
E N (t) = λt and therefore λ = E N (1) .
A Motivating Argument: The property that under the stated conditions, N (t) (the number of
events that has occurred in [0, t]) is a Poisson random variable is not accidental. It is related
to the fact that a Poisson random variable is approximately binomial for large n, small p, and
moderate np. To see this, divide the interval [0, t] into n subintervals of equal length. Then,
as n → ∞, the probability of two or more events in any of these subintervals is 0. Therefore,
N (t), the number of events in [0, t], is equal to the number of subintervals in which an event
occurs. If a subinterval in which an event occurs is called a success, then N (t) is the number of
successes. Moreover, the stationarity and independent-increments properties imply that N (t)
is a binomial random variable with parameters (n, p), where p is the probability of success
(i.e., the probability that an event occurs in a subinterval). Now let λ be the expected number of
events in an interval of unit length. Because of stationarity, events occur at a uniform rate over
the entire time period. Therefore, the expected number of events in any period of length t is λt.
Hence, in particular, the expected number of the events in [0, t] is λt. But by the formula for
the expected value of binomial random variables, the expected number of the events in [0, t]
is np. Thus np = λt or p = (λt)/n. Since n is extremely large (n → ∞), we have that p
is extremely small while np = λt is of moderate size. Therefore, N (t) is a Poisson random
variable with parameter λt. In the study of sequences of random
events occurring in time, suppose that N (0) = 0
and, for all t > 0, 0 < P N (t) = 0 < 1. Furthermore, suppose that the events occur in a
way that the preceding conditions—stationarity, independent increments, and orderliness—are
always satisfied. We argued that for each value of t, the discrete random variable N (t), the
number of events in [0, t] and hence in any other time interval of length t, is a Poisson random
variable with parameter λt. Any
process with this property is called a Poisson process with
rate λ and is often denoted by N (t), t ≥ 0 .
Theorem 5.2 is astonishing because it shows how three simple and natural physical conditions on N (t) characterize the probability mass functions of random variables N (t), t > 0.
Moreover, λ, the only unknown parameter of the probability mass functions of N (t)’s, equals
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 215 — #231
✐
✐
Section 5.2
Poisson Random Variables
215
E N (1) . That is, it is the average number of events that occur in one unit of time. Hence in
practice it can be measured readily. Historically, Theorem 5.2 is the most elegant and evolved
form of several previous theorems. It was discovered in 1955 by the Russian mathematician
Aleksandr Khinchin (1894–1959). The first major work in this direction was done by Thornton
Fry (1892–1992) in 1929. But before Fry, Albert Einstein (1879–1955) and Roman Smoluchowski (1910–1996) had also discovered important results in connection with their work on
the theory of Brownian motion.
Example 5.16 (Standard Birthday Problem Revisited) Consider a class of size k
of unrelated students. Assuming that the birth rates are constant throughout the year and each
year has 365 days, with probability p = 1/365, each pair of the students has the same birthday.
Let X be the number of such pairs. Argue that X is approximately a Poisson random variable
and find the probability of at least one birthday match for k = 23, 30, 50, and 60. Compare
your results with the exact probabilities that were calculated up to three decimal places in
Example 2.8.
!
k
Solution: There are n =
pairs of birthdays. Call a pair success if its elements are
2
identical and failure otherwise. Clearly X is a binomial random variable with parameters n
and p. Since n is large and p is small,!X is approximately a Poisson random variable with the
k 1
. The probability of at least one birthday match is
appreciable parameter λ = np =
2 365
e−λ · λ0
= 1 − e−λ .
0!
The following table shows how good this approximation is
k
23
30
50
60
P (X ≥ 1) = 1 − P (X = 0) = 1 −
P (X ≥ 1) by Poisson approximation
0.500
0.696
0.965
0.992
P (X ≥ 1) obtained in Example 2.8
0.507
0.706
0.970
0.995
Using Poisson random variables, similarly to the method used in this exercise, we can easily
find excellent approximations for the probability that two students have, say, birthdays within
one day of each other or within two days of each other and so forth. For having birthdays
within one day, in the solution above, all we need to do is to change p from 1/365 to 3/365.
Similarly, for having birthdays within two days, it suffices to change p from 1/365 to 5/365.
Such birthday problems are not easy to solve. Poisson approximation comes in handy and gives
excellent results. Example 5.17 Suppose that children are born at a Poisson rate of five per day in a certain
hospital. What is the probability that (a) at least two babies are born during the next six hours;
(b) no babies are born during the next two days?
Solution: Let N (t) denote the number of babies born at or prior to t. The assumption that
N (t), t ≥ 0 is a Poisson process is reasonable because it is stationary, it has independent
increments, N (0) = 0, and simultaneous births are impossible. Thus N (t), t ≥ 0 is a
Poisson process. If we choose one day as time unit, then λ = E N (1) = 5. Therefore,
(5t)n e−5t
P N (t) = n =
.
n!
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 216 — #232
✐
✐
216
Chapter 5
Special Discrete Distributions
Hence the probability that at least two babies are born during the next six hours is
P N (1/4) ≥ 2 = 1 − P N (1/4) = 0 − P N (1/4) = 1
=1−
(5/4)0 e−5/4 (5/4)1 e−5/4
−
≈ 0.36,
0!
1!
where 1/4 is used since 6 hours is 1/4 of a day. The probability that no babies are born during
the next two days is
(10)0 e−10
P N (2) = 0 =
≈ 4.54 × 10−5 . 0!
Example 5.18 Suppose that earthquakes occur in a certain region of California, in accordance with a Poisson process, at a rate of seven per year.
(a)
What is the probability of no earthquakes in one year?
(b)
What is the probability that in exactly three of the next eight years no earthquakes will
occur?
Solution:
(a)
Let
N (t) be the number of earthquakes in this region at or prior to t. We are given that
N (t), t ≥ 0 is a Poisson process. If we choose one year as the unit of time, then
λ = E N (1) = 7. Thus
(7t)n e−7t
,
P N (t) = n =
n!
(b)
n = 0, 1, 2, 3, . . . .
Let p be the probability of no earthquakes in one year; then
p = P N (1) = 0 = e−7 ≈ 0.00091.
Suppose that a year is called a success if during its course no earthquakes occur. Of the
next eight years, let X be the number of years in which no earthquakes will occur. Then
X is a binomial random variable with parameters (8, p). Thus
!
8
P (X = 3) ≈
(0.00091)3 (1 − 0.00091)5 ≈ 4.2 × 10−8 . 3
Example 5.19 A fisherman catches fish at a Poisson rate of two per hour from a large lake
with lots of fish. Yesterday, he went fishing at 10:00 A.M. and caught just one fish by 10:30 and
a total of three by noon. What is the probability that he can duplicate this feat tomorrow?
Solution: Label the time the fisherman starts fishing tomorrow at t = 0. Let N (t) denote
the total number of fish caught at or prior to t. Clearly, N (0) = 0. It is reasonable to assume
that catching two or more fish simultaneously is impossible. It is also reasonable to assume
that N (t), t ≥ 0 is stationary and has independent increments. Thus the assumption that
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 217 — #233
✐
✐
Section 5.2
Poisson Random Variables
217
N (t), t ≥ 0 is a Poisson process is well grounded. Choosing 1 hour as the unit of time, we
have λ = E N (1) = 2. Thus
(2t)n e−2t
,
n = 0, 1, 2, . . . .
P N (t) = n =
n!
We want to calculate the probability of the event N (1/2) = 1 and N (2) = 3 . But this event
is the same as N (1/2) = 1 and N (2) − N (1/2) = 2 . Thus by the independent-increments
property,
P N (1/2) = 1 and N (2) − N (1/2) = 2 = P N (1/2) = 1 · P N (2) − N (1/2) = 2 .
Since stationarity implies that P N (2) − N (1/2) = 2 = P N (3/2) = 2 , the desired
probability equals
11 e−1 32 e−3
P N (1/2) = 1 · P N (3/2) = 2 =
·
≈ 0.082. 1!
2!
Example 5.20 Let
N (t) be the number of earthquakes that occur at or prior to time t worldwide. Suppose that N (t) : t ≥ 0 is a Poisson process with rate λ and the probability that
the magnitude of an earthquake on the Richter scale is 5 or more is p. Find the probability of k
earthquakes of such magnitudes at or prior to t worldwide.
Solution: Let X(t) be the number of earthquakes of magnitude
5 or more on the Richter
scale at or prior to t worldwide. Since the sequence of events N (t) = n , n = 0, 1, 2, . . . is
S∞ mutually exclusive and n=0 N (t) = n is the sample space, by the law of total probability,
Theorem 3.4,
∞
X
P X(t) = k =
P X(t) = k | N (t) = n P N (t) = n .
n=0
Now clearly, P X(t) = k | N (t) = n = 0 if n < k . If n > k, the conditional probability
mass function of X(t) given that N (t) = n is binomial with parameters n and p. Thus
!
∞
X
n k
e−λt (λt)n
P X(t) = k =
p (1 − p)n−k
k
n!
n=k
=
∞
X
n!
e−λt (λt)k (λt)n−k
pk (1 − p)n−k
k! (n − k)!
n!
n=k
∞
n−k
e−λt (λtp)k X
1
=
λt(1 − p)
k!
(n − k)!
n=k
=
∞
j
e−λt (λtp)k X 1 λt(1 − p)
k!
(j)!
j=0
e−λt (λtp)k λt(1−p)
e
k!
e−λtp (λtp)k
=
.
k!
=
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 218 — #234
✐
✐
218
Chapter 5
Special Discrete Distributions
It should be clear that X(t) : t ≥ 0 is stationary, orderly, and possesses independent in
crement property. Therefore, X(t) : t ≥ 0 is itself a Poisson process with parameter λp.
Remark 5.1 We only considered sequences of random events that occur in time. However,
the restriction to time is not necessary. Random events that occur on the real line, the plane, or
in space and satisfy the stationarity and orderliness conditions and possess independent increments also form Poisson processes. For example, suppose that a wire manufacturing company
produces a wire that has various fracture sites where the wire will fail intension. Let N (t) be
the number of fracture sites in the first t meters of wire. The process N (t), t ≥ 0 may
be modeled as a Poisson process. As another example, suppose that in a certain region S, the
numbers of trees that grow in nonoverlapping subregions are independent of each other, the
distributions of the number of trees in subregions of equal area are identical, and the probability of two or more trees in a very small subregion is negligible. Let λ be the expected number
of trees in a region of area 1 and A(R) be the area of a region R. Then N (R), the number
of
trees in a subregion R is a Poisson random variable with parameter λA(R), and the set
N (R), R ⊆ S is a two-dimensional Poisson process. EXERCISES
A
1.
Jim buys 60 lottery tickets every week. If only 5% of the lottery tickets win, what is the
probability that he wins next week?
2.
Suppose that 3% of the families in a large city have an annual income of over $60,000.
What is the probability that, of 60 random families, at most three have an annual income
of over $60,000?
3.
Suppose that 2.5% of the population of a border town are illegal immigrants. Find the
probability that, in a theater of this town with 80 random viewers, there are at least two
illegal immigrants.
4.
By Example 2.24, the probability that a poker hand is a full house is 0.0014. What is the
probability that in 500 random poker hands there are at least two full houses?
5.
On a random day, the number of vacant rooms of a big hotel in New York City is 35,
on average. What is the probability that next Saturday this hotel has at least 30 vacant
rooms?
6.
On average, there are three misprints in every 10 pages of a particular book. If every
chapter of the book contains 35 pages, what is the probability that Chapters 1 and 5 have
10 misprints each?
7.
Suppose that X is a Poisson random variable with P (X = 1) = P (X = 3). Find
P (X = 5).
8.
Suppose that n raisins have been carefully mixed with a batch of dough. If we bake
k (k > 4) raisin buns of equal size from this mixture, what is the probability that two
out of four randomly selected buns contain no raisins?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 219 — #235
✐
✐
Section 5.2
Poisson Random Variables
219
Hint: Note that, by Example 5.15, the number of raisins in a given bun is approximately Poisson with parameter n/k .
9.
The children in a small town all own slingshots. In a recent contest, 4% of them were
such poor shots that they did not hit the target even once in 100 shots. If the number
of times a randomly selected child has hit the target is approximately a Poisson random
variable, determine the percentage of children who have hit the target at least twice.
10.
At a large mall, on a major holiday, they give a gift to the first, and then to every 100th
customer who enters the mall. If customers arrive at the mall according to a Poisson
process with rate λ, find the probability mass function of Xt , the number of gifts given
by time t.
11.
The department of mathematics of a state university has 26 faculty members. For i =
0, 1, 2, 3, find pi , the probability that i of them were born on Independence Day (a) using
the binomial distribution; (b) using the Poisson distribution. Assume that the birth rates
are constant throughout the year and that each year has 365 days.
12.
Suppose that on a summer evening, shooting stars are observed at a Poisson rate of one
every 12 minutes. What is the probability that three shooting stars are observed in 30
minutes?
13.
Let X be the number of times, within the next five years, that a policyholder of an
insurance company gets injured in a car accident. An actuary has shown that X is a
Poisson random variable, and the probability that, within 5 years, a policyholder gets
injured in exactly two accidents is 1/5th of the probability that he or she gets injured in
exactly one accident. Find the mean and the standard deviation of X .
14.
It is 6:00 P.M. now and, for t > 0, the number of orders that are placed, between now
and t minutes past 6:00 P.M. , through an electronic commerce company’s website, is a
Poisson process. Suppose that the first order placed through the company’s website will
arrive within the next minute with probability 0.09. Find the expected number of all the
orders that will be placed within the next hour.
15.
Suppose that in Japan earthquakes occur at a Poisson rate of three per week. What is the
probability that the next earthquake occurs after two weeks?
16.
Shocks occur independently to a system at a Poisson rate of λ. The probability that the
system survives a shock is p independently of other shocks. Find the probability that the
system has survived all shocks by time t.
17.
A pharmaceutical company has developed an Ebola vaccine, which is 98.5% of the time
effective. If 5000 people have received the vaccine, and sooner or later all of them will be
exposed to the Ebola virus, what is the probability that at most 70 of them will become
sick with Ebola? Assume that a person reacts to the Ebola vaccine independently of
other people.
18.
Suppose that, for a telephone subscriber, the number of wrong numbers is Poisson, at a
rate of λ = 1 per week. A certain subscriber has not received any wrong numbers from
Sunday through Friday. What is the probability that he receives no wrong numbers on
Saturday either?
19.
In a certain town, crimes occur at a Poisson rate of five per month. What is the probability
of having exactly two months (not necessarily consecutive) with no crimes during the
next year?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 220 — #236
✐
✐
220
Chapter 5
20.
Accidents occur at an intersection at a Poisson rate of three per day. What is the probability that during January there are exactly three days (not necessarily consecutive) without
any accidents?
21.
Customers arrive at a bookstore at a Poisson rate of six per hour. Given that the store
opens at 9:30 A.M., what is the probability that exactly one customer arrives by 10:00 A.M.
and 10 customers by noon?
22.
A company is located in a region that is prone to dust storms. The company has purchased an insurance policy that pays $200,000 for each dust storm in a year except the
first one. If the number of the dust storms per year is a Poisson random variable with
mean 0.83, what is the expected amount of money that the company is paid in a random
year under this insurance policy?
23.
A wire manufacturing company has inspectors to examine the wire for fractures as it
comes out of a machine. The number of fractures is distributed in accordance with a
Poisson process, having one fracture on the average for every 60 meters of wire. One
day an inspector has to take an emergency phone call and is missing from his post for
ten minutes. If the machine turns out 7 meters of wire per minute, what is the probability
that the inspector will miss more than one fracture?
24.
Let X be a Poisson random variable with parameter λ. Let
0
if X is zero or even
Y =
1
If X is odd.
Special Discrete Distributions
Find the probability mass function of Y .
B
25.
On a certain two-lane north-south highway, there is a T junction. Cars arrive at the
junction according to a Poisson process, on the average four per minute. For cars to turn
left onto the side street, the highway is widened by the addition of a left-turn lane that
is long enough to accommodate three cars. If four or more cars are trying to turn left,
the fourth car will effectively block north-bound traffic. At the junction for the left-turn
lane there is a left-turn signal that allows cars to turn left for one minute and prohibits
such turns for the next three minutes. The probability of a randomly selected car turning
left at this T junction is 0.22. Suppose that during a green light for the left-turn lane all
waiting cars were able to turn left. What is the probability that during the subsequent red
light for the left-turn lane, the north-bound traffic will be blocked?
26.
Suppose that, on the Richter scale, earthquakes of magnitude 5.5 or higher have probability 0.015 of damaging certain types of bridges. Suppose that such intense earthquakes
occur at a Poisson rate of 1.5 per ten years. If a bridge of this type is constructed to last
at least 60 years, what is the probability that it will be undamaged by earthquakes for
that period of time?
27.
According to the United States Postal Service, http:www.usps.gov, May 15, 1998,
Dogs have caused problems for letter carriers for so long that the situation
has become a cliché. In 1983, more than 7000 letter carriers were bitten
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 221 — #237
✐
✐
Section 5.2
Poisson Random Variables
221
by dogs. . . . However, the 2,795 letter carriers who were bitten by dogs
last year represent less than one-half of 1 percent of all reported dog-bite
victims.
Suppose that during a year 94% of the letter carriers are not bitten by dogs. Assuming
that dogs bite letter carriers randomly, what percentage of those who sustained one bite
will be bitten again?
28.
Suppose that in Maryland, on a certain day, N lottery tickets are sold and M win. To
have a probability of at least α of winning on that day, approximately how many tickets
should be purchased?
29.
Balls numbered 1,2, . . . , and n are randomly placed into cells numbered 1, 2, . . . , and
n. Therefore, for 1 ≤ i ≤ n and 1 ≤ j ≤ n, the probability that ball i is in cell j is
1/n. For each i, 1 ≤ i ≤ n, if ball i is in cell i, we say that a match has occurred at
cell i.
30.
31.
(a)
What is the probability of exactly k matches?
(b)
Let n → ∞. Show that the probability mass function of the number of matches
is Poisson with mean 1.
Let N (t), t ≥ 0 be a Poisson process. What is the probability of (a) an even number
of events in (t, t + α); (b) an odd number of events in (t, t + α)?
Let N (t), t ≥ 0 be a Poisson process with rate λ. Suppose that N (t) is the total
number of two types of events that have occurred in [0, t]. Let N1 (t) and N2 (t) be
the total number of events of type 1 and events of type 2 that have occurred in [0, t],
respectively. If events of type 1 and
with probabilities p and
type 2 occur independently
1 − p, respectively, prove that N1 (t), t ≥ 0 and N2 (t), t ≥ 0 are Poisson
processes with respective rates λp and λ(1 − p).
Hint: First calculate P N1 (t) = n and N2 (t) = m using the relation
P N1 (t) = n and N2 (t) = m
∞
X
=
P N1 (t) = n and N2 (t) = m | N (t) = i P N (t) = i .
i=0
(This is true because of Theorem 3.4.) Then use the relation
∞
X
P N1 (t) = n =
P N1 (t) = n and N2 (t) = m .
m=0
32.
Customers arrive at a grocery store at a Poisson rate of one per minute. If 2/3 of the
customers are female and 1/3 are male, what is the probability that 15 females enter the
store between 10:30 and 10:45?
Hint: Use the result of Exercise 31.
33.
In a forest, the number of trees that grow in a region of area R has a Poisson distribution
with mean λR, where λ is a given positive number.
(a)
Find the probability that the distance from a certain tree to the nearest tree is more
than d.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 222 — #238
✐
✐
222
Chapter 5
(b)
34.
Special Discrete Distributions
Find the probability that the distance from a certain tree to the nth nearest tree is
more than d.
Let X be a Poisson random variable with parameter λ. Show that the maximum of
P (X = i) occurs at [λ], where [λ] is the greatest integer less than or equal to λ.
Hint: Let p be the probability mass function of X . Prove that
p(i) =
λ
p(i − 1).
i
Use this to find the values of i at which p is increasing and the values of i at which it is
decreasing.
Self-Quiz on Section 5.2
Time allotted: 15 Minutes
Each problem is worth 5 points.
1.
Suppose that due to an exhaust leak in 0.3% of cars manufactured by a car company,
their passenger side exhaust manifolds must be replaced. What is the probability that in
the next 5000 cars manufactured by this company, at least 12 but no more than 15 need
such a repair?
2.
Patients arrive at a pharmacy for flu shots at a Poisson rate of 3 per hour. Given that
the pharmacy opens at 8:00 A.M., what is the probability that exactly 2 patients arrive by
9:00 A.M. and 12 patients by noon?
5.3
OTHER DISCRETE RANDOM VARIABLES
Geometric Random Variables
Consider an experiment in which independent Bernoulli trials are performed until the first
success occurs. The sample space for such an experiment is
S = {s, f s, f f s, f f f s, . . . , f f · · · f s, . . .}.
Now, suppose that a sequence of independent Bernoulli trials, each with probability of success
p, 0 < p < 1, are performed. Let X be the number of experiments until the first success
occurs. Then X is a discrete random variable called geometric. It is defined on S, its set of
possible values is {1, 2, . . .}, and
P (X = n) = (1 − p)n−1 p,
n = 1, 2, 3, . . . .
This equation follows since (a) the first (n − 1) trials are all failures, (b) the nth trial is a
success, and (c) the successive Bernoulli trials are all independent.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 223 — #239
✐
✐
Section 5.3
Other Discrete Random Variables
223
Let p(x) = (1 − p)x−1 p for x = 1, 2, 3, . . . , and 0 elsewhere. Then, for all values of x in
R , p(x) ≥ 0 and
∞
X
p(x) =
x=1
∞
X
x=1
(1 − p)x−1 p =
p
= 1,
1 − (1 − p)
by the geometric series theorem. Hence p(x) is a probability mass function.
Definition 5.4
The probability mass function
(
(1 − p)x−1 p
0 < p < 1,
p(x) =
0
elsewhere
x = 1, 2, 3, . . . ,
is called geometric.
Let X be a geometric random variable with parameter p; then
E(X) =
∞
X
x=1
=
x−1
xp(1 − p)
∞
p X
=
x(1 − p)x
1 − p x=1
p
1
1−p
= ,
1 − p 1 − (1 − p) 2
p
P∞
where the third equality follows from the relation x=1 xr x = r/(1 − r)2 , |r| < 1. E(X) =
1/p indicates that toP
get the first success,
on average, 1/p independent Bernoulli trials are
∞
needed. The relation x=1 x2 r x = r(r + 1) /(1 − r)3 , |r| < 1, implies that
E(X 2 ) =
∞
X
x=1
x2 p(1 − p)x−1 =
∞
2−p
p X 2
.
x (1 − p)x =
1 − p x=1
p2
Hence
2
2 − p 1 2 1 − p
Var(X) = E(X 2 ) − E(X) =
−
.
=
p2
p
p2
We have established the following formulas:
Let X be a geometric random variable with parameter p, 0 < p < 1. Then
√
1
1−p
1−p
E(X) = , Var(X) =
.
,
σ
=
X
p
p2
p
Example 5.21 From an ordinary deck of 52 cards we draw cards at random, with replacement, and successively until an ace is drawn. What is the probability that at least 10 draws are
needed?
Solution: Let X be the number of draws until the first ace. The random variable X is geometric with the parameter p = 1/13. Thus
P (X = n) =
12 n−1 1 ,
13
13
n = 1, 2, 3, . . . ,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 224 — #240
✐
✐
224
Chapter 5
Special Discrete Distributions
and so the probability that at least 10 draws are needed is
1 X 12 n−1
13
13
13 n=10 13
n=10
9
1 (12/13)
12 9
=
≈ 0.49.
·
=
13 1 − 12/13
13
P (X ≥ 10) =
∞ X
12 n−1 1 =
∞
Remark: There is a shortcut to the solution of this problem: The probability that at least 10
draws are needed to get an ace is the same as the probability that in the first nine draws there
are no aces. This is equal to (12/13)9 ≈ 0.49. Let X be a geometric random variable with parameter p, 0 < p < 1. Then, for all positive
integers n and m,
P (X > n + m|X > m) =
P (X > n + m)
(1 − p)n+m
=
= (1 − p)n = P (X > n).
P (X > m)
(1 − p)m
This is called the memoryless property of geometric random variables. It means that
In successive independent Bernoulli trials, the probability that the next n
outcomes are all failures does not change if we are given that the previous
m successive outcomes were all failures.
This is obvious by the independence of the trials. Interestingly enough, in the following sense,
geometric random variable is the only memoryless discrete random variable.
Let X be a discrete random variable with the set of possible values
{1, 2, 3 . . .}. If for all positive integers n and m,
P (X > n + m | X > m) = P (X > n),
then X is a geometric random variable. That is, there exists a number p,
0 < p < 1, such that
P (X = n) = p(1 − p)n−1 ,
n ≥ 1.
We leave the proof of this theorem as an exercise (see Exercise 25).
Example 5.22 A father asks his sons to cut their backyard lawn. Since he does not specify
which of the three sons is to do the job, each boy tosses a coin to determine the odd person,
who must then cut the lawn. In the case that all three get heads or tails, they continue tossing
until they reach a decision. Let p be the probability of heads and q = 1 − p, the probability of
tails.
(a)
Find the probability that they reach a decision in less than n tosses.
(b)
If p = 1/2, what is the minimum number of tosses required to reach a decision with
probability 0.95?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 225 — #241
✐
✐
Section 5.3
Other Discrete Random Variables
225
Solution:
(a)
The probability that they reach a decision on a certain round of coin tossing is
!
!
3 2
3 2
p q+
q p = 3pq(p + q) = 3pq.
2
2
The probability that they do not reach a decision on a certain round is 1 − 3pq. Let X be
the number of tosses until they reach a decision; then X is a geometric random variable
with parameter 3pq. Therefore,
P (X < n) = 1 − P (X ≥ n) = 1 − (1 − 3pq)n−1 ,
where the second equality follows since X ≥ n if and only if none of the first n − 1
tosses results in a success.
(b)
We want to find the minimum n so that P (X ≤ n) ≥ 0.95. This gives
1 − P (X > n) ≥ 0.95 or P (X > n) ≤ 0.05. But
P (X > n) = (1 − 3pq)n = (1 − 3/4)n = (1/4)n .
Therefore, we must have (1/4)n ≤ 0.05, or n ln 1/4 ≤ ln 0.05. This gives n ≥ 2.16;
hence the smallest n is 3. Negative Binomial Random Variables
Negative binomial random variables are generalizations of geometric random variables. Suppose that a sequence of independent Bernoulli trials, each with probability of success p,
0 < p < 1, is performed. Let X be the number of experiments until the r th success occurs.
Then X is a discrete random variable called a negative binomial. Its set of possible values is
{r, r + 1, r + 2, r + 3, . . .} and
!
n−1 r
P (X = n) =
p (1 − p)n−r ,
n = r, r + 1, . . . .
(5.4)
r−1
This equation follows since if the outcome of the nth trial is the r th success, then in the first
(n − 1) trials exactly (r − 1) successes have occurred and the nth trial is a success. The
probability of the former event is
!
!
n − 1 r−1
n − 1 r−1
(n−1)−(r−1)
p (1 − p)
=
p (1 − p)n−r ,
r−1
r−1
and the probability of the latter is p. Therefore, by the independence of the trials, (5.4) follows.
Definition 5.5
p(x) =
The probability mass function
!
x−1 r
p (1 − p)x−r , 0 < p < 1,
r−1
x = r, r + 1, r + 2, r + 3, . . . ,
is called negative binomial with parameters (r, p).
Note that a negative binomial probability mass function with parameters (1, p) is geometric. In Chapter 10, Examples 10.7 and 10.15, we will show that
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 226 — #242
✐
✐
226
Chapter 5
Special Discrete Distributions
If X is a negative binomial random variable with parameters (r, p), then
p
r
r(1 − p)
r(1 − p)
E(X) = , Var(X) =
, σX =
.
2
p
p
p
Example 5.23 Sharon and Ann play a series of backgammon games until one of them wins
five games. Suppose that the games are independent and the probability that Sharon wins a
game is 0.58.
(a)
Find the probability that the series ends in seven games.
(b)
If the series ends in seven games, what is the probability that Sharon wins?
Solution: (a) Let X be the number of games until Sharon wins five games. Let Y be the
number of games until Ann wins five games. The random variables X and Y are negative
binomial with parameters (5, 0.58) and (5, 0.42), respectively. The probability that the series
ends in seven games is
!
!
6
6
P (X = 7) + P (Y = 7) =
(0.58)5 (0.42)2 +
(0.42)5 (0.58)2
4
4
≈ 0.17 + 0.066 ≈ 0.24.
(b) Let A be the event that Sharon wins and B be the event that the series ends in seven
games. Then the desired probability is
P (A | B) =
P (X = 7)
0.17
P (AB)
=
≈
≈ 0.71. P (B)
P (X = 7) + P (Y = 7)
0.24
The following example, given by Kaigh in January 1979 issue of Mathematics Magazine,
is a modification of the gambler’s ruin problem, Example 3.15.
Example 5.24 (Attrition Ruin Problem) Two gamblers play a game in which in each
play gambler A beats B with probability p, 0 < p < 1, and loses to B with probability
q = 1 − p. Suppose that each play results in a forfeiture of $1 for the loser and in no change for
the winner. If player A initially has a dollars and player B has b dollars, what is the probability
that B will be ruined?
Solution: Let Ei be the event that, in the first b + i plays, B loses b times. Let A∗ be the event
that A wins. Then
P (A∗ ) =
a−1
X
P (Ei ).
i=0
If every time that A wins is called a success, Ei is the event that the bth success occurs on the
(b + i)th play. Using the negative binomial distribution, we have
!
i+b−1 b i
P (Ei ) =
pq.
b−1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 227 — #243
✐
✐
Section 5.3
Other Discrete Random Variables
227
Therefore,
!
a−1
X
i+b−1 b i
P (A ) =
pq.
b−1
i=0
∗
As a numerical illustration, let a = b = 4, p = 0.6, and q = 0.4. We get P (A∗ ) = 0.710208
and P (B ∗ ) = 0.289792 exactly. The following example is due to Hugo Steinhaus (1887–1972), who brought it up in a conference honoring Stefan Banach (1892–1945), a smoker and one of the greatest mathematicians
of the twentieth century.
Example 5.25 (Banach Matchbox Problem) A smoking mathematician carries two
matchboxes, one in his right pocket and one in his left pocket. Whenever he wants to smoke,
he selects a pocket at random and takes a match from the box in that pocket. If each matchbox
initially contains N matches, what is the probability that when the mathematician for the first
time discovers that one box is empty, there are exactly m matches in the other box, m =
0, 1, 2, . . . , N ?
Solution: Every time that the left pocket is selected we say that a success has occurred. When
the mathematician discovers that the left box is empty, the right one contains m matches if and
only if the (N + 1)st success occurs on the (N − m) + (N + 1) = (2N − m + 1)st trial.
The probability of this event is
!
!
2N − m 1 2N −m+1
(2N − m + 1) − 1 1 N +1 1 (2N −m+1)−(N +1)
=
.
2
2
2
N
(N + 1) − 1
By symmetry,
! when the mathematician discovers that the right box is empty, with probability
2N − m 1 2N −m+1
, the left box contains m matches. Therefore, the desired probability
N
2
is
!
!
2N − m 1 2N −m
2N − m 1 2N −m+1
=
. 2
2
2
N
N
Hypergeometric Random Variables
Suppose that, from a box containing D defective and N − D nondefective items, n are drawn
at random and without replacement. Furthermore, suppose that the number of items drawn does
not exceed the number of defective or the number of nondefective items. That is, suppose that
n ≤ min(D, N − D). Let X be the number of defective items drawn. Then X is a discrete
random variable with the set of possible values {0, 1, . . . n}, and a probability mass function
!
!
D
N −D
x
n−x
!
,
x = 0, 1, 2, . . . , n.
p(x) = P (X = x) =
N
n
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 228 — #244
✐
✐
228
Chapter 5
Special Discrete Distributions
Any random variable X with such a probability mass function is called a hypergeometric
random variable. The fact that
Pnp(x) is a probability mass function is readily verified. Clearly,
p(x) ≥ 0, ∀x. To prove that x=0 p(x) = 1, note that this is equivalent to
!
!
!
n
X
D
N −D
N
=
,
x
n−x
n
x=0
which can be shown by a simple combinatorial argument. (See Exercise 58 of Section 2.4; in
that exercise let m = D, n = N − D, and r = n.)
Let N, D, and n be positive integers with n ≤ min(D, N − D). Then
!
!
D
N
−
D
n−x
x
!
if x ∈ {0, 1, 2, . . . n}
p(x) = P (X = x) =
N
n
0
elsewhere
Definition 5.6
is said to be a hypergeometric probability mass function.
For the following important results, see Example 10.8, and Exercise 25, Section 10.2.
For the hypergeometric random variable X, defined above,
nD
nD(N − D) n−1
E(X) =
,
Var(X) =
1
−
.
N
N2
N −1
Note that if the experiment of drawing n items from a box containing D defective and N − D
nondefective items is performed with replacement, then X is binomial with parameters n and
D/N . Hence
nD
D
D nD(N − D)
E(X) =
, Var(X) = n
1−
=
.
N
N
N
N2
These show that if items are drawn with replacement, then the expected value of X will not
change, but the variance will increase. However, if n is much smaller than N, then, as the
variance formulae confirm, drawing with replacement is a good approximation for drawing
without replacement.
Example 5.26 In 500 independent calculations a scientist has made 25 errors. If a second
scientist checks seven of these calculations randomly, what is the probability that he detects
two errors? Assume that the second scientist will definitely find the error of a false calculation.
Solution: Let X be the number of errors found by the second scientist. Then X is hypergeometric with N = 500, D = 25, and n = 7. We are interested in P (X = 2), which is given
by
!
!
25
500 − 25
2
7−2
!
p(2) =
≈ 0.04. 500
7
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 229 — #245
✐
✐
Section 5.3
Other Discrete Random Variables
229
Example 5.27 In a community of a + b potential voters, a are pro-choice and b (b < a)
are pro-life. Suppose that a vote is taken to determine the will of the majority with regard to
legalizing abortion. If n (n < b) random persons of these a + b potential voters do not vote,
what is the probability that the pro-life advocates will win?
Solution: Let X be the number of those who do not vote and are pro-choice. The pro-life
advocates will win if and only if
a − X < b − (n − X).
But this is true if and only if X > (a−b+n)/2. Since X is a hypergeometric random variable,
we have that
!
!
a
b
n
n
X
X
i
n−i
a − b + n
! ,
=
P X>
P (X = i) =
2
a
+
b
a−b+n
a−b+n
i=[
]+1
]+1
i=[
2
2
n
where
ha − b + ni
2
is the greatest integer less than or equal to
a−b+n
.
2
Example 5.28 Every month, a large lot of 2000 fuses manufactured by a certain company
is shipped to a major retail store. The policy of the store is to return a lot if more than 2% of
its fuses are defective. Since 2% of 50 fuses is 1 fuse, the store selects 50 of the 2000 fuses at
random and tests them all. If it finds more than one defective fuse among them, it will return
the entire lot. Does this sampling strategy, with probability of 0.95 or higher, guarantee that the
store’s return policy is satisfied?
Solution: Let X be the number of defective items among the 50 randomly selected fuses
from a lot. According to the store’s sampling strategy, P (X > 1) is the probability that the
store will return the lot. Note that, based on the store’s policy, the lot must contain at least
41 defective fuses for the store to return it. Suppose that it actually has exactly 41 such fuses.
Then X is hypergeometric with parameters N = 2000, D = 41, and n = 50. In this case, the
probability of returning the lot, based on the sampling strategy, is
P (X > 1) = 1 − P (X = 0) − P (X = 1)
!
!
!
!
41
2000 − 41
41
2000 − 41
1
49
0
50
!
!
−
≈ 0.274.
=1−
2000
2000
50
50
Thus the probability of returning a lot that must be returned is too low, much lower than the
required 0.95. This alone shows that the store’s sampling strategy does not guarantee that its
return policy is satisfied. Furthermore, by similar calculations, the following table shows that
even if there are greater than 41 defective fuses in the lot, still the store’s sampling strategy fails
to guarantee that its policy is satisfied with the required probability.
Number of defective items in the lot
50
70
100
150
P (X > 1)
0.357
0.580
0.724
0.900
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 230 — #246
✐
✐
230
Chapter 5
Special Discrete Distributions
It is interesting to note that, with probability of 0.95 or higher, the store’s sampling strategy
works if its policy is to return a lot in which more than 9% of the fuses are defective. That is,
a lot with at least 181 defective fuses. For exactly 181 such fuses, P (X > 1) = 0.95, and if
there are more than 181 defective items in the lot, then P (X > 1) is greater than 0.95. Example 5.29 Professors Davidson and Johnson from the University of Victoria in Canada
gave the following problem to their students in a finite math course:
An urn contains N balls of which B are black and N − B are white; n balls are
chosen at random and without replacement from the urn. If X is the number
of black balls chosen, find P (X = i).
Using the hypergeometric formula, the solution to this problem is
!
!
B
N −B
i
n−i
!
P (X = i) =
,
i ≤ min(n, B).
N
n
However, by mistake, some students interchanged B with n and came up with the following
answer:
!
!
n
N −n
i
B−i
!
,
P (X = i) =
N
B
which can be verified immediately as the correct answer for all values of i. Explain if this is
accidental or if there is a probabilistic reason for its validity.
Solution: In their paper “Interchanging Parameters of Hypergeometric Distributions,” Mathematics Magazine, December 1993, Volume 66, Davidson and Johnson indicate that what happened is not accidental and that there is a probabilistic reason for it. Their justification is elegant
and is as follows:
We begin with an urn that contains N white balls. When Ms. Painter arrives,
she picks B balls at random without replacement from the urn, then paints
each of the drawn balls black with instant dry paint, and finally returns these
B balls to the urn. When Mr. Carver arrives, he chooses n balls at random
without replacement from the urn, then engraves each of the chosen balls with
the letter C and, finally, returns these n balls to the urn. Let the random variable
X denote the number of black balls chosen (i.e., painted and engraved) when
both Painter and Carver have finished their jobs. Since the tasks of Painter
and Carver do not depend on which task is done first, the distribution of X
is the same whether Painter does her job before or after Carver does his. If
Painter goes first, Carver chooses n balls from an urn that contains B black
balls and N − B white balls, so
B
N −B
i
n−i
for i = 0, 1, 2, . . . , min(n, B).
P (X = i) =
N
n
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 231 — #247
✐
✐
Section 5.3
Other Discrete Random Variables
231
On the other hand, if Carver goes first, Painter draws B balls from an urn that
contains n balls engraved with C and N − n balls not engraved, so
n N −n
i
B−i
for i = 0, 1, 2, . . . , min(n, B).
P (X = i) =
N
B
Thus changing the order of who goes first provides an urn-drawing explanation
of why the distribution of X remains the same when B is interchanged with n.
Let X be a hypergeometric random variable with parameters n, D, and N . Calculation of
E(X) and Var(X) directly from p(x), the probability mass function of X, is tedious. By the
methods developed in Sections 10.1 and 10.2, calculation of these quantities is considerably
easier. In Section 10.1 we will prove that E(X) = nD/N ; it can be shown that
n−1
nD(N − D) 1
−
.
Var(X) =
N2
N −1
(See Exercise 25, Section 10.2.) Using these formulas, for example, we have that the average
number of face cards in a bridge hand is (13 × 12)/52 = 3 with variance
13 − 1 13 × 12(52 − 12) 1
−
= 1.765.
522
52 − 1
Remark 5.2 If the n items that are selected at random from the D defective and N − D
nondefective items are chosen with replacement rather than without replacement, then X, the
number of defective items, is a binomial random variable with parameters n and D/N . Thus
!
D n−x
n D x 1−
,
x = 0, 1, 2, . . . , n.
P (X = x) =
N
x N
Now, if N is very large, it is not that important whether a sample is taken with or without replacement. For example, if we have to choose 1000 persons from a population of 200,000,000,
replacing persons that are selected back into the population will change the chance of other
persons to be selected only insignificantly. Therefore, for large N, the binomial probability
mass function is an excellent approximation for the hypergeometric probability mass function,
which can be stated mathematically as follows.
!
!
D
N −D
!
x
n−x
n x
!
lim
=
p (1 − p)n−x .
N →∞
x
N
D→∞
D/N →p
n
To prove this, note that, by Example 5.29,
!
!
D
N −D
x
n−x
!
=
N
n
!
N −n
!
D−x
n
! .
x
N
D
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 232 — #248
✐
✐
232
Chapter 5
Special Discrete Distributions
Therefore,
(N − n)!
D
N −D
N −n
(D
−
x)!
(N
− n − D + x)!
n
n
x
n−x
D−x
=
=
N
N
x
x
N!
n
D
D! (N − D)!
(N − D)!
(N − n)!
n
D!
·
·
=
N!
x (D − x)! (N − D − n + x)!
=
n D(D − 1) · · · (D − x + 1)(N − D)(N − D − 1) · · · (N − D − n + x + 1)
N (N − 1) · · · (N − n + 1)
x
n
=
x
D(D − 1) · · · (D − x + 1)(N − D)(N − D − 1) · · · (N − D − n + x + 1)
Nn
N (N − 1) · · · (N − n + 1)
Nn
D D − 1 D − x + 1 N − D N − D − 1 N − D − n + x + 1 ···
···
n
N
N
N
N
N
N
=
N − n + 1
N N − 1 N − 2 x
···
N
N
N
N
DD
1 D
x − 1 1 n − x − 1
D D
D
−
−
−
−
···
1−
1−
··· 1 −
n N N
N
N
N
N
N
N
N
N
=
x
2 n − 1
1 1−
··· 1 −
1−
N
N
N
Now as N → ∞, D → ∞, and D/N
! → p, the
! denominator approaches 1, and the numerator
D
N −D
!
x
n−x
n
!
approaches px (1 − p)n−x . So
approaches
px (1 − p)n−x .
x
N
n
EXERCISES
A
1.
Define a sample space for the experiment that, from a box containing two defective and
five nondefective items, three items are drawn at random and without replacement.
2.
Define a sample space for the experiment of performing independent Bernoulli trials
until the second success occurs.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 233 — #249
✐
✐
Section 5.3
3.
4.
Other Discrete Random Variables
233
An absentminded professor does not remember which of his 12 keys will open his office
door. If he tries them at random and with replacement:
(a)
On average, how many keys should he try before his door opens?
(b)
What is the probability that he opens his office door after only three tries?
The probability is p that Marty hits target M when he fires at it. The probability is q that
Alvie hits target A when he fires at it. Marty and Alvie fire one shot each at their targets.
If both of them hit their targets, they stop; otherwise, they will continue.
(a)
What is the probability that they stop after each has fired r times?
(b)
What is the expected value of the number of the times that each of them has fired
before stopping?
5.
Suppose that 20% of a group of people have hazel eyes. What is the probability that
the eighth passenger boarding a plane is the third one having hazel eyes? Assume that
passengers boarding the plane form a randomly chosen group.
6.
A certain basketball player makes a foul shot with probability 0.45. What is the probability that (a) his first basket occurs on the sixth shot; (b) his first and second baskets
occur on his fourth and eighth shots, respectively?
7.
A store has 50 light bulbs available for sale. Of these, five are defective. A customer buys
eight light bulbs randomly from this store. What is the probability that he finds exactly
one defective light bulb among them?
8.
The probability is p that a randomly chosen light bulb is defective. We screw a bulb into
a lamp and switch on the current. If the bulb works, we stop; otherwise, we try another
and continue until a good bulb is found. What is the probability that at least n bulbs are
required?
9.
Suppose that independent Bernoulli trials with parameter p are performed successively.
Let N be the number of trials needed to get x successes, and X be the number of
successes in the first n trials. Show that
P (N = n) =
x
P (X = x).
n
Remark: By this relation, in coin tossing, for example, we can state that the probability
of getting a fifth head on the seventh toss is 5/7 of the probability of five heads in seven
tosses.
10.
In rural Ireland, a century ago, the students had to form a line. The student at the front
of the line would be asked to spell a word. If he spelled it correctly, he was allowed to
sit down. If not, he received a whack on the hand with a switch and was sent to the end
of the line. Suppose that a student could spell correctly 70% of the words in the lesson.
What is the probability that the student would be able to sit down before receiving four
whacks on the hand? Assume that the master chose the words to be spelled randomly
and independently.
11.
The digits after the decimal point of a random number between 0 and 1 are numbers
selected at random, with replacement, independently, and successively from the set
{0, 1, . . . , 9}. In a random number from (0, 1), on the average, how many digits are
there before the fifth 3?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 234 — #250
✐
✐
234
Chapter 5
12.
On average, how many games of bridge are necessary before a player is dealt three aces?
A bridge hand is 13 randomly selected cards from an ordinary deck of 52 cards.
13.
Solve the Banach matchbox problem (Example 5.25) for the case where the matchbox
in the right pocket contains M matches, the matchbox in his left pocket contains N
matches, and m ≤ min(M, N ).
14.
For certain software, independently of other users, the probability is 0.07 that a user
encounters a fault. What are the chances that the 30th user is the 5th person encountering
a fault?
15.
In recent months, an electric utility company has received complaints from its customers
who think that, due to their unusually high electricity bills, their electric meters are
defective. The company begins to inspect the meters one by one. If only 3.6% of them
are defective, what is the probability that it takes at least 50 inspections until the 5th
defective electric meter is found?
16.
Suppose that 15% of the population of a town are senior citizens. Let X be the number
of nonsenior citizens who enter a mall before the tenth senior citizen arrives. Find the
probability mass function of X . Assume that each customer who enters the mall is a
random person from the entire population.
17.
Florence is moving and wishes to sell her package of 100 computer diskettes. Unknown
to her, 10 of those diskettes are defective. Sharon will purchase them if a random sample
of 10 contains no more than one defective disk. What is the probability that she buys
them?
1 .
Let X be a geometric random variable with parameter p, 0 < p < 1. Find E
1+X
In an annual charity drive, 35% of a population of 560 make contributions. If, in a
statistical survey, 15 people are selected at random and without replacement, what is the
probability that at least two persons have contributed?
18.
19.
Special Discrete Distributions
20.
The probability is p that a message sent over a communication channel is garbled. If
the message is sent repeatedly until it is transmitted reliably, and if each time it takes 2
minutes to process it, what is the probability that the transmission of a message takes
more than t minutes?
21.
A vending machine contains cans of grapefruit juice that cost 75 cents each, but it is not
working properly. The probability that it accepts a coin is 10%. Angela has a quarter
and five dimes. Determine the probability that she should try the coins at least 50 times
before she gets a can of grapefruit juice.
22.
A computer network consists of several stations connected by various media (usually
cables). There are certain instances when no message is being transmitted. At such “suitable instances,” each station will send a message with probability p, independently of
the other stations. However, if two or more stations send messages, a collision will corrupt the messages, and they will be discarded. These messages will be retransmitted
until they reach their destination. If the network consists of N stations, on average, how
many times should a certain station transmit and retransmit a message until it reaches its
destination?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 235 — #251
✐
✐
Section 5.3
Other Discrete Random Variables
235
B
23.
A fair coin is flipped repeatedly. What is the probability that the fifth tail occurs before
the tenth head?
24.
On average, how many independent games of poker are required until a preassigned
player is dealt a straight? (See Exercise 35 of Section 2.4 for a definition of a straight.
The cards have distinct consecutive values that are not of the same suit: for example, 3
of hearts, 4 of hearts, 5 of spades, 6 of hearts, and 7 of hearts.)
25.
Let X be a geometric random variable with parameter p, and n and m be nonnegative
integers.
(a)
For what values of n is P (X = n) maximum?
(b)
What is the probability that X is even?
(c)
Show that the geometric is the only distribution on the positive integers with the
memoryless property:
P (X > n + m | X > m) = P (X > n).
26.
In data communication, messages are usually combinations of characters, and each character consists of a number of bits. A bit is the smallest unit of information and is either 1
or 0. Suppose that the length of a character (in bits) is a geometric random variable with
parameter p. Furthermore, suppose that the lengths of characters are independent random variables. What is the distribution of the total number of the bits forming a message
of k random characters?
27.
Twelve hundred eggs, of which 200 are rotten, are distributed randomly in 100 cartons,
each containing a dozen eggs. These cartons are then sold to a restaurant. How many
cartons should we expect the chef of the restaurant to open before finding one without
rotten eggs?
28.
In the Banach matchbox problem, Example 5.25, find the probability that when the first
matchbox is emptied (not found empty) there are exactly m matches in the other box.
29.
In the Banach matchbox problem, Example 5.25, find the probability that the box which
is emptied first is not the one that is first found empty.
30.
Adam rolls a well-balanced die until he gets a 6. Andrew rolls the same die until he rolls
an odd number. What is the probability that Andrew rolls the die more than Adam does?
31.
To estimate the number of trout in a lake, we caught 50 trout, tagged and returned them.
Later we caught 50 trout and found that four of them were tagged. From this experiment
estimate n, the total number of trout in the lake.
Hint: Let pn be the probability of four tagged trout among the 50 trout caught. Find
the value of n that maximizes pn .
32.
Suppose that, from a box containing D defective and N − D nondefective items, n
(n ≤ D ) are drawn one by one, at random and without replacement.
(a)
Find the probability that the k th item drawn is defective.
(b)
If the (k − 1)st item drawn is defective, what is the probability that the k th item
is also defective?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 236 — #252
✐
✐
236
Chapter 5
Special Discrete Distributions
Self-Quiz on Section 5.3
Time allotted: 15 Minutes
Each problem is worth 5 points.
1.
A maximum of m (m ≥ 1) independent Bernoulli trials, each with probability of success p, 0 < p < 1, are performed successively. However, we stop the trials if the first
success occurs before the mth trial. Let Z be the number of experiments performed.
Find the probability mass function of Z .
2.
At the Antonio Car dealership, the probability that during a month at least one consumer
returns a car for warranty work is 0.45. Suppose that the numbers of cars returned for
warranty work in different months are independent. Consider the period until the fifth
month in which at least one car was returned for warranty work. What is the probability
that, during that period, there were exactly three months with no car returned for warranty work?
Hint: Call a month “success” if, in that month, at least one car is returned for warranty
work.
CHAPTER 5 SUMMARY
◮ Bernoulli Random Variables
Bernoulli trials are experiments that have only two possible outcomes. One outcome is usually called a success, denoted by s. The other outcome is
called a failure, denoted by f . The random variable defined by X(s) = 1 and X(f ) = 0 is
called a Bernoulli random variable. Therefore, a Bernoulli random variable takes on the value
1 when the outcome of the Bernoulli trial is a success and 0 when it is a failure. If p is the
probability of a success, then 1 − p (sometimes denoted q ) is the probability of a failure. Hence
the probability mass function of X is p(x) = 1 − p ≡ q if x = 0, p(x) = p if x = 1,
2
and p(x)
p = 0, otherwise. We have that E(X) = E(X ) = p, Var(X) = p(1 − p), and
σX = p(1 − p).
◮ Binomial Random Variables If n Bernoulli trials all with probability of success p are
performed independently, then X, the number of successes, is said to be a binomial random
variable with parameters n and p. The set of possible values
of X is {0, 1, 2, . . . , n}, and its
!
n x
probability mass function is p(x) = P (X = x) =
p (1 − p)n−x , if x = 0, 1, 2, . . . , n,
x
and p(x) = 0, elsewhere. For a binomial random
variable X with parameters n and p,
p
E(X) = np, Var(X) = np(1 − p), and σX = np(1 − p).
◮ Poisson Random Variables For a binomial random variable X with parameters n and
p, when the number of trials is large (n → ∞), the probability of success is small (p → 0), and
the average number of successes remains a fixed quantity of moderate value (np = λ for some
e−λ λi
. This approximation is called Poisson approximation
constant λ), then P (X = i) →
i!
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 237 — #253
✐
✐
Chapter 5
Summary
237
to binomial. Poisson’s approximation by itself is a probability mass function, and a discrete
random variable X with possible values 0, 1, 2, 3, . . . is called Poisson with parameter λ,
λ > 0, if its probability mass function is given by
P (X = i) =
e−λ λi
,
i!
i = 0, 1, 2, 3, . . . .
For a Poisson random variable with parameter λ, λ > 0, E(X) = Var(X) = λ and σX =
√
λ.
◮ Poisson Processes The Poisson distribution also appears in connection with the study
of sequences of random events occurring over time. Suppose that, starting from a time point
labeled t = 0, we begin counting the number of events. Then for each value of t, we obtain
a number denoted by N (t), which is the number of events that have occurred during [0, t]. If
random events occur in time in a way that the conditions stationarity, independent increments,
and orderliness are always satisfied, N (0) = 0 and, for all t > 0, 0 < P N (t) = 0 < 1,
then there exists a positive number λ such that
(λt)n e−λt
.
P N (t) = n =
n!
That is, for all t > 0, N (t) is a Poisson random variable with parameter λt. Hence E N (t) =
λt and therefore λ = E N (1) .
◮ Geometric Random Variables Suppose that a sequence of independent Bernoulli tri-
als, each with probability of success p, 0 < p < 1, are performed. Let X be the number
of experiments until the first success occurs. Then X is a discrete random variable called
geometric. Its set of possible values is {1, 2, . . .}, and its probability mass function is given
by p(x) = (1 − p)x−1 p, 0 < p < 1, x = 1, 2, 3, . . . , p(x) = 0, elsewhere. For X, a geomet2
ric random
√ variable with parameter p, 0 < p < 1, E(X) = 1/p, Var(X) = (1 − p)/p , and
σX = 1 − p/p.
◮ Negative Binomial Random Variables Suppose that a sequence of independent
Bernoulli trials, each with probability of success p, 0 < p < 1, is performed. Let X be
the number of experiments until the r th success occurs. Then X is a discrete random variable
called a negative binomial with parameters (r, p). Clearly, {r, r + 1, r + 2, r + 3, . . .} is the
set of possible values of X and its probability mass function is given by
!
x−1 r
p(x) =
p (1 − p)x−r , 0 < p < 1, x = r, r + 1, r + 2, r + 3, . . . .
r−1
For a negative binomial random variable X with parameters (r, p),
p
r(1 − p)
r
r(1 − p)
E(X) = , Var(X) =
.
, σX =
p
p2
p
◮ Hypergeometric Random Variables Suppose that, from a box containing D defective
and N − D nondefective items, n are drawn at random and without replacement. Furthermore,
suppose that the number of items drawn does not exceed the number of defective or the number
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 238 — #254
✐
✐
238
Chapter 5
Special Discrete Distributions
of nondefective items. That is, suppose that n ≤ min(D, N − D). Let X be the number of
defective items drawn. Then X is a discrete random variable with the set of possible values
{0, 1, . . . n}, and a probability mass function
!
!
D
N −D
x
n−x
!
p(x) = P (X = x) =
,
x = 0, 1, 2, . . . , n.
N
n
Any random variable X with such a probability mass function is called a hypergeometric random variable. For the hypergeometric random variable X, defined above,
E(X) =
nD
,
N
Var(X) =
n−1
nD(N − D) 1
−
.
N2
N −1
REVIEW PROBLEMS
1.
Of police academy applicants, only 25% will pass all the examinations. Suppose that 12
successful candidates are needed. What is the probability that, by examining 20 candidates, the academy finds all of the 12 persons needed?
2.
The time between the arrival of two consecutive customers at a post office is 3 minutes,
on average. Assuming that customers arrive in accordance with a Poisson process, find
the probability that tomorrow during the lunch hour (between noon and 12:30 P.M.) fewer
than seven customers arrive.
3.
A restaurant serves 8 fish entrées, 12 beef, and 10 poultry. If customers select from these
entrées randomly, what is the expected number of fish entrées ordered by the next four
customers?
4.
A university has n students, 70% of whom will finish an aptitude test in less than 25
minutes. If 12 students are selected at random, what is the probability that at most two
of them will not finish in 25 minutes?
5.
A doctor has five patients with migraine headaches. He prescribes for all five a drug that
relieves the headaches of 82% of such patients. What is the probability that the medicine
will not relieve the headaches of two of these patients?
6.
In a community, the chance of a set of triplets is 1 in 1000 births. Determine the probability that the second set of triplets in this community occurs before the 2000th birth.
7.
From a panel of prospective jurors, 12 are selected at random. If there are 200 men and
160 women on the panel, what is the probability that more than half of the jury selected
are women?
8.
Suppose that 10 trains arrive independently at a station every day, each at a random time
between 10:00 A.M. and 11:00 A.M.. What is the expected number and the variance of
those that arrive between 10:15 A.M. and 10:28 A.M.?
9.
Suppose that a certain bank returns bad checks at a Poisson rate of three per day. What is
the probability that this bank returns at most four bad checks during the next two days?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 239 — #255
✐
✐
Chapter 5
Review Problems
239
10.
The policy of the quality control division of a certain corporation is to reject a shipment
if more than 5% of its items are defective. A shipment of 500 items is received, 30 of
them are randomly tested, and two have been found defective. Should that shipment be
rejected?
11.
A fair coin is tossed successively until a head occurs. If N is the number of tosses
required, what are the expected value and the variance of N ?
12.
What is the probability that the sixth toss of a die is the first 6?
13.
In data communication, one method for error control makes the receiver send a positive acknowledgment for every message received with no detected error and a negative
acknowledgment for all the others. Suppose that (i) acknowledgments are transmitted
error free, (ii) a negative acknowledgment makes the sender transmit the same message
over, (iii) there is no upper limit on the number of retransmissions of a message, and
(iv) errors in messages are detected with probability p, independently of other messages.
On average, how many times is a message retransmitted?
14.
Past experience shows that 30% of the customers entering Harry’s Clothing Store will
make a purchase. Of the customers who make a purchase, 85% use credit cards. Let X
be the number of the next six customers who enter the store, make a purchase, and use a
credit card. Find the probability mass function, the expected value, and the variance of
X.
15.
Of the 28 professors in a certain department, 18 drive foreign and 10 drive domestic
cars. If five of these professors are selected at random, what is the probability that at
least three of them drive foreign cars?
16.
A bowl contains 10 red and six blue chips. What is the expected number of blue chips
among five randomly selected chips?
17.
A certain type of seed when planted fails to germinate with probability 0.06. If 40 of
such seeds are planted, what is the probability that at least 36 of them germinate?
18.
Suppose that 6% of the claims received by an insurance company are for damages from
vandalism. What is the probability that at least three of the 20 claims received on a
certain day are for vandalism damages?
19.
At the Antonio Car dealership, the probability that during a given month at least one
consumer returns a car for warranty work is 0.45. Suppose that the numbers of cars
returned for warranty work in different months are independent. Consider the period
until the fifth month in which at least one car was returned for warranty work. What
is the probability that, during that period, there were at least three months with no car
returned for warranty work?
20.
Passengers are making reservations for a particular flight on a small commuter plane 24
hours a day at a Poisson rate of 3 reservations per 8 hours. If 24 seats are available for
the flight, what is the probability that by the end of the second day all the plane seats are
reserved?
21.
An insurance company claims that only 70% of the drivers regularly use seat belts. In a
statistical survey, it was found that out of 20 randomly selected drivers, 12 regularly used
seat belts. Is this sufficient evidence to conclude that the insurance company’s claim is
false?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 240 — #256
✐
✐
240
Chapter 5
22.
From the set {x : 0 ≤ x ≤ 1} numbers are selected at random and independently and
rounded to three decimal places. What is the probability that 0.345 is obtained (a) for
the first time on the 1000th selections; (b) for the third time on the 3000th selections?
23.
The probability that a child of a certain family inherits a certain disease is 0.23 independently of other children inheriting the disease. If the family has five children and the
disease is detected in one child, what is the probability that exactly two more children
have the disease as well?
24.
A bowl contains w white and b blue chips. Chips are drawn at random and with replacement until a blue chip is drawn. What is the probability that (a) exactly n draws are
required; (b) at least n draws are required?
25.
Experience shows that 75% of certain kinds of seeds germinate when planted under
normal conditions. Determine the minimum number of seeds to be planted so that the
chance of at least five of them germinating is more than 90%.
26.
Suppose that n babies were born at a county hospital last week. Also suppose that the
probability of a baby having blonde hair is p. If k of these n babies are blondes, what is
the probability that the ith baby born is blonde?
27.
A farmer, plagued by insects, seeks to attract birds to his property by distributing seeds
over a wide area. Let λ be the average number of seeds per unit area, and suppose that
the seeds are distributed in a way that the probability of having any number of seeds in
a given area depends only on the size of the area and not on its location. Argue that X,
the number of seeds that fall on the area A, is approximately a Poisson random variable
with parameter λA.
28.
Show that if all three of n, N, and D → ∞ so that n/N → 0, D/N converges to a
small number, and nD/N → λ, then for all x,
!
!
D
N −D
x
n−x
e−λ λx
!
→
.
x!
N
Special Discrete Distributions
n
This formula shows that the Poisson distribution is the limit of the hypergeometric distribution.
Self-Test on Chapter 5
Time allotted: 120 Minutes
1.
Each problem is worth 10 points.
Fifty identical balls are numbered 1 through 50 and are mixed in a box. Five balls are
drawn randomly, one by one, and without replacement. Let X be the number of balls
drawn that are numbered at least 40. Let Y be the number on the 3rd ball drawn. Find
the probability mass functions of X and Y .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 241 — #257
✐
✐
Chapter 5
Self-Test Problems
241
2.
A company produces liquid silicone rubber that is designed specifically for mold making. If in the molding process, defective items are produced at a rate of 0.6%, find the
probability that, in a sample of 200 items, there are at most 3 defective ones.
3.
For certain software, independently of other users, the probability is 0.07 that a user
encounters a fault. Let X be the number of users who do not encounter a fault before
the 12th user who encounters a fault. Find the probability mass function of X .
Hint: Note that X + 12 is a negative binomial random variable.
4.
Ginny has begun a new website that aggregates reviews of movies and TV shows. Suppose that, all over the world, every day, 10 million people search for film review websites, and Ginny’s site is one of the last sites out of hundreds that a search turns up.
If independently of the others, the probability is 5 × 10−8 that a person visits Ginny’s
site via a search, find the probability that tomorrow at least 4 people visit her site by
searching the Web.
5.
To estimate the number of kangaroos living in a particular region of Western Victoria in
Australia, Liam captured 25 kangaroos, marked them and then let them go free. After
a few days he captured 30 kangaroos and observed that 5 of them were marked. Using
this experiment, estimate n, the total number of kangaroos in the area.
Hint: Let pn be the probability of 5 marked kangaroos among the 30 captured kangaroos. By analyzing pn /pn−1 , find the value of n that makes pn maximum.
6.
There are two fair coins on a table. The first coin is flipped successively and independently until a heads appears. Then the second coin is flipped repeatedly and independently until a heads appears. Let X be the number of flips until the first coin lands heads,
and let Y be the number of flips until the second coin lands heads. Find P (X = Y ).
7.
To reward improved product quality, a small manufacturing company that has only 27
employees, has developed a cash incentive program. Every year, the company pays out
a $7000 bonus on top of the salaries of the individuals who have performed at or above
certain thresholds. Suppose that, independently of other employees, each individual has
an 8% chance of achieving that level of performance. Find the minimum amount of
bonus money that the company should fund annually so that the probability is less than
1% that the amount is inadequate.
8.
Suppose that, independently of other tornado-force winds, with probability 0.063, such
a wind even damages tornado-proof fortified rooms that are constructed with cinder
blocks and are filled with mortar and rebar. Such a room is constructed in a town in the
plains located in a Tornado alley. If in this town, tornado-force winds occur at a Poisson
rate of 0.3 per year, what is the probability that this room can remain undamaged from
tornadoes for at least 70 years?
Hint: Let X be the number of tornado-force winds during the next 70 years.Observe
that X is a Poisson random variable with parameter λ = (70)(0.3) = 21. We will
show this fact rigorously in Chapter 11 (see Theorem 11.5).
9.
A gambler has two dice, an unbiased one, and a loaded one which, when tossed, lands
on 6 with probability of 4/9 and lands on any of the other faces with probability of 1/9.
The gambler does not know which die is the loaded one. He chooses one of the two dice
at random and tosses that die ten times in a row. If it lands on 6 exactly four times, what
is the probability that it is the loaded one?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 242 — #258
✐
✐
242
Chapter 5
10.
A company has insurance coverage to pay for its indirect cost each time it closes down
from a devastating calamity such as hurricane, tornado, flood, and snowstorm. Suppose
that the policy pays $50,000 for each calamity that causes the company to shut down
temporarily, except for the first calamity, for which it pays nothing. Furthermore, suppose that such calamities occur at a Poisson rate of 4/3 per year. Find the expected
amount paid to the company each year under this policy.
Special Discrete Distributions
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 243 — #259
✐
✐
Chapter 6
C ontinuous R andom
Variables
6.1
PROBABILITY DENSITY FUNCTIONS
As discussed in Section 4.2, the distribution function of a random variable X is a function F from (−∞, +∞) to R defined by F (t) = P (X ≤ t). From the definition of F
we deduced that it is nondecreasing, right continuous, and satisfies limt→∞ F (t) = 1 and
limt→−∞ F (t) = 0. Furthermore, we showed that, for discrete random variables, distributions
are step functions. We also proved that if X is a discrete random variable with set of possible values {x1 , x2 , . . .}, probability mass function p, and distribution function F, then F has
jump discontinuities at x1 , x2 , . . . , where the magnitude of the jump at xi is p(xi ) and for
xn−1 ≤ t < xn ,
n−1
X
F (t) = P (X ≤ t) =
p(xi ).
(6.1)
i=1
In the case of discrete random variables a very small change in t may cause relatively large
changes in the values of F . For example,
Pnfrom xn −ε to xn , ε > 0 being arbitrarily
Pn−1 if t changes
small, then F (x) changes from i=1 p(xi ) to i=1 p(xi ), a change of magnitude p(xn ),
which might be large. In cases such as the lifetime of a random light bulb, the arrival time of
a train at a station, and the weight of a random watermelon grown in a certain field, where the
set of possible values of X is uncountable, small changes in x produce correspondingly small
changes in the distribution of X . In such cases we expect that F, the distribution function of
X, will be a continuous function. Random variables that have continuous distributions can be
studied under general conditions. However, for practical reasons and mathematical simplicity,
we restrict ourselves to a class of random variables that are called absolutely continuous and
are defined as follows:
Definition 6.1
Let X be a random variable. Suppose that there exists a nonnegative realvalued function f : R → [0, ∞) such that for any subset of real numbers A that can be
constructed from intervals by a countable number of set operations,
Z
P (X ∈ A) =
f (x) dx.
(6.2)
A
Then X is called absolutely continuous or, in this book, for simplicity, continuous. Therefore,
whenever we say that X is continuous, we mean that it is absolutely continuous and hence
243
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 244 — #260
✐
✐
244
Chapter 6
Continuous Random Variables
satisfies (6.2). The function f is called the probability density function, or simply the density
function of X .
Let f be the probability density function of a random variable X with distribution function
F . Some immediate properties of f are as follows:
Rt
(a) F (t) = −∞ f (x) dx.
Let A = (−∞, t]; then using (6.2), we can write
Z
Z t
F (t) = P (X ≤ t) = P (X ∈ A) =
f (x) dx =
f (x) dx.
A
(b)
(c)
(d)
−∞
Comparing (a) with (6.1), we see that f, the probability density function of a continuous
random variable X, can be regarded as an extension of the idea of p, the probability mass
function of a discrete random variable. As we study f further, the resemblance between
f and p becomes clearer. For example, recall that if X is a discrete random
variable with
P
the set of possible values A and probability mass function p, then x∈A p(x) = 1. In
the continuous case, the relation analogous to this is the following:
R∞
f (x) dx = 1.
−∞
This is true by (a) and the fact that F (∞) = 1 i.e., limt→∞ F (t) = 1 . It also follows
from (6.2) with RA = R. Because of this property, in general, if a function g : R →
∞
[0, ∞) satisfies −∞ g(x) dx = 1, we say that g is a probability density function, or
simply a density function.
If f is continuous, then F ′ (x) = f (x). This follows from (a) and the fundamental
Z
d t
f (x) dx = f (t). Note that even if f is not continuous, still
theorem of calculus:
dt a
′
F (x) = f (x) for every x at which f is continuous. Therefore, the distribution functions
of continuous random variables are integrals of their derivatives.
Rb
For real numbers a ≤ b, P (a ≤ X ≤ b) = a f (t) dt.
Let A = [a, b] in (6.2); then
P (a ≤ X ≤ b) = P (X ∈ A) =
Z
A
f (x) dx =
Z b
f (x) dx.
a
Property (d) states that the probability of X being between a and b is equal to the area
under the graph of f from a to b. Letting a = b in (d), we obtain
Z a
P (X = a) =
f (x) dx = 0.
a
This means that for any real number a, P (X = a) = 0. That is, the probability that
a continuous random variable assumes a certain value is 0. From P (X = a) = 0,
∀a ∈ R, we have that the probability of X lying in an interval does not depend on the
endpoints of the interval. Therefore,
(e)
P (a < X < b) = P (a ≤ X < b) = P (a < X ≤ b) = P (a ≤ X ≤ b) =
Rb
a f (t) dt.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 245 — #261
✐
✐
Section 6.1
Probability Density Functions
245
Another implication of P (X = a) = 0, ∀a ∈ R, is that the value of the probability
density function f at no point represents a probability. The probabilistic significance of f is
that its integral over any subset of real numbers B gives the probability that X lies in B . In
particular,
The area over an interval I under the graph of f represents the probability
that the random variable X will belong to I . The area under f to the left of
a given point t is F (t), the value of the distribution function of X at t.
As an example, let X be a continuous random variable with probability density function f the
graph of which is sketched in Figure 6.1. Then the shaded area under f is the probability that
X is between a and b.
f
a
Figure 6.1
b
The shaded area under f is the probability that X ∈ I = (a, b).
Intuitively, f (a) is a measure that determines how likely it is for X to be close to a. To see
this, note that if ε > 0 is very small, then P (a − ε < X < a + ε) is the probability that X is
close to a. Now
Z a+ε
P (a − ε < X < a + ε) =
f (t) dt
a−ε
is the area under f (t) from a − ε to a + ε. This area is almost equal to the area of a rectangle
with sides of lengths (a + ε) − (a − ε) = 2ε and f (a). Thus
Z a+ε
P (a − ε < X < a + ε) =
f (t) dt ≈ 2εf (a).
(6.3)
a−ε
Similarly, for any other real number b in the domain of f,
P (b − ε < X < b + ε) ≈ 2εf (b).
(6.4)
Relations (6.3) and (6.4) show that for small fixed ε > 0, if f (a) < f (b), then
P (a − ε < X < a + ε) < P (b − ε < X < b + ε).
That is, if f (a) < f (b), the probability that X is close to b is higher than the probability that
X is close to a. Thus the larger f (x), the higher the probability will be that X is close to x.
Example 6.1 Experience has shown that while walking in a certain park, the time X, in
minutes, between seeing two people smoking has a probability density function of the form
(
λxe−x
x>0
f (x) =
0
otherwise.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 246 — #262
✐
✐
246
Chapter 6
(a)
Calculate the value of λ.
(b)
Find the distribution function of X .
(c)
What is the probability that Jeff, who has just seen a person smoking, will see another
person smoking in 2 to 5 minutes? In at least 7 minutes?
Continuous Random Variables
Solution:
(a)
To determine λ, we use the property
Z ∞
f (x) dx =
−∞
Z ∞
R∞
f (x) dx = 1:
−∞
−x
λxe
dx = λ
0
Now, by integration by parts,
Z
Thus
Z ∞
xe−x dx.
0
xe−x dx = −(x + 1)e−x .
h
i∞
λ − (x + 1)e−x
= 1.
(6.5)
0
But as x → ∞, using l’Hôpital’s rule, we get
lim (x + 1)e−x = lim
x→∞
(b)
x→∞
x+1
1
= lim x = 0.
x→∞ e
ex
Therefore, (6.5) implies that λ 0 − (−1) = 1 or λ = 1.
To find F, the distribution function of X, note that F (t) = 0 if t < 0. For t ≥ 0,
F (t) =
Z t
−∞
h
it
f (x) dx = − (x + 1)e−x = −(t + 1)e−t + 1.
0
Thus
F (t) =
(c)
(
0
if t < 0
−t
−(t + 1)e
+1
if t ≥ 0.
The desired probabilities P (2 < X < 5) and P (X ≥ 7) are calculated as follows:
P (2 < X < 5) = P (2 < X ≤ 5) = P (X ≤ 5) − P (X ≤ 2) = F (5) − F (2)
= (1 − 6e−5 ) − (1 − 3e−2 ) = 3e−2 − 6e−5 ≈ 0.37.
P (X ≥ 7) = 1 − P (X < 7) = 1 − P (X ≤ 7) = 1 − F (7)
= 8e−7 ≈ 0.007.
Note: To calculate
R ∞ these quantities, we can use P (2 < X < 5) =
P (X ≥ 7) = 7 f (t) dt as well. R5
2
f (t) dt and
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 247 — #263
✐
✐
Section 6.1
Probability Density Functions
247
Example 6.2
(a)
Sketch the graph of the function
1 − 1 |x − 3|
f (x) = 2 4
0
1≤x≤5
otherwise,
and show that it is the probability density function of a random variable X .
(b)
Find F, the distribution function of X, and sketch its graph.
(c)
Show that F is continuous.
Solution:
(a)
Note that
0
1 1
+ (x − 3)
2
4
f (x) =
1 1
− (x − 3)
2 4
0
x<1
1≤x<3
3≤x<5
x ≥ 5.
Therefore, the graph of f is as shown in Figure 6.2. Now since f (x) ≥ 0 and the area
under f from 1 to 5, being the area of the triangle ABC (as seen from Figure 6.2) is
(1/2)(4 × 1/2) = 1, f is a probability density function of some random variable X .
f (x)
B
1/2
Figure 6.2
(b)
A
C
1
5
2
3
4
x
Probability density function of Example 6.2.
To calculate the distribution function of X, we use the formula F (t) =
For t < 1,
Z t
F (t) =
0 dx = 0;
Rt
−∞ f (x) dx.
−∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 248 — #264
✐
✐
248
Chapter 6
Continuous Random Variables
for 1 ≤ t < 3,
F (t) =
Z t
f (x) dx =
−∞
1
for 3 ≤ t < 5,
Z t
Z th
i
1
1
1
1 1
+ (x − 3) dx = t2 − t + ;
2 4
8
4
8
Z 3h
Z th
i
i
1 1
1 1
+ (x − 3) dx +
− (x − 3) dx
2 4
2 4
−∞
1
3
1
1
5
21
1
5
17
= + − t2 + t −
= − t2 + t − ;
2
8
4
8
8
4
8
F (t) =
f (x) dx =
and for t ≥ 5,
Z t
Z 3h
Z 5h
i
i
1 1
1 1
F (t) =
f (x) dx =
+ (x − 3) dx +
− (x − 3) dx
2 4
2 4
−∞
1
3
1 1
= + = 1.
2 2
Therefore,
0
1
1
1
t2 − t +
4
8
F (t) = 8
1
5
17
− t2 + t −
8
4
8
1
t<1
1≤t<3
3≤t<5
t ≥ 5.
The graph of F is shown in Figure 6.3.
F(x)
1
1/2
x
1
Figure 6.3
2
3
4
5
Distribution function of Example 6.2.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 249 — #265
✐
✐
Section 6.1
(c)
Probability Density Functions
249
F is continuous because
lim F (t) = 0 = F (1),
t→1−
1
1
1
1
1
5
17
lim F (t) = (3)2 − (3) + = = − (3)2 + (3) −
= F (3),
8
4
8
2
8
4
8
5
17
1
= 1 = F (5). lim F (t) = − (5)2 + (5) −
t→5−
8
4
8
t→3−
Remark 6.1 Under what conditions is a distribution function F the distribution
function of a continuous random variable? The answer is that if F is a continuous function and if it is differentiable everywhere except possibly at a finite number of points, then it
is the distribution Z
function of a continuous random variable. To show this, let f (x) = F ′ (x).
x
f (t) dt; (ii) since F is a distribution function and hence nondecreasing,
Z ∞
f (t) dt = 1; (iv) for any
we must have f (x) ≥ 0; (iii) limt→∞ F (t) = 1 implies that
Then, (i) F (x) =
−∞
−∞
interval I = (a, b), [a, b), (a, b], or [a, b],
P (X ∈ I) = F (b) − F (a) =
Z b
f (x) dx,
a
and, as a result of (iv), if A is a subset of R that can be constructed from intervals by a countable
number of set operations, then
Z
P (X ∈ A) =
f (x) dx.
A
EXERCISES
1.
2.
When a certain car breaks down, the time that it takes to fix it (in hours) is a random
variable with the probability density function
(
ce−3x
if 0 ≤ x < ∞
f (x) =
0
otherwise.
(a)
Calculate the value of c.
(b)
Find the probability that when this car breaks down, it takes at most 30 minutes
to fix it.
The distribution function for the duration of a certain soap opera (in tens of hours) is
1 − 16
x≥4
x2
F (x) =
0
x < 4.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 250 — #266
✐
✐
250
3.
4.
5.
6.
Chapter 6
Continuous Random Variables
(a)
Calculate f, the probability density function of the soap opera.
(b)
Sketch the graphs of F and f .
(c)
What is the probability that the soap opera takes at most 50 hours? At least 60
hours? Between 50 and 70 hours? Between 10 and 35 hours?
The time it takes for a student to finish an aptitude test (in hours) has a probability
density function of the form
(
c(x − 1)(2 − x)
if 1 < x < 2
f (x) =
0
elsewhere.
(a)
Determine the constant c.
(b)
Calculate the distribution function of the time it takes for a randomly selected
student to finish the aptitude test.
(c)
What is the probability that a student will finish the aptitude test in less than 75
minutes? Between 1 12 and 2 hours?
The lifetime of a tire selected randomly from a used tire shop is 10, 000X miles, where
X is a random variable with the probability density function
2/x2
if 1 < x < 2
f (x) =
0
elsewhere.
(a)
What percentage of the tires of this shop last fewer than 15,000 miles?
(b)
What percentage of those having lifetimes fewer than 15,000 miles last between
10,000 and 12,500 miles?
The probability density function of a random variable X is given by
√ c
if −1 < x < 1
1 − x2
f (x) =
0
elsewhere.
(a)
Calculate the value of c.
(b)
Find the distribution function of X .
At a grocery store in Idaho, the demand for honey, in kilograms, per week, is a random
variable with probability density function
5
(4 − x)(4 + x)3
0<x<4
6656
f (x) =
0
otherwise.
How many kilograms of honey must the grocery store carry per week so that the stock
out probability for honey is at most 0.05?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 251 — #267
✐
✐
Section 6.1
7.
Probability Density Functions
251
Let X be a continuous random variable with probability density and distribution functions f and F, respectively. Assuming that α ∈ R is a point at which P (X ≤ α) < 1,
prove that
f (x)
h(x) = 1 − F (α)
0
if x ≥ α
if x < α
is also a probability density function.
8.
An insurance company has sold a large number of homeowner insurance policies in
an affluent area, where the insured value of homes is above $750,000. An actuary has
calculated that X, the insured value of a randomly selected home, in millions of dollars,
is a continuous random variable with probability density function
81
x > 3/4
5
f (x) = 64x
0
otherwise.
Of all homes that are insured for at least 1.5 million dollars, what percentage are insured
for less than 2 million dollars?
9.
Let X be a continuous random variable with probability density function f . We say that
X is symmetric about α if for all x,
P (X ≥ α + x) = P (X ≤ α − x).
(a)
Prove that X is symmetric about α if and only if for all x, we have f (α − x) =
f (α + x).
(b)
Show that X is symmetric about α if and only if f (x) = f (2α − x) for all x.
(c)
Let X be a continuous random variable with probability density function
2
1
f (x) = √ e−(x−3) /2 ,
2π
x ∈ R,
and Y be a continuous random variable with probability density function
1
,
g(x) = π 1 + (x − 1)2
x ∈ R.
Find the points about which X and Y are symmetric.
10.
The function
0
F (x) = cos x
1
x<π
π ≤ x < 2π
x ≥ 2π
is continuous, differentiable, nondecreasing, F (−∞) = 0, and F (∞) = 1. Is F the
distribution function of a continuous random variable?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 252 — #268
✐
✐
252
Chapter 6
11.
Suppose that the loss in a certain investment, in thousands of dollars, is a continuous
random variable X that has a probability density function of the form
(
k(2x − 3x2 )
−1 < x < 0
f (x) =
0
elsewhere.
12.
Continuous Random Variables
(a)
Calculate the value of k .
(b)
Find the probability that the loss is at most $500.
Let X denote the lifetime of a radio, in years, manufactured by a certain company. The
probability density function of X is given by
1 e−x/15
if 0 ≤ x < ∞
f (x) = 15
0
elsewhere.
What is the probability that, of eight such radios, at least four last more than 15 years?
13.
Let X be a continuous random variable with probability density function
1 − |x|
−1 < x < 1
f (x) =
0
otherwise.
Find the distribution function of X .
14.
Prove that if f and g are two probability density functions, then for α ≥ 0, β ≥ 0, and
α + β = 1, αf + βg is also a probability density function.
15.
The distribution function of a random variable X is given by
x
F (x) = α + β arctan ,
2
−∞ < x < ∞.
Determine the constants α and β and the probability density function of X .
Self-Quiz on Section 6.1
Time allotted: 15 Minutes
1.
Each problem is worth 5 points.
For what value of α is the following the probability density function of a random variable
X ? For that value of α, find P (1.5 < X < 2.5).
α/x4
1<x<3
f (x) =
0
otherwise.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 253 — #269
✐
✐
Section 6.2
2.
Density Function of a Function of a Random Variable
253
The distribution for the duration of a certain soap opera, in tens of hours, is
16
1 −
x≥4
x2
F (x) =
0
otherwise.
If the soap opera has been running for 110 hours, what is the probability that it will run
for at least another 110 hours?
6.2
DENSITY FUNCTION OF A FUNCTION OF A RANDOM VARIABLE
In the theory and applications of probability, we are often faced with situations in which f,
the probability density function of a random variable X, is known but the probability density function h(X), for some function of X, is needed. In such cases there are two common
methods for the evaluation of the probability density function of h(X) that we will explain in
this section. One method is to find the probability density of h(X) by calculating its distribution function. Another method is to find it directly from the probability density function of
X . We now explain the first method, called the method of distribution functions. To calculate the probability
density function of h(X), we calculate the distribution function of h(X),
P h(X) ≤ t , by finding a set A for which h(X) ≤ t if and only if X ∈ A and then
computing P (X ∈ A) from the distribution function of X . When P h(X) ≤ t is found, the
probability density function of h(X), if it exists, is obtained by differentiation. Some examples
follow.
Example 6.3
Let X be a continuous random variable with the probability density function
2/x2
if 1 < x < 2
f (x) =
0
elsewhere.
Find the distribution and the probability density functions of Y = X 2 .
Solution: Let G and g be the distribution function and the probability density function of Y,
respectively. By definition,
√
√
G(t) = P (Y ≤ t) = P (X 2 ≤ t) = P (− t ≤ X ≤ t )
0
t<1
√
= P (1 ≤ X ≤ t )
1≤t<4
1
t ≥ 4.
Now
P (1 ≤ X ≤
√
t) =
Z √t
1
h 2 i√t
2
2
=2− √ .
dx = −
2
x
x 1
t
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 254 — #270
✐
✐
254
Chapter 6
Continuous Random Variables
Therefore,
0
2
G(t) = 2 − √
t
1
t<1
1≤t<4
t ≥ 4.
The probability density function of Y, g, is found by differentiation of G:
1
√
if 1 ≤ t ≤ 4
′
g(t) = G (t) = t t
0
elsewhere. Example 6.4 Let X be a continuous random variable with distribution function F and probability density function f . In terms of f, find the distribution and the probability density functions of Y = X 3 .
Solution: Let G and g be the distribution function and the probability density function of Y .
Then
√
√
3
3
G(t) = P (Y ≤ t) = P (X 3 ≤ t) = P (X ≤ t ) = F ( t ).
Hence
√
√
1
1
3
3
F ′( t ) = √
g(t) = G′ (t) = √
f ( t ),
3 2
3 2
3 t
3 t
where g is defined at all points t 6= 0 at which f is defined.
Example 6.5
The error of a measurement has the probability density function
(
1/2
if −1 < x < 1
f (x) =
0
otherwise.
Find the distribution and the probability density functions of the magnitude of the error.
Solution: Let X be the error of the measurement. We want to find G, the distribution, and g,
the probability density functions of |X|, the magnitude of the error. By definition,
G(t) = P |X| ≤ t = P (−t ≤ X ≤ t) =
0
Z t
1
dx = t
2
−t
1
t<0
0≤t<1
t ≥ 1.
The probability density function of Y, g, is obtained by differentiation of G:
(
1
if 0 ≤ t ≤ 1
′
g(t) = G (t) =
0
elsewhere. ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 255 — #271
✐
✐
Section 6.2
Density Function of a Function of a Random Variable
255
In Examples 6.3, 6.4, and 6.5, the method used to find the probability density function
of h(X), a function of the random variable X, is to determine the distribution of h(X) and
then differentiate it. It is also possible to find the probability density function of h(X) without
obtaining its distribution function. We now explain this method, called the method of transformations.
Recall that a real-valued function h : R → R defined by y = h(x) is invertible if and only
if when solving y = h(x) for x, we find that x = h−1 (y) is itself a function. That is, y = h(x)
if and only if x = h−1 (y), where x = h−1 (y) is a function. For example, the function
√
y = h(x) = x3 is invertible because if we solve it for x, we obtain x = h−1 (y) = 3 y, which
is itself a function. The function y = h(x) = x2 is not invertible if its domain is (−∞, ∞).
√
When solving for x we obtain x = ± y, which is not a function. The function y = x2 will be
invertible if we restrict its domain to [0, ∞), the set of nonnegative numbers. This follows since
√
solving y = x2 for x gives x = y, which is a function. Other examples of invertible functions
x
are h(x) = e , x ∈ R; h(x) = ln x, x ∈ (0, ∞); and h(x) = sin x, x ∈ (0, π/2). Examples
of functions that are not invertible are h(x) = sin x, x ∈ R; h(x) = x2 + 5, x ∈ R; and
h(x) = exp(x2 ), x ∈ R.
In calculus, we have seen that a continuous function is invertible if and only if it is either
strictly increasing or strictly decreasing. If it is strictly increasing, its inverse is also strictly
increasing. If it is strictly decreasing, its inverse is strictly decreasing as well. We use these
to prove the next theorem, which enables us to find the probability density function of h(X)
mainly by finding the inverse of h. For this application, all we need to know is how to find the
inverse of a function, if it exists. Studying different criteria for invertibility does not serve our
purpose.
Theorem 6.1 (Method of Transformations)
Let X be a continuous random variable
with probability density function fX and the set of possible values A. For the invertible function
h
B = h(A) =
: A → R, let Y = h(X) be a random variable with the set of possible values
h(a) : a ∈ A . Suppose that the inverse of y = h(x) is the function x = h−1 (y), which is
differentiable for all values of y ∈ B . Then fY , the probability density function of Y, is given
by
fY (y) = fX h−1 (y) (h−1 )′ (y) ,
y ∈ B.
Proof: Let FX and FY be distribution functions of X and Y = h(X), respectively. Differentiability of h−1 implies that it is continuous. Since a continuous invertible function is strictly
monotone, h−1 is either strictly increasing or strictly decreasing. If it is strictly increasing,
(h−1 )′ (y) > 0, and hence (h−1 )′ (y) = (h−1 )′ (y) . Moreover, in this case h is also strictly
increasing, so
FY (y) = P h(X) ≤ y = P X ≤ h−1 (y) = FX h−1 (y) .
Differentiating this by chain rule, we obtain
FY′ (y) = (h−1 )′ (y)FX′ h−1 (y) = (h−1 )′ (y) fX h−1 (y) ,
which gives the theorem. If h−1 is strictly decreasing, (h−1 )′ (y) < 0 and hence
(h−1 )′ (y) = −(h−1 )′ (y). In this case h is also strictly decreasing and we get
FY (y) = P h(X) ≤ y = P X ≥ h−1 (y) = 1 − FX h−1 (y) .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 256 — #272
✐
✐
256
Chapter 6
Continuous Random Variables
Differentiating this by chain rule, we find that
FY′ (y) = −(h−1 )′ (y)FX′ h−1 (y) = (h−1 )′ (y) fX h−1 (y) ,
showing that the theorem is valid in this case as well.
Example 6.6
Let X be a random variable with the probability density function
(
2e−2x
if x > 0
fX (x) =
0
otherwise.
Using the method of transformations, find the probability density function of Y =
√
X.
Solution: The set of possible values of X is A = (0, ∞). Let h : (0, ∞) → R
√be defined by
√
h(x) = x. We want to find the probability
density function of Y = h(X) = X . The set of
possible values of h(X) is B = h(a) : a ∈ A = (0, ∞). The function h is invertible with
the inverse x = h−1 (y) = y 2 , which is differentiable, and its derivative is (h−1 )′ (y) = 2y .
Therefore, by Theorem 6.1,
2
fY (y) = fX h−1 (y) (h−1 )′ (y) = 2e−2y |2y|,
y ∈ B = (0, ∞).
Since y ∈ (0, ∞), |2y| = 2y and we find that
(
2
4ye−2y
fY (y) =
0
√
is the probability density function of Y = X .
Example 6.7
if y > 0
otherwise
Let X be a continuous random variable with the probability density function
(
4x3
if 0 < x < 1
fX (x) =
0
otherwise.
Using the method of transformations, find the probability density function of the random
variable Y = 1 − 3X 2 .
Solution: The set of possible values of X is A = (0, 1). Let h : (0, 1) → R be defined by
h(x) = 1 − 3x2 . We want to find the probability
density function of Y = h(X) = 1 − 3X 2 .
The set of possible values of h(X) is B = h(a) : a ∈ A = (−2, 1). Since the domain of
2
the function h is (0,
p 1), h is invertible, and its inverse is found by solving 1 − 3x = y for x,
which gives x = (1 − y)/3. Therefore, the inverse of h is
q
x = h−1 (y) = (1 − y)/3,
p
which is differentiable, and its derivative is (h−1 )′ (y) = −1 2 3(1 − y) . Using
Theorem 6.1, we find that
r 1 − y 3
−1 ′
1
2
−1
fY (y) = fX h (y) (h ) (y) = 4
− p
= (1 − y)
3
9
2 3(1 − y)
is the probability density function of Y when y ∈ (−2, 1).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 257 — #273
✐
✐
Section 6.2
Density Function of a Function of a Random Variable
257
EXERCISES
A
1.
Let X be a continuous random variable with the probability density function
(
1/4
if x ∈ (−2, 2)
f (x) =
0
otherwise.
Using the method of distribution functions, find the probability density functions of
Y = X 3 and Z = X 4 .
2.
A manufacturer produces metal disks with radii being random variables with the identical probability density function
5
5x4
x
−
if 3 < x < 4
f (x) = 431 2586
0
otherwise.
What is the probability that the area of a random disk is greater than 39?
3.
Let X be a continuous random variable with distribution function F and probability
density function f . Calculate the probability density function of the random variable
Y = eX .
4.
Let the probability density function of X be
(
e−x
if x > 0
f (x) =
0
elsewhere.
Using
√ the method of transformations, find the probability density functions of Y =
X X and Z = e−X .
5.
Let X be a continuous random variable with the probability density function
(
3e−3x
if x > 0
f (x) =
0
otherwise.
Using the method of transformations, find the probability density function of the random
variable Y = log2 X.
6.
Let the probability density function of X be
(
λe−λx
f (x) =
0
if x ≥ 0
otherwise,
for some λ > 0. Using the
method of distribution functions, calculate the probability
√
3
density function of Y = X 2 .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 258 — #274
✐
✐
258
Chapter 6
Continuous Random Variables
7.
Let f be the probability density function of a random variable X . In terms of f, calculate
the probability density function of X 2 .
8.
Let X be a random variable with the probability density function
f (x) =
1
,
π(1 + x2 )
−∞ < x < ∞.
(X is called a Cauchy random variable.) Find the probability density function of Z =
arctan X.
B
9.
Let X be a random variable with the probability density function given by
(
e−x
if x ≥ 0
f (x) =
0
elsewhere.
Let
Y =
X
1/X
if X ≤ 1
if X > 1.
Find the probability density function of Y .
Self-Quiz on Section 6.2
Time allotted: 20 Minutes
Each problem is worth 5 points.
1.
Let X be a random point from the interval (−π/2, π/2). Find the probability density
function of Y = tan X .
2.
Let X be a random variable with probability density function
(
2
xe−x /2
if x > 0
f (x) =
0
otherwise.
Using the method of transformations, find the probability density function of Y = X 2 .
Then calculate E(Y ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 259 — #275
✐
✐
Section 6.3
6.3
Expectations and Variances
259
EXPECTATIONS AND VARIANCES
Expectations of Continuous Random Variables
Let X be a continuous random variable with probability density and distribution functions f
and F, respectively. To define E(X), the average or expected value of X, first suppose that X
only takes values from the interval [a, b], and divide [a, b] into n subintervals of equal lengths.
Let h = (b − a)/n, x0 = a, x1 = a + h, x2 = a + 2h, . . . , xn = a + nh = b. Then
a = x0 < x1 < x2 < · · · < xn = b is a partition of [a, b]. Since F is continuous, let us
assume that it is differentiable on (a, b). By the mean-value theorem (of calculus), there exists
ti ∈ (xi−1 , xi ) such that
F (xi ) − F (xi−1 ) = F ′ (ti )(xi − xi−1 ),
1 ≤ i ≤ n,
or, equivalently,
P (xi−1 < X ≤ xi ) = f (ti )h,
1 ≤ i ≤ n.
(6.6)
If n is sufficiently large (n → ∞), then h, the width of the intervals, is sufficiently small
(h → 0), and f (x) does not vary appreciably over any subinterval of (xi−1 , xi ], 1 ≤ i ≤ n.
Thus ti is an approximate value of X in the interval (xi−1 , xi ]. Now
n
X
i=1
ti P (xi−1 < X ≤ xi )
finds the product of an approximate value of X when it is in (xi−1 , xi ] and the probability
that it is in (xi−1 , xi ], and then sums over all these intervals. From the concept of expectation
of a discrete
Pn random variable, it is clear that, as the lengths of these intervals get smaller and
smaller, i=1 ti P (xi−1 < X ≤ xi ) gets closer and closer to the “average” value of X . So it
is desirable to define E(X) as
lim
n→∞
n
X
i=1
ti P (xi−1 < X ≤ xi ).
Pn
But by (6.6) this is the same as limn→∞ i=1 ti f (ti )h, where this limit is equal to
Rb
xf (x) dx as known from calculus. If X is not restricted to an interval [a, b], a definition
a
for E(X) is motivated in the same way, but at the end the limit is taken as a → −∞ and
b → ∞.
Definition 6.2
If X is a continuous random variable with probability density function f,
the expected value of X is defined by
Z ∞
E(X) =
xf (x) dx.
−∞
The expected value of X is also called the mean, or mathematical expectation, or simply the
expectation of X, and as in the discrete case, sometimes it is denoted by EX, E[X], µ, or
µX .
To get a geometric feeling for mathematical expectation, consider a piece of cardboard of
uniform density on which the graph of the probability density function f of a random variable
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 260 — #276
✐
✐
260
Chapter 6
Continuous Random Variables
X is drawn. Suppose that the cardboard is cut along the graph of f and we are asked to balance
it on a given edge perpendicular to the x-axis. Then, to have it in equilibrium, we must balance
the cardboard on the given edge at the point x = E(X) on the x-axis.
Example 6.8 In a group of adult males, the difference between the uric acid value and 6, the
standard value, is a random variable X with the following probability density function:
27
(3x2 − 2x)
if 2/3 < x < 3
490
f (x) =
0
elsewhere.
Calculate the mean of these differences for the group.
Solution: By definition,
E(X) =
Z ∞
xf (x) dx =
−∞
=
Z 3
27
(3x3 − 2x2 ) dx
2/3 490
27 h 3 4 2 3 i3
283
=
x − x
= 2.36. 490 4
3
120
2/3
Remark 6.2 If X is a continuous random variable with probability density function f, X is
said to have a finite expected value if
Z ∞
|x|f (x) dx < ∞;
−∞
that is, X has a finite expected value if the integral of xf (x) converges absolutely. Otherwise,
we say that the expected value of X is not finite. We now justify why the absolute convergence
of the integral is required. Note that
E(X) =
Z ∞
xf (x) dx =
−∞
−∞
=−
(−x)f (x) dx +
−∞
xf (x) dx +
Z ∞
Z ∞
xf (x) dx
0
xf (x) dx,
0
R∞
(−x)f (x) dx ≥ 0 and 0 xf (x) dx ≥ 0. Thus E(X) is well defined if the
−∞
R0
R∞
integrals −∞ (−x)f (x) dx and 0 xf (x) dx are not both ∞. Moreover, E(X) < ∞ if
where
R0
Z 0
Z 0
neither of these two integrals is +∞. Hence a necessary and sufficient condition for E(X) to
R0
R∞
exist and to be finite is that −∞ (−x)f (x) dx < ∞ and 0 xf (x) dx < ∞. Since both of
these integrals are finite if and only if
Z ∞
−∞
|x|f (x) dx =
Z 0
−∞
(−x)f (x) dx +
Z ∞
xf (x) dx
0
R∞
is finite, we have that E(X) is well defined and finite if and only if the integral −∞ xf (x) dx
is absolutely convergent. ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 261 — #277
✐
✐
Section 6.3
Example 6.9
Expectations and Variances
261
A random variable X with probability density function
f (x) =
c
,
1 + x2
−∞ < x < ∞,
is called a Cauchy random variable.
(a)
Find c.
(b)
Show that E(X) does not exist.
Solution:
(a)
Since f is a probability density function
Z ∞
R∞
−∞
c
dx = c
2
−∞ 1 + x
f (x) dx = 1. Thus
Z ∞
dx
= 1.
2
−∞ 1 + x
Now
Z
dx
= arctan x.
1 + x2
Since the range of arctan x is (−π/2, +π/2), we get
Z ∞
h π π i
h
i∞
dx
=
c
=
c
arctan
x
− −
= cπ.
1=c
2
2
2
−∞
−∞ 1 + x
Thus c = 1/π .
(b)
To show that E(X) does not exist, note that
Z ∞
−∞
Z ∞
Z ∞
|x| dx
x dx
=
2
2
π(1 + x2 )
−∞ π(1 + x )
0
i∞
1h
= ∞. =
ln(1 + x2 )
π
0
|x|f (x) dx =
Remark 6.3 In this book, unless otherwise specified, it is implicitly assumed that the expected value of a random variable is finite. The following theorem directly relates the distribution function of a random variable to its
expected value. It enables us to find the expected value of a continuous random variable without
calculating its probability density function. It also has important theoretical applications.
Theorem 6.2 For any continuous random variable X with distribution function F and
probability density function f,
Z ∞
Z ∞
F (−t) dt.
1 − F (t) dt −
E(X) =
0
0
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 262 — #278
✐
✐
262
Chapter 6
Proof:
Note that
Continuous Random Variables
E(X) =
Z ∞
xf (x) dx =
−∞
=−
=−
Z 0 Z −x
−∞
0
Z ∞ Z −t
0
−∞
Z 0
−∞
Z ∞
Z ∞ Z ∞
xf (x) dx +
xf (x) dx
0
Z ∞ Z x dt f (x) dx +
dt f (x) dx
0
f (x) dx dt +
0
0
t
f (x) dx dt,
whereRthe last equality is obtainedRby changing the order of integration. The theorem follows
−t
∞
since −∞ f (x) dx = F (−t) and t f (x) dx = P (X > t) = 1 − F (t). Remark 6.4 In the proof of this theorem we assumed that the random variable X is continuous. Even without this condition the theorem is still valid. Also note that, since 1 − F (t) =
P (X > t), this theorem may be stated as follows.
For any random variable X,
Z ∞
Z ∞
E(X) =
P (X > t) dt −
P (X ≤ −t) dt.
0
0
In particular, if X is nonnegative, that is, P (X < 0) = 0, this theorem
states that
Z ∞
Z ∞
E(X) =
1 − F (t) dt =
P (X > t) dt. 0
0
As an important application of Theorem 6.2, we now prove the law of the unconscious
statistician, Theorem 4.2, for continuous random variables.
Theorem 6.3 Let X be a continuous random variable with probability density function
f (x); then for any function h : R → R,
Z ∞
E h(X) =
h(x)f (x) dx.
−∞
Proof: Let
h−1 (t, ∞) = x : h(x) ∈ (t, ∞) = x : h(x) > t
with similar representation for h−1 (−∞, −t). Notice
that we are not claiming that h has an
inverse function. We are simply considering the set x : h(x) ∈ (t, ∞) , which is called the
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 263 — #279
✐
✐
Section 6.3
Expectations and Variances
263
inverse image of (t, ∞) and is denoted by h−1 (t, ∞). By Theorem 6.2,
Z ∞
Z ∞
E h(X) =
P h(X) > t dt −
P h(X) ≤ −t dt
0
=
Z ∞
0
=
0
P X ∈ h−1 (t, ∞) dt −
Z ∞ Z
0
−
{x : x ∈ h−1 (t, ∞)}
Z ∞Z
Z ∞
0
P X ∈ h−1 (−∞, −t] dt
f (x) dx dt
f (x) dx dt
{x : x ∈ h−1 (−∞, −t]}
Z ∞ Z
Z ∞ Z
=
f (x) dx dt −
f (x) dx dt.
0
{x : h(x) > t}
0
{x : h(x) ≤ −t}
0
Now we change the order of integration for both of these double integrals. Since
(t, x) : 0 < t < ∞, h(x) > t = (t, x) : h(x) > 0, 0 < t < h(x) ,
and
we get
(t, x) : 0 < t < ∞, h(x) ≤ −t = (t, x) : h(x) < 0, 0 < t ≤ −h(x) ,
E h(X) =
Z
Z h(x)
dt f (x) dx
{x : h(x) > 0} 0
Z
Z −h(x) −
dt f (x) dx
{x : h(x) < 0} 0
Z
Z
=
h(x)f (x) dx +
h(x)f (x) dx
{x : h(x) > 0}
{x : h(x) < 0}
Z ∞
=
h(x)f (x) dx.
−∞
Note that the last equality follows because
Z
{x : h(x) = 0}
h(x)f (x) dx = 0.
Corollary Let X be a continuous random variable with probability density function f (x).
Let h1 , h2 , . . . , hn be real-valued functions, and α1 , α2 , . . . , αn be real numbers. Then
E α1 h1 (X) + α2 h2 (X) + · · · + αn hn (X)
= α1 E h1 (X) + α2 E h2 (X) + · · · + αn E hn (X) .
Proof: In the discrete case, in the proof of the corollary of Theorem 4.2, replace
and p(x) by f (x) dx. P
by
R
,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 264 — #280
✐
✐
264
Chapter 6
Continuous Random Variables
Just as in the discrete case, by this corollary we can write that, for example,
E(3X 4 + cos X + 3eX + 7) = 3E(X 4 ) + E(cos X) + 3E(eX ) + 7.
Moreover, this corollary implies that if α and β are constants, then
E(αX + β) = αE(X) + β.
Example 6.10 A point X is selected from the interval (0, π/4) randomly. Calculate
E(cos 2X) and E(cos2 X).
Solution: First we calculate the distribution function of X . Clearly,
0
t<0
π
4t
t−0
0≤t<
=
F (t) = P (X ≤ t) = π
π
4
−0
4
π
1
t≥ .
4
Thus f, the probability density function of X, is
4
π
0<t<
π
4
f (t) =
0
otherwise.
Now, by Theorem 6.3,
E(cos 2X) =
Z π/4
0
2
h2
iπ/4
4
2
cos 2x dx =
sin 2x
= .
π
π
π
0
To calculate E(cos X), note that cos2 X = (1 + cos 2X)/2. So by the corollary of
Theorem 6.3,
1 1
1 1
1 1 2
1 1
+ cos 2X = + E(cos 2X) = + · = + . E(cos2 X) = E
2 2
2 2
2 2 π
2 π
Variances of Continuous Random Variables
Definition 6.3
If X is a continuous random variable with E(X) = µ, then Var(X) and
σX , called the variance and standard deviation of X, respectively, are defined by
Var(X) = E (X − µ)2 ,
σX =
q E (X − µ)2 .
Therefore, if f is the probability density function of X, then by Theorem 6.3,
Z ∞
Var(X) = E (X − µ)2 =
(x − µ)2 f (x) dx.
−∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 265 — #281
✐
✐
Section 6.3
Expectations and Variances
265
Also, as before, we have the following important relations whose proofs are analogous to those
in the discrete case.
2
Var(X) = E(X 2 ) − E(X) ,
Var(aX + b) = a2 Var(X), σaX+b = |a|σX , a and b being constants.
Example 6.11 The time elapsed, in minutes, between the placement of an order of pizza
and its delivery is random with the probability density function
if 25 < x < 40
1/15
f (x) =
0
otherwise.
(a)
Determine the mean and standard deviation of the time it takes for the pizza shop to
deliver pizza.
(b)
Suppose that it takes 12 minutes for the pizza shop to bake pizza. Determine the mean
and standard deviation of the time it takes for the delivery person to deliver pizza.
Solution:
(a)
(b)
Let the time elapsed between the placement of an order and its delivery be X minutes.
Then
Z 40
1
E(X) =
x dx = 32.5,
15
25
Z 40
1
E(X 2 ) =
x2 dx = 1075.
15
25
p
Therefore, Var(X) = 1075 − (32.5)2 = 18.75, and hence σX = Var(X) = 4.33.
The time it takes for the delivery person to deliver pizza is X −12. Therefore, the desired
quantities are
E(X − 12) = E(X) − 12 = 32.5 − 12 = 20.5
σX−12 = |1|σX = σX = 4.33. Remark 6.5 For a continuous random variable X, the moments, absolute moments, moments about a constant c, and central moments are all defined in a manner similar to those of
Section 4.5 (the discrete case). ⋆ Remark 6.6 In Remark 4.2, we showed that, for a discrete random variable X and
positive integer n, if E(X n+1 ) exists, then E(X n ) also exists. That is, the existence of higher
moments of a discrete random variable implies the existence of its lower moments. In particular,
the existence of E(X 2 ) implies the existence of E(X) and, hence, Var(X). These facts are
also true for continuous random variables, and their proofs are similar. Let X be a continuous
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 266 — #282
✐
✐
266
Chapter 6
Continuous Random Variables
random variable with probability density function f (x). Let n be a positive integer. We will
show that if E(X n+1 ) exists, then E(X n ) also exists. Clearly,
Z
Z 1
−1
|x|>1
n
|x| f (x) dx ≤
|x|n f (x) dx ≤
Z 1
Z
−1
f (x) dx ≤
|x|>1
Z ∞
f (x) dx = 1;
−∞
|x|n+1 f (x) dx ≤
Z ∞
−∞
|x|n+1 f (x) dx.
These inequalities yield
Z ∞
−∞
n
|x| f (x) dx =
Therefore,
Z ∞
−∞
Z 1
−1
n
|x| f (x) dx +
Z
n
|x|>1
|x| f (x) dx ≤ 1 +
|x|n+1 f (x) dx < ∞ implies that
Z ∞
−∞
Z ∞
−∞
|x|n+1 f (x) dx.
|x|n f (x) dx < ∞. By Remark 6.2,
this shows that the existence of E(X n+1 ) implies the existence of E(X n ).
⋆ Remark 6.7 Random Variables that are neither Discrete nor Continuous
Here we discuss one such random variable. Examples 4.7 and 4.8 present two more. At a bank,
when customers arrive, they take their place in a waiting line for service if all of the servers are
busy. If there is no queue and a server is free, then customers will be served immediately upon
arrival. Otherwise, the customer waits in the line until his or her turn for service. Let W be the
waiting time of a randomly selected customer. Then W is not a continuous random variable,
since P (W = 0) > 0 and is not a discrete random variable, since its set of possible values is
the interval [0, ∞), a set, which is not countable. W can take any value in [0, ∞). Suppose that
the probability is p that a customer who arrives at the bank finds no queue and a free server.
Furthermore, suppose that if the customer has to wait to be served, then for some λ > 0, his or
her waiting time is a continuous random variable, X, with probability density function
λe−λx
if x ≥ 0
f (x) =
0
otherwise.
Clearly,
W =
0
X
with probability p
with probability 1 − p.
Random variables such as W are called mixtures of discrete and continuous random
variables. In general, let
(
Y with probability of p
Z=
X with probability 1 − p.
Then, by the methods of Section 10.4, we can show that
E(Z) = pE(Y ) + (1 − p)E(X).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 267 — #283
✐
✐
Section 6.3
Expectations and Variances
267
(See Exercise 4, Section 10.4.) Applying this to W, we obtain
Z ∞
E(W ) = p · 0 + (1 − p)E(X) = (1 − p)
x · λe−λx dx.
0
Using integration by parts with u = x and dv = λe−λx dx, we have
Z ∞
i∞
i∞ Z ∞
h1
h
1
−λx
−λx
= .
+
e−λx dx = 0 − e−λx
x · λe
dx = − xe
λ
λ
0
0
0
0
Therefore, E(W ) =
1−p
.
λ
Let us now calculate P (W ≤ 5); that is, the probability that a customer’s waiting time at
the bank is at most 5. By the law of total probability,
P (W ≤ 5) = P (W ≤ 5 | W = 0)P (W = 0) + P (W ≤ 5 | W > 0)P (W > 0)
Z 5
=1·p+
λe−λx dx · (1 − p)
0
= p + (1 − p)(1 − e−5λ ).
It should be noted that there are also random variables that are neither discrete nor continuous, and they cannot be expressed as a mixture of discrete and continuous random variables.
Discussion of such random variables is beyond the scope of this book. EXERCISES
A
1.
2.
The distribution function for the duration of a certain soap opera (in tens of hours) is
16
1 −
if x ≥ 4
x2
F (x) =
0
if x < 4.
(a)
Find E(X).
(b)
Show that Var(X) does not exist.
The time it takes for a student to finish an aptitude test (in hours) has the probability
density function
(
6(x − 1)(2 − x)
if 1 < x < 2
f (x) =
0
otherwise.
Determine the mean and standard deviation of the time it takes for a randomly selected
student to finish the aptitude test.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 268 — #284
✐
✐
268
Chapter 6
Continuous Random Variables
3.
The mean and standard deviation of the lifetime of a car muffler manufactured by company A are 5 and 2 years, respectively. These quantities for car mufflers manufactured
by company B are, respectively, 4 years and 18 months. Brian buys one muffler from
company A and one from company B. That of company A lasts 4 years and 3 months,
and that of company B lasts 3 years and 9 months. Determine which of these mufflers
has performed relatively better.
Hint: Find the standardized lifetimes of the mufflers and compare (see Section 4.6).
4.
At a university, which does not qualify as a tax-exempt entity under federal and state
laws, each year, without taxes, it costs an average of $56,000 with a variance of 81
million dollars to repair and maintain all the dormitories. If all items associated with
the maintenance and repair of the dormitories are subject to 8.875 percent tax, find the
standard deviation of the total cost of maintaining and repairing the dormitories, in a
random year, after all the taxes are paid.
5.
An actuary of an insurance company has discovered that the company’s total monthly
claims, in hundred thousands of dollars, is a continuous random variable proportional to
1
, for 0 < x < ∞. Find the expected value of the total amount of next
(1 + x)(1 + x2 )
month’s claims.
6.
A random variable X has the probability density function
(
3e−3x
if 0 ≤ x < ∞
f (x) =
0
otherwise.
Calculate E(eX ).
7.
8.
Find the expected value of a random variable X with the probability density function
√ 1
if −1 < x < 1
f (x) = π 1 − x2
0
otherwise.
Let Y be a continuous random variable with distribution function
(
e−k(α−y)/A
−∞ < y ≤ α
F (y) =
1
y > α,
where A, k, and α are positive constants. (Such distribution functions arise in the study
of local computer network performance.) Find E(Y ).
9.
Let the probability density function of tomorrow’s Celsius temperature be h. In terms
of h, calculate the corresponding probability density function and its expected value for
Fahrenheit temperature.
Hint: Let C and F be tomorrow’s temperature in Celsius and Fahrenheit, respectively.
Then F = 1.8C + 32.
10.
Find E(ln X) if X is a continuous random variable with probability density function
2/x2
if 1 < x < 2
f (x) =
0
elsewhere.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 269 — #285
✐
✐
Section 6.3
11.
Expectations and Variances
269
A right triangle has a hypotenuse of length 9. If the probability density function of one
side’s length is given by
x/6
if 2 < x < 4
f (x) =
0
otherwise,
what is the expected value of the length of the other side?
12.
Let X be a random variable with probability density function
1
f (x) = e−|x| ,
2
−∞ < x < ∞.
Calculate Var(X).
B
13.
Let X be a random variable with the probability density function
f (x) =
14.
1
,
π(1 + x2 )
−∞ < x < ∞.
Prove that E |X|α converges if 0 < α < 1 and diverges if α ≥ 1.
Suppose that X, the interarrival time between two customers entering a certain post
office, satisfies
P (X > t) = αe−λt + βe−µt ,
t ≥ 0,
where α + β = 1, α ≥ 0, β ≥ 0, λ > 0, µ > 0. Calculate the expected value of X .
Hint: For a fast calculation, use Remark 6.4.
15.
For n ≥ 1, let Xn be a continuous random variable with the probability density function
c
n
if x ≥ cn
n+1
x
fn (x) =
0
otherwise.
Xn ’s are called Pareto random variables and are used to study income distributions.
(a)
(b)
(c)
(d)
16.
Calculate cn , n ≥ 1.
Find E(Xn ), n ≥ 1.
Determine the probability density function of Zn = ln Xn , n ≥ 1.
For what values of m does E(Xnm+1 ) exist?
Let X be a continuous random variable with the probability density function
1
x sin x
if 0 < x < π
f (x) = π
0
otherwise.
Prove that
E(X n+1 ) + (n + 1)(n + 2)E(X n−1 ) = π n+1 .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 270 — #286
✐
✐
270
Chapter 6
17.
Let X be a continuous random variable with probability density function f . A number
t is said to be the median of X if
Continuous Random Variables
P (X ≤ t) = P (X ≥ t) =
1
.
2
By Exercise 9, Section 6.1, X is symmetric about α if and only if for all x we have
f (α − x) = f (α + x). Show that if X is symmetric about α, then
E(X) = Median(X) = α.
18.
19.
Let X be a continuous random variable with
probability density function f (x). Determine the value of y for which E |X − y| is minimum.
Let X be a nonnegative random variable with distribution function F . Define
(
1
if X > t
I(t) =
0
otherwise.
R∞
I(t) dt = X .
(a)
Prove that
(b)
By calculating the expected value of both sides of part (a), prove that
Z ∞
E(X) =
1 − F (t) dt.
0
0
This is a special case of Theorem 6.2.
(c)
For r > 0, use part (b) to prove that
Z ∞
r
E(X ) = r
tr−1 1 − F (t) dt.
0
20.
Let X be a continuous random variable. Prove that
∞
X
n=1
∞
X
P |X| ≥ n ≤ E |X| ≤ 1 +
P |X| ≥ n .
n=1
These important inequalities show that E |X| < ∞ if and only if the series
P∞
n=1 P |X| ≥ n converges.
Hint: By Exercise 19,
E |X| =
21.
Z ∞
0
∞
X
P |X| > t dt =
n=0
Z n+1
n
P |X| > t dt.
Note that on the interval [n, n + 1),
P |X| ≥ n + 1 < P |X| > t ≤ P |X| ≥ n .
Let X be the random variable introduced in Exercise 14. Applying the results of
Exercise 19, calculate Var(X ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 271 — #287
✐
✐
Section 6.3
22.
Expectations and Variances
271
Suppose that X is the lifetime of a randomly selected fan used in certain types of diesel
engines. Let Y be a randomly selected competing fan for the same type of diesel engines manufactured by another company. To compare the lifetimes X and Y, it is not
sufficient to compare E(X) and E(Y ). For example, E(X) > E(Y ) does not necessarily imply that the first manufacturer’s fan outlives the second manufacturer’s fan.
Knowing Var(X) and Var(Y ) will help, but variance is also a crude measure. One of the
best tools for comparing random variables in such situations is stochastic comparison.
Let X and Y be two random variables. We say that X is stochastically larger than Y,
denoted by X ≥st Y, if for all t,
P (X > t) ≥ P (Y > t).
Show that if X ≥st Y, then E(X) ≥ E(Y ), but not conversely.
Hint: Use Theorem 6.2.
23.
Let X be a continuous random
R ∞ variable with probability density function f . Show that
if E(X) exists; that is, if −∞ |x|f (x) dx < ∞, then
lim xP (X ≤ x) = lim xP (X > x) = 0.
x→−∞
x→∞
Self-Quiz on Section 6.3
Time allotted: 20 Minutes
1.
2.
Each problem is worth 5 points.
Suppose that the weekly profit of a certain store, in thousands of dollars, is a random
variable X with probability density function
3
(x + x2 )
if 0 < x < 2
f (x) = 14
0
otherwise.
(i)
What is the expected weekly profit of this store?
(ii)
What is the probability that for 8 consecutive weeks, every week, the store’s profit
is less than $500? Assume that the profit in one week is independent of the profits
in other weeks.
Let X be the value of the claims made, in one year, in millions of dollars, for partial
losses from fires that damage homes but do not destroy them. Suppose that the probability density function of X is given by
1
(2 − x)3
if 0 < x < 2
f (x) = 4
0
otherwise.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 272 — #288
✐
✐
272
Chapter 6
Continuous Random Variables
(i)
Find E(X), Var(X), and σX .
(ii)
Calculate the probability that the claims made in a certain year exceed 1,200,000,
if we know that they exceed 800,000.
CHAPTER 6 SUMMARY
◮ Continuous Random Variables Let X be a random variable. Suppose that there exists
a nonnegative real-valued function f : R → [0, ∞) such that for any subset of real numbers A
that can be constructed from intervals by a countable number of set operations,
Z
P (X ∈ A) =
f (x) dx.
A
Then X is called continuous. The function f is called the probability density function, or simply
the density function of X . Let F be the distribution function
properties
R t of X . Some immediate
R∞
of the probability density function f are: (i) F (t) = −∞ f (x) dx; (ii) −∞ f (x) dx = 1;
(iii) If f is continuous, then F ′ (x) = f (x); (iv) For real numbers a ≤ b, P (a ≤ X ≤ b) =
Rb
f (t) dt; (v) for any real number a, P (X = a) = 0; and
a
Rb
(vi) P (a < X < b) = P (a ≤ X < b) = P (a < X ≤ b) = P (a ≤ X ≤ b) = a f (t) dt.
◮ The area over an interval I under the graph of f represents the probability that the random
variable X will belong to I . The area under f to the left of a given point t is F (t), the value of
the distribution function of X at t.
◮ Density Function of a Function of a Random Variable Suppose that f, the probability density function of a random variable X, is known but the probability density function
h(X), for some function of X, is needed. In such cases there are two common methods for
the evaluation of the probability density function of h(X). In a method called the method of
distribution functions, to calculate the probability
density function of h(X), we calculate the
distribution function of h(X), P h(X) ≤ t , by finding a set E for which h(X) ≤ t if and
only if X ∈ E
and then computing P (X ∈ E) from the distribution function of X . When
P h(X) ≤ t is found, the probability density function of h(X), if it exists, is obtained by
differentiation.
In the second method, called the method of transformations, we find the probability density
function of h(X) without obtaining its distribution function. Let fX be the probability density
function of X, and let its set of possible values be A. For the invertible functionh : A → R, let
Y = h(X) be a random variable with the set of possible values B = h(A) = h(a) : a ∈ A .
Suppose that the inverse of y = h(x) is the function x = h−1 (y), which is differentiable for
all values of y ∈ B . Then fY , the probability density function of Y, is given by
fY (y) = fX h−1 (y) (h−1 )′ (y) ,
y ∈ B.
◮ Expectations of Continuous Random Variables If X is a continuous random variable with probability density function f, the expected value of X is defined by
Z ∞
xf (x) dx.
E(X) =
−∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 273 — #289
✐
✐
Chapter 6
Summary
273
R∞
X is said to have a finite expected value if −∞ |x|f (x) dx < ∞; that is, X has a finite
expected value if the integral of xf (x) converges absolutely.
◮ For any continuous random variable X with distribution function F and probability density function f,
E(X) =
Z ∞
0
E(X) =
Z ∞
0
1 − F (t) dt −
P (X > t) dt −
Z ∞
F (−t) dt,
0
Z ∞
0
P (X ≤ −t) dt.
◮ If X is a nonnegative continuous random variable with distribution function F, that is, if
P (X < 0) = 0, then
Z ∞
Z ∞
P (X > t) dt.
1 − F (t) dt =
E(X) =
0
0
◮ Let X be a continuous random variable with probability density function f (x); then for
any function h : R → R,
E h(X) =
Z ∞
h(x)f (x) dx.
−∞
◮ Let h1 , h2 , . . . , hn be real-valued functions, and α1 , α2 , . . . , αn be real numbers. Then
E α1 h1 (X) + α2 h2 (X) + · · · + αn hn (X)
= α1 E h1 (X) + α2 E h2 (X) + · · · + αn E hn (X) .
This relation implies the linearity property of continuous random variables. That is, if α and β
are constants, then E(αX + β) = αE(X) + β .
◮ Variances of Continuous Random Variables If X is a continuous random variable
with E(X) = µ, then Var(X) and σX , called the variance and standard deviation of X,
respectively, are defined by
Var(X) = E (X − µ)2 ,
σX =
q E (X − µ)2 .
As in discrete case, we have the following relations:
2
Var(X) = E(X 2 ) − E(X) ,
Var(aX + b) = a2 Var(X), σaX+b = |a|σX , a and b being constants.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 274 — #290
✐
✐
274
Chapter 6
Continuous Random Variables
REVIEW PROBLEMS
1.
Let X be a random number from (0, 1). Find the probability density function of the
random variable Y = 1/X .
2.
Let X be a continuous random variable with the probability density function
2/x3
if x > 1
f (x) =
0
otherwise.
Find E(X) and Var(X) if they exist.
3.
Let X be a continuous random variable with probability density function
f (x) = 6x(1 − x),
0 < x < 1.
What is the probability that X is within two standard deviations of the mean?
4.
Let X be a random variable with probability density function
f (x) =
e−|x|
,
2
−∞ < x < ∞.
Find P (−2 < X < 1).
5.
6.
Does there exist a constant c for which the following is a probability density function?
c
if x > 0
f (x) = 1 + x
0
otherwise.
Let X be a random variable with probability density function
4x3 /15
1≤x≤2
f (x) =
0
otherwise.
Find the probability density functions of Y = eX , Z = X 2 , and W = (X − 1)2 .
7.
8.
The probability density function of a continuous random variable X is
(
30x2 (1 − x)2
if 0 < x < 1
f (x) =
0
otherwise.
Find the probability density function of Y = X 4 .
n
Pn
Prove or disprove: If i=1 αi = 1, αi ≥ 0, ∀i, and fi i=1 is a sequence of probability
Pn
density functions, then i=1 αi fi is a probability density function.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 275 — #291
✐
✐
Chapter 6
9.
Self-Test Problems
275
Let X be a continuous random variable with set of possible values {x : 0 < x < α}
(where α < ∞), distribution function F, and probability density function f . Using
integration by parts, prove the following special case of Theorem 6.2.
Z α
E(X) =
1 − F (t) dt.
0
10.
The lifetime (in hours) of a light bulb manufactured by a certain company is a random
variable with probability density function
if x ≤ 500
0
f (x) =
5
5 × 10
if x > 500.
x3
Suppose that, for all nonnegative real numbers a and b, the event that any light bulb lasts
at least a hours is independent of the event that any other light bulb lasts at least b hours.
Find the probability that, of six such light bulbs selected at random, exactly two last over
1000 hours.
Self-Test on Chapter 6
Time allotted: 120 Minutes
Each problem is worth 10 points.
1.
Let X be a strictly positive and continuous random variable with distribution function
F and probability density function f . Find the distribution and the probability density
functions of Y = log2 X.
2.
When Kara throws a dart at a dartboard, X, the distance, in centimeters, between the
impact point of the dart and O, the center of the dartboard, is a continuous random
variable with probability density function
3
x(10 − x)
0 < x < 10
f (x) = 500
0
otherwise.
If the bull’s eye of the dartboard is a circle of diameter 6 centimeters centered at O, and
the probability of hitting the bull’s eye in one throw is independent of hitting it in other
throws, what is the probability that Kara misses the bull’s eye in exactly 7 of her next 10
throws?
3.
An actuary has calculated that the probability density function of the cost to repair a
vehicle after a certain type of car accident, in thousands of dollars, is
x2 /9
if 0 < x < 3
f (x) =
0
otherwise.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 276 — #292
✐
✐
276
Chapter 6
Continuous Random Variables
The actuary has also calculated that, for all auto insurance policies with the same deductible amount d, in 7.24% of the time, the cost of such a repair to the insurance
company is less than $1000. Calculate d.
4.
5.
6.
There is a spring inside pendulum clocks that is an important component of its mechanism. Suppose that, for some c, the lifetime of such a spring, in tens of thousands of
hours, is a random variable that has a probability density function of the form
c
8 ≤ x ≤ 15
(1 − x)2
f (x) =
0
otherwise.
(i)
Determine the constant c.
(ii)
Find the probability that the spring of a randomly selected pendulum clock lasts
at least 120,000 hours.
(iii)
Calculate the distribution function of the spring of a randomly selected pendulum
clock with the probability density function f given above.
At a local hardware store, the demand, per week, for propane gas, primarily for barbecue
propane tanks, in 100’s of gallons, is a random variable with probability density function
1
(6 − x2 )2
1≤x≤6
875
f (x) =
0
otherwise.
(a)
How many gallons of propane gas must the hardware store carry, per week, so
that the stock out probability for the gas is at most 0.05?
(b)
Suppose that on a Sunday morning, when the store opens, it has 590 gallons of
propane gas. What is the expected value of its inventory of gas on the following
Saturday evening when the store closes, given that no propane gas delivery has
been made all week?
Let X be a continuous random variable with probability density function
1
cos x
−π/2 ≤ x ≤ π/2
f (x) = 2
0
otherwise.
Find Var(X).
7.
Let F, the distribution of a random variable X, be defined by
0
x < −1
1 arcsin x
F (x) =
+
−1 ≤ x < 1
2
π
1
x ≥ 1,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 277 — #293
✐
✐
Chapter 6
Self-Test Problems
277
where arcsin x lies between −π/2 and π/2. Find f, the probability density function of
X and E(X).
8.
Let X be a continuous random variable with probability density function
1
π
π
cos x
− <x<
2
2
f (x) = 2
0
otherwise.
Find E(X), Var(X), and E(cos X).
1 + cos 2x
.
2
Let X be a random variable with probability density function
(2/3)x
if 1 < x < 2
f (x) =
0
otherwise.
Hint: Note that cos2 x =
9.
10.
(i)
Determine the probability density function of eX .
(ii)
Calculate E(eX ).
Each year, a business loses some money because the market prices of some merchandise
in its inventory decline. Let X be the amount of loss in a random year, in thousands of
dollars, and suppose that the probability density function of X is given by
1
(10 − x)
0 < x < 10
f (x) = 50
0
otherwise.
Of all the years that the company loses at least $4000 due to decline of market prices of
some of its merchandise, in what percentage does its total loss exceed $8000?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 278 — #294
✐
✐
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 279 — #295
✐
✐
Chapter 7
Special C ontinuous
Distributions
In this chapter we study some examples of continuous random variables. These random
variables appear frequently in theory and applications of probability, statistics, and branches of
science and engineering.
7.1
UNIFORM RANDOM VARIABLES
In Sections 1.6 and 1.7 we explained that in random selection of a point from an interval (a, b),
the probability of the occurrence of any particular point is zero. As a result, we stated that
if [α, β] ⊆ (a, b), the events that the point falls in [α, β], (α, β), [α, β), and (α, β] are all
equiprobable. Moreover, we said that a point is randomly selected from an interval (a, b) if any
two of its subintervals that have the same length are equally likely to include the point. We also
mentioned that the probability associated with the event that the subinterval (α, β) includes
the point is defined to be (β − α)/(b − a). Applications of these facts have been discussed
throughout the book. Therefore, their significance should be clear by now. In particular, in
Chapter 13 we show that the core of computer simulations is selection of random points from
intervals. In this section we introduce the concept of a uniform random variable. Then we
study its properties and applications. As we will see now, uniform random variables are directly
related to random selection of points from intervals.
Suppose that X is the value of the random point selected from an interval (a, b). Then X
is called a uniform random variable over (a, b). Let F and f be distribution and probability
density functions of X, respectively. Clearly,
0
t<a
t − a
F (t) =
a≤t<b
b
−
a
1
t ≥ b.
Therefore,
1
f (t) = F ′ (t) = b − a
0
if a < t < b
(7.1)
otherwise.
279
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 280 — #296
✐
✐
280
Chapter 7
Special Continuous Distributions
Definition 7.1
A random variable X is said to be uniformly distributed over an interval
(a, b) if its probability density function is given by (7.1).
Another way of reaching this definition is to note that f (x) is a measure that determines
how likely it is for X to be close to x. Since for all x ∈ (a, b) the probability that X is close
to x is the same, f should be a nonzero constant on (a, b); zero, elsewhere. Therefore,
c
if a < x < b
f (x) =
0
elsewhere.
Rb
Rb
Now a f (x) dx = 1 implies that a c dx = c(b − a) = 1. Thus c = 1/(b − a).
Figure 7.1 represents the graphs of f and F, the probability density function and the distribution function of a uniform random variable over the interval (a, b).
F(x)
f(x)
1
_
b a
1
a
b
Figure 7.1
x
a
b
x
Density and distribution functions of a uniform random
variable.
In random selections of a large number of points from (a, b), we expect that the average of
the values of the points will be approximately (a + b)/2, the midpoint of (a, b). This can be
shown by calculating E(X), the expected value of a uniform random variable X over (a, b).
Z b
1
1 h 1 2 ib
1 1 2 1 2
E(X) =
x
dx =
x
=
b − a
b−a
b−a 2
a
b−a 2
2
a
=
(b − a)(b + a)
a+b
=
.
2(b − a)
2
To find Var(X), note that
Z b
E(X 2 ) =
x2
a
1
1 b3 − a3
1
dx =
= (a2 + ab + b2 ).
b−a
3 b−a
3
Hence
a + b 2
2
1
(b − a)2
Var(X) = E(X 2 ) − E(X) = (a2 + ab + b2 ) −
=
.
3
2
12
We have shown that
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 281 — #297
✐
✐
Section 7.1
Uniform Random Variables
281
If X is uniformly distributed over an interval (a, b), then
E(X) =
a+b
,
2
(b − a)2
,
12
Var(X) =
b−a
σX = √ .
12
It is interesting to note that the expected value and the variance of a randomly selected integer
Y from the set {1, 2, 3, . . . , N } are very similar to E(X) and Var(X) obtained above. By
Exercise 6 of Section 4.5, E(Y ) = (1 + N )/2 and Var(Y ) = (N 2 − 1)/12.
Example 7.1 Starting at 5:00 A.M., every half hour there is a flight from San Francisco airport
to Los Angeles International airport. Suppose that none of these planes is completely sold out
and that they always have room for passengers. A person who wants to fly to L.A. arrives at
the airport at a random time between 8:45 A.M. and 9:45 A.M. Find the probability that she waits
(a) at most 10 minutes; (b) at least 15 minutes.
Solution: Let the passenger arrive at the airport X minutes past 8:45. Then X is a uniform
random variable over the interval (0, 60). Hence the probability density function of X is given
by
(
1/60
if 0 < x < 60
f (x) =
0
elsewhere.
Now the passenger waits at most 10 minutes if she arrives between 8:50 and 9:00 or 9:20 and
9:30; that is, if 5 < X < 15 or 35 < X < 45. So the answer to (a) is
P (5 < X < 15) + P (35 < X < 45) =
Z 15
5
1
dx +
60
Z 45
35
1
1
dx = .
60
3
The passenger waits at least 15 minutes if she arrives between 9:00 and 9:15 or 9:30 and 9:45;
that is, if 15 < X < 30 or 45 < X < 60. Thus the answer to (b) is
P (15 < X < 30) + P (45 < X < 60) =
Z 30
15
1
dx +
60
Z 60
45
1
1
dx = . 60
2
Example 7.2 A person arrives at a bus station every day at 7:00 A.M. If a bus arrives at a
random time between 7:00 A.M. and 7:30 A.M., what is the average time spent waiting?
Solution: If the bus arrives X minutes past 7:00 A.M., then X is a uniform random variable
over the interval (0, 30). Hence the average waiting time is
E(X) =
0 + 30
= 15 minutes. 2
The uniform distribution is often used to solve elementary problems in geometric
probability. In Chapter 8 we solve several such problems. Here we discuss only a famous
problem introduced by the French mathematician Joseph Bertrand (1822–1900) in 1889. It is
called Bertrand’s paradox. Bertrand seriously doubted that probability could be defined on
infinite sample spaces. To make his point, he posed this problem:
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 282 — #298
✐
✐
282
Chapter 7
Special Continuous Distributions
Example 7.3 What is the probability that a random chord of a circle is longer than a side of
an equilateral triangle inscribed into the circle?
Solution: Since the exact meaning of a random chord is not given, we cannot solve this problem as stated. We interpret the expression random chord in three different ways and solve the
problem in each case. Let the center of the circle be at C and its radius be r .
First interpretation: To draw a random chord, by considerations of symmetry, first choose
a random point A on the circle and connect it to C, the center; then choose a random number d
from (0, r) and place M on AC so that CM = d. Finally, from M draw a chord perpendicular
to the radius AC (see Figure 7.2). Since d < r/2 if and only if the chord is longer than a side
of an equilateral triangle inscribed into the circle and d is uniformly distributed over (0, r), the
desired probability is
r r/2
1
=
= .
P d<
2
r
2
C
M
A
Figure 7.2
First interpretation of Example 7.3.
Second interpretation: To draw a random chord, by considerations of symmetry, first
choose a random point A and then another random point D on the circle and connect AD. Let
B and E be the points on the circle that make ABE an equilateral triangle. The random chord
AD is longer than a side of ABE if and only if D lies on the arc BE . Since the length of the
arc BE is one-third of the length of the whole circle and D is a random point on the circle, D
lies on the arc BE with probability 1/3. Thus the desired probability is 1/3 as well (see Figure
7.3).
E
D
A
B
Figure 7.3
Second interpretation of Example 7.3.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 283 — #299
✐
✐
Section 7.1
Uniform Random Variables
283
Third interpretation: Since a chord is perpendicular to the radius connecting its midpoint
to the center of the circle, every chord is uniquely determined by its midpoint. To draw a
random chord, choose a random point M inside the circle, connect it to C, and draw a chord
perpendicular to M C from M . It is clear that the chord is longer than a side of the equilateral
triangle inscribed into the circle if and only if its midpoint M lies inside the circle centered at
C with radius r/2 (see Figure 7.4). Since each choice of M uniquely determines one choice of
a chord, the desired probability is the area of the small circle divided by the area of the original
circle (a fact discussed in detail in Section 8.1). It is equal to
1
π(r/2)2
= .
2
πr
4
M
C
Figure 7.4
Third interpretation of Example 7.3.
As we showed, three different interpretations of random chord resulted in three different
answers. Because of this, Bertrand’s problem was once considered a paradox. At that time,
one did not pay attention to the fact that the three interpretations correspond to three different
experiments concerning the selection of a random chord. In this process we are dealing with
three different probability functions defined on the same set of events.
EXERCISES
A
1.
It takes a professor a random time between 20 and 27 minutes to walk from his home to
school every day. If he has a class at 9:00 A.M. and he leaves home at 8:37 A.M., find the
probability that he reaches his class on time.
2.
Suppose that 15 points are selected at random and independently from the interval (0, 1).
How many of them can be expected to be greater than 3/4?
3.
The time at which a bus arrives at
√a station is uniform over an interval (a, b) with mean
2:00 P.M. and standard deviation 12 minutes. Determine the values of a and b.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 284 — #300
✐
✐
284
Chapter 7
Special Continuous Distributions
4.
Suppose that b is a random number from the interval (−3, 3). What is the probability
that the quadratic equation x2 + bx + 1 = 0 has at least one real root?
5.
The radius of a sphere is a random number between 2 and 4. What is the expected value
of its volume? What is the probability that its volume is at most 36π ?
6.
A point is selected at random on a line segment of length ℓ. What is the probability that
the longer segment is at least twice as long as the shorter segment?
7.
The loss for an insurance policy, in thousands of dollars, denoted by X, is a random
amount between 0 and 10. The insurance company wants to determine the amount of
deductible under this policy so that the expected value of the payment is (0.81)E(X),
which is 81% of the expected value of the loss if there was no deductible. What should
the amount of the deductible be?
8.
From the class of all triangles one is selected at random. What is the probability that it
is obtuse?
Hint: The largest angle of a triangle is less than 180 degrees but greater than or equal
to 60 degrees.
9.
A farmer who has two pieces of lumber of lengths a and b (a < b) decides to build a pen
in the shape of a triangle for his chickens. He sends his foolish son out to cut the longer
piece and the boy, without taking any thought as to the ultimate purpose, makes a cut on
the lumber of length b, at a point selected randomly. What are the chances that the two
resulting pieces and the piece of length a can be used to form a triangular pen?
Hint: Three segments form a triangle if and only if the length of any one of them is
less than the sum of the lengths of the remaining two.
10.
Let θ be a random number between −π/2 and π/2. Find the probability density function
of X = tan θ .
11.
Let X be a random number from [0, 1]. Find the probability mass function of [nX], the
greatest integer less than or equal to nX .
B
12.
13.
14.
15.
16.
17.
Let X be a random number from (0, 1). Find the probability density functions of the
random variables Y = − ln(1 − X) and Z = X n , n 6= 0.
Let X be a uniform random variable over the interval (0, 1 +
0 < θ < 1 is a
θ), where
given parameter. Find a function of X, say g(X), so that E g(X) = θ 2 .
Let X be a continuous random variable with distribution function F . Prove that F (X)
is uniformly distributed over (0, 1).
g be a nonnegative real-valued function on R that satisfies the relation
RLet
∞
−∞ g(t) dt = 1. Show that if, for a random variable X, the random variable
RX
Y = −∞ g(t) dt is uniform, then g is the probability density function of X .
RThe sample space of an experiment is S = (0, 1), and for every subset A of S, P (A) =
dx. Let X be a random variable defined on S by X(ω) = 5ω − 1. Prove that X is a
A
uniform random variable over the interval (−1, 4).
√
Let Y be a random number from (0, 1). Let X be the second digit of Y . Prove that
for n = 0, 1, 2, . . . , 9, P (X = n) increases as n increases. This is remarkable because
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 285 — #301
✐
✐
Section 7.2
Normal Random Variables
285
it shows that P (X = n), n = 1, 2, 3, . . . is not constant. That is, Y is uniform but X is
not.
Self-Quiz on Section 7.1
Time allotted: 20 Minutes
Each problem is worth 5 points.
1.
A point is selected at random on a line segment of length ℓ. What is the probability that
none of the two segments is smaller than ℓ/3?
2.
Suppose that X, the error of a measurement, is a random number between −1 and 1.
Find the expected value and the variance of |X|, the magnitude of the error.
Hint: By calculating the distribution function of |X|, show that it is a uniform random
variable.
7.2
NORMAL RANDOM VARIABLES
In search of formulas to approximate binomial probabilities, Poisson was not alone. Other
mathematicians had also realized the importance of such investigations. In 1718, before
Poisson, De Moivre had discovered the following approximation, which is completely different
from Poisson’s.
De Moivre’s Theorem Let X be a binomial random variable with parameters n and 1/2.
Then for any numbers a and b, a < b,
Z b
2
X − (1/2)n
1
√
lim P a <
e−t /2 dt.
<b = √
n→∞
(1/2) n
2π a
√
Note that in this formula (1/2)n = E(X) and (1/2) n = σX .
De Moivre’s theorem was appreciated by Laplace, who recognized its importance. In 1812
he generalized it to binomial random variables with parameters n and p. He showed the following theorem, now called the De Moivre–Laplace theorem.
Theorem 7.1 (De Moivre–Laplace Theorem) Let X be a binomial random variable
with parameters n and p. Then for any numbers a and b, a < b,
Z b
2
X − np
1
lim P a < p
<b = √
e−t /2 dt.
n→∞
np(1 − p)
2π a
Note that np and
p
np(1 − p) appearing in this formula are, respectively, E(X) and σX .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 286 — #302
✐
✐
286
Chapter 7
Special Continuous Distributions
Poisson’s approximation, as discussed, is good when n is large, p is relatively small, and
np is appreciable. The De Moivre–Laplace formula yields excellent approximations for values
of n and p for which np(1 − p) ≥ 10.
To find his approximation, Poisson first considered a binomial random variable X with
parameters (n, p). Then he showed that if λ = np, P (X = i) is approximately (e−λ λi )/i!
for large n. Finally, he proved that (e−λ λi )/i! itself is a probability mass function. The same
procedure can be followed for the De Moivre–Laplace approximation as well. By this theorem,
if X is a binomial random variable with parameters (n, p), the sequence of probabilities
1
converges to √
2π
X − np
≤t ,
P p
np(1 − p)
n = 1, 2, 3, 4, . . . ,
Z t
2
1
e−x /2 dx, where the function Φ(t) = √
2π
−∞
Z t
2
e−x /2 dx is a dis-
−∞
tribution function itself. To prove that Φ is a distribution function, note that it is increasing,
continuous, and Φ(−∞) = 0. The proof of Φ(∞) = 1 is tricky. We use the following ingenious technique introduced by Gauss to show it. Let
Z ∞
2
e−x /2 dx.
I=
−∞
Then
I2 =
Z ∞
2
e−x /2 dx
Z ∞
−∞
−∞
Z ∞Z ∞
2
2
2
e−(x +y )/2 dx dy.
e−y /2 dy =
−∞
−∞
To evaluate this double integral, we change the variables to polar coordinates. That is, we let
x = r cos θ, y = r sin θ. We get dx dy = r dθ dr and
I2 =
Z ∞ Z 2π
0
= 2π
2
e−r /2 r dθ dr =
0
Z ∞
0
h
2
Z ∞
2
e−r /2 r
0
2
re−r /2 dr = 2π − e−r /2
Thus
I=
Z ∞
i∞
0
Z 2π
0
= 2π.
2
√
Z ∞
e−x /2 dx = 1.
e−x /2 dx =
dθ dr
2π,
−∞
and hence
1
Φ(∞) = √
2π
2
−∞
Therefore, Φ is a distribution function.
Definition 7.2
is Φ, that is, if
A random variable X is called standard normal if its distribution function
1
P (X ≤ t) = Φ(t) ≡ √
2π
Z t
2
e−x /2 dx.
−∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 287 — #303
✐
✐
Section 7.2
Normal Random Variables
287
By the fundamental theorem of calculus, f, the probability density function of a standard
normal random variable, is given by
1
2
f (x) = Φ′ (x) = √ e−x /2 .
2π
The standard normal density function is a bell-shaped curve that is symmetric about the y -axis
(see Figure 7.5).
f(x)
1/
x
0
Figure 7.5
Graph of the standard normal density function.
Since Φ is the distribution function of the standard normal random variable, Φ(t) is the area
under this curve from −∞ to t. Because Φ(∞) = 1 and the curve is symmetric about the
y -axis, Φ(0) = 1/2. Moreover,
Φ(−t) = 1 − Φ(t).
To see this, note that
1
Φ(−t) = √
2π
Z −t
2
e−x /2 dx.
−∞
Substituting u = −x, we obtain
Z t
Z ∞
2
−1
1
−u2 /2
Φ(−t) = √
e
du = √
e−u /2 du
2π ∞
2π t
Z ∞
Z t
2
1
1
−u2 /2
=√
e
du − √
e−u /2 du
2π −∞
2π −∞
= 1 − Φ(t).
Example 7.4 Let Z be a standard normal random variable and α be a given constant. Find
the real number x that maximizes P (x < Z < x + α).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 288 — #304
✐
✐
288
Chapter 7
Special Continuous Distributions
Solution: It should be intuitively clear that P (x < Z < x + α) is maximum at x = −α/2.
Z x+α
2
1
To show this mathematically, let g(x) = P (x < Z < x + α) = √
e−y /2 dy. The
2π x
number x that maximizes P (x < Z < x + α) is the root of g ′ (x) = 0. Therefore, by the
fundamental theorem of calculus, it is the solution of
which is x = −α/2.
2
1 −(x+α)2 /2
e
− e−x /2 = 0,
g ′ (x) = √
2π
Correction for Continuity
Thus far we have shown that the De Moivre–Laplace theorem approximates the distribution of
a discrete random variable by that of a continuous one. But we have not demonstrated how this
approximation works in practice. To do so, first we explain how, in general, a probability based
on a discrete random variable is approximated by a probability based on a continuous one. Let
X be a discrete random variable with probability mass function p(x), and suppose that we want
to find P (i ≤ X ≤ j), i < j . Consider the histogram of X, as sketched in Figure 7.6 from i
to j . In that figure the base of each rectangle equals 1, and the height (and therefore the area) of
the rectangle with the base midpoint k is p(k), i ≤ k ≤ j . Thus the sum of the areas of all rectPj
angles is k=i p(k), which is the exact value of P (i ≤ X ≤ j). Now suppose that f (x), the
probability density function of a continuous random variable, sketched in Figure 7.7, is a good
approximation to p(x). Then, as this figure shows, P (i ≤ X ≤ j), the sum of the areas of all
rectangles of the figure, is approximately the area under f (x) from i − 1/2 to j + 1/2 rather
than from i to j . That is,
P (i ≤ X ≤ j) ≈
Z j+1/2
f (x) dx.
i−1/2
p(k)
i i+1i+2
Figure 7.6
j_1 j
k
Histogram of X from i to j.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 289 — #305
✐
✐
Section 7.2
Normal Random Variables
289
f
j_1 j
i i+1i+2
Figure 7.7
Histogram of X and the probability density function f .
This adjustment is called correction for continuity and is necessary for approximation of the
distribution of a discrete random variable with that of a continuous one. Similarly, the following
corrections for continuity are made to calculate the given probabilities.
P (X = k) ≈
P (X ≥ i) ≈
P (X ≤ j) ≈
Z k+1/2
k−1/2
Z ∞
f (x) dx,
f (x) dx,
i−1/2
Z j+1/2
f (x) dx.
−∞
In real-world problems, or even sometimes in theoretical ones, to apply the De Moivre–
Rb
2
Laplace theorem, we need to calculate the numerical values of a e−x /2 dx for some real
2
numbers a and b. Since e−x /2 has no antiderivative in terms of elementary functions, such
integrals are approximated by numerical techniques. Tables for such approximations have
been available since 1799. Nowadays, most of the scientific calculators give excellent approximate values
integrals. Table 1 of the Appendix Tables gives the values of
√ for
R xthese
−y 2 /2
Φ(x) = (1/ 2π ) −∞ e
dy for x = −3.89 to x = 0. For x ≤ −3.90, Φ(x) ≈ 0.
Table 2 of the Appendix Tables gives the values of Φ(x) for x = 0 to x = 3.89. For x > 3.89,
Φ(x) ≈ 1. Note that, using the relation Φ(x) = 1 − Φ(−x), we can also use Table 1 to find
Φ(x) for x > 0 and Table 2 to find Φ(x) for x < 0. Therefore, one of the tables 1 or 2 is
sufficient to find Φ(x) for all values of x. However, for convenience, we have included a table
for negative values of x and a separate table for positive values of x. The following example
shows how the De Moivre–Laplace theorem is applied.
Example 7.5 Suppose that of all the clouds that are seeded with silver iodide, 58% show
splendid growth. If 60 clouds are seeded with silver iodide, what is the probability that exactly
35 show splendid growth?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 290 — #306
✐
✐
290
Chapter 7
Special Continuous Distributions
Solution: Let X be the p
number of clouds that show splendid growth. Then E(X) = 60(0.58)
= 34.80 and σX =
60(0.58)(1 − 0.58) = 3.82. By correction for continuity and
De Moivre–Laplace theorem,
P (X = 35) ≈ P (34.5 < X < 35.5)
34.5 − 34.80
X − 34.80
35.5 − 34.80 =P
<
<
3.82
3.82
3.82
Z 0.18
2
X − 34.8
1
= P − 0.08 <
< 0.18 ≈ √
e−x /2 dx
3.82
2π −0.08
1
=√
2π
Z 0.18
2
1
e−x /2 dx − √
2π
−∞
Z −0.08
2
e−x /2 dx
−∞
= Φ(0.18) − Φ(−0.08) = 0.5714 − 0.4681 = 0.1033.
!
60
The exact value of P (X = 35) is
(0.58)35 (0.42)25 , which up to four decimal points
35
equals 0.1039. Therefore, the answer obtained by the De Moivre–Laplace approximation is
very close to the actual probability. ⋆ Remark 7.1 Using
powerful scientific calculators, calculating the numerical
√ today’s
Rx
2
value of Φ(x) = (1/ 2π ) −∞ e−y /2 dy up to several decimal points of accuracy is almost
as easy as addition or subtraction. However, both for theoretical and computational reasons,
approximating Φ(x) by simpler functions has been the subject of many studies. One such approximation was given by A. K. Shah in his paper “A Simpler Approximation for Areas Under
the Standard Normal Curve” (The American Statistician, 39, 80, 1985). According to this approximation, for 0 ≤ x ≤ 2.2,
Φ(x) ≈ 0.5 +
x(4.4 − x)
,
10
where the error of approximation is at most 0.005. That is,
Φ(x) − 0.5 −
x(4.4 − x)
≤ 0.005. 10
Now that the importance of the standard normal random variable is clarified, we will
calculate its expected value and variance. The graph of the standard normal density function
(Figure 7.5) suggests that its expected value is zero. To prove this, let X be a standard normal
random variable. Then
Z ∞
2
1
E(X) = √
xe−x /2 dx = 0,
2π −∞
2
because the integrand, xe−x /2 , is a finite odd function, and the integral is taken from −∞ to
+∞. To calculate Var(X), note that
Z ∞
2
1
x2 e−x /2 dx.
E(X 2 ) = √
2π −∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 291 — #307
✐
✐
Section 7.2
Normal Random Variables
291
2
Using integration by parts, we get (let u = x, dv = xe−x /2 dx)
Z ∞
Z ∞
h
i∞
√
√
2
2
2
xxe−x /2 dx = − xe−x /2
+
e−x /2 dx = 0 + 2π = 2π.
−∞
−∞
−∞
p
2
Therefore, E(X 2 ) = 1, Var(X) = E(X 2 ) − E(X) = 1, and σX = Var(X) = 1. We
have shown that
The expected value of a standard normal random variable is 0. Its standard
deviation is 1.
Although some natural phenomena obey a standard normal distribution, when it comes
to the analysis of data, due to the lack of parameters in the standard normal distribution, it
cannot be used. To overcome this difficulty, mathematicians generalized the standard normal
distribution by introducing the following probability density function.
Definition 7.3
A random variable X is called normal, with parameters µ and σ, if its
probability density function is given by
h −(x − µ)2 i
1
,
f (x) = √ exp
2σ 2
σ 2π
−∞ < x < ∞.
f is a probability density function because by the change of variable y = (x − µ)/σ :
Z ∞
Z ∞
h −(x − µ)2 i
2
1
1
√ exp
√
e−y /2 dy = Φ(∞) = 1.
dx
=
2
2σ
2π −∞
−∞ σ 2π
If X is a normal random variable with parameters µ and σ, we write X ∼ N (µ, σ 2 ).
One of the first applications of N (µ, σ 2 ) was given by Gauss in 1809. Gauss used
N (µ, σ 2 ) to model the errors of observations in astronomy. For this reason, the normal distribution is sometimes called the Gaussian distribution. Larsen and Marx, in their An Introduction to Mathematical Statistics and Its Applications (Prentice Hall, Upper Saddle River, N.J.,
1986), explain that N (µ, σ 2 ) was popularized by the Belgian scholar Lambert Quetelet (1796–
1874), who used it successfully in data analysis in many situations.† The chest measurement of Scottish soldiers is one of the studies of N (µ, σ 2 ) by Quetelet. He measured Xi ,
i = 1, 2, . . . , 5738, the respective sizes of the chests of 5738 Scottish soldiers, and, using
statistical methods, found that their average and standard deviation were 39.8 and 2.05 inches,
respectively. Then he counted the number of soldiers who had a chest of size i, i = 33, 34,
. . . , 48, and calculated the relative frequencies of sizes 33 through 48. He showed that for i =
33, . . . , 48 the relative frequency of the size i is very close to P (i − 1/2 < X < i + 1/2),
where X is N (µ, σ 2 ) with µ = 39.8 and σ = 2.05. Hence he concluded that the measure of
the chest of a Scottish soldier is approximately a normal random variable with parameters 39.8
and 2.05.
The following lemma shows that, by a simple change of variable, N (µ, σ 2 ) can be transformed to N (0, 1).
†
An excellent treatise on the work of Quetelet concerning the applications of probability to the measurement
of uncertainty in the social sciences can be found in Chapter 5 of The History of Statistics by Stephen M. Stigler,
The Belknap Press of Harvard University Press, 1986.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 292 — #308
✐
✐
292
Chapter 7
Special Continuous Distributions
Lemma 7.1 If X ∼ N (µ, σ 2 ), then Z = (X −µ)/σ is N (0, 1). That is, if X ∼ N (µ, σ 2 ),
the standardized X is N (0, 1).
Proof:
√
Rx
2
We show that the distribution function of Z is (1/ 2π ) −∞ e−y /2 dy. Note that
≤ x = P (X ≤ σx + µ)
σ
Z σx+µ
h −(t − µ)2 i
1
= √
exp
dt.
2σ 2
σ 2π −∞
P (Z ≤ x) = P
X − µ
Let y = (t − µ)/σ; then dt = σ dy and we get
Z x
Z x
2
1
1
−y 2 /2
P (Z ≤ x) = √
e
σ dy = √
e−y /2 dy.
σ 2π −∞
2π −∞
Let X ∼ N (µ, σ 2 ). Quetelet’s analysis of the sizes of the chests of Scottish soldiers is an
indication that µ and σ are the expected value and the standard deviation of X, respectively. To
prove this, note that the random variable Z = (X − µ)/σ is N (0, 1) and X = σZ + µ. Hence
E(X) = E(σZ + µ) = σE(Z) + µ = µ and Var(X) = Var(σZ + µ) = σ 2 Var(Z) = σ 2 .
Therefore,
The parameters µ and σ that appear in the formula of the probability density function of X are its expected value and standard deviation, respectively.
h −(t − µ)2 i
1
√ exp
is a bell-shaped curve symmetric about
2σ 2
σ 2π √
x = µ with the maximum at (µ, 1/σ 2π ) and inflection points at µ ± σ (see Figures 7.8 and
7.9).
The transformation Z = (X − µ)/σ enables us to use Tables 1 and 2 of the Appendix
Tables to calculate the probabilities concerning X . Some examples follow.
The graph of f (x) =
Example 7.6 Suppose that a Scottish soldier’s chest size is normally distributed with mean
39.8 and standard deviation 2.05 inches, respectively. What is the probability that of 20 randomly selected Scottish soldiers, five have a chest of at least 40 inches?
Solution: Let p be the probability that a randomly selected Scottish soldier has a chest of 40
or more inches. If X is the normal random variable with mean 39.8 and standard deviation
2.05, then
X − 39.8
X − 39.8
40 − 39.8 p = P (X ≥ 40) = P
≥
=P
≥ 0.10
2.05
2.05
2.05
= P (Z ≥ 0.10) = 1 − Φ(0.1) ≈ 1 − 0.5398 ≈ 0.46.
Therefore, the probability that of 20 randomly selected Scottish soldiers, five have a chest of at
least 40 inches is
!
20
(0.46)5 (0.54)15 ≈ 0.03. 5
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 293 — #309
✐
✐
Section 7.2
Normal Random Variables
293
f (x)
1__
____
x
0
Figure 7.8
Density of N (µ, σ 2 ).
f(x)
=5
=8
=15
10
Figure 7.9
20
x
30
Different normal densities with specified parameters.
Example 7.7 Let X, the grade of a randomly selected student in a test of a probability
course, be a normal random variable. A professor is said to grade such a test on the curve if he
finds the average µ and the standard deviation σ of the grades and then assigns letter grades
according to the following table.
Range of
the grade
X ≥ µ+σ
µ ≤ X < µ+σ
µ−σ ≤ X < µ
µ−2σ ≤ X < µ−σ
X < µ−2σ
Letter
grade
A
B
C
D
F
Suppose that the professor of the probability course grades the test on the curve. Determine
the percentage of the students who will get A, B, C, D, and F, respectively.
Solution: By the fact that (X − µ)/σ is standard normal,
P (X ≥ µ + σ) = P
X − µ
≥ 1 = 1 − Φ(1) ≈ 0.1587,
σ
X −µ
< 1 = Φ(1) − Φ(0) ≈ 0.3413,
P (µ ≤ X < µ + σ) = P 0 ≤
σ
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 294 — #310
✐
✐
294
Chapter 7
Special Continuous Distributions
X −µ
P (µ − σ ≤ X < µ) = P − 1 ≤
< 0 = Φ(0) − Φ(−1)
σ
= 0.5 − 0.1587 ≈ 0.3413,
X −µ
P (µ − 2σ ≤ X < µ − σ) = P − 2 ≤
< −1 = Φ(−1) − Φ(−2)
σ
= 0.1587 − 0.0228 ≈ 0.1359,
X − µ
< −2 = Φ(−2) ≈ 0.0228.
P (X < µ − 2σ) = P
σ
Therefore, approximately 16% should get A, 34% B, 34% C, 14% D, and 2% F. If an instructor
grades a test on the curve, instead of calculating µ and σ, he or she may assign A to the top
16%, B to the next 34%, and so on. Example 7.8 The scores on an achievement test given to 100,000 students are normally
distributed with mean 500 and standard deviation 100. What should the score of a student be
to place him among the top 10% of all students?
Solution: Letting X be a normal random variable with mean 500 and standard deviation 100,
we must find x so that P (X ≥ x) = 0.10 or P (X < x) = 0.90. This gives
X − 500
x − 500 <
= 0.90.
P
100
100
Thus
Φ
x − 500 100
= 0.90.
From Table 2 of the Appendix Tables, we have that Φ(1.28) ≈ 0.8997, implying that
(x − 500)/100 ≈ 1.28. This gives x ≈ 628; therefore, a student should earn 628 or more to
be among the top 10% of the students. Example 7.9 The annual rate of return for a share of a specific stock is a normal random
variable with mean 10% and standard deviation 12%. Ms. Couture buys 100 shares of the
stock at a price of $60 per share. What is the probability that after a year her net profit from
that investment is at least $750? Ignore transaction costs and assume that there is no annual
dividend.
Solution: Let r be the rate of return of this stock. The random variable r is normal with µ =
0.10 and σ = 0.12. Let X be the price of the total shares of the stock that Ms. Couture buys
this year. We are given that X = 6000. Let Y be the total value of the shares next year. The
desired probability is
Y − X
750 750 P Y − X ≥ 750 = P
≥
=P r≥
X
X
6000
0.125 − 0.10 = P (r ≥ 0.125) = P Z ≥
0.12
= P (Z ≥ 0.21) = 1 − P (Z < 0.21)
= 1 − Φ(0.21) = 1 − 0.5832 = 0.4168. ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 295 — #311
✐
✐
Section 7.2
Normal Random Variables
295
It is important to note that a large family of random variables in mathematics and statistics
have distributions that are either normal or approximately normal. One of the reasons for this
is the celebrated central limit theorem which we will study in Chapter 11. To express that
theorem, we will use Richard von Mises’ words from his Probability, Statistics and the Truth
(page 171, Dover Publications, 1981):
The normal curve represents the distribution in all cases where a final collective is formed by combination of a very large number of initial collectives, the
attribute in the final collective being the sum of the results in the initial collectives. The original collectives are not necessarily simple alternatives as they
are in Bernoulli’s problem. It is not even necessary for them to have the same
attributes or the same distributions. The only conditions are that a very great
number of collectives are combined and that the attributes are mixed in such
a way that the final attribute is the sum of all the original ones. Under these
conditions the final distribution is always represented by a normal curve.
For example, consider the annual profit of an insurance company that sells thousands of policies
each year. This is a normal random variable since the total profit is the sum of profits (losses
being treated as negative profits) obtained from all policies sold. Some policies might be very
complicated (i.e., those sold to the banks or oil companies). Moreover, the profit from each
policy is itself a random variable which depends on many factors: the premium, risk involved,
damages to be paid, quit claims, profit or loss from investments, the level of competition in the
industry, and so on.
EXERCISES
A
1.
2.
Let X ∼ N (−2, 5). Find P |X| < 4 .
Suppose that 90% of the patients with a certain disease can be cured with a certain drug.
What is the approximate probability that, of 50 such patients, at least 45 can be cured
with the drug?
3.
A small college has 1095 students. What is the approximate probability that more than
five students were born on Christmas day? Assume that the birthrates are constant
throughout the year and that each year has 365 days.
4.
Let Ψ(x) = 2Φ(x) − 1. The function Ψ is called the positive normal distribution.
Prove that if Z is standard normal, then |Z| is positive normal.
5.
6.
Let X be a standard normal random variable. Calculate E(X cos X), E(sin X), and
X E
.
1 + X2
The ages of subscribers to a certain newspaper are normally distributed with mean 35.5
years and standard deviation 4.8. What is the probability that the age of a random subscriber is (a) more than 35.5 years; (b) between 30 and 40 years?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 296 — #312
✐
✐
296
Chapter 7
Special Continuous Distributions
7.
The grades for a certain exam are normally distributed with mean 67 and variance 64.
What percent of students get A(≥ 90), B(80 − 90), C(70 − 80), D(60 − 70), and
F(< 60)?
8.
Suppose that the distribution of the diastolic blood pressure for a randomly selected
person in a certain population is normal with mean 80 mm Hg and standard deviation 7
mm Hg. If people with diastolic blood pressures 95 or above are considered hypertensive
and people with diastolic blood pressures above 89 and below 95 are considered to have
mild hypertension, what percent of that population have mild hypertension and what
percent are hypertensive? Assume in that population no one has abnormal systolic blood
pressure.
9.
The length of an aluminum-coated steel sheet manufactured by a certain factory is approximately normal with mean 75 centimeters and standard deviation 1 centimeter. Find
the probability that a randomly selected sheet manufactured by this factory is between
74.5 and 75.8 centimeters.
10.
Suppose that the IQ of a randomly selected student from a university is normal with
mean 110 and standard deviation 20. Determine the interval of values that is centered at
the mean and includes 50% of the IQ’s of the students at that university.
11.
The amount of cereal in a box is normal with mean 16.5 ounces. If the packager is
required to fill at least 90% of the cereal boxes with 16 or more ounces of cereal, what
is the largest standard deviation for the amount of cereal in a box?
12.
Suppose that the scores on a certain manual dexterity test are normal with mean 12 and
standard deviation 3. If eight randomly selected individuals take the test, what is the
probability that none will make a score less than 14?
13.
A number t is said to be the median of a continuous random variable X if
P (X ≤ t) = P (X ≥ t) = 1/2.
14.
Calculate the median of the normal random variable with parameters µ and σ 2 .
Let X ∼ N (µ, σ 2 ). Prove that P |X − µ| > kσ does not depend on µ or σ .
15.
Suppose that lifetimes of light bulbs produced by a certain company are normal random variables with mean 1000 hours and standard deviation 100 hours. Is this company
correct when it claims that 95% of its light bulbs last at least 900 hours?
16.
Suppose that lifetimes of light bulbs produced by a certain company are normal random
variables with mean 1000 hours and standard deviation 100 hours. Suppose that lifetimes
of light bulbs produced by a second company are normal random variables with mean
900 hours and standard deviation 150 hours. Howard buys one light bulb manufactured
by the first company and one by the second company. What is the probability that at
least one of them lasts 980 or more hours?
17.
The annual rate of return for a share of a specific stock is a normal random variable
with mean 0.12 and standard deviation of 0.06. The current price of the stock is $35 per
share. Mrs. Lovotti would like to purchase enough shares of this stock to make at least
$1000 profit with a probability of at least 90% in one year. Find the minimum number of
shares that she should buy. Ignore transaction costs and assume that there are no annual
dividends.
18.
Find the expected value and the variance of a random variable with the probability
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 297 — #313
✐
✐
Section 7.2
Normal Random Variables
297
density function
f (x) =
19.
r
2 −2(x−1)2
e
,
π
−∞ < x < ∞.
Let X ∼ N (µ, σ 2 ). Find the distribution function of |X − µ| and its expected value.
20.
Determine the value(s) of k for which the following is the probability density function
of a normal random variable.
√
2 2
−∞ < x < ∞.
f (x) = ke−k x −2kx−1 ,
21.
The viscosity of a brand of motor oil is normal with mean 37 and standard deviation 10.
What is the lowest possible viscosity for a specimen that has viscosity higher than at
least 90% of that brand of motor oil?
22.
Let α ∈ (−∞, ∞) and Z ∼ N (0, 1); find E(eαZ ).
23.
24.
25.
Let X ∼ N (0, σ 2 ). Calculate the probability density function of Y = X 2 .
Let X ∼ N (µ, σ 2 ). Calculate the probability density function of Y = eX .
p
Let X ∼ N (0, 1). Calculate the probability density function of Y = |X|.
B
26.
Suppose that the odds are 1 to 5000 in favor of a customer of a particular bookstore
buying a certain fiction bestseller. If 800 customers enter the store every day, how many
copies of that bestseller should the store stock every month so that, with a probability of
more than 98%, it does not run out of this book? For simplicity, assume that a month is
30 days.
27.
To examine the accuracy of an algorithm that selects random numbers from the set
{1, 2, . . . , 40}, 100,000 numbers are selected and there are 3500 ones. Given that the
expected number of ones is 2500, is it fair to say that this algorithm is not accurate?
28.
Prove that for some constant k, f (x) = ka−x , a ∈ (0, ∞), is a normal probability
density function.
29.
(a)
2
Prove that for all x > 0,
2
1
1 2
1 √
1 − 2 e−x /2 < 1 − Φ(x) < √ e−x /2 .
x
x 2π
x 2π
Hint:
Integrate the following inequalities:
2
2
2
(1 − 3y −4 )e−y /2 < e−y /2 < (1 + y −2 )e−y /2 .
(b)
Use part (a) to prove that 1 − Φ(x) ∼
ratio of the two sides approaches 1.
2
1
√ e−x /2 . That is, as x → ∞, the
x 2π
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 298 — #314
✐
✐
298
Chapter 7
30.
Let Z be a standard normal random variable. Show that for x > 0,
x
lim P Z > t + | Z ≥ t = e−x .
t→∞
t
Hint:
Special Continuous Distributions
Use part (b) of Exercise 29.
31.
The amount of soft drink in a bottle is a normal random variable. Suppose that in 7%
of the bottles containing this soft drink there are fewer than 15.5 ounces, and in 10% of
them there are more than 16.3 ounces. What are the mean and standard deviation of the
amount of soft drink in a randomly selected bottle?
32.
At an archaeological site 130 skeletons are found and their heights are measured and
found to be approximately normal with mean 172 centimeters and variance 81 centimeters. At a nearby site, five skeletons are discovered and it is found that the heights of
exactly three of them are above 185 centimeters. Based on this information is it reasonable to assume that the second group of skeletons belongs to the same family as the first
group of skeletons?
33.
In a forest, the number of trees that grow in a region of area R has a Poisson distribution
with mean λR, where λ is a positive real number. Find the expected value of the distance
from a certain tree to its nearest neighbor.
R∞
2
Let I = 0 e−x /2 dx; then
34.
2
I =
Z ∞hZ ∞
0
0
i
2
2
e−(x +y )/2 dy dx.
Let y/x = s and change the order of integration to show that I 2 = π/2. This gives
an alternative proof of the fact that Φ is a distribution function. The advantage of this
method is that it avoids polar coordinates.
Self-Quiz on Section 7.2
Time allotted: 20 Minutes
Each problem is worth 5 points.
1.
The grades of students in a calculus-based probability course are normal with mean 72
and standard deviation 7. If 90, 80, 70, and 60 are the respective lowest, A, B, C, and D,
what percent of students in this course get A’s, B’s, C’s, D’s, and F’s?
2.
Every day a factory produces 5000 light bulbs, of which 2500 are type I and 2500 are
type II. If a sample of 40 light bulbs is selected at random to be examined for defects,
what is the approximate probability that this sample contains at least 18 light bulbs of
each type?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 299 — #315
✐
✐
Section 7.3
7.3
Exponential Random Variables
299
EXPONENTIAL RANDOM VARIABLES
Let N (t) : t ≥ 0 be a Poisson process. Then, as discussed in Section 5.2, N (t) is the number
of “events” that have occurred at or prior to time t. Let X1 be the time of the first event, X2
be the elapsed time between the first and the second events, X3 be the elapsed time between
the second and third events, and so on. The sequence of random variables
{X1 , X2 , X3 , . . .}
is called the sequence of interarrival times of the Poisson process N (t) : t ≥ 0 . Let
λ = E N (1) ; then
e−λt (λt)n
P N (t) = n =
.
n!
This enables us to calculate the distribution functions of the random variables Xi , i ≥ 1. For
t ≥ 0,
P (X1 > t) = P N (t) = 0 = e−λt .
Therefore,
P (X1 ≤ t) = 1 − P (X1 > t) = 1 − e−λt .
Since a Poisson process is stationary and possesses independent increments, at any time t, the
process probabilistically starts all over again. Hence the interarrival time of any two consecutive
events has the same distribution as X1 ; that is, the sequence {X1 , X2 , X3 , . . .} is identically
distributed. Therefore, for all n ≥ 1,
(
1 − e−λt
t≥0
P (Xn ≤ t) = P (X1 ≤ t) =
0
t < 0.
Let
F (t) =
(
1 − e−λt
t≥0
0
t<0
for some λ > 0. Then F is the distribution function of Xn for all n ≥ 1. It is called exponential
distribution and is one of the most important distributions of pure and applied probability.
Since
(
λe−λt
t≥0
′
f (t) = F (t) =
(7.2)
0
t<0
and
Z ∞
0
λe−λt dt = lim
b→∞
Z b
λe−λt dt = lim
0
f is a probability density function.
b→∞
h
− e−λt
ib
0
= lim (1 − e−λb ) = 1,
b→∞
Definition 7.4
A continuous random variable X is called exponential with parameter
λ > 0 if its probability density function is given by (7.2).
Because the interarrival times of a Poisson process are exponential, the following are
examples of random variables that might be exponential.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 300 — #316
✐
✐
300
Chapter 7
Special Continuous Distributions
1.
The interarrival time between two customers at a post office
2.
The duration of Jim’s next telephone call
3.
The time between two consecutive earthquakes in California
4.
The time between two accidents at an intersection
5.
The time until the next baby is born in a hospital
6.
The time until the next crime in a certain town
7.
The time to failure of the next fiber segment in a large group of such segments when all
of them are initially fault free
8.
The time interval between the observation of two consecutive shooting stars on a summer
evening
9.
The time between two consecutive fish caught by a fisherman from a large lake with lots
of fish
From
Section
5.2 we know that λ is the average number of the events in one time unit; that
is, E N (1) = λ. Therefore, we should expect an average time 1/λ between two consecutive
events. To prove this, let X be an exponential random variable with parameter λ; then
Z ∞
Z ∞
E(X) =
xf (x) dx =
x(λe−λx ) dx.
−∞
0
Using integration by parts with u = x and dv = λe−λx dx, we obtain
i∞
h
i∞ Z ∞
h1
1
= .
E(X) = − xe−λx
+
e−λx dx = 0 − e−λx
λ
0
λ
0
0
A similar calculation shows that
2
E(X ) =
Z ∞
2
x f (x) dx =
−∞
Hence
Z ∞
x2 (λe−λx ) dx
0
h i∞
2
2
2
= − x2 + x + 2 e−λx
= 2.
λ
λ
λ
0
2
1
1
2
Var(X) = E(X 2 ) − E(X) = 2 − 2 = 2 ,
λ
λ
λ
and therefore σX = 1/λ. We have shown that
For an exponential random variable with parameter λ,
E(X) = σX =
1
λ
,
Var(X) =
1
λ2
.
Figures 7.10 and 7.11 represent the graphs of the exponential density and exponential distribution functions, respectively.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 301 — #317
✐
✐
Section 7.3
Exponential Random Variables
301
f (x)
4/
1/
Figure 7.10
x
Exponential density function with parameter λ.
F(x)
1
1/2
x
4/
Figure 7.11
Exponential distribution function.
Example 7.10 Suppose that every three months, on average, an earthquake occurs in
California. What is the probability that the next earthquake occurs after three but before seven
months?
Solution: Let X be the time (in months) until the next earthquake; it can be assumed that X
is an exponential random variable with 1/λ = 3 or λ = 1/3. To calculate P (3 < X < 7),
note that since F, the distribution function of X, is given by
F (t) = P (X ≤ t) = 1 − e−t/3
for t > 0,
we can write
P (3 < X < 7) = F (7) − F (3) = (1 − e−7/3 ) − (1 − e−1 ) ≈ 0.27. Example 7.11 At an intersection there are two accidents per day, on average. What is the
probability that after the next accident there will be no accidents at all for the next two days?
Solution: Let X be the time (in days) between the next two accidents. It can be assumed that
X is exponential with parameter λ, satisfying 1/λ = 1/2, so that λ = 2. To find P (X > 2),
note that F, the distribution function of X, is given by F (t) = 1 − e−2t , t > 0. Hence
P (X > 2) = 1 − P (X ≤ 2) = 1 − F (2) = e−4 ≈ 0.02. ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 302 — #318
✐
✐
302
Chapter 7
Special Continuous Distributions
An important feature of exponential distribution is its memoryless property. A nonnegative
random variable X is called memoryless if, for all s, t ≥ 0,
P (X > s + t | X > t) = P (X > s).
(7.3)
If, for example, X is the lifetime of some type of instrument, then (7.3) means that there is no
deterioration with age of the instrument. The probability that a new instrument will last more
than s years is the same as the probability that a used instrument that has lasted more than t
years will last at least another s years. In other words, the probability that such an instrument
will deteriorate in the next s years does not depend on the age of the instrument.
To show that an exponential distribution is memoryless, note that (7.3) is equivalent to
P (X > s + t, X > t)
= P (X > s)
P (X > t)
and
P (X > s + t) = P (X > s)P (X > t).
(7.4)
Now since
P (X > s + t) = 1 − [1 − e−λ(s+t) ] = e−λ(s+t) ,
P (X > s) = 1 − (1 − e−λs ) = e−λs ,
and
P (X > t) = 1 − (1 − e−λt ) = e−λt ,
we have that (7.4) follows. Hence X is memoryless. It can be shown that exponential is the
only continuous distribution which possesses a memoryless property (see Exercise 18).
Example 7.12 The lifetime of a TV tube (in years) is an exponential random variable with
mean 10. If Jim bought his TV set 10 years ago, what is the probability that its tube will last
another 10 years?
Solution: Let X be the lifetime of the tube. Since X is an exponential random variable, there
is no deterioration with age of the tube. Hence
P (X > 20 | X > 10) = P (X > 10) = 1 − [1 − e(−1/10)10 ] ≈ 0.37. Example 7.13 Suppose that, on average, two earthquakes occur in San Francisco and two
in Los Angeles every year. If the last earthquake in San Francisco occurred 10 months ago and
the last earthquake in Los Angeles occurred two months ago, what is the probability that the
next earthquake in San Francisco occurs after the next earthquake in Los Angeles?
Solution: It can be assumed that the number of earthquakes in San Francisco and Los Angeles
are both Poisson processes with common rate λ = 2. Hence the times between two consecutive
earthquakes in Los Angeles and two consecutive earthquakes in San Francisco are both exponentially distributed with the common mean 1/λ = 1/2. Because of the memoryless property
of the exponential distribution, it does not matter when the last earthquakes in San Francisco
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 303 — #319
✐
✐
Section 7.3
Exponential Random Variables
303
and Los Angeles have occurred. The times between now and the next earthquake in San Francisco and the next earthquake in Los Angeles both have the same distribution. Since these time
periods are exponentially distributed with the same parameter, by symmetry, the probability
that the next earthquake in San Francisco occurs after that in Los Angeles is 1/2. Relationship between Exponential and Geometric: Recall that if a Bernoulli trial is performed successively and independently, then the number of trials until the first success occurs
is geometric. Furthermore, the number of trials between two consecutive successes is also geometric. Sometimes exponential is considered to be the continuous analog of geometric because,
for a Poisson process, the time it will take until the first event occurs is exponential, and the time
between two consecutive events is also exponential. Moreover, exponential is the only memoryless continuous distribution, and geometric is the only memoryless discrete distribution. It is
also interesting to know that if X is an exponential random variable, then [X], the integer part
of X; i.e., the greatest integer less than or equal to X, is geometric (see Exercise 17). Remark 7.2 In this section, we showed that if {N (t) : t ≥ 0} is a Poisson process with rate
λ, then the interarrival times of the process form an independent sequence of identically distributed exponential random variables with mean 1/λ. Using the tools of an area of probability
called renewal theory, we can prove that the converse of this fact is also true:
If, for some process, N (t) is the number of “events” occurring in [0, t], and
if the times between consecutive events form a sequence of independent
and identically
distributed exponential random variables with mean 1/λ,
then N (t) : t ≥ 0 is a Poisson process with rate λ.
EXERCISES
A
1.
Customers arrive at a post office at a Poisson rate of three per minute. What is the
probability that the next customer does not arrive during the next 3 minutes?
2.
Find the median of an exponential random variable with rate λ. Recall that for a continuous distribution F, the median Q0.5 is the point at which F (Q0.5 ) = 1/2.
3.
Let X be an exponential random variable with mean 1. Find the probability density
function of Y = − ln X.
4.
The time between the first and second heart attacks for a certain group of people is an
exponential random variable. If 50% of those who have had a heart attack will have
another one within the next five years, what is the probability that a person who had one
heart attack five years ago will not have another one in the next five years?
5.
Guests arrive at a hotel, in accordance with a Poisson process, at a rate of five per hour.
Suppose that for the last 10 minutes no guest has arrived. What is the probability that
(a) the next one will arrive in less than 2 minutes; (b) from the arrival of the tenth to the
arrival of the eleventh guest takes no more than 2 minutes?
6.
Let X be an exponential random variable with parameter λ. Find
P X − E(X) ≥ 2σX .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 304 — #320
✐
✐
304
Chapter 7
Special Continuous Distributions
7.
Suppose that losses under an insurance policy are distributed exponentially with mean
$6000. If the amount of deductible for each loss is $500, what is the median of the insurance company’s claim payment under this policy? Note that the median of a continuous
random variable Y is a number m, which satisfies P (Y ≤ m) = P (Y ≥ m) = 0.5.
8.
Suppose that, at an Italian restaurant, the time, in minutes, between two customers ordering pizza is exponential with parameter λ. What is the probability that (a) no customer
orders pizza during the next t minutes; (b) the next pizza order is placed in at least t
minutes but no later than s minutes (t < s)?
9.
Suppose that the time it takes for a novice secretary to type a document is exponential
with mean 1 hour. If at the beginning of a certain eight-hour working day the secretary
receives 12 documents to type, what is the probability that she will finish them all by the
end of the day?
10.
The time between consecutive arrival times of two buses at a station is exponential with
mean 20 minutes. Every day, during the operating hours of the buses, Kayla arrives at
the station at a random time and waits for a bus to arrive. However, while she expects
that her average waiting time be only 10 minutes, she complains that, on the average,
she must wait much longer for a bus to arrive. What is the fallacy in Kayla’s argument?
11.
The profit is $350 for each computer assembled by a certain person. Suppose that the
assembler guarantees his computers for one year and the time between two failures of a
computer is exponential with mean 18 months. If it costs the assembler $40 to repair a
failed computer, what is the expected profit per computer?
Hint: Let N (t) be the number of times that the computer fails in [0, t]. Then
N (t) : t ≥ 0 is a Poisson process with parameter λ = 1/18.
12.
Mr. Jones is waiting to make a phone call at a train station. There are two public telephone booths next to each other, occupied by two persons, say A and B . If the duration
of each telephone call is an exponential random variable with λ = 1/8, what is the
probability that among Mr. Jones, A, and B, Mr. Jones will not be the last to finish his
call?
13.
The service times at a two-teller bank are independent exponential random variables
each with parameter λ. At a random time Jim is being served by teller 2, Jeff is being
served by teller 1, Judy and Julie are in the queue waiting to be served, and Judy is ahead
of Julie. What is the probability that when Julie’s service is completed Jim is still being
served?
14.
In a factory, a certain machine operates for a period which is exponentially distributed
with parameter λ. Then it breaks down and will be in the repair shop for a period, which
is also exponentially distributed with mean 1/λ. The operating and the repair times are
independent. For this machine, we say that a change of “state” occurs each time that
it breaks down, or each time that it is fixed. In a time interval of length t, find the
probability mass function of the number of times a change of state occurs.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 305 — #321
✐
✐
Section 7.3
Exponential Random Variables
305
B
15.
In data communication, messages are usually combinations of characters, and each character consists of a number of bits. A bit is the smallest unit of information and is either
1 or 0. Suppose that L, the length of a character (in bits) is a geometric random variable
with parameter p. If a sender emits messages at the rate of 1000 bits per second, what is
the distribution of T, the time it takes the sender to emit a character?
16.
The random variable X is said to be a Laplace random variable or double
exponentially distributed if its probability density function is given by
f (x) = ce−|x| ,
−∞ < x < +∞.
(a)
Find the value of c.
(b)
Prove that E(X 2n ) = (2n)! and E(X 2n+1 ) = 0.
17.
Let X, the lifetime (in years) of a radio tube, be exponentially distributed with mean
1/λ. Prove that [X], the integer part of X, which is the complete number of years that
the tube works, is a geometric random variable.
18.
Prove that if X is a positive, continuous, memoryless random variable with distribution
function F, then F (t) = 1 − e−λt for some λ > 0. This shows that the exponential is
the only distribution on (0, ∞) with the memoryless property.
Self-Quiz on Section 7.3
Time allotted: 20 Minutes
Each problem is worth 5 points.
1.
When solid state hard drives fail to function, it is not possible to repair them. They should
be replaced. Such drives, manufactured by Seaportal Technology, have lifetimes that are
exponentially distributed with mean 8 years. An online company, which sells this brand
of hard drives, offers its own warranty that fully refunds the price of the hard drive if
it fails within the first year, refunds half of its price if it fails during the second year,
and offers no refund if it fails thereafter. If the online company sells each hard drive for
$250, what is the expected value of its loss due to warranty claims per hard drive?
2.
Patients arrive at a pharmacy for flu shots at a Poisson rate of 3 per hour. If the last patient, who had a flu shot at the pharmacy, arrived 45 minutes ago, what is the probability
that the next patient does not arrive for a flu shot within the next 15 minutes?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 306 — #322
✐
✐
306
Chapter 7
7.4
GAMMA DISTRIBUTIONS
Special Continuous Distributions
Let N (t) : t ≥ 0 be a Poisson process, X1 be the time of the first event, and for n ≥ 2,
let Xn be the time between the (n − 1)st and nth events. As we explained in Section 7.3,
{X1 , X2 , . . .} is a sequenceof identically distributed exponential random variables with mean
1/λ, where λ is the rate of N (t) : t ≥ 0 . For this Poisson process let X be the time of the
nth event. Then X is said to have a gamma distribution with parameters (n, λ). Therefore,
exponential is the time we will wait for the first event to occur, and gamma is the time we will
wait for the nth event to occur. Clearly, a gamma distribution with parameters (1, λ) is identical
with an exponential distribution with parameter λ.
Let X be a gamma random variable with parameters (n, λ). To find f, the probability
density function of X, note that {X ≤ t} occurs if the time of the nth event is in [0, t], that
is, if the number of events occurring in [0, t] is at least n. Hence F, the distribution function of
X, is given by
∞
X
e−λt (λt)i
.
F (t) = P (X ≤ t) = P N (t) ≥ n =
i!
i=n
Differentiating F, the probability density function f is obtained:
∞ h
X
(λt)i
(λt)i−1 i
+ λe−λt
i!
(i − 1)!
i=n
∞
∞
h
i
n−1
X
X
(λt)i−1 i
−λt (λt)
−λt (λt)
=
−λe
+ λe
+
λe−λt
i!
(n − 1)! i=n+1
(i − 1)!
i=n
∞
∞
h X
(λt)n−1 h X −λt (λt)i i
(λt)i i
+ λe−λt
+
λe
= −
λe−λt
i!
(n − 1)!
i!
i=n
i=n
f (t) =
− λe−λt
= λe−λt
(λt)n−1
.
(n − 1)!
The probability density function
(λx)n−1
λe−λx
(n − 1)!
f (x) =
0
if x ≥ 0
elsewhere
is called the gamma (or n-Erlang) density with parameters (n, λ).
Now we extend the definition of the gamma density from parameters (n, λ) to (r, λ),
where r > 0 is not necessarily a positive integer. As we shall see later, this extension has useful
applications in probability and statistics. In the formula of the gamma density function, the term
(n − 1)! is defined only for positive integers. So the only obstacle in such an extension is to find
a function of r that has the basic property of the factorial function, namely, n! = n · (n − 1)!,
and coincides with (n − 1)! when n is a positive integer. The function with these properties is
Γ : (0, ∞) → R defined by
Z ∞
Γ(r) =
tr−1 e−t dt.
0
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 307 — #323
✐
✐
Section 7.4
Gamma Distributions
307
The property analogous to n! = n · (n − 1)! is
Γ(r + 1) = rΓ(r),
r > 1,
which is obtained by an integration by parts applied to Γ(r + 1) with u = tr and dv = e−t dt:
Z ∞
Z ∞
h
i∞
r −t
r −t
tr−1 e−t dt
+r
t e dt = − t e
Γ(r + 1) =
0
0
0
Z ∞
tr−1 e−t dt = rΓ(r).
=r
0
To show that Γ(n) coincides with (n − 1)! when n is a positive integer, note that
Z ∞
e−t dt = 1.
Γ(1) =
0
Therefore,
Γ(2) = (2 − 1)Γ(2 − 1) = 1 = 1!,
Γ(3) = (3 − 1)Γ(3 − 1) = 2 · 1 = 2!,
Γ(4) = (4 − 1)Γ(4 − 1) = 3 · 2 · 1 = 3!.
Repetition of this process or a simple induction implies that Γ(n+1) = n!. Hence Γ(r+1) is
the natural generalization of n! for a noninteger r > 0. This motivates the following definition.
Definition 7.5
A random variable X with probability density function
−λx
(λx)r−1
λe
if x ≥ 0
Γ(r)
f (x) =
0
elsewhere
is said to have a gamma distribution with parameters (r, λ), λ > 0, r > 0.
Let X be a gamma random variable with parameters (r, λ). To find E(X) and Var(X),
note that for all n ≥ 0,
E(X n ) =
Z ∞
0
xn
λe−λx (λx)r−1
λr
dx =
Γ(r)
Γ(r)
Z ∞
xn+r−1 e−λx dx.
0
Let t = λx; then dt = λ dx, so
Z ∞
Z ∞ n+r−1
λr
λr
Γ(n + r)
t
n
−t 1
E(X ) =
dt =
e
.
tn+r−1 e−t dt =
n+r−1
n+r
Γ(r) 0 λ
λ
Γ(r)λ
Γ(r)λn
0
For n = 1, this gives
E(X) =
rΓ(r)
r
Γ(r + 1)
=
= .
Γ(r)λ
λΓ(r)
λ
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 308 — #324
✐
✐
308
Chapter 7
Special Continuous Distributions
For n = 2,
E(X 2 ) =
(r + 1)Γ(r + 1)
Γ(r + 2)
(r + 1)rΓ(r)
r2 + r
=
.
=
=
Γ(r)λ2
λ2 Γ(r)
λ2 Γ(r)
λ2
Thus
Var(X) =
We have shown that
r 2 + r r 2
r
−
= 2.
2
λ
λ
λ
For a gamma random variable with parameters r and λ,
E(X) =
r
λ
,
Var(X) =
r
λ2
,
σX =
√
r
λ
.
Figures 7.12 and 7.13 demonstrate the shape of the gamma density function for several
values of r and λ.
Example 7.14 Suppose that, on average, the number of β -particles emitted from a radioactive substance is four every second. What is the probability that it takes at least 2 seconds before
the next two β -particles are emitted?
Solution: Let N (t) denote the number
of β -particles emitted from a radioactive substance in
[0, t]. It is reasonable to
assume that N (t) : t ≥ 0 is a Poisson process. Let 1 second be the
time unit; then λ = E N (1) = 4. X, the time between now and when the second β -particle
is emitted, has a gamma distribution with parameters (2, 4). Therefore,
Z ∞
Z ∞ −4x
4e (4x)2−1
dx =
16xe−4x dx
P (X ≥ 2) =
Γ(2)
2
2
h
i∞ Z ∞
= − 4xe−4x
−
−4e−4x dx = 8e−8 + e−8 ≈ 0.003.
2
2
Note that an alternative solution to this problem is
P (X ≥ 2) = P N (2) ≤ 1 = P N (2) = 0 + P N (2) = 1
=
e−8 (8)1
e−8 (8)0
+
= 9e−8 ≈ 0.003. 0!
1!
Example 7.15 There are 100 questions in a test. Suppose that, for all s > 0 and t > 0, the
event that it takes t minutes to answer one question is independent of the event that it takes s
minutes to answer another one. If the time that it takes to answer a question is exponential with
mean 1/2, find the distribution, the average time, and the standard deviation of the time it takes
to do the entire test.
Solution: Let Xbe the time to answer a question and N (t) the number of questions answered
by time t. Then N (t) : t ≥ 0 is a Poisson process at the rate of λ = 1/E(X) = 2 per
minute. Therefore, the time that it takes to complete all the questions is gamma with parameters
(100, 2). p
The average
ptime to finish the test is r/λ = 100/2 = 50 minutes with standard
deviation r/λ2 = 100/4 = 5. ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 309 — #325
✐
✐
Section 7.4
Gamma Distributions
309
f(x)
0.1
r=2
r=
3
r=4
x
0
30
Figure 7.12
Gamma densities for λ = 1/4.
f(x)
= 0.6
= 0.2
= 0.5
30
Figure 7.13
x
Gamma densities for r = 4.
Relationship between Gamma and Negative Binomial: Recall that if a Bernoulli trial is
performed successively and independently, then the number of trials until the r th success occurs is negative binomial. Sometimes gamma is viewed as the continuous analog of negative
binomial, one reason being that, for a Poisson process, the time it will take until the r th event
occurs is gamma. There are more serious relationships between the two distributions that we
will not discuss. For example, in some sense, in limit, negative binomial’s distribution approaches gamma distribution. EXERCISES
A
1.
Show that the gamma density function with parameters (r, λ) has a unique maximum at
(r − 1)/λ.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 310 — #326
✐
✐
310
Chapter 7
Special Continuous Distributions
2.
Let X be a gamma random variable with parameters (r, λ). Find the distribution function of cX, where c is a positive constant.
3.
In a hospital, babies are born at a Poisson rate of 12 per day. What is the probability that
it takes at least seven hours before the next three babies are born?
4.
Let X be the time to failure of a mobile robot that is used as an automated guide vehicle.
Suppose that E(X) = 7 years, and X has a gamma distribution with r = 4. Find
Var(X) and P (5 ≤ X ≤ 7).
5.
Let f be the probability
R ∞density function of a gamma random variable X, with parameters (r, λ). Prove that −∞ f (x) dx = 1.
6.
Customers arrive at a restaurant at a Poisson rate of 12 per hour. If the restaurant makes
a profit only after 30 customers have arrived, what is the expected length of time until
the restaurant starts to make a profit?
7.
Experience has shown that the bugs of a certain smartphone app cause battery drainage.
The time that it will take for the app developers to locate and remove the bugs is a gamma
random variable with mean 300 and standard deviation 100 days. What is the probability
that it will take less than 250 days for the app developers to get rid of its bugs?
8.
Suppose that the lifetime of a certain wheel bearing manufactured by a company is a
gamma random variable with mean 75,000 and standard deviation 15,000 miles. If the
manufacturer offers a free repair warranty of 60,000 miles on its wheel bearings, what
percentage of its wheel bearings must be repaired free of charge?
9.
A manufacturer produces light bulbs at a Poisson rate of 200 per hour. The probability
that a light bulb is defective is 0.015. During production, the light bulbs are tested one
by one, and the defective ones are put in a special can that holds up to a maximum of 25
light bulbs. On average, how long does it take until the can is filled?
10.
The lifetime of the alarm system installed in Herman Hall at a private university is
exponential with mean 1/λ. Beginning at time T, the facility management crew inspects
the system every T units of time. Assuming that without inspection there is no way for
the facility management to know whether the system is up or down, find the expected
value of the time until the system is found to be down.
B
11.
For n = 0, 1, 2, 3, . . . , calculate Γ(n + 1/2).
12.
(a)
13.
Let X be a normal random variable with mean µ and standard deviation σ . Find
X − µ 2
.
the distribution function of W =
σ
Howard enters a bank that has n tellers. All the tellers are busy serving customers, and
there is exactly one queue being served by all tellers, with one customer ahead of Howard
waiting to be served. If the service time of a customer is exponential with parameter λ,
find the distribution of the waiting time for Howard in the queue.
Let Z be a standard normal random variable. Show that the random variable Y =
Z 2 is gamma and find its parameters.
(b)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 311 — #327
✐
✐
Section 7.5
14.
Beta Distributions
311
In data communication, messages are usually combinations of characters, and each character consists of a number of bits. A bit is the smallest unit of information and is either 1
or 0. Suppose that the length of a character (in bits) is a geometric random variable with
parameter p. Suppose that a sender emits messages at the rate of 1000 bits per second.
What is the distribution of T, the time it takes the sender to emit a message combined of
k characters of independent lengths?
Hint: Let N (t) be the number of characters emitted at or prior to t. First argue that
N (t) : t ≥ 0 is a Poisson process and find its parameter.
Self-Quiz on Section 7.4
Time allotted: 15 Minutes
Each problem is worth 5 points.
1.
A vacuum tube is a device that facilitates the free passage of electric current between
electrodes in an evacuated container. The lifetime, in thousands of hours, of a vacuum
tube manufactured by a certain company is a random variable with probability density
function
1 3 −x/2
x e
,
x ≥ 0.
f (x) =
96
Find the mean and the standard deviation of a randomly selected vacuum tube manufactured by this company.
2.
A company is behind in manufacturing upper fuser rolling bearings for its photocopiers
and needs to produce 500 of them right away. If the company produces these items at a
Poisson rate of 5 per hour, what is the expected period and the standard deviation until
it manufactures the needed items?
7.5
BETA DISTRIBUTIONS
A random variable X is called beta with parameters (α, β), α > 0, β > 0 if f, its probability
density function, is given by
1
xα−1 (1 − x)β−1
if 0 < x < 1
B(α, β)
f (x) =
0
otherwise,
where
B(α, β) =
Z 1
0
xα−1 (1 − x)β−1 dx.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 312 — #328
✐
✐
312
Chapter 7
Special Continuous Distributions
B(α, β) is related to the gamma function by the relation
B(α, β) =
Γ(α)Γ(β)
.
Γ(α + β)
Beta distributions are often appropriate models for random variables that vary between
two finite limits—an upper and a lower. For this reason the following are examples of random
variables that might be beta.
1.
The fraction of people in a community who use a certain product in a given period of
time
2.
The percentage of total farm acreage that produces healthy watermelons
3.
The distance from one end of a tree to that point where it breaks in a severe storm
In these three instances the random variables are restricted between 0 and 1, 0 and 100, and 0
and the length of the tree, respectively.
Beta density also occurs in a natural way in the study of the median of a sample of random
points from (0, 1). Let X(1) be the smallest of these numbers, X(2) be the second smallest,
. . . , X(i) be the ith smallest, . . . , and X(n) be the largest of these numbers. If n = 2k + 1
is odd, X(k+1) is called the median of these n random numbers, whereas if n = 2k is even,
[X(k) + X(k+1) ]/2 is called the median. It can be shown that the median of (2n + 1) random
numbers from the interval (0, 1) is a beta random variable with parameters (n + 1, n + 1).
As Figures 7.14 and 7.15 show, by changing the values of parameters α and β, beta densities
cover a wide range of different shapes. If α = β, the median is x = 1/2, and the probability
density function of the beta random variable is symmetric about the median. In particular, if
α = β = 1, the uniform density over the interval (0, 1) is obtained.
To find the expected value and the variance of a beta random variable X, with parameters
(α, β), note that for n ≥ 1,
Z 1
B(α + n, β)
Γ(α + n)Γ(α + β)
1
n
xα+n−1 (1 − x)β−1 dx =
=
.
E(X ) =
B(α, β) 0
B(α, β)
Γ(α)Γ(α + β + n)
Letting n = 1 and n = 2 in this relation, we find that
E(X) =
Γ(α + 1)Γ(α + β)
αΓ(α)Γ(α + β)
α
=
=
,
Γ(α)Γ(α + β + 1)
Γ(α) (α + β)Γ(α + β)
α+β
E(X 2 ) =
(α + 1)α
Γ(α + 2)Γ(α + β)
=
.
Γ(α)Γ(α + β + 2)
(α + β + 1)(α + β)
Thus
2
Var(X) = E(X 2 ) − E(X) =
αβ
.
(α + β + 1)(α + β)2
We have established the following formulas:
For a beta random variable with parameters (α, β),
E(X) =
α
α+β
,
Var(X) =
αβ
(α + β + 1)(α + β)2
.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 313 — #329
✐
✐
Section 7.5
Beta Distributions
313
f (x)
<1
>1
<1
<1
x
Figure 7.14
Beta densities for α < 1, β < 1 and α < 1, β > 1.
f (x)
>1
>1
>1
<1
= 2,
=1
x
Figure 7.15
Beta densities for the indicated values of α and β.
There is also an interesting relation between beta and binomial distributions that the following theorem explains.
Theorem 7.2 Let α and β be positive integers, X be a beta random variable with parameters α and β, and Y be a binomial random variable with parameters α + β − 1 and p,
0 < p < 1. Then
P (X ≤ p) = P (Y ≥ α).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 314 — #330
✐
✐
314
Chapter 7
Proof:
Special Continuous Distributions
From the definition,
(α + β − 1)!
(α − 1)! (β − 1)!
Z p
xα−1 (1 − x)β−1 dx
0
!Z
p
α+β−2
= (α + β − 1)
xα−1 (1 − x)β−1 dx
α−1
0
!
α+β−1
X
α+β−1 i
p (1 − p)α+β−1−i
=
i
i=α
P (X ≤ p) =
= P (Y ≥ α),
where the third equality follows from using integration by parts (α − 1) times.
Example 7.16 The proportion of the stocks that will increase in value tomorrow is a beta
random variable with parameters α and β . Today, these parameters were determined to be
α = 5 and β = 4, by a financial analyst, after she assessed the current political, social,
economical, and financial factors. What is the probability that tomorrow the values of at least
70% of the stocks will move up?
Solution: Let X be the proportion of the stocks that will move up in value tomorrow; X is beta
with parameters α = 5 and β = 4. Hence the desired probability is
Z 1
Z 1
1
4
3
x
(1
−
x)
dx
=
280
x4 (1 − x)3 dx = 0.194. P (X ≥ 0.70) =
B(5,
4)
0.70
0.70
Example 7.17 A beam of length ℓ, rigidly supported at both ends, is hit suddenly at a
random point. This leads to a break in the beam at a position X units from the right end. If X/ℓ
is beta with parameters α = β = 3, find E(X), Var(X), and P (ℓ/5 < X < ℓ/4).
Solution: Since
E(X/ℓ) =
α
3
1
= = ,
α+β
6
2
Var(X/ℓ) =
9
1
αβ
,
=
=
2
2
(α + β + 1)(α + β)
7(6)
28
we have that E(X) = ℓ/2, Var(X) = ℓ2 /28, and
P (ℓ/5 < X < ℓ/4) = P (1/5 < X/ℓ < 1/4) =
Z 1/4
1/5
= 30
Z 1/4
1/5
x2 (1 − x)2 dx =
1
x2 (1 − x)2 dx
B(3, 3)
72953
= 0.046. 1, 600, 000
Example 7.18 Suppose that all 25 passengers of a commuter plane, departing at 2:00 P.M.,
arrive at random times between 12:45 and 1:45 P.M. Find the probability that the median of the
arrival times of the passengers is at or prior to 1:12 P.M.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 315 — #331
✐
✐
Section 7.5
Beta Distributions
315
Solution: Let the arrival time of a randomly selected passenger be W hours past 12:45.
Then W is a uniform random variable over the interval (0, 1). Let X be the median of the
arrival times of the 25 passengers. Since the median of (2n + 1) random numbers from (0, 1)
is beta with parameters (n + 1, n + 1), we have that X is beta with parameters (13, 13).
Since the length of the time interval from 12:45 to 1:12 is 27 minutes, the desired quantity is
P (X ≤ 27/60). To calculate this, let Y be a binomial random variable with parameters 25
and p = 27/60 = 0.45. Then, by Theorem 7.2,
!
25
X
25
P (X ≤ 0.45) = P (Y ≥ 13) =
(0.45)i (0.55)25−i ≈ 0.306. i
i=13
EXERCISES
A
1.
Is the following the probability density function of some beta random variableX ? If so,
find E(X) and Var(X).
(
12x(1 − x)2 0 < x < 1
f (x) =
0
otherwise.
2.
Is the following a probability density function? Why or why not?
(
120x2 (1 − x)4 0 < x < 1
f (x) =
0
otherwise.
3.
For what value of c is the following a probability density function of some random
variable X ? Find E(X) and Var(X).
(
cx4 (1 − x)5 0 < x < 1
f (x) =
0
otherwise.
4.
Suppose that new blood pressure medicines introduced are effective on 100p% of the
patients, where p is a beta random variable with parameters α = 20 and β = 13. What
is the probability that a new blood pressure medicine is effective on at least 60% of the
hypertensive population?
5.
The proportion of resistors a procurement office of an engineering firm orders every
month, from a specific vendor, is a beta random variable with mean 1/3 and variance
1/18. What is the probability that next month, the procurement office orders at least
7/12th of its purchase from this vendor?
6.
At a certain university, the fraction of students who get a C in any section of a certain
course is uniform over (0, 1). Find the probability that the median of these fractions for
the 13 sections of the course that are offered next semester is at least 0.40.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 316 — #332
✐
✐
316
Chapter 7
Special Continuous Distributions
7.
Suppose that while daydreaming, the fraction X of the time that one commits brave
deeds is beta with parameters (5, 21). What is the probability that next time Jeff is
daydreaming, he commits brave deeds at least 1/4 of the time?
8.
A beam of length ℓ is rigidly supported at both ends. Experience shows that whenever
the beam is hit at a random point, it breaks at a position X units from the right end,
where X/ℓ is a beta random variable. If E(X) = 3ℓ/7 and Var(X) = 3ℓ2 /98, find
P (ℓ/7 < X < ℓ/3).
9.
For complicated projects such as construction of spacecrafts, project managers estimate
two quantities, a and b, the minimum and maximum lengths of time it will take for a
project to be completed, respectively. In estimating b, they consider all of the possible
complications that might delay the completion date. Experience shows that the actual
length of time it takes for a project to be completed is a random variable defined by
Y = a + (b − a)X,
where X is beta with parameters α and β (α > 0, β > 0) that can be determined for
the project.
(a)
Find E(Y ) and Var(Y ).
(b)
Find the probability density function of Y .
(c)
Suppose that it takes Y years to complete a specific project, where Y = 2 + 4X,
and X is beta with parameters α = 2 and β = 3. What is the probability that it
takes less than 3 years to complete the project?
B
10.
Under what conditions and about which point(s) is the probability density function of a
beta random variable symmetric?
11.
For α, β > 0, show that
B(α, β) = 2
Z ∞
t2α−1 (1 + t2 )−(α+β) dt.
0
Hint: Make the substitution x = t2 /(1 + t2 ) in
Z 1
xα−1 (1 − x)β−1 dx.
B(α, β) =
0
12.
Prove that
B(α, β) =
13.
Γ(α)Γ(β)
.
Γ(α + β)
For an integer n ≥ 3, let X be a random variable with the probability density function
n + 1
Γ
2 −(n+1)/2
2 1 + x
, −∞ < x < ∞.
f (x) = √
n
n
nπ Γ
2
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 317 — #333
✐
✐
Section 7.6
Survival Analysis and Hazard Functions
317
Such random variables have significant applications in statistics. They are called
t-distributed with n degrees of freedom. Using the previous two exercises, find E(X)
and Var(X).
Self-Quiz on Section 7.5
Time allotted: 15 Minutes
Each problem is worth 5 points.
1.
Suppose that each time the price of a commodity decreases, X, the proportion of change
in the price, is a beta random variable with parametrs α = 1 and β = 10. That is, if,
the price of commodity changes, say, from q1 to q2 , q2 < q1 , then X = (q1 − q2 )/q1
is beta with parameters 1 and 10. Find the probability that next time the price falls, the
proportion of change is at least 0.12.
2.
Let X be a random number between 0 and 1. Show that Y = X 2 is beta with parameters
1/2 and 1.
⋆ 7.6
SURVIVAL ANALYSIS AND HAZARD FUNCTIONS
In this section, we will study the risk or rate of failure, per unit of time, of lifetimes that have
already survived a certain length of time. The term lifetime is broad and applies to appropriate
quantities in various models in science, engineering, and business. Depending on the context
of a study, by a lifetime, we might mean the lifetime of a machine, an electrical component,
a living organism, an organism that is under an experimental medical treatment, a financial
contract such as a mortgage, the waiting time until a customer arrives at a bank, or the time it
takes until a customer is served at a post office. The failure for a living organism is its death.
For a mortgage, it is when the mortgage is paid off. For the time until a customer arrives at a
bank, failure occurs when a customer enters the bank.
In the following discussion, for convenience, we will talk about the lifetime of a system.
However, our definitions and results are general and apply to other examples such as the ones
just stated. Let the lifetime of a system be X, where X is a nonnegative, continuous random
variable with distribution function F and probability density function f . The function defined
by
F̄ (t) = 1 − F (t) = P (X > t)
is called the survival function of X . For t > 0, F̄ (t) is the probability that the system has
already survived at least t units of time. Furthermore, the following relation, shown in Remark 6.4, enables us to calculate E(X), the expected value of the lifetime of the system, using
F̄ (t):
Z
∞
E(X) =
F̄ (t) dt.
0
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 318 — #334
✐
✐
318
Chapter 7
Special Continuous Distributions
The probability that a system, which has survived at least t units of time, fails on the time
interval (t, t + ∆t ] is given by
P (X ≤ t + ∆t | X > t) =
F (t + ∆t ) − F (t)
P (t < X ≤ t + ∆t )
.
=
P (X > t)
F̄ (t)
To find the instantaneous failure rate of a system of age t at time t, note that during the interval
(t, t + ∆t ], the system fails at a rate of
1
F (t + ∆t) − F (t)
1
·
P (X ≤ t + ∆t | X > t) =
∆t
∆t
F̄ (t)
per unit of time. As ∆t → 0, this quantity approaches the instantaneous failure rate of the
system at time t, given that it has already survived t units of time. Let
λ(t) = lim
1
∆t →0 F̄ (t)
·
F (t + ∆t ) − F (t)
∆t
=
1
F (t + ∆t ) − F (t)
· lim
∆t
F̄ (t) ∆t →0
=
F ′ (t)
f (t)
=
.
F̄ (t)
F̄ t)
Then λ(t) is called the hazard function of the random variable X . It is the instantaneous
failure rate at t, per unit of time, given that the system has already survived until time t. Note
that λ(t) ≥ 0, but it is not a probability density function.
Remark 7.3 An alternate term used for F̄ (t), the survival function of X, is the reliability
function. Other terms used for λ(t), the hazard function, are hazard rate, failure rate function,
failure rate, intensity rate, and conditional failure rate; sometimes actuarial scientists call it
force of mortality. We know that if ∆t > 0 is very small,
P (t < X ≤ t + ∆t ) =
Z t+∆t
f (x) dx
t
is the area under f from t to t + ∆t . This area is almost equal to the area of a rectangle with
sides of length ∆t and f (t). Thus
P (t < X ≤ t + ∆t ) ≈ f (t)∆t .
The smaller ∆t , the closer f (t)∆t is to the probability that the system fails in (t, t + ∆t ]. This
approximation implies that, for infinitesimal ∆t > 0,
λ(t)∆t =
f (t)∆t
P (t < X ≤ t + ∆t )
= P (X ≤ t + ∆t | X > t).
≈
P (X > t)
F̄ (t)
Therefore,
For very small values of ∆t > 0, the quantity λ(t)∆t is approximately the
conditional probability that the system fails in (t, t + ∆t ), given that it has
lasted at least until t. That is,
P (X ≤ t + ∆t | X > t) ≈ λ(t)∆t .
(7.5)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 319 — #335
✐
✐
Section 7.6
Survival Analysis and Hazard Functions
319
Example 7.19 Dr. Hirsch has informed one of his employees, Dr. Kizanis, that she will
receive her next year’s employment contract at a random time between 10:00 A.M. and 3:00 P.M.
Suppose that Dr. Kizanis will receive her contract X minutes past 10:00 A.M. Then X is a
uniform random variable over the interval (0, 300). Hence the probability density function of
X is given by
(
1/300 if 0 < t < 300
f (t) =
0
otherwise.
Straightforward calculations show that the survival function of X is given by
1
if t < 0
F̄ (t) = P (X > t) = 300 − t if 0 ≤ t < 300
300
0
if t ≥ 300.
The hazard function, λ(t) = f (t)/F̄ (t), is defined only for t < 300. It is unbounded for
t ≥ 300. We have
if t < 0
0
λ(t) =
1
if 0 ≤ t < 300.
300 − t
Since λ(0) = 0.003333, at 10:00 A.M. the instantaneous arrival rate of the contract (the failure
rate in this context) is 0.003,333 per minute. At noon this rate will increase to λ(120) =
0.0056, at 2:00 P.M. to λ(240) = 0.017, at 2:59 P.M. it reaches λ(299) = 1. One second before
3:00 P.M., the instantaneous arrival of the contract is, approximately, λ(299.983) = 58.82.
This shows that if Dr. Kizanis has not received her contract by one second before 3:00 P.M.,
the instantaneous arrival rate at that time is very high. That rate approaches ∞, as the time
approaches 3:00 P.M. To translate all these into probabilities, let ∆t = 1/60. Then f (t)∆t is
approximately the probability that the contract will arrive within one second after t, whereas
λ(t)∆t is approximately the probability that the contract will arrive within one second after t,
given that it has not yet arrived by time t. The following table shows these probabilities at the
indicated times:
t
0
120
240
299
299.983
f (t)∆t
0.000,056
0.000,056
0.000,056
0.000,056
0.000,056
λ(t)∆t
0.000,056
0.000,093
0.000,283
0.016,7
0.98
The fact that f (t)∆t is constant is expected because X is uniformly distributed over (0, 300),
and all of the intervals under consideration are subintervals of (0, 300) of equal lengths.
Let X be a nonnegative continuous random variable with distribution function F, probability density function f, survival function F̄ , and hazard function λ(t). We will now calculate
F̄ and f in terms
of λ(t). The formulas obtained are very useful in various applications. Let
G(t) = − ln F̄ (t) . Then
f (t)
= λ(t).
G′ (t) =
F̄ (t)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 320 — #336
✐
✐
320
Chapter 7
Consequently,
Special Continuous Distributions
Z t
′
G (u) du =
0
Hence
Z t
λ(u) du.
Z t
λ(u) du.
0
G(t) − G(0) =
0
Since X is nonnegative, G(0) = − ln 1 − F (0) = − ln 1 = 0. So
implies that
− ln F̄ (t) − 0 = G(t) − G(0) =
F̄ (t) = exp −
By this equation,
Z t
h
0
F (t) = 1 − exp −
Z t
λ(u) du
0
λ(u) du .
Z t
0
(7.6)
i
λ(u) du .
Differentiating both sides of this relation, with respect to t, yields
h Z t
i
f (t) = λ(t) exp −
λ(u) du .
(7.7)
0
This demonstrates that the hazard function uniquely determines the probability density function.
In reliability theory, a branch of engineering, it is often observed that λ(t), the hazard
function of the lifetime of a manufactured machine, is initially large due to undetected defective
components during testing. Later, λ(t) will decrease and remains more or less the same until
a time when it increases again due to aging, which makes worn-out components more likely
to fail. A random variable X is said to have an increasing failure rate if λ(t) is increasing.
It is said to have a decreasing failure rate if λ(t) is decreasing. In the next example, we will
show that if λ(t) is neither increasing nor decreasing; that is, if λ(t) is a constant, then X is an
exponential random variable. In such a case, the fact that aging does not change the failure rate
is consistent with the memoryless property of exponential random variables. To summarize, the
lifetime of a manufactured machine is likely to have a decreasing failure rate in the beginning,
a constant failure rate later on, and an increasing failure rate due to wearing out after an aging
process. Similarly, lifetimes of living organisms, after a certain age, have increasing failure
rates. However, for a newborn baby, the longer he or she survives, the chances of surviving is
higher. This means that as t increases, λ(t) decreases. That is, λ(t) is decreasing.
Example 7.20
In this example, we will prove the following important theorem:
Let λ(t), the hazard function of a continuous, nonnegative random variable X, be a constant λ. Then X is an exponential random variable with
parameter λ.
To show this theorem, let f be the probability density function of X . By (7.7),
f (t) = λ exp −
Z t
0
λ du = λe−λt .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 321 — #337
✐
✐
Section 7.6
Survival Analysis and Hazard Functions
321
This is the probability density function of an exponential random variable with parameter λ.
Thus a random variable with constant hazard function is exponential. As mentioned previously,
this result is not surprising. The fact that an exponential random variable is memoryless implies
that age has no effect on the distribution of the remaining lifetimes of exponentially distributed
random variables. For systems that have exponential lifetime distribution, failures are not due
to aging and wearing out. In fact, such systems do not wear out at all. Failures occur abruptly.
There is no transition from the previous “state” of the system and no preparation for, or gradual
approach to, failure. EXERCISES
1.
Suppose that the lifetime of a device, in hundreds of hours, is a random variable with
probability density function
3t2
0≤t≤1
f (t) =
0
otherwise.
(a)
Find λ(t), the hazard function of the lifetime of the device.
(b)
Find the instantaneous failure rates of such a device if it has survived 50 hours, 90
hours, and 99 hours, respectively. Using these, calculate the approximate probabilities that the device fails within one hour if it has survived 50 hours, 90 hours,
and 99 hours, respectively.
2.
The hazard function of a random variable X with the set of possible values [1, 3), is
given by
3t2
,
1 ≤ t < 3.
λ(t) =
27 − t3
Find the probability density function of X .
3.
Suppose that the lifetime of a cell phone chip manufactured by a wireless company, in
years, is gamma with parameters r = 4 and λ = 1. Find λ(t), the hazard function
of such a chip, and its instantaneous failure rate if it has survived for four years. Using
this, find an approximate value for the probability that an exactly four-year-old chip fails
within one week. Assume that a year is 52 weeks.
4.
Experience shows that the failure rate of a certain electrical component is a linear function. Suppose that after two full days of operation, the failure rate is 10% per hour and
after three full days of operation, it is 15% per hour.
5.
(a)
Find the probability that the component operates for at least 30 hours.
(b)
Suppose that the component has been operating for 30 hours. What is the probability that it fails within the next hour?
One of the most popular distributions used to model the lifetimes of electric components
is the Weibull distribution, whose probability density function is given by
α
f (t) = αtα−1 e−t ,
t > 0, α > 0.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 322 — #338
✐
✐
322
Chapter 7
Special Continuous Distributions
Determine for which values of α the hazard function of a Weibull random variable is
increasing, for which values it is decreasing, and for which values it is constant.
Self-Quiz on Section 7.6
Time allotted: 15 Minutes
Each problem is worth 5 points.
1.
A batch of cell phones manufactured by a wireless company’s plant in the U.S.A. has
been ordered by the company’s main facility in Portugal. Suppose that the hazard function of the delivery lead time, in weeks, is λ(t) = 1/8, t ≥ 0. If the order for the batch
of cell phones was placed four weeks ago, what is the probability that it will arrive in
less than two weeks?
2.
The hazard function of a random variable X with the set of possible values (0, 1), is
given by
2
0<t<1
λ(t) = 1 − t
0
otherwise.
Find the probability density function of X .
CHAPTER 7 SUMMARY
◮ Uniform Random Variables Suppose that X is the value of the random point selected
from an interval (a, b). Then X is called a uniform random variable over (a, b). f, the probability density functions of X, E(X), Var(X), and σX are given by f (t) = 1/(b − a) if
a < t < b and√f (t) = 0, otherwise. E(X) = (a + b)/2, Var(X) = (b − a)2 /12, and
σX = (b − a)/ 12.
◮ De Moivre–Laplace Theorem Let X be a binomial random variable with parameters
n and p. Then for any numbers a and b, a < b,
Z b
2
X − np
1
lim P a < p
e−t /2 dt.
<b = √
n→∞
np(1 − p)
2π a
By this theorem, if X is a binomial random variable with parameters (n, p), the sequence of
probabilities
X − np
≤t ,
n = 1, 2, 3, . . . ,
P p
np(1 − p)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 323 — #339
✐
✐
Chapter 7
Summary
323
Z t
Z t
2
2
1
1
converges to √
e−x /2 dx, where the function Φ(t) = √
e−x /2 dx is a
2π −∞
2π −∞
distribution function itself.
◮ Standard Normal Random Variables A random variable X is called standard normal
if its distribution function is Φ. The probability density function of a standard normal random
variable, is given by
2
1
f (x) = Φ′ (x) = √ e−x /2 .
2π
The standard normal density function is a bell-shaped curve that is symmetric about the y -axis,
and E(X) = 0, Var(X) = σX = 1.
◮ Φ(t), the distribution function of the standard normal random variable, is the area under
the curve of f from −∞ to t. Because Φ(∞) = 1 and the curve is symmetric about the y -axis,
Φ(0) = 1/2, and Φ(−t) = 1 − Φ(t).
◮ Normal Random Variables A random variable X is called normal, with parameters µ
and σ, if its probability density function is given by
h −(x − µ)2 i
1
,
f (x) = √ exp
2σ 2
σ 2π
−∞ < x < ∞.
The parameters µ and σ that appear in the formula of the probability density function of X are
its expected value and standard deviation, respectively.
◮ If X ∼ N (µ, σ 2 ), then Z = (X − µ)/σ is N (0, 1). That is, if X ∼ N (µ, σ 2 ), the
standardized X is N (0, 1).
◮ Exponential Random Variables A continuous random variable X is called exponential
with parameter λ > 0 if its probability density function is given by f (t) = λe−λt , t ≥ 0, and
f (t) = 0 if t < 0. For an exponential random variable with parameter λ, E(X) = σX = 1/λ
and Var(X) = 1/λ2 .
◮ In a Poisson process, the time between two consecutive events is an exponential random
variable.
◮ An important feature of exponential distribution is its memoryless property. A nonnegative
random variable X is called memoryless if, for all s, t ≥ 0,
P (X > s + t | X > t) = P (X > s).
This shows that if, for example, X is the lifetime of some type of instrument, then there is no
deterioration with age of the instrument.
◮ Gamma Distributions
We call a random variable X with probability density function
−λx
λe (λx)r−1
if x ≥ 0
Γ(r)
f (x) =
0
elsewhere
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 324 — #340
✐
✐
324
Chapter 7
Special Continuous Distributions
R∞
a gamma distribution with parameters (r, λ), λ > 0, r > 0, where, Γ(r) = 0 tr−1 e−t dt.
For a gamma
random variable with parameters r and λ, E(X) = r/λ, Var(X) = r/λ2 , and
√
σX = r/λ.
◮ For a Poisson process with rate λ, the time of the nth event is gamma with parameters
(n, λ).
◮ Beta Distributions
A random variable X is called beta with parameters (α, β), α > 0,
β > 0 if f, its probability density function, is given by
f (x) =
1
xα−1 (1 − x)β−1 ,
B(α, β)
0 < x < 1,
R1
and f (x) = 0, otherwise, where B(α, β) = 0 xα−1 (1 − x)β−1 dx. B(α, β) is related to the
gamma function by the relation B(α, β) = Γ(α)Γ(β)/Γ(α + β). For a beta random variable
with parameters (α, β),
E(X) =
α
,
α+β
Var(X) =
αβ
.
(α + β + 1)(α + β)2
◮ Let α and β be positive integers, X be a beta random variable with parameters α and β,
and Y be a binomial random variable with parameters α + β − 1 and p, 0 < p < 1. Then
P (X ≤ p) = P (Y ≥ α).
◮ Survival Analysis and Hazard Functions
Let X be a continuous random variable
with probability density function f and distribution function F . The function defined by
F̄ (t) = 1 − F (t) = P (X > t) is called the survival function of X . Let λ(t) = f (t)/F̄ (t).
Then λ(t) is called the hazard function of the random variable X . If X is the lifetime of a
“system,” then λ(t) is the instantaneous failure rate at t, per unit of time, given that the system
has already survived until time t. For very small values of ∆t > 0, the quantity λ(t)∆t is
approximately the conditional probability that the system fails in (t, t + ∆t ), given that it has
lasted at least until t. That is, P (X ≤ t + ∆t | X > t) ≈ λ(t)∆t .
◮ Let X be a nonnegative continuous random variable with hazard function λ(t). Then the
hazard function uniquely determines the probability density function:
i
h Z t
λ(u) du .
f (t) = λ(t) exp −
0
In particular, by this formula, if λ(t), the hazard function of a continuous, nonnegative random
variable X, be a constant λ. Then X is an exponential random variable with parameter λ.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 325 — #341
✐
✐
Chapter 7
Review Problems
325
REVIEW PROBLEMS
1.
For a restaurant, the time it takes to deliver pizza (in minutes) is uniform over the interval
(25, 37). Determine the proportion of deliveries that are made in less than half an hour.
2.
It is known that the weight of a random woman from a community is normal with mean
130 pounds and standard deviation 20. Of the women in that community who weigh
above 140 pounds, what percent weigh over 170 pounds?
3.
One thousand random digits are generated. What is the probability that digit 5 is generated at most 93 times?
4.
Let X, the lifetime of a light bulb, be an exponential random variable with parameter λ.
Is it possible that X satisfies the following relation?
P (X ≤ 2) = 2P (2 < X ≤ 3).
If so, for what value of λ?
5.
The time that it takes for a computer system to fail is exponential with mean 1700 hours.
If a lab has 20 such computer systems, what is the probability that at least two fail before
1700 hours of use?
6.
Let X be a uniform random variable over the interval (0, 1). Calculate E(− ln X).
7.
Suppose that the diameter of a randomly selected disk produced by a certain manufacturer is normal with mean 4 inches and standard deviation 1 inch. Find the distribution
function of the diameter of a randomly chosen disk, in centimeters.
8.
Let X be an exponential random variable with parameter λ. Prove that
P (α ≤ X ≤ α + β) ≤ P (0 ≤ X ≤ β).
9.
The time that it takes for a calculus student to answer all the questions on a certain exam
is an exponential random variable with mean 1 hour and 15 minutes. If all 10 students
of a calculus class are taking that exam, what is the probability that at least one of them
completes it in less than one hour?
10.
Determine the value(s) of k for which the following is a probability density function.
2
f (x) = ke−x +3x+2 ,
11.
−∞ < x < ∞.
The number of minutes that a train from Milan to Rome
is late is an exponential random
variable X with parameter λ. Find P X > E(X) .
12.
Suppose that the weights of passengers taking an elevator in a certain building are normal
with mean 175 pounds and standard deviation 22. What is the minimum weight for a
passenger who outweighs at least 90% of the other passengers?
13.
The breaking strength of a certain type of yarn produced by a certain vendor is normal
with mean 95 and standard deviation 11. What is the probability that, in a random sample
of size 10 from the stock of this vendor, the breaking strengths of at least two are over
100?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 326 — #342
✐
✐
326
Chapter 7
14.
The number of phone calls to a specific exchange is a Poisson process with rate 23 per
hour. Calculate the probability that the time until the 91st call is at least 4 hours.
15.
Let X be a uniform random variable over the interval (1 − θ, 1 +
θ), where
0 < θ < 1 is a given parameter. Find a function of X, say g(X), so that E g(X) = θ 2 .
Special Continuous Distributions
Self-Test on Chapter 7
Time allotted: 120 Minutes
Each problem is worth 10 points.
1.
In a city, a passenger transit bus operates two routes that both stop at the Downtown
Center and at the Embarcadero stations. Route B1 operates every day between 6:00 A.M.
and 9:30 P.M. and arrives at Embarcadero every 15 minutes starting at 6:15 A.M. Route
B26 operates during the same period but arrives at Embarcadero every 20 minutes starting at 6:20 A.M. To go to the Downtown Center station, Nolan arrives at the Embarcadero
station at a random time between 1:05 P.M. and 1:42 P.M., and he takes either B1 or B26,
whichever arrives first. What is the probability that he ends up taking bus route B26?
2.
While in flight, the fraction X of the time that Nolan sleeps is beta with parameters
(6, 10). What is the probability that next time Nolan flies from Chicago to Shanghai, he
sleeps longer than 6 hours during the 14-hour flight?
3.
In a certain town the length of residence of a family in a home is normal with mean 80
months and variance 900. What is the probability that of 12 independent families, living
on a certain street of that town, at least three will have lived there more than eight years?
4.
In a huge office building, the alarm system has n sensors, the lifetime of each being
exponential with mean 1/λ, independently of the other ones. If for the next t units of
time the alarm system is not checked, what is the probability that at time t only k sensors
are working? Assume that without inspection there is no way for the technical crew to
know that a sensor is down.
5.
Suppose that buses arrive at a station every 15 minutes. During the operating hours of the
buses, Kayla arrives at the station at a random time. (a) If she has been waiting for the
bus for 9 minutes, what is the probability that she has to wait at least another 4 minutes?
(b) Calculate the same probability if the time between consecutive arrival times of two
buses is exponential with mean 15 minutes.
6.
Suppose that, every day, all over the world, 10 million people search the Web for film
review websites. Suppose that, among many other sites, the search always turns up a
specific website that aggregates reviews of movies and TV shows, independently of all
the other sites. If 0.2% of such people visit this site, find the probability that, on a given
day, at least 20,200 people visit this website.
7.
Let X be the weight of a randomly selected adult from a certain community, and assume
that X is a normal random variable. If it is known that, in that community, the mean of
the weights of the adults is 160 pounds, and 85% of the adult population has a weight
between 120 and 200 pounds, find Var(X).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 327 — #343
✐
✐
Chapter 7
8.
Self-Test Problems
327
To improve the reliability of a system, sometimes manufacturers add one or more operating components in parallel so that if one of the components fails, the system still
continues to function. Therefore, a parallel system functions if and only if at least one
of its components functions. Suppose that the probability density function of X, the
operating period, in tens of thousands of hours, of a parallel system is given by
f (x) = 2e−2x + 4e−4x − 6e−6x ,
x ≥ 0.
(a)
Find λ(t), the hazard function of an operating period of this system.
(b)
Find the instantaneous failure rate of such a parallel system if it has survived
6000 hours. Using this, calculate an approximate probability for the system to
fail within 100 hours if it has survived 6000 hours.
9.
In a mayoral election in a small island in Greece, of 5000 voters only 1000 voted for
Tasoula. What is the probability that, in a random sample of 100 voters, more than 15
voted for this mayoral candidate?
10.
A wireless phone company manufactures cell phone chips for 40% of the cell phones it
produces and purchases chips for the remaining 60% of its cell phone production from a
supplier. Suppose that the lifetime of a chip manufactured by the company’s own plant is
gamma with mean 4 and variance 4 years whereas the lifetime of a chip purchased from
the supplier is gamma with mean 6 and standard deviation 6 years. The company offers
its customers a one year replacement warranty. Find the percentage of the cell phones
that must be replaced by the company due to chip failure.
Hint: Let X be the lifetime of the chip of a randomly selected cell phone manufactured
by the wireless phone company. Let A be the event that the chip was manufactured by
the company’s own plant. Use the law of total probability to find P (X ≤ 1).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 328 — #344
✐
✐
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 329 — #345
✐
✐
Chapter 8
Bivariate Distributions
8.1
JOINT DISTRIBUTIONS OF TWO RANDOM VARIABLES
Joint Probability Mass Functions
Thus far we have studied probability mass functions of single discrete random variables and
probability density functions of single continuous random variables. We now consider two
or more random variables that are defined simultaneously on the same sample space. In this
section we consider such cases with two variables. Cases of three or more variables are studied
in Chapter 9.
Definition 8.1 Let X and Y be two discrete random variables defined on the same sample
space. Let the sets of possible values of X and Y be A and B, respectively. The function
p(x, y) = P (X = x, Y = y)
is called the joint probability mass function of X and Y .
Note that p(x, y) ≥ 0. If x 6∈ A or y 6∈ B, then p(x, y) = 0. Also,
X X
p(x, y) = 1.
(8.1)
x∈A y∈B
Let X and Y have joint probability mass function p(x, y). Let pX be the probability mass
function of X . Then
pX (x) = P (X = x) = P (X = x, Y ∈ B)
X
X
P (X = x, Y = y) =
p(x, y).
=
y∈B
y∈B
Similarly, pY , the probability mass function of Y, is given by
pY (y) =
X
p(x, y).
x∈A
These relations motivate the following definition.
Definition 8.2
Let X and Y have joint probability mass function p(x, y). Let A be the
set of possible values of X and B be the set of possible values of Y . Then the functions
329
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 330 — #346
✐
✐
330
Chapter 8
Bivariate Distributions
P
P
pX (x) = y∈B p(x, y) and pY (y) = x∈A p(x, y) are called, respectively, the marginal
probability mass functions of X and Y .
Example 8.1 A small college has 90 male and 30 female professors. An ad hoc committee
of five is selected at random to write the vision and mission of the college. Let X and Y be the
number of men and women on this committee, respectively.
(a)
Find the joint probability mass function of X and Y .
(b)
Find pX and pY , the marginal probability mass functions of X and Y .
Solution:
(a)
(b)
The set of possible values for both X and Y is {0, 1, 2, 3, 4, 5}. The joint probability
mass function of X and Y, p(x, y), is given by
!
!
90
30
x
y
!
if x, y ∈ {0, 1, 2, 3, 4, 5}, x + y = 5
120
p(x, y) =
5
0
otherwise.
To find pX and pY , the
P5marginal probability mass
P5functions of X and Y, respectively,
note that pX (x) =
p(x, y). Since p(x, y) = 0 if
y=0 p(x, y), pY (y) =
P5
Px=0
5
x + y 6= 5, y=0 p(x, y) = p(x, 5 − x) and x=0 p(x, y) = p(5 − y, y). Therefore,
!
!
90
30
x
5−x
!
,
x ∈ {0, 1, 2, 3, 4, 5},
pX (x) =
120
5
!
!
90
30
5−y
y
!
pY (y) =
,
y ∈ {0, 1, 2, 3, 4, 5}.
120
5
Note that, as expected, pX and pY are hypergeometric. Example 8.2 Roll a balanced die and let the outcome be X . Then toss a fair coin X times
and let Y denote the number of tails. What is the joint probability mass function of X and Y
and the marginal probability mass functions of X and Y ?
Solution: Let p(x, y) be the joint probability mass function of X and Y . Clearly, X ∈
{1, 2, 3, 4, 5, 6} and Y ∈ {0, 1, 2, 3, 4, 5, 6}. Now if X = 1, then Y = 0 or 1; we have
1
1 1
p(1, 0) = P (X = 1, Y = 0) = P (X = 1)P (Y = 0 | X = 1) = · = ,
6 2
12
1
1 1
p(1, 1) = P (X = 1, Y = 1) = P (X = 1)P (Y = 1 | X = 1) = · = .
6 2
12
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 331 — #347
✐
✐
Section 8.1
Joint Distributions of Two Random Variables
331
If X = 2, then y = 0, 1, or 2, where
p(2, 0) = P (X = 2, Y = 0) = P (X = 2)P (Y = 0 | X = 2) =
1 1
1
· = .
6 4
24
Similarly, p(2, 1) = 1/12, p(2, 2) = 1/24. If X = 3, then y = 0, 1, 2, or 3, and
p(3, 0) = P (X = 3, Y = 0) = P (X = 3)P (Y = 0 | X = 3)
!
1
1 3 1 0 1 3
,
=
=
6 0 2
2
48
p(3, 1) = P (X = 3, Y = 1) = P (X = 3)P (Y = 1 | X = 3)
!
1 3 1 2 1 1
3
=
.
=
6 1 2
2
48
Similarly, p(3, 2) = 3/48, p(3, 3) = 1/48. Similar calculations will yield the following table
for p(x, y).
y
x
0
1
2
3
4
5
6
pX (x)
1
2
3
4
5
6
1/12
1/24
1/48
1/96
1/192
1/384
1/12
2/24
3/48
4/96
5/192
6/384
0
1/24
3/48
6/96
10/192
15/384
0
0
1/48
4/96
10/192
20/384
0
0
0
1/96
5/192
15/384
0
0
0
0
1/192
6/384
0
0
0
0
0
1/384
1/6
1/6
1/6
1/6
1/6
1/6
pY (y)
63/384
120/384
99/384
64/384
29/384
8/384
1/384
Note that pX (x) = P (X = x) and pY (y) = P (Y = y), the probability mass functions
of X and Y, are obtained by summing up the rows and the columns of this table, respectively.
Let X and Y be discrete random variables with joint probability mass function p(x, y).
Let the sets of possible values of X and Y be A and B, respectively. To find E(X) and E(Y ),
first we calculate pX and pY , the marginal probability mass functions of X and Y, respectively.
Then we will use the following formulas.
X
X
E(X) =
xpX (x);
E(Y ) =
ypY (y).
x∈A
Example 8.3
by
y∈B
Let the joint probability mass function of random variables X and Y be given
1
x(x + y) if x = 1, 2, 3,
p(x, y) = 70
0
elsewhere.
y = 3, 4
Find E(X) and E(Y ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 332 — #348
✐
✐
332
Chapter 8
Bivariate Distributions
Solution: To find E(X) and E(Y ), first we need to calculate pX (x) and pY (y), the marginal
probability mass functions of X and Y, respectively. Note that
pX (x) = p(x, 3) + p(x, 4)
=
1
1
1
1
x(x + 3) + x(x + 4) = x2 + x,
70
70
35
10
x = 1, 2, 3;
pY (y) = p(1, y) + p(2, y) + p(3, y)
=
1
2
3
1
3
(1 + y) + (2 + y) + (3 + y) = + y,
70
70
70
5 35
y = 3, 4.
Therefore,
E(X) =
3
X
xpX (x) =
3
1
X
1 17
x
x2 + x =
≈ 2.43;
35
10
7
x=1
ypY (y) =
4
1
X
3 124
y
+ y =
≈ 3.54.
5 35
35
y=3
x=1
E(Y ) =
4
X
y=3
We will now state the following generalization of Theorem 4.2 from one dimension to two.
As an immediate application of this important generalization, we will show that the expected
value of the sum of two random variables is equal to the sum of their expected values. We will
prove a generalization of this theorem in Chapter 10 and discuss some of its many applications
in that chapter.
Theorem 8.1 Let p(x, y) be the joint probability mass function of discrete random variables X and Y . Let A and B be the set of possible values of X and Y, respectively. If h is a
function of two variables from R2 to R , then h(X, Y ) is a discrete random variable with the
expected value given by
X X
E h(X, Y ) =
h(x, y)p(x, y),
x∈A y∈B
provided that the sum is absolutely convergent.
Letting h(x, y) = x + y in this theorem, it follows that
X X
(x + y)p(x, y)
E(X + Y ) =
x∈A y∈B
=
X X
xp(x, y) +
x∈A y∈B
XX
yp(x, y)
x∈A y∈B
= E(X) + E(Y ).
We have shown the following important corollary.
Corollary
For discrete random variables X and Y,
E(X + Y ) = E(X) + E(Y ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 333 — #349
✐
✐
Section 8.1
Example 8.4
Y be given by
Find E(XY ).
Joint Distributions of Two Random Variables
333
Let the joint probability mass function of the discrete random variables X and
2 x
p(x, y) = 11 y
0
if x = 1, 2,
y = 1, 2, 3
otherwise.
Solution: By Theorem 8.1,
E(XY ) =
2 X
3
X
x=1 y=1
=
xy ·
2 x
11 y
2
3
2
2 XX 2
2 X 2 30
x =
3x =
≈ 2.73. 11 x=1 y=1
11 x=1
11
Joint Probability Density Functions
To define the concept of joint probability density function of two continuous random variables
X and Y, recall from Section 6.1 that a single random variable X is called continuous if there
exists a nonnegative real-valued function f : R → [0, ∞) such that for any subset A of real
numbers that can be constructed from intervals by a countable number of set operations,
Z
P (X ∈ A) =
f (x) dx.
A
This definition is generalized in the following obvious way:
Definition 8.3 Two random variables X and Y, defined on the same sample space, have a
continuous joint distribution if there exists a nonnegative function of two variables, f (x, y) on
R × R , such that for any region R in the xy -plane that can be formed from rectangles by a
countable number of set operations,
ZZ
P (X, Y ) ∈ R =
f (x, y) dx dy.
(8.2)
R
The function f (x, y) is called the joint probability density function of X and
Y.
Note that if in (8.2) the region R is a plane curve, then P (X, Y ) ∈ R = 0. That is, the
probability that (X, Y ) lies on any curve (in particular, a circle, an ellipse, a straight line, etc.)
is 0.
Let R = (x, y) : x ∈ A, y ∈ B , where A and B are any subsets of real numbers that
can be constructed from intervals by a countable number of set operations. Then (8.2) gives
Z Z
P (X ∈ A, Y ∈ B) =
f (x, y) dx dy.
(8.3)
B
A
Letting A = (−∞, ∞), B = (−∞, ∞), (8.3) implies the relation
Z ∞ Z ∞
f (x, y) dx dy = 1,
−∞
−∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 334 — #350
✐
✐
334
Chapter 8
Bivariate Distributions
which is the continuous analog of (8.1). The relation (8.3) also implies that
P (X = a, Y = b) =
Z bZ a
b
f (x, y) dx dy = 0.
a
Hence, for a < b and c < d,
P (a < X ≤ b, c ≤ Y ≤ d) = P (a < X < b, c < Y < d)
= P (a ≤ X < b, c ≤ Y < d) = · · ·
Z d Z b
=
f (x, y) dx dy.
c
a
Because, for real numbers a and b, P (X = a, Y = b) = 0, in general, f (a, b) is not equal to
P (X = a, Y = b). At no point do the values of f represent probabilities. Intuitively, f (a, b)
is a measure that determines how likely it is that X is close to a and Y is close to b. To see this,
note that if ε and δ are very small positive numbers, then
P (a − ε < X < a + ε, b − δ < Y < b + δ)
is the probability that X is close to a and Y is close to b. Now
P (a − ε < X < a + ε, b − δ < Y < b + δ) =
Z b+δ Z a+ε
b−δ
f (x, y) dx dy
a−ε
is the volume under the surface z = f (x, y), above the region (a−ε, a+ε)×(b−δ, b+δ). For
infinitesimal ε and δ, this volume is approximately equal to the volume of a rectangular paral
lelepiped with sides of lengths 2ε, 2δ, and height f (a, b) i.e., (2ε)(2δ)f (a, b) = 4εδf (a, b) .
Therefore,
P (a − ε < X < a + ε, b − δ < Y < b + δ) ≈ 4εδf (a, b).
Hence for fixed small values of ǫ and δ, we observe that a larger value of f (a, b) makes
P (a − ε < X < a + ε, b − δ < Y < b + δ) larger. That is, the larger the value of
f (a, b), the higher the probability that X and Y are close to a and b, respectively. A sketch of
a joint probability density function of two continuous random variables is shown in Figure 8.1.
Note that f (x, y) is the height of the surface at each point (x, y) in its domain.
Figure 8.1
An example of a joint probability density function.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 335 — #351
✐
✐
Section 8.1
Joint Distributions of Two Random Variables
335
Let X and Y have joint probability density function f (x, y). Let fY be the probability
density function of Y . To find fY in terms of f, note that, on the one hand, for any subset B of
R,
Z
P (Y ∈ B) =
fY (y) dy,
(8.4)
B
and, on the other hand, using (8.3),
P (Y ∈ B) = P X ∈ (−∞, ∞), Y ∈ B =
Comparing this with (8.4), we can write
fY (y) =
Z Z ∞
B
−∞
f (x, y) dx dy.
Z ∞
f (x, y) dx.
(8.5)
Z ∞
f (x, y) dy.
(8.6)
−∞
Similarly,
fX (x) =
−∞
Therefore, it is reasonable to make the following definition:
Definition 8.4
Let X and Y have joint probability density function f (x, y); then the
functions fX and fY , given by (8.6) and (8.5), are called, respectively, the marginal probability
density functions of X and Y .
Note that while from the joint probability density function of X and Y we can find the
marginal probability density functions of X and Y, it is not possible, in general, to find the joint
probability density function of two random variables from their marginals. This is because the
marginal probability density function of a random variable X gives information about X without looking at the possible values of other random variables. However, sometimes with more
information the marginal probability density functions enable us to find the joint probability
density function of two random variables. An example of such a case, studied in Section 8.2, is
when X and Y are independent random variables.
Let X and Y be two random variables (discrete, continuous, or mixed). The joint
distribution function, or joint cumulative distribution function of X and Y , is defined by
F (t, u) = P (X ≤ t, Y ≤ u)
for all −∞ < t, u < ∞. The marginal distribution function of X, FX , can be found from
F as follows:
FX (t) = P (X ≤ t) = P (X ≤ t, Y < ∞) = P lim {X ≤ t, Y ≤ n}
n→∞
= lim P (X ≤ t, Y ≤ n) = lim F (t, n) ≡ F (t, ∞).
n→∞
n→∞
To justify the fourth equality, note that the sequence of events {X ≤ t, Y ≤ n}, n ≥ 1, is an
increasing sequence. Therefore, by continuity of probability function (Theorem 1.8),
P lim {X ≤ t, Y ≤ n} = lim P (X ≤ t, Y ≤ n).
n→∞
n→∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 336 — #352
✐
✐
336
Chapter 8
Bivariate Distributions
Similarly, FY , the marginal distribution function of Y , is
FY (u) = P (Y ≤ u) = lim F (n, u) ≡ F (∞, u).
n→∞
Now suppose that the joint probability density function of X and Y is f (x, y). Then
F (x, y) = P (X ≤ x, Y ≤ y) = P X ∈ (−∞, x], Y ∈ (−∞, y]
Z y Z x
=
f (t, u) dt du.
(8.7)
−∞
−∞
Assuming that the partial derivatives of F exist, by differentiation of (8.7), we get
f (x, y) =
∂2
∂x ∂y
F (x, y).
Moreover, from (8.7), we obtain
FX (x) = F (x, ∞) =
=
Z x Z ∞
−∞
Z ∞ Z x
−∞
−∞
f (t, u) dt du
f (t, u) du dt =
−∞
Z x
fX (t) dt,
−∞
and, similarly,
FY (y) =
Z y
fY (u) du.
−∞
These relations show that if X and Y have joint probability density function f (x, y), then
X and Y are continuous random variables with probability density functions fX and fY , and
distribution functions FX and FY , respectively. Therefore, FX′ (x) = fX (x) and FY′ (y) =
fY (y). We also have
Z ∞
Z ∞
E(X) =
xfX (x) dx;
E(Y ) =
yfY (y) dy.
−∞
Example 8.5
by
−∞
The joint probability density function of random variables X and Y is given
(
λxy 2
f (x, y) =
0
0≤x≤y≤1
otherwise.
(a)
Determine the value of λ.
(b)
Find the marginal probability density functions of X and Y .
(c)
Calculate E(X) and E(Y ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 337 — #353
✐
✐
Section 8.1
Joint Distributions of Two Random Variables
337
Solution:
(a)
To find λ, note that
R∞ R∞
−∞
−∞ f (x, y) dx dy = 1 gives
Z 1 Z 1
0
x
λxy 2 dy dx = 1.
Therefore,
Z 1Z 1
Z 1h
1 3 i1
1=
y dy λx dx =
y λx dx
3 x
0
x
0
Z 1
Z
Z
1 1 3
λ 1
λ 1
=λ
− x x dx =
(1 − x3 )x dx =
(x − x4 ) dx
3
3
3
3
0
0
0
h
i
1
λ
λ 1 2 1 5
x − x
=
,
=
3 2
5
10
0
2
and hence λ = 10.
(b)
To find fX and fY , the respective marginal probability density functions of X and Y,
we use (8.6) and (8.5):
Z ∞
Z 1
h 10
i1
10
fX (x) =
f (x, y) dy =
10xy 2 dy =
xy 3 =
x(1 − x3 ), 0 ≤ x ≤ 1;
3
3
x
−∞
x
Z ∞
Z y
h
iy
fY (y) =
f (x, y) dx =
10xy 2 dx = 5x2 y 2 = 5y 4 , 0 ≤ y ≤ 1.
−∞
(c)
0
0
To find E(X) and E(Y ), we use the results obtained in part (b):
E(X) =
Z 1
0
E(Y ) =
Example 8.6
For λ > 0, let
(
F (x, y) =
x·
Z 1
5
y · 5y 4 dy = . 6
0
1 − λe−λ(x+y)
0
10
5
x(1 − x3 ) dx = ;
3
9
if x > 0,
y>0
otherwise.
Determine if F is the joint distribution function of two random variables X and Y .
Solution: If F is the joint distribution function of two random variables X and Y, then
∂2
F (x, y) is the joint probability density function of X and Y . But
∂x ∂y
(
−λ3 e−λ(x+y)
if x > 0, y > 0
∂2
F (x, y) =
∂x ∂y
0
otherwise.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 338 — #354
✐
✐
338
Chapter 8
Bivariate Distributions
∂2
F (x, y) < 0, it cannot be a joint probability density function. Therefore, F is not
∂x ∂y
a joint distribution function. Since
Example 8.7 A circle of radius 1 is inscribed in a square with sides of length 2. A point is
selected at random from the square. What is the probability that it is inside the circle? Note that
by a point being selected at random from the square we mean that the point is selected in a way
that all the subsets of equal areas of the square are equally likely to contain the point.
y
x
Figure 8.2
Geometric model of Example 8.7.
Solution: Let the square and the circle be situated in the coordinate system as shown in
Figure 8.2. Let the coordinates of the point selected at random be (X, Y ); then X and Y are
random variables. By definition, regions inside the square with equal areas are equally likely to
contain (X, Y ). Hence, for all (a, b) inside the square, the probability that (X, Y ) is “close”
to (a, b) is the same. Let f (x, y) be the joint probability density function of X and Y . Since
f (x, y) is a measure that determines how likely it is that X is close to x and Y is close to y,
f (x, y) must be constant for the points inside the square, and 0 elsewhere. Therefore, for some
c > 0,
(
c
if 0 < x < 2, 0 < y < 2
f (x, y) =
0
otherwise,
where
R∞ R∞
−∞
−∞
f (x, y) dx dy = 1 gives
Z 2Z 2
0
c dx dy = 1,
0
implying that c = 1/4. Now let R be the region inside the circle. Then the desired probability
is
ZZ
dx dy
ZZ
ZZ
1
dx dy = R
.
f (x, y) dx dy =
4
4
R
R
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 339 — #355
✐
✐
Section 8.1
Note that
ZZ
Joint Distributions of Two Random Variables
339
dx dy is the area of the circle, and 4 is the area of the square. Thus the desired
R
probability is
area of the circle
π(1)2
π
=
= . area of the square
4
4
What we showed in Example 8.7 is true, in general. Let S be a bounded region in the
Euclidean plane and suppose that R is a region inside S . Fix a coordinate system, and let the
coordinates of a point selected at random from S be (X, Y ). By an argument similar to that of
Example 8.7, we have that, for some c > 0, the joint probability density function of X and Y,
f (x, y), is given by
f (x, y) =
(
c
0
if (x, y) ∈ S
otherwise,
where
ZZ
f (x, y) dx dy = 1.
S
This gives c
hence
ZZ
S
dx dy = 1, or, equivalently, c × area(S) = 1. Therefore, c = 1/area(S) and
f (x, y) =
Thus
P (X, Y ) ∈ R =
ZZ
R
1
area(S)
0
if (x, y) ∈ S
(8.8)
otherwise.
1
f (x, y) dx dy =
area(S)
ZZ
R
dx dy =
area(R)
.
area(S)
Based on these observations, we make the following definition.
Definition 8.5 Let S be a subset of the plane with area A(S). A point is said to be randomly
selected from S if for any subset R of S with area A(R), the probability that R contains the
point is A(R)/A(S).
This definition is essential in the field of geometric probability. By the following examples, we will show how it can help to solve problems readily.
Example 8.8 A man invites his fiancée to a fine hotel for a Sunday brunch. They decide to
meet in the lobby of the hotel between 11:30 A.M. and 12 noon. If they arrive at random times
during this period, what is the probability that they will meet within 10 minutes?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 340 — #356
✐
✐
340
Chapter 8
Bivariate Distributions
Solution: Let X and Y be the minutes past 11:30 A.M. that the man and his fiancée arrive at
the lobby, respectively. Let
S = (x, y) : 0 ≤ x ≤ 30, 0 ≤ y ≤ 30 , and R = (x, y) ∈ S : |x − y| ≤ 10 .
Then the desired probability, P |X − Y | ≤ 10 , is given by
area of R
area(R)
area of R
=
=
.
P |X − Y | ≤ 10 =
area of S
30 × 30
900
y
30
y
_
x=
10
10
x
10
Figure 8.3
=
_ y
10
30
x
Geometric model of Example 8.8.
R = (x, y) ∈ S : x − y ≤ 10 and y − x ≤ 10 is the shaded region of Figure 8.3, and
its area is the area of the square minus the areas of the two unshaded triangles: (30)(30) −
2(1/2 × 20 × 20) = 500. Hence the desired probability is 500/900 = 5/9.
Example 8.9 A farmer decides to build a pen in the shape of a triangle for his chickens.
He sends his son out to cut the lumber and the boy, without taking any thought as to the ultimate purpose, makes two cuts at two points selected at random. What are the chances that the
resulting three pieces of lumber can be used to form a triangular pen?
Solution: Suppose that the length of the lumber is ℓ. Let A and B be the random points placed
on the lumber; let the distances of A and B from the left end of the lumber be denoted by X
and Y, respectively. If X < Y, the lumber is divided into three parts of lengths X, Y − X, and
ℓ − Y ; otherwise, it is divided into three parts of lengths Y, X − Y, and ℓ − X . Because of
symmetry, we calculate the probability that X < Y, and X, Y − X, and ℓ − Y form a triangle.
Then we multiply the result by 2 to obtain the desired probability. We know that three segments
form a triangle if and only if the length of any one of them is less than the sum of the lengths
of the remaining two. Therefore, we must have
X<Y
X < (Y − X) + (ℓ − Y )
(Y − X) < X + (ℓ − Y )
ℓ − Y < X + (Y − X),
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 341 — #357
✐
✐
Section 8.1
Joint Distributions of Two Random Variables
341
or, equivalently,
X < Y,
X<
ℓ
,
2
ℓ
Y <X+ ,
2
and Y >
ℓ
.
2
y
x
y=
x
Figure 8.4
Geometric model of Example 8.9.
Since (X, Y ) is a random point from the square (0, ℓ) × (0, ℓ), the probability that it satisfies
these inequalities is the area of R, where
n
ℓ
ℓ
ℓo
R = (x, y) : x < y, x < , y < x + , y >
,
2
2
2
divided by ℓ2 , the area of the square. Note that R is the shaded region of Figure 8.4 and its area
is 1/8 of the area of the square. Thus the probability that (X, Y ) lies in R is 1/8 and hence the
desired probability is 1/4. Before closing this section, we state the continuous analog of Theorem 8.1, which is a
generalization of Theorem 6.3 from one dimension to two. As an immediate application of this
important generalization, we will show the continuous version of the following theorem that
we proved above for discrete random variables: The expected value of the sum of two random
variables is equal to the sum of their expected values.
Theorem 8.2 Let f (x, y) be the joint probability density function of random variables X
and Y . If h is a function of two variables from R2 to R , then h(X, Y ) is a random variable
with the expected value given by
Z ∞ Z ∞
E h(X, Y ) =
h(x, y)f (x, y) dx dy,
−∞
−∞
provided that the integral is absolutely convergent.
Corollary
For random variables X and Y,
E(X + Y ) = E(X) + E(Y ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 342 — #358
✐
✐
342
Chapter 8
Proof:
In Theorem 8.2 let h(x, y) = x + y. Then
Z ∞Z ∞
E(X + Y ) =
(x + y)f (x, y) dx dy
Bivariate Distributions
−∞
=
−∞
Z ∞Z ∞
−∞
xf (x, y) dx dy +
−∞
Z ∞Z ∞
−∞
yf (x, y) dx dy
−∞
= E(X) + E(Y ). Let X and Y have joint probability density function
3
(x2 + y 2 )
if 0 < x < 1, 0 < y < 1
f (x, y) = 2
0
otherwise.
Example 8.10
Find E(X 2 + Y 2 ).
Solution: By Theorem 8.2,
Z ∞Z ∞
Z 1Z 1
3 2
2
2
2
2
(x + y 2 )2 dx dy
E(X + Y ) =
(x + y )f (x, y) dx dy =
−∞ −∞
0
0 2
Z Z
14
3 1 1 4
(x + 2x2 y 2 + y 4 ) dx dy =
=
. 2 0 0
15
EXERCISES
A
1.
Let the joint probability mass function of discrete random variables X and Y be given
by
x
if x = 1, 2, y = 1, 2
k
y
p(x, y) =
0
otherwise.
Determine (a) the value of the constant k, (b) the marginal probability mass functions of
X and Y, (c) P (X > 1 | Y = 1), (d) E(X) and E(Y ).
2.
Let the joint probability mass function of discrete random variables X and Y be given
by
(
c(x + y)
if x = 1, 2, 3, y = 1, 2
p(x, y) =
0
otherwise.
Determine (a) the value of the constant c, (b) the marginal probability mass functions of
X and Y, (c) P (X ≥ 2 | Y = 1), (d) E(X) and E(Y ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 343 — #359
✐
✐
Section 8.1
3.
Joint Distributions of Two Random Variables
343
Let the joint probability mass function of discrete random variables X and Y be given
by
(
k(x2 + y 2 )
if (x, y) = (1, 1), (1, 3), (2, 3)
p(x, y) =
0
otherwise.
Determine (a) the value of the constant k, (b) the marginal probability mass functions of
X and Y, and (c) E(X) and E(Y ).
4.
Let the joint probability mass function of discrete random variables X and Y be given
by
1
if x = 1, 2, y = 0, 1, 2
(x2 + y 2 )
p(x, y) = 25
0
otherwise.
Find P (X > Y ), P (X + Y ≤ 2), and P (X + Y = 2).
5.
Thieves stole four animals at random from a farm that had seven sheep, eight goats, and
five burros. Calculate the joint probability mass function of the number of sheep and
goats stolen.
6.
Two dice are rolled. The sum of the outcomes is denoted by X and the absolute value of
their difference by Y . Calculate the joint probability mass function of X and Y and the
marginal probability mass functions of X and Y .
7.
In a community 30% of the adults are Republicans, 50% are Democrats, and the rest are
independent. For a randomly selected person, let
(
1
if he or she is a Republican
X=
0
otherwise,
(
1
if he or she is a Democrat
Y =
0
otherwise.
Calculate the joint probability mass function of X and Y .
8.
In an area prone to flood, an insurance company covers losses only up to three separate
floods per year. Let X be the total number of floods in a random year and Y be the
number of the floods that each cause over $74 million in insured damage. Suppose that,
for some constant c, the joint probability mass function of X and Y is
c(3x + 2y)
0 ≤ x ≤ 3, 0 ≤ y ≤ x
p(x, y) =
0
otherwise.
Find the expected value of the number of floods, in a random year, that cause $74 million
or less in insured damage.
9.
From an ordinary deck of 52 cards, seven cards are drawn at random and without replacement. Let X and Y be the number of hearts and the number of spades drawn,
respectively.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 344 — #360
✐
✐
344
10.
11.
12.
Chapter 8
Bivariate Distributions
(a)
Find the joint probability mass function of X and Y .
(b)
Calculate P (X ≥ Y ).
Let the joint probability density function of random variables X and Y be given by
(
2
if 0 ≤ y ≤ x ≤ 1
f (x, y) =
0
elsewhere.
(a)
Calculate the marginal probability density functions of X and Y, respectively.
(b)
Find E(X) and E(Y ).
(c)
Calculate P (X < 1/2), P (X < 2Y ), and P (X = Y ).
Let the joint probability density function of random variables X and Y be given by
(
8xy
if 0 ≤ y ≤ x ≤ 1
f (x, y) =
0
elsewhere.
(a)
Calculate the marginal probability density functions of X and Y, respectively.
(b)
Calculate E(X) and E(Y ).
Let the joint probability density function of random variables X and Y be given by
1
if x > 0, 0 < y < 2
ye−x
f (x, y) = 2
0
elsewhere.
Find the marginal probability density functions of X and Y .
13.
Let X and Y have the joint probability density function
(
1
if 0 ≤ x ≤ 1, 0 ≤ y ≤ 1
f (x, y) =
0
elsewhere.
Calculate P (X + Y ≤ 1/2),
P (X 2 + Y 2 ≤ 1).
P (X − Y ≤ 1/2),
P (XY ≤ 1/4), and
14.
Let X be the proportion of customers of an insurance company who bundle their auto
and home insurance policies. Let Y be the proportion of customers who insure at least
their car with the insurance company. An actuary has discovered that for, 0 ≤ x ≤ y ≤
1, the joint distribution function of X and Y is F (x, y) = x(y 2 + xy − x2 ). Find the
expected value of the proportion of the customers of the insurance company who bundle
their auto and home insurance policies.
15.
Let F be the joint distribution function of random variables X and Y . For x1 < x2 and
y1 < y2 , in terms of F, calculate P (x1 < X ≤ x2 , y1 < Y ≤ y2 ).
16.
Let F be the joint distribution function of the random variables X and Y . In terms of F,
calculate P (X > t, Y > u).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 345 — #361
✐
✐
Section 8.1
17.
Joint Distributions of Two Random Variables
345
Let R be the bounded region between y = x and y = x2 . A random point (X, Y ) is
selected from R.
(a)
Find the joint probability density function of X and Y .
(b)
Calculate the marginal probability density functions of X and Y .
(c)
Find E(X) and E(Y ).
18.
A man invites his fiancée to an elegant hotel for a Sunday brunch. They decide to meet
in the lobby of the hotel between 11:30 A.M. and 12 noon. If they arrive at random times
during this period, what is the probability that the first to arrive has to wait at least 12
minutes?
19.
A farmer makes cuts at two points selected at random on a piece of lumber of length ℓ.
What is the expected value of the length of the middle piece?
20.
On a line segment AB of length ℓ, two points C and D are placed at random and
independently. What is the probability that C is closer to D than to A?
21.
Two points X and Y are selected at random and independently from the interval (0, 1).
Calculate P (Y ≤ X and X 2 + Y 2 ≤ 1).
22.
For λ > 0, let
F (x, y) =
1 − λe−λ(x+y)
0
if x > 0, y > 0
otherwise.
Determine if F is the joint distribution function of two random variables X and Y .
23.
Let
0
1
xy(x + y)
2
1
F (x, y) =
x(x + 1)
2
1
2 y(y + 1)
1
x ≤ 0 or y ≤ 0
0 < x ≤ 1, 0 < y ≤ 1
0 ≤ x ≤ 1, y ≥ 1
x ≥ 1, 0 < y ≤ 1
x ≥ 1, y ≥ 1.
Determine if F is the joint distribution function of two random variables X and Y .
24.
Let X and Y have the joint probability density function
3
(x2 + y 2 )
0 < x < 1 and 0 < y < 1
f (x, y) = 2
0
otherwise.
Find the joint distribution function, the marginal distribution functions, and the marginal
probability density functions of X and Y .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 346 — #362
✐
✐
346
Chapter 8
Bivariate Distributions
B
25.
Suppose that h is the probability density function of a continuous random variable. Let
the joint probability density function of two random variables X and Y be given by
f (x, y) = h(x)h(y),
x ∈ R,
y ∈ R.
Prove that P (X ≥ Y ) = 1/2.
26.
27.
Let X and Y be random variables with finite expected values. Show that if
P (X ≤ Y ) = 1, then E(X) ≤ E(Y ).
Let g and h be two probability density functions with distribution functions G and H,
respectively. Show that for −1 ≤ α ≤ 1, the function
f (x, y) = g(x)h(y) 1 + α[2G(x) − 1][2H(y) − 1]
is a joint probability density function of two random variables. Moreover, prove that g
and h are the marginal probability density functions of f .
Note: This exercise gives an infinite family of joint probability density functions all with
the same marginal probability density function of X and of Y .
28.
Three points M, N, and L are placed on a circle at random and independently. What is
the probability that M N L is an acute angle?
29.
Two numbers x and y are selected at random from the interval (0, 1). For i = 0, 1, 2,
determine the probability that the integer nearest to x + y is i.
Note: This problem was given by Hilton and Pedersen in the paper “A Role for Untraditional Geometry in the Curriculum,” published in the December 1989 issue of Kolloquium Mathematik-Didaktik der Universität Bayreuth. It is not hard to solve, but it has
interesting consequences in arithmetic.
30.
A farmer who has two pieces of lumber of lengths a and b (a < b) decides to build a pen
in the shape of a triangle for his chickens. He sends his son out to cut the lumber and
the boy, without taking any thought as to the ultimate purpose, makes two cuts, one on
each piece, at randomly selected points. He then chooses three of the resulting pieces at
random and takes them to his father. What are the chances that they can be used to form
a triangular pen?
31.
Two points are placed on a segment of length ℓ independently and at random to divide
the line into three parts. What is the probability that the length of none of the three parts
exceeds a given value α, ℓ/3 ≤ α ≤ ℓ?
32.
A point is selected at random and uniformly from the region
R = (x, y) : |x| + |y| ≤ 1 .
Find the probability density function of the x-coordinate of the point selected at random.
33.
Let X and Y be continuous random variables with joint probability density function
f (x, y). Let Z = Y /X, X 6= 0. Prove that the probability density function of Z is
given by
Z ∞
|x|f (x, xz) dx.
fZ (z) =
−∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 347 — #363
✐
✐
Section 8.1
34.
35.
347
Joint Distributions of Two Random Variables
Consider a disk centered at O with radius R. Suppose that n ≥ 3 points P1 , P2 , . . . , Pn
are independently placed at random inside the disk. Find the probability that all these
points are contained in a closed semicircular disk.
Hint: For each 1 ≤ i ≤ n, let Ai be the endpoint of the radius through Pi and
let Bi be the corresponding antipodal point. Let Di be the closed semicircular disk
OAi Bi with negatively (clockwise) oriented boundary. Note that there is at most one
Di , 1 ≤ i ≤ n, that contains all the Pi ’s.
For α > 0, β > 0, and γ > 0, the following function is called the bivariate Dirichlet
probability density function
f (x, y) =
Γ(α + β + γ) α−1 β−1
x y (1 − x − y)γ−1
Γ(α)Γ(β)Γ(γ)
if x ≥ 0, y ≥ 0, and x + y ≤ 1; f (x, y) = 0, otherwise. Prove that fX , the marginal
probability density function of X, is beta with parameters (α, β + γ); and fY is beta
with the parameters (β, α + γ).
Hint: Note that
h Γ(α + β + γ) ih Γ(β + γ) i
1
1
Γ(α + β + γ)
=
=
.
Γ(α)Γ(β)Γ(γ)
Γ(α)Γ(β + γ) Γ(β)Γ(γ)
B(α, β + γ) B(β, γ)
36.
As Liu Wen from Hebei University of Technology in Tianjin, China, has noted in the
April 2001 issue of The American Mathematical Monthly, in some reputable probability
and statistics texts it has been asserted that “if a two-dimensional distribution function
F (x, y) has a continuous density of f (x, y), then
f (x, y) =
∂ 2 F (x, y)
.
∂x ∂y
(8.9)
Furthermore, some intermediate textbooks in probability and statistics even assert that
at a point of continuity for f (x, y), F (x, y) is twice differentiable, and (8.9) holds at
that point.” Let the joint probability density function of random variables X and Y be
given by
if y ≥ 0, (−1/2)e−y ≤ x ≤ 0
1 + 2xey
1 − 2xey
if y ≥ 0, 0 ≤ x ≤ (1/2)e−y
f (x, y) = 1 + 2xe−y if y ≤ 0, (−1/2)ey ≤ x ≤ 0
1 − 2xe−y if y ≤ 0, 0 ≤ x ≤ (1/2)ey
0
otherwise.
Show that even though f is continuous everywhere, the partial derivatives of its distribution function F do not exist at (0, 0). This counterexample is constructed based on a
general example given by Liu Wen in the aforementioned paper.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 348 — #364
✐
✐
348
Chapter 8
Bivariate Distributions
Self-Quiz on Section 8.1
Time allotted: 20 Minutes
1.
2.
8.2
Each problem is worth 5 points.
A device with two components functions only if both components function. Suppose
that the joint probability density function of X and Y, the lifetimes of the components,
in years, is
1
(x + y)
if 0 < x < 3 and 0 < y < 5
f (x, y) = 60
0
otherwise.
(a)
Find the probability that the device functions for at least 2 years.
(b)
Find the marginal distribution function of X .
Let X be a car insurance company’s annual losses under medical coverage insurance.
Let Y be its annual losses under comprehensive and collisions losses. Suppose that the
joint probability density function of X and Y, in millions of dollars, is
1
(2x + 3y)
1 < x < 5, 1 < y < 7
f (x, y) = 432
0
otherwise.
(a)
Find the marginal probability density function of Y .
(b)
Find the expected value of the sum of the annual losses of the company under
medical coverage and under comprehensive and collision.
INDEPENDENT RANDOM VARIABLES
Two random variables X and Y are called independent if, for arbitrary subsets A and B of
real numbers, the events {X ∈ A} and {Y ∈ B} are independent, that is, if
P (X ∈ A, Y ∈ B) = P (X ∈ A)P (Y ∈ B).
(8.10)
Using the axioms of probability, we can prove that X and Y are independent if and only if for
any two real numbers a and b,
P (X ≤ a, Y ≤ b) = P (X ≤ a)P (Y ≤ b).
(8.11)
Hence (8.10) and (8.11) are equivalent.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 349 — #365
✐
✐
Section 8.2
Independent Random Variables
349
Relation (8.11) states that X and Y are independent random variables if and only if their
joint distribution function is the product of their marginal distribution functions. The following
theorem states this fact.
Theorem 8.3 Let X and Y be two random variables defined on the same sample space. If
F is the joint distribution function of X and Y, then X and Y are independent if and only if
for all real numbers t and u,
F (t, u) = FX (t)FY (u).
Independence of Discrete Random Variables
If X and Y are discrete random variables with sets of possible values E and F, respectively,
and joint probability mass function p(x, y), then the definition of independence is satisfied if
for all x ∈ E and y ∈ F, the events {X = x} and {Y = y} are independent; that is,
P (X = x, Y = y) = P (X = x)P (Y = y).
(8.12)
Relation (8.12) shows that X and Y are independent if and only if their joint probability mass
function is the product of the marginal probability mass functions of X and Y . Therefore, we
have the following theorem.
Theorem 8.4 Let X and Y be two discrete random variables defined on the same sample space. If p(x, y) is the joint probability mass function of X and Y, then X and Y are
independent if and only if for all real numbers x and y,
p(x, y) = pX (x)pY (y).
Let X and Y be discrete independent random variables with sets of possible values A and
B, respectively. Then (8.12) implies that for all x ∈ A and y ∈ B,
P (X = x | Y = y) = P (X = x)
and
P (Y = y | X = x) = P (Y = y).
Hence X and Y are independent if knowing the value of one of them does not change the
probability mass function of the other.
Example 8.11 Suppose that 4% of the bicycle fenders, produced by a stamping machine
from the strips of steel, need smoothing. What is the probability that, of the next 13 bicycle
fenders stamped by this machine, two need smoothing and, of the next 20, three need smoothing?
Solution: Let X be the number of bicycle fenders among the first 13 that need smoothing.
Let Y be the number of those among the next 7 that need smoothing. We want to calculate
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 350 — #366
✐
✐
350
Chapter 8
Bivariate Distributions
P (X = 2, Y = 1). Since X and Y are independent binomial random variables with parameters (13, 0.04) and (7, 0.04), respectively, we can write
P (X = 2, Y = 1) = P (X = 2)P (Y = 1)
!
!
13
7
=
(0.04)2 (0.96)11
(0.04)1 (0.96)6 ≈ 0.0175. 2
1
The idea of Example 8.11 can be stated in general terms: Suppose that n + m independent
Bernoulli trials, each with parameter p, are performed. If X and Y are the number of successes
in the first n and in the last m trials, then X and Y are binomial random variables with parameters (n, p) and (m, p), respectively. Furthermore, they are independent because knowing the
number of successes in the first n trials does not change the probability mass function of the
number of successes in the last m trials. Hence
P (X = i, Y = j) = P (X = i)P (Y = j)
!
!
n i
m
=
p (1 − p)n−i
pj (1 − p)m−j
i
j
!
!
n
m i+j
=
p (1 − p)(n+m)−(i+j) .
i
j
We now prove that functions of independent random variables are also independent.
Theorem 8.5 Let X and Y be independent random variables. Then for the real-valued
functions g : R → R and h : R → R, g(X) and h(Y ) are also independent random variables.
Proof: To show that g(X) and h(Y ) are independent, by (8.11) it suffices to prove that, for
any two real numbers a and b,
P g(X) ≤ a, h(Y ) ≤ b = P g(X) ≤ a P h(Y ) ≤ b .
Let A = x : g(x) ≤ a and B = y : h(y) ≤ b . Clearly, x ∈ A if and only if g(x) ≤ a,
and y ∈ B if and only if h(y) ≤ b. Therefore,
P g(X) ≤ a, h(Y ) ≤ b = P (X ∈ A, Y ∈ B) = P (X ∈ A)P (Y ∈ B)
= P g(X) ≤ a P h(Y ) ≤ b . By this theorem, if X and Y are independent random variables, then sets such as {X 2 , Y },
{sin X, eY }, {X 2 − 2X, Y 3 + 3Y } are sets of independent random variables.
Another important property of independent random variables is that the expected value of
their product is equal to the product of their expected values. To prove this, we use Theorem 8.1.
Theorem 8.6 Let X and Y be independent random variables. Then for all real-valued
functions g : R → R and h : R → R ,
E g(X)h(Y ) = E g(X) E h(Y ) ,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 351 — #367
✐
✐
Section 8.2
Independent Random Variables
351
where, as usual, we assume that E g(X) and E h(Y ) are finite.
Solution: Let A be the set of possible values of X, and B be the set of possible values for Y .
Let p(x, y) be the joint probability mass function of X and Y . Then
XX
XX
E g(X)h(Y ) =
g(x)h(y)p(x, y) =
g(x)h(y)pX (x)pY (y)
x∈A y∈B
=
Xh
x∈A y∈B
g(x)pX (x)
x∈A
X
y∈B
i X
h(y)pY (y) =
g(x)pX (x)E h(Y )
x∈A
X
= E h(Y )
g(x)pX (x) = E h(Y ) E g(X) . x∈A
By this theorem, if X and Y are independent, then E(XY ) = E(X)E(Y ), and
relations such as the following are valid.
E X 2 |Y | = E(X 2 )E |Y | ,
E[(sin X)eY ] = E(sin X)E(eY ),
E (X 2 − 2X)(Y 3 + 3Y ) = E(X 2 − 2X)E(Y 3 + 3Y ).
However, the converse of Theorem 8.6 is not necessarily true. That is, two random variables X
and Y might be dependent while E(XY ) = E(X)E(Y ). Here is an example:
Example 8.12 Let X be a random variable with the set of possible values {−1, 0, 1} and
probability mass function p(−1) = p(0) = p(1) = 1/3. Letting Y = X 2 , we have
1
1
1
+ 0 · + 1 · = 0,
3
3
3
1
1
1
2
E(Y ) = E(X 2 ) = (−1)2 · + 02 · + (1)2 · = ,
3
3
3
3
3 1
3 1
3
3 1
E(XY ) = E(X ) = (−1) · + 0 · + (1) · = 0.
3
3
3
E(X) = −1 ·
Thus E(XY ) = E(X)E(Y ) while, clearly, X and Y are dependent.
Independence of Continuous Random Variables
For jointly continuous independent random variables X and Y having joint distribution function F (x, y) and joint probability density function f (x, y), Theorem 8.3 implies that, for any
two real numbers x and y,
F (x, y) = FX (x)FY (y).
(8.13)
Differentiating (8.13) with respect to x, we get
∂F
= fX (x)FY (y),
∂x
which upon differentiation with respect to y yields
∂2F
= fX (x)fY (y),
∂y∂x
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 352 — #368
✐
✐
352
Chapter 8
Bivariate Distributions
or, equivalently,
f (x, y) = fX (x)fY (y).
(8.14)
We leave it as an exercise that, if (8.14) is valid, then X and Y are independent. We have the
following theorem:
Theorem 8.7 Let X and Y be jointly continuous random variables with joint probability
density function f (x, y). Then X and Y are independent if and only if f (x, y) is the product
of their marginal densities fX (x) and fY (y).
Example 8.13
A point is selected at random from the rectangle
R = (x, y) ∈ R2 : 0 < x < a, 0 < y < b .
Let X be the x-coordinate and Y be the y -coordinate of the point selected. Determine if X and
Y are independent random variables.
Solution: From Section 8.1 we know that f (x, y), the joint probability density function of X
and Y, is given by
1
1
=
if (x, y) ∈ R
area(R)
ab
f (x, y) =
0
elsewhere.
Now the marginal probability density functions of X and Y, fX and fY , are given by
Z b
1
1
fX (x) =
dy = ,
x ∈ (0, a),
a
0 ab
Z a
1
1
fY (y) =
dx = ,
y ∈ (0, b).
b
0 ab
Therefore, f (x, y) = fX (x)fY (y), ∀x, y ∈ R, and hence X and Y are independent. Example 8.14 Stores A and B, which belong to the same owner, are located in two different
towns. If the probability density function of the weekly profit of each store, in thousands of
dollars, is given by
(
x/4
if 1 < x < 3
f (x) =
0
otherwise,
and the profit of one store is independent of the other, what is the probability that next week
one store makes at least $500 more than the other store?
Solution: Let X and Y denote next week’s profits of A and B, respectively. The desired
probability is P (X > Y + 1/2) + P (Y > X + 1/2). Since X and Y have the same
probability density function, by symmetry, this sum equals 2P (X > Y + 1/2). To calculate
this, we need to know f (x, y), the joint probability density function of X and Y . Since X and
Y are independent,
f (x, y) = fX (x)fY (y),
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 353 — #369
✐
✐
Section 8.2
Independent Random Variables
353
where
fX (x) =
(
x/4
if 1 < x < 3
0
otherwise,
fY (y) =
(
y/4
if 1 < y < 3
0
otherwise.
Thus
f (x, y) =
(
xy/16
0
if 1 < x < 3, 1 < y < 3
otherwise.
The desired probability, as seen from Figure 8.5, is therefore equal to
n
3
1 o
1
= 2P (X, Y ) ∈ (x, y) : < x < 3, 1 < y < x −
2P X > Y +
2
2
2
Z 3 Z x−1/2
Z 3 h 2 ix−1/2
xy
1
xy
=2
dy dx =
dx
16
8
2 1
3/2
1
3/2
Z
i
1 3 h
1 2
=
x x−
− 1 dx
16 3/2
2
Z 3 1
3 =
x3 − x2 − x dx
16 3/2
4
h
1 1 4 1 3 3 2 i3
549
=
x − x − x
=
≈ 0.54. 16 4
3
8
1024
3/2
/2
_ 1
y=
x
y
3
5/2
1
1 3/2
Figure 8.5
3
x
Figure of Example 8.14.
We now explain one of the most interesting problems of geometric probability,
Buffon’s needle problem. In Chapter 13 of the Companion Website for this book, we will
show how the solution of this problem and the Monte Carlo method can be used to find estimations for π by simulation. Georges Louis Buffon (1707–1784), who proposed and solved
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 354 — #370
✐
✐
354
Chapter 8
Bivariate Distributions
this famous problem, was a French naturalist who, in the eighteenth century, used probability
to study natural phenomena. In addition to the needle problem, his studies of the distribution
and expected value of the remaining lifetimes of human beings are famous among mathematicians. These works, together with many more, are published in his gigantic 44-volume, Histoire
Naturelle (Natural History; 1749–1804).
Example 8.15 (Buffon’s Needle Problem) A plane is ruled with parallel lines a distance d apart. A needle of length ℓ, ℓ < d, is tossed at random onto the plane. What is the
probability that the needle intersects one of the parallel lines?
Solution: Let X denote the distance from the center of the needle to the closest line, and let
Θ denote the angle between the needle and the line. The position of the needle is completely
determined by the coordinates X and Θ.
Figure 8.6
Buffon’s needle problem.
As Figure 8.6 shows, the needle intersects the line if and only if the length of the hypotenuse
of the triangle, X/ sin Θ, is less than ℓ/2. Since X and Θ are independent uniform random
variables over the RR
intervals (0, d/2) and (0, π), respectively, the probability that the needle
intersects a line is
f (x, θ) dx dθ, where f (x, θ) is the joint probability density function of
R
X and Θ and is given by
2
πd
f (x, θ) = fX (x)fΘ (θ) =
0
and
n
R = (x, θ) :
Thus the desired probability is
ZZ
R
f (x, θ) dx dθ =
d
.
2
elsewhere
o
ℓo n
ℓ
x
<
= (x, θ) : x < sin θ .
sin θ
2
2
Z π Z (ℓ/2) sin θ
0
if 0 ≤ θ ≤ π, 0 ≤ x ≤
0
2
ℓ
dx dθ =
πd
πd
Z π
0
sin θ dθ =
2ℓ
. πd
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 355 — #371
✐
✐
Section 8.2
Independent Random Variables
355
Example 8.16 Prove that two random variables X and Y with the following joint probability density function are not independent.
(
8xy
0≤x≤y≤1
f (x, y) =
0
otherwise.
Solution: To check the validity of (8.14), we first calculate fX and fY .
fX (x) =
Z 1
8xy dy = 4x(1 − x2 ),
0 ≤ x ≤ 1,
Z y
8xy dx = 4y 3 ,
0 ≤ y ≤ 1.
x
fY (y) =
0
Now since f (x, y) =
6 fX (x)fY (y), X and Y are dependent. This is expected because of the
range 0 ≤ x ≤ y ≤ 1, which is not a Cartesian product of one-dimensional regions. The results of Theorems 8.5 and 8.6 are valid for continuous random variables as well:
Let X and Y be independent continuous random variables and g : R → R
and h : R → R be real-valued functions; then g(X) and h(Y ) are also
independent random variables.
The proof of this fact is identical to the proof of Theorem 8.5. The continuous analog of Theorem 8.6 and its proof are as follows:
Let X and Y be independent continuous random variables. Then for all
real-valued functions g : R → R and h : R → R ,
E g(X)h(Y ) = E g(X) E h(Y ) ,
where, as usual, we assume that E g(X) and E h(Y ) are finite.
Proof: Let f (x, y) be the joint probability density function of X and Y . Then
Z ∞Z ∞
E g(X)h(Y ) =
g(x)h(y)f (x, y) dx dy
=
=
−∞
−∞
−∞
Z ∞
−∞
Z ∞Z ∞
−∞
=
Z ∞
h(y)fY (y)
g(x)fX (x) dx dy
Z ∞
g(x)h(y)fX (x)fY (y) dx dy
−∞
−∞
g(x)fX (x) dx
Z ∞
= E g(X) E h(Y ) . −∞
h(y)fY (y) dy
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 356 — #372
✐
✐
356
Chapter 8
Bivariate Distributions
Again, by this theorem, if X and Y are independent, then E(XY ) = E(X)E(Y ).
As we know from the discrete case, the converse of this fact is not necessarily true. (See
Example 8.12.)
EXERCISES
A
1.
Let the joint probability mass function of random variables X and Y be given by
1
if x = 1, 2, y = 0, 1, 2
(x2 + y 2 )
p(x, y) = 25
0
elsewhere.
Are X and Y independent? Why or why not?
2.
Let the joint probability mass function of random variables X and Y be given by
1
if (x, y) = (1, 1), (1, 2), (2, 1)
x2 y
p(x, y) = 7
0
elsewhere.
Are X and Y independent? Why or why not?
3.
Let X and Y be independent random variables each having the probability mass function
p(x) =
1 2 x
,
2 3
x = 1, 2, 3, . . . .
Find P (X = 1, Y = 3) and P (X + Y = 3).
4.
From an ordinary deck of 52 cards, eight cards are drawn at random and without replacement. Let X and Y be the number of clubs and spades, respectively. Are X and Y
independent?
5.
What is the probability that there are exactly two girls among the first seven and exactly
four girls among the first 15 babies born in a hospital in a given week? Assume that the
events that a child born is a girl or is a boy are equiprobable.
6.
Suppose that the number of claims received by an insurance company in a given week
is independent of the number of claims received in any other week. An actuary has calculated that the probability mass function of the number of claims received in a random
week is
1 4 n
,
n ≥ 0.
p(n) =
5 5
Find the probability that the company will receive exactly 8 claims during the next two
weeks.
7.
Let X and Y be two independent random variables with distribution functions F and
G, respectively. Find the distribution functions of max(X, Y ) and min(X, Y ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 357 — #373
✐
✐
Section 8.2
Independent Random Variables
357
8.
A fair coin is tossed n times by Adam and n times by Andrew. What is the probability
that they get the same number of heads?
9.
The joint probability mass function p(x, y) of the random variables X and Y is given
by the following table. Determine if X and Y are independent.
y
10.
x
0
1
2
3
0
1
2
3
0.1681
0.1804
0.0574
0.0041
0.1804
0.1936
0.0616
0.0044
0.0574
0.0616
0.0196
0.0014
0.0041
0.0044
0.0014
0.0001
Let the joint probability density function of random variables X and Y be given by
(
2
if 0 ≤ y ≤ x ≤ 1
f (x, y) =
0
elsewhere.
Are X and Y independent? Why or why not?
11.
Suppose that the amount of cholesterol in a certain type of sandwich is 100X milligrams, where X is a random variable with the following probability density function:
2x + 3
if 2 < x < 4
18
f (x) =
0
otherwise.
Find the probability that two such sandwiches made independently have the same
amount of cholesterol.
12.
Let the joint probability density function of random variables X and Y be given by
(
x2 e−x(y+1)
if x ≥ 0, y ≥ 0
f (x, y) =
0
elsewhere.
Are X and Y independent? Why or why not?
13.
Let the joint probability density function of X and Y be given by
(
8xy
if 0 ≤ x < y ≤ 1
f (x, y) =
0
otherwise.
Determine if E(XY ) = E(X)E(Y ).
14.
Let the joint probability density function of X and Y be given by
(
2e−(x+2y)
if x ≥ 0, y ≥ 0
f (x, y) =
0
otherwise.
Find E(X 2 Y ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 358 — #374
✐
✐
358
Chapter 8
15.
Let X and Y be two independent random variables with the same probability density
function given by
(
e−x
if 0 < x < ∞
f (x) =
0
elsewhere.
16.
17.
18.
19.
Bivariate Distributions
Show that g, the probability density function of X/Y, is given by
1
if 0 < t < ∞
(1 + t)2
g(t) =
0
t ≤ 0.
Let
X and Y be
independent exponential random variables both with mean 1. Find
E max(X, Y ) .
Let
X and Y be independent random points from the interval (−1, 1). Find
E max(X, Y ) .
Let X and Y be independent random points from the interval (0, 1). Find the probability
density function of the random variable XY .
A point is selected at random from the disk
R = (x, y) ∈ R2 : x2 + y 2 ≤ 1 .
Let X be the x-coordinate and Y be the y -coordinate of the point selected. Determine
if X and Y are independent random variables.
20.
Six brothers and sisters who are all either under 10 or in their early teens are having
dinner with their parents and four grandparents. Their mother unintentionally feeds the
entire family (including herself) a type of poisonous mushrooms that makes 20% of
the adults and 30% of the children sick. What is the probability that more adults than
children get sick?
21.
The lifetimes of mufflers manufactured by company A are random with the following
probability density function:
1
if x > 0
e−x/6
f (x) = 6
0
elsewhere.
The lifetimes of mufflers manufactured by company B are random with the following
probability density function:
2
if y > 0
e−2y/11
g(y) = 11
0
elsewhere.
Elizabeth buys two mufflers, one from company A and the other one from company B
and installs them on her cars at the same time. What is the probability that the muffler
of company B outlasts that of company A?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 359 — #375
✐
✐
Section 8.2
22.
23.
Independent Random Variables
359
Let X and Y be two independent random integers from the set {0, 1, 2, . . . , n}. Find
P (X = Y ) and P (X ≤ Y ).
Let the joint probability density function of X and Y be given by
1
if |x| < y, 0 < y < 1
f (x, y) =
0
otherwise.
Show that X and Y are dependent but E(XY ) = E(X)E(Y ).
24.
Andy buys an MP3 player and six new AAA batteries. The player is operated with two
AAA batteries, and each time a battery dies, Andy replaces that battery with a new one.
If the lifetimes of the batteries are independent exponential random variables each with
parameter λ, find the expected length of time Andy can listen to his MP3 player with
the six batteries.
Hint: Let X1 and X2 be the lifetimes of the first two batteries Andy will use to operate
his MP3. The first time a battery
dies is after min(X1 , X2 ) units of time. Begin by
calculating E min(X1 , X2 ) .
B
25.
Let E be an event; the random variable IE , defined as follows, is called the indicator of
E:
(
1
if E occurs
IE =
0
otherwise.
Show that A and B are independent events if and only if IA and IB are independent
random variables.
26.
Let B and C be two independent random variables both having the following probability
density function:
2
3x
if 1 < x < 3
f (x) = 26
0
otherwise.
What is the probability that the quadratic equation X 2 + BX + C = 0 has two real
roots?
27.
Let the joint probability density function of two random variables X and Y satisfy
f (x, y) = g(x)h(y),
−∞ < x < ∞,
−∞ < y < ∞,
where g and h are two functions from R to R. Show that X and Y are independent.
28.
Let X and Y be two independent random points from the interval (0, 1). Calculate the
distribution function and the probability density function of max(X, Y )/ min(X, Y ).
29.
Suppose that X and Y are independent, identically distributed exponential random variables with mean 1/λ. Prove that X/(X + Y ) is uniform over (0, 1).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 360 — #376
✐
✐
360
Chapter 8
30.
Let f (x, y) be the joint probability density function of two
√ continuous random variables;
f is called circularly symmetrical if it is a function of x2 + y 2 , the distance
√ of (x, y)
from the origin; that is, if there exists a function ϕ so that f (x, y) = ϕ( x2 + y 2 ).
Prove that if X and Y are independent random variables, their joint probability density
function is circularly symmetrical if and only if they are both normal with mean 0 and
equal variance.
Hint: Suppose that f is circularly symmetrical; then
p
fX (x)fY (y) = ϕ( x2 + y 2 ).
Bivariate Distributions
Differentiating this relation with respect to x yields
√
ϕ′ ( x2 + y 2 )
fX′ (x)
√ 2
√
=
.
xfX (x)
ϕ( x + y 2 ) x2 + y 2
This implies that both sides are constants, so that
fX′ (x)
=k
xfX (x)
for some constant k . Solve this and use the fact that fX is a probability density function
to show that fX is normal. Repeat the same procedure for fY .
Self-Quiz on Section 8.2
Time allotted: 15 Minutes
1.
Each problem is worth 5 points.
The joint probability density function of the continuous random variables X and Y is
given by
1
y 3 e−2x
x ≥ 0, 0 ≤ y ≤ 2
f (x, y) = 2
0
otherwise.
Are X and Y independent? Why or why not?
2.
8.3
Let X and Y be independent exponential random variables with parameters λ and µ,
respectively. What is the probability that X < Y ?
CONDITIONAL DISTRIBUTIONS
Let X and Y be two discrete or two continuous random variables. In this section we study the
distribution function and the expected value of the random variable X given that Y = y .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 361 — #377
✐
✐
Section 8.3
Conditional Distributions
361
Conditional Distributions: Discrete Case
Let X be a discrete random variable with set of possible values A, and let Y be a discrete random variable with set of possible values B . Let p(x, y) be the joint probability mass function
of X and Y, and let pX and pY be the marginal probability mass functions of X and Y . When
no information is given about the value of Y,
pX (x) = P (X = x) =
X
P (X = x, Y = y) =
y∈B
X
p(x, y)
(8.15)
y∈B
is used to calculate the probabilities of events concerning X . However, if the value of Y is
known, then instead of pX (x), the conditional probability mass function of X given that
Y = y is used. This function, denoted by pX|Y (x|y), is defined as follows:
pX|Y (x|y) = P (X = x | Y = y) =
p(x, y)
P (X = x, Y = y)
=
,
P (Y = y)
pY (y)
where x ∈ A, y ∈ B, and pY (y) > 0. Hence, for x ∈ A, y ∈ B, and pY (y) > 0,
pX|Y (x|y) =
Note that
X
pX|Y (x|y) =
x∈A
X p(x, y)
x∈A
pY (y)
=
p(x, y)
pY (y)
.
(8.16)
1 X
1
pY (y) = 1.
p(x, y) =
pY (y) x∈A
pY (y)
Hence, for any fixed y ∈ B, pX|Y (x|y) is itself a probability mass function with the set of
possible values A. If X and Y are independent, pX|Y coincides with pX because
pX|Y (x|y) =
p(x, y)
P (X = x, Y = y)
P (X = x)P (Y = y)
=
=
pY (y)
P (Y = y)
P (Y = y)
= P (X = x) = pX (x).
Note that, for x ∈ A, y ∈ B, and pX (x) > 0,
pY |X (y|x) =
p(x, y)
.
pX (x)
(8.17)
Similar to pX|Y (x|y), the conditional distribution function of X, given that Y = y is
defined as follows:
X
X
FX|Y (x|y) = P (X ≤ x | Y = y) =
P (X = t | Y = y) =
pX|Y (t|y).
t≤x
Example 8.17
t≤x
Let the joint probability mass function of X and Y be given by
1
if x = 0, 1, 2, y = 1, 2
(x + y)
p(x, y) = 15
0
otherwise.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 362 — #378
✐
✐
362
Chapter 8
Bivariate Distributions
Find pX|Y (x|y) and P (X = 0 | Y = 2).
Solution: To use pX|Y (x|y) = p(x, y)/pY (y), we must first calculate pY (y).
pY (y) =
2
X
p(x, y) =
x=0
1+y
.
5
Therefore,
pX|Y (x|y) =
(x + y)/15
x+y
=
,
(1 + y)/5
3(1 + y)
In particular, P (X = 0 | Y = 2) =
x = 0, 1, 2 when y = 1 or 2.
0+2
2
= . 3(1 + 2)
9
Example 8.18 Let N (t) be the number of males who enter a certain post office at or prior
to time t. Let M (t) be the number of females who enter a certain post office at or prior to t.
Suppose that N (t) : t ≥ 0 and M (t) : t ≥ 0 are independent Poisson processes with
rates λ and µ, respectively. So for all t > 0 and s > 0, N (t) is independent of M (s). If at
some instant t, N (t) + M (t) = n, what is the conditional probability mass function of N (t)?
Solution: For simplicity, let K(t) = N (t) + M (t) and p(x, n) be the joint probability mass
function of N (t) and K(t). Then pN (t)|K(t) (x|n), the desired probability mass function, is
found as follows:
P N (t) = x, K(t) = n
p(x, n)
=
pN (t)|K(t) (x|n) =
pK(t) (n)
P K(t) = n
P N (t) = x, M (t) = n − x
=
.
P K(t) = n
Since N (t) and M (t) are independent random variables and K(t) = N (t)+M (t) is a Poisson
random variable with rate λt + µt (accept this for now; we will prove it in Theorem 11.5),
pN (t)|K(t) (x|n) =
P N (t) = x P M (t) = n − x
=
P K(t) = n
e−λt (λt)x e−µt (µt)n−x
x!
(n − x)!
e−(λt+µt) (λt + µt)n
n!
!
n!
λx µn−x
n λ x µ n−x
=
=
x! (n − x)! (λ + µ)n
x λ+µ
λ+µ
!
n λ x λ n−x
1−
=
,
x = 0, 1, . . . , n.
x λ+µ
λ+µ
This shows that the conditional probability mass function of the number of males who have
entered the post office at or prior to t, given that altogether n persons have entered the post
office during this period, is binomial with parameters n and λ/(λ + µ). ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 363 — #379
✐
✐
Section 8.3
Conditional Distributions
363
We now generalize the concept of mathematical expectation to the conditional case. Let
X and Y be discrete random variables, and let the set of possible values of X be A. The
conditional expectation of the random variable X given that Y = y is as follows:
E(X | Y = y) =
X
x∈A
xP (X = x | Y = y) =
X
xpX|Y (x|y),
x∈A
where pY (y) > 0. Hence, by definition, for pY (y) > 0,
E(X | Y = y) =
X
xpX|Y (x|y).
(8.18)
X
ypY |X (y|x).
(8.19)
x∈A
Similarly, for pX (x) > 0,
E(Y | X = x) =
y∈B
Note that conditional expectations are simply ordinary expectations computed relative to
conditional distributions. For this reason, they satisfy the same properties that ordinary expectations do. For example, if h is an ordinary function from R to R , then for the discrete random
variables X and Y with set of possible values A for X, the expected value of h(X) is obtained
from
X
E h(X) | Y = y =
h(x)pX|Y (x|y).
(8.20)
x∈A
Example 8.19 Calculate the expected number of aces in a randomly selected poker hand
that is found to have exactly two jacks.
Solution: Let X and Y be the number of aces and jacks in a random poker hand, respectively.
Then
E(X | Y = 2) =
=
3
X
x=0
x
3
X
x=0
3
X
p(x, 2)
x
pY (2)
x=0
!
xpX|Y (x|2) =
! !
4
4
44
2
x
3−x
!
52
5
!
!
4
48
2
3
!
52
5
=
3
X
x=0
x
4
x
!
44
3−x
!
48
3
!
≈ 0.25. ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 364 — #380
✐
✐
364
Chapter 8
Bivariate Distributions
Example 8.20 While rolling a balanced die successively, the first 6 occurred on the third
roll. What is the expected number of rolls until the first 1?
Solution: Let X and Y be the number of rolls until the first 1 and the first 6, respectively. The
required quantity, E(X | Y = 3), is calculated from
∞
X
∞
X
p(x, 3)
E(X | Y = 3) =
xpX|Y (x|3) =
x
.
pY (3)
x=1
x=1
Clearly,
5 2 1 25
,
6
6
216
4 1 1 4
p(2, 3) =
,
=
6 6 6
216
pY (3) =
=
p(1, 3) =
1 5 1 6
6
6
=
5
,
216
p(3, 3) = 0,
and for x > 3,
p(x, 3) =
Therefore,
E(X | Y = 3) =
4 2 1 5 x−4 1 6
6
6
6
=
1 5 x−4
.
81 6
∞
∞
X
p(x, 3)
216 X
x
=
xp(x, 3)
pY (3)
25 x=1
x=1
∞
X
4
1 5 x−4 i
5
216 h
+2·
+
x·
1·
=
25
216
216 x=4
81 6
∞
8 X 5 x−4
13
+
x
=
25 75 x=4 6
∞
5 x−4
13
8 X
=
+
(x − 4 + 4)
25 75 x=4
6
∞
∞ x−4 i
5 x−4
X
13
8h X
5
=
+
(x − 4)
+4
25 75 x=4
6
6
x=4
∞
∞ x i
X
8 h X 5 x
5
13
+
x
+4
=
25 75 x=0 6
6
x=0
h
13
1 i
5/6
8
=
+
+
4
= 6.28. 25 75 (1 − 5/6)2
1 − 5/6
Example 8.21 (The Box Problem: To Switch or Not to Switch) Suppose that there
are two identical closed boxes, one containing twice as much money as the other. David is
asked to choose one of these boxes for himself. He picks a box at random and opens it. If, after
observing the content of the box, he has the option to exchange it for the other box, should he
exchange or not?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 365 — #381
✐
✐
Section 8.3
Conditional Distributions
365
Solution: Suppose that David finds $x in the box. Then the other box has $(2x) with probability 1/2 and $(x/2) with probability 1/2. Therefore, the expected value of the dollar amount
in the other box is
1
x 1
· + 2x · = 1.25x.
2 2
2
This implies that David should switch no matter how much he finds in the box he picks at
random. But as Steven J. Brams and D. Marc Kilgour observe in their paper “The Box Problem:
To Switch or Not to Switch” (Mathematics Magazine, February 1995, Volume 68, Number 1),
It seems paradoxical that it would always be better to switch to the second
box, no matter how much money you found in the first one.
They go on to explain,
What is needed to determine whether or not switching is worthwhile is some
prior notion of the amount of money to be found in each box. For particular, knowing the prior distribution would enable you to calculate whether the
amount found in the first box is less than the expected value of what is in the
other box—and, therefore, whether a switch is profitable.
Let L be the event that David picked the box with the larger amount. Let X be the dollar
amount David finds in the box he picks. According to a general exchange condition, discovered
by Brams and Kilgour, on the average, an exchange is profitable if and only if for the observed
value x,
P (L | X = x) < 2/3.
That is, David should switch if and only if the conditional probability that he picked the box
with the larger amount given that he found $x is less than 2/3. To show this, let Y be the amount
in the box David did not pick, and let S be the event that David picked the box with the smaller
amount. Then
E(Y | X = x) =
x
· P (L | X = x) + 2x · P (S | X = x).
2
On average, exchange is profitable if and only if E(Y | X = x) > x, or, equivalently, if and
only if
x
· P (L | X = x) + 2x · P (S | X = x) > x.
2
Given that P (S | X = x) = 1 − P (L | X = x), this relation yields
1
P (L | X = x) + 2 1 − P (L | X = x) > 1,
2
which implies that P (L | X = x) < 2/3.
As an example, suppose that the box with the larger amount contains $1, $2, $4, or $8 with
equal probabilities. Then, for x = 1/2, 1, 2, and 4, David should switch , for x = 8 he should
not. This is because, for example, by Bayes’ formula (Theorem 3.5),
1 1
×
P (X = 2 | L)P (L)
1
2
4 2
P (L | X = 2) =
=
= < ,
P (X = 2 | L)P (L) + P (X = 2 | S)P (S)
2
3
1 1 1 1
× + ×
4 2 4 2
whereas P (L | X = 8) = 1 > 2/3. ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 366 — #382
✐
✐
366
Chapter 8
Bivariate Distributions
Conditional Distributions: Continuous Case
Now let X and Y be two continuous random variables with the joint probability density
function
f (x, y). Again, when no information is given about the value of Y, fX (x) =
R∞
f
(x,
y) dy is used to calculate the probabilities of events concerning X . However, when
−∞
the value of Y is known, to find such probabilities, fX|Y (x|y), the conditional probability
density function of X given that Y = y is used. Similar to the discrete case, fX|Y (x|y) is
defined as follows:
fX|Y (x|y) =
f (x, y)
,
fY (y)
(8.21)
provided that fY (y) > 0. Note that
Z ∞
fX|Y (x|y) dx =
−∞
Z ∞
f (x, y)
1
dx =
f
(y)
f
Y
Y (y)
−∞
Z ∞
f (x, y) dx =
−∞
1
fY (y) = 1,
fY (y)
showing that for a fixed y, fX|Y (x|y) is itself a probability density function. If X and Y are
independent, then fX|Y coincides with fX because
fX|Y (x|y) =
fX (x)fY (y)
f (x, y)
=
= fX (x).
fY (y)
fY (y)
Similarly, the conditional probability density function of Y given that X = x is defined
by
fY |X (y|x) =
f (x, y)
,
fX (x)
(8.22)
provided that fX (x) > 0. Also, as we expect, FX|Y (x|y), the conditional distribution
function of X given that Y = y is defined as follows:
Z x
fX|Y (t|y) dt.
FX|Y (x|y) = P (X ≤ x | Y = y) =
−∞
Therefore,
d
FX|Y (x|y) = fX|Y (x|y).
dx
Example 8.22
function
(8.23)
Let X and Y be continuous random variables with joint probability density
3
(x2 + y 2 )
f (x, y) = 2
0
if 0 < x < 1,
0<y<1
otherwise.
Find fX|Y (x|y).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 367 — #383
✐
✐
Section 8.3
Solution: By definition, fX|Y (x|y) =
fY (y) =
Z ∞
Conditional Distributions
367
f (x, y)
, where
fY (y)
f (x, y) dx =
−∞
Z 1
0
3
1
3 2
(x + y 2 ) dx = y 2 + .
2
2
2
Thus
fX|Y (x|y) =
3/2(x2 + y 2 )
3(x2 + y 2 )
=
(3/2)y 2 + 1/2
3y 2 + 1
for 0 < x < 1 and 0 < y < 1. Everywhere else, fX|Y (x|y) = 0.
Figure 8.7
Conditional probability density function of Y given X = x.
Figure 8.7 shows the conditional probability density function of Y given that X = x. Note
that fY |X (y|x) is a vertical slice taken from the surface of f (x, y) at the given point x. Thus
the graph of fY |X (y|x) is obtained from the intersection of the plane X = x and the surface
f (x, y).
Example 8.23 First, a point Y is selected at random from the interval (0, 1). Then another
point X is chosen at random from the interval (0, Y ). Find the probability density function of
X.
Solution: Let f (x, y) be the joint probability density function of X and Y . Then
Z ∞
fX (x) =
f (x, y) dy,
−∞
where from fX|Y (x|y) =
f (x, y)
, we obtain
fY (y)
f (x, y) = fX|Y (x|y)fY (y).
Therefore,
fX (x) =
Z ∞
fX|Y (x|y)fY (y) dy.
−∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 368 — #384
✐
✐
368
Chapter 8
Bivariate Distributions
Since Y is uniformly distributed over (0, 1),
(
1
fY (y) =
0
if 0 < y < 1
elsewhere.
Since given Y = y, X is uniformly distributed over (0, y),
(
1/y
if 0 < y < 1, 0 < x < y
fX|Y (x|y) =
0
elsewhere.
Thus
fX (x) =
Z ∞
fX|Y (x|y)fY (y) dy =
−∞
Z 1
dy
= ln 1 − ln x = − ln x.
y
x
Therefore,
fX (x) =
Example 8.24
(
− ln x
if 0 < x < 1
0
elsewhere.
Let the conditional probability density function of X, given that Y = y, be
fX|Y (x|y) =
x + y −x
e ,
1+y
0 < x < ∞,
0 < y < ∞.
Find P (X < 1 | Y = 2).
Solution: The probability density function of X given Y = 2 is
fX|Y (x|2) =
x + 2 −x
e ,
3
0 < x < ∞.
Therefore,
Z 1
Z
Z 1
i
x + 2 −x
1 h 1 −x
P (X < 1 | Y = 2) =
e dx =
xe dx +
2e−x dx
3
3 0
0
0
h
h
i
i
1
1
2
4
1
− xe−x − e−x − e−x = 1 − e−1 ≈ 0.509.
=
3
3
3
0
0
Note that while we calculated P (X < 1 | Y = 2), it is not possible to calculate probabilities such as P (X < 1 | 2 < Y < 3) using fX|Y (x|y). This is because the probability density
function of X, given that 2 < Y < 3, cannot be found from the conditional probability density
function of X, given Y = y . Similar to the case where X and Y are discrete, for continuous random variables X and Y
with joint probability density function f (x, y), the conditional expectation of X given that
Y = y is as follows:
Z ∞
xfX|Y (x|y) dx,
(8.24)
E(X | Y = y) =
−∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 369 — #385
✐
✐
Section 8.3
Conditional Distributions
369
where fY (y) > 0. Similarly, the conditional expectation of Y given that X = x is given
by
Z ∞
yfY |X (y|x) dy,
(8.25)
E(Y | X = x) =
−∞
where fX (x) > 0.
As explained in the discrete case, conditional expectations are simply ordinary expectations
computed relative to conditional distributions. For this reason, they satisfy the same properties
that ordinary expectations do. For example, if h is an ordinary function from R to R , then, for
continuous random variables X and Y, with joint probability density function f (x, y),
E h(X) | Y = y =
Z ∞
h(x)fX|Y (x|y) dx.
(8.26)
−∞
In particular, this implies that the conditional variance of X given that Y = y is given by
Z ∞
2
2
x − E(X | Y = y) fX|Y (x|y) dx.
(8.27)
σX|Y =y =
−∞
Example 8.25
function
Let X and Y be continuous random variables with joint probability density
f (x, y) =
(
e−y
0
if y > 0,
0<x<1
elsewhere.
Find E(X | Y = 2).
Solution: From the definition,
Z 1
Z ∞
Z 1
e−2
f (x, 2)
dx =
dx.
E(X | Y = 2) =
x
xfX|Y (x|2) dx =
x
fY (2)
fY (2)
0
−∞
0
But fY (2) =
R1
0
f (x, 2) dx =
R1
0
e−2 dx = e−2 ; therefore,
E(X | Y = 2) =
Z 1
1
e−2
x −2 dx = . e
2
0
Example 8.26 The lifetimes of batteries manufactured by a certain company are identically
distributed with distribution and probability density functions F and f, respectively. In terms
of F, f, and s, find the expected value of the lifetime of an s-hour-old battery.
Solution: Let X be the lifetime of the s-hour-old battery. We want to calculate
E(X | X > s). Let
FX|X>s (t) = P (X ≤ t | X > s),
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 370 — #386
✐
✐
370
Chapter 8
Bivariate Distributions
′
and fX|X>s (t) = FX|X>s (t). Then
E(X | X > s) =
Z ∞
tfX|X>s(t) dt.
0
Now
P (X ≤ t, X > s)
FX|X>s (t) = P (X ≤ t | X > s) =
P (X > s)
0
if t ≤ s
= P (s < X ≤ t)
if t > s.
P (X > s)
Therefore,
FX|X>s (t) =
0
F (t) − F (s)
1 − F (s)
Differentiating FX|X>s (t) with respect to t, we obtain
0
fX|X>s(t) =
f (t)
1 − F (s)
if t ≤ s
if t > s.
if t ≤ s
if t > s.
This yields
E(X | X > s) =
Z ∞
0
1
tfX|X>s(t) dt =
1 − F (s)
Z ∞
tf (t) dt. s
⋆ Remark 8.1 Suppose that, in Example 8.26, a battery manufactured by the company is
installed at time 0 and begins to operate. If at time s an inspector finds the battery dead, then
the expected lifetime of the dead battery is E(X | X < s). Similar to the calculations in that
example, we can show that
Z s
1
E(X | X < s) =
tf (t) dt.
F (s) 0
(See Exercise 24.)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 371 — #387
✐
✐
Section 8.3
Conditional Distributions
371
EXERCISES
A
1.
Let the joint probability mass function of discrete random variables X and Y be given
by
1
(x2 + y 2 )
if x = 1, 2, y = 0, 1, 2
p(x, y) = 25
0
otherwise.
Find pX|Y (x|y), P (X = 2 | Y = 1), and E(X | Y = 1).
2.
Regions A and B are prone to dust storms. An actuary has calculated that the number
of dust storms that strike region A, in a year, is binomial with parameters 4 and 0.7.
Furthermore, she has observed that if, in a year, n dust storms strike region A, then
the number of dust storms striking region B is n with probability 2/3 and n + 1 with
probability 1/3. What is the expected number of dust storms in region A in a year in
which 3 dust storms strike region B.
3.
Let the joint probability density function of continuous random variables X and Y be
given by
(
2
if 0 < x < y < 1
f (x, y) =
0
elsewhere.
Find fX|Y (x|y).
4.
An unbiased coin is flipped until the sixth head is obtained. If the third head occurs on
the fifth flip, what is the probability mass function of the number of flips?
5.
Let the conditional probability density function of X given that Y = y be given by
fX|Y (x|y) =
3(x2 + y 2 )
,
3y 2 + 1
0 < x < 1,
0 < y < 1.
Find P (1/4 < X < 1/2 | Y = 3/4).
6.
7.
Let X and Y be independent discrete random variables. Prove that for all y,
E(X | Y = y) = E(X). Do the same for continuous random variables X and Y .
Let X and Y be continuous random variables with joint probability density function
(
x+y
if 0 ≤ x ≤ 1, 0 ≤ y ≤ 1
f (x, y) =
0
elsewhere.
Calculate fX|Y (x|y).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 372 — #388
✐
✐
372
8.
Chapter 8
Bivariate Distributions
Let X and Y be continuous random variables with joint probability density function
given by
(
e−x(y+1)
if x ≥ 0, 0 ≤ y ≤ e − 1
f (x, y) =
0
elsewhere.
Calculate E(X | Y = y).
9.
A box contains 5 blue, 10 green, and 5 red chips. We draw 4 chips at random and without
replacement. If exactly one of them is blue, what is the probability mass function of the
number of green balls drawn?
10.
First a point Y is selected at random from the interval (0, 1). Then another point X is
selected at random from the interval (Y, 1). Find the probability density function of X .
11.
Let (X, Y ) be a random point from a unit disk centered at the origin. Find
P (0 ≤ X ≤ 4/11 | Y = 4/5).
12.
An actuary working for an insurance company has calculated that the time that it will
take for an insured driver to report an accident he or she has been involved in, in days,
is a continuous random variable X, with the probability density function
196 − x2
579
f (x) =
0
0<x<3
otherwise.
Furthermore, the actuary has discovered that if the accident is reported x days after its
occurrence, and if the insured is entitled to claim payment, the time that it will take
until he or she receives payment is a uniform random variable between x + 7 and 21
days. What is the probability that a driver involved in an accident and entitled to claim
payment will receive his or her payment within 14 days from the time of the accident?
13.
14.
15.
The joint probability density function of X and Y is given by
(
c e−x
if x ≥ 0, |y| < x
f (x, y) =
0
otherwise.
(a)
Determine the constant c.
(b)
Find fX|Y (x|y) and fY |X (y|x).
(c)
Calculate E(Y | X = x) and Var(Y | X = x).
Leon leaves his office every day at a random time between 4:30 P.M. and 5:00 P.M. If he
leaves t minutes past 4:30, the time it will take him to reach home is a random number
between 20 and 20 + (2t)/3 minutes. Let Y be the number of minutes past 4:30 that
Leon leaves his office tomorrow and X be the number of minutes it takes him to reach
home. Find the joint probability density function of X and Y .
Show that if N (t) : t ≥ 0 is a Poisson process, the conditional distribution of the first
arrival time given N (t) = 1 is uniform on (0, t).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 373 — #389
✐
✐
Section 8.3
16.
Conditional Distributions
373
In a sequence of independent Bernoulli trials, let X be the number of successes in the
first m trials and Y be the number of successes in the first n trials, m < n. Show
that the conditional distribution of X, given Y = y, is hypergeometric. Also, find the
conditional distribution of Y given X = x.
B
17.
18.
A point is selected at random and uniformly from the region
R = (x, y) : |x| + |y| ≤ 1 .
Find the conditional probability density function of X given Y = y .
Let N (t) : t ≥ 0 be a Poisson process. For s < t show that the conditional distribution of N (s) given N (t) = n is binomial with parameters n and p = s/t. Also find the
conditional distribution of N (t) given N (s) = k.
19.
Cards are drawn from an ordinary deck of 52, one at a time, randomly and with replacement. Let X and Y denote the number of draws until the first ace and the first king are
drawn, respectively. Find E(X | Y = 5).
20.
A box contains 10 red and 12 blue chips. Suppose that 18 chips are drawn, one by one,
at random and with replacement. If it is known that 10 of them are blue, show that the
expected number of blue chips in the first nine draws is five.
21.
Let X and Y be continuous random variables with joint probability density function
(
n(n − 1)(y − x)n−2
if 0 ≤ x ≤ y ≤ 1
f (x, y) =
0
otherwise.
Find the conditional expectation of Y given that X = x.
22.
23.
A point (X, Y ) is selected randomly from the triangle with vertices (0, 0), (0, 1), and
(1, 0).
(a)
Find the joint probability density function of X and Y .
(b)
Calculate fX|Y (x|y).
(c)
Evaluate E(X | Y = y).
Let X and Y be discrete random variables with joint probability mass function
p(x, y) =
1
,
e2 y! (x − y)!
x = 0, 1, 2, . . . ,
y = 0, 1, 2, . . . , x,
p(x, y) = 0, elsewhere. Find E(Y | X = x).
24.
The lifetimes of batteries manufactured by a certain company are identically distributed
with distribution and probability density functions F and f, respectively. Suppose that
a battery manufactured by this company is installed at time 0 and begins to operate. If
at time s an inspector finds the battery dead, in terms of F, f, and s, find the expected
lifetime of the dead battery.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 374 — #390
✐
✐
374
Chapter 8
Bivariate Distributions
Self-Quiz on Section 8.3
Time allotted: 20 Minutes
Each problem is worth 5 points.
1.
There are 16 pea seeds in a packet available to a gardener. Seven of the peas are Wando,
4 are Maestro, and 5 are Lincoln. To test the soil of a specific garden for growing peas,
the gardener picks 8 of these seeds at random to plant in that garden. Let X be the
number of Lincoln seeds and Y be the number of Wando seeds chosen. Find the joint
probability mass function of X and Y, the conditional probability mass function of X
given that Y = y, and P (X = 3 | Y = 3).
2.
A bivariate probability density function is called copula if its marginal probability density functions are uniform random variables over the interval (0, 1). For a constant α
between −1 and 1, let X and Y be continuous random variables with the joint probability density function given by
f (x, y) = 1 + α(1 − 2x)(1 − 2y),
0 < x < 1, 0 < y < 1.
The function f is an example of a copula bivariate probability density function. Copula
distributions have significant applications in portfolio optimization and financial risk
assessment. Find fX|Y (x|y) and E(X | Y = y).
8.4
TRANSFORMATIONS OF TWO RANDOM VARIABLES
In our preceding discussions of random variables, cases have arisen where we have calculated
the distribution and the probability density functions of a function of a random variable X .
Functions such as X 2 , eX , cos X, X 3 + 1, and so on. In particular, in Section 6.2 we explained
how, in general, probability density functions and distribution functions of such functions can
be obtained. In this section we demonstrate a method for finding the joint probability density
function of functions of two random variables. The key is the following, which is the analog of
the change of variable theorem for functions of several variables.
Theorem 8.8 Let X and Y be continuous random variables with joint probability density
function f (x, y). Let h1 and h2 be real-valued functions of two variables, U = h1 (X, Y ) and
V = h2 (X, Y ). Suppose that
(a)
u = h1 (x, y) and v = h2 (x, y) defines a one-to-one transformation of a set R in
the xy -plane onto a set Q in the uv -plane. That is, for (u, v) ∈ Q, the system of two
equations in two unknowns,
(
h1 (x, y) = u
(8.28)
h2 (x, y) = v,
has a unique solution x = w1 (u, v) and y = w2 (u, v) for x and y, in terms of u and
v; and
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 375 — #391
✐
✐
Section 8.4 Transformations of Two Random Variables
(b)
375
the functions w1 and w2 have continuous partial derivatives, and the Jacobian of the
transformation x = w1 (u, v) and y = w2 (u, v) is nonzero at all points (u, v) ∈ Q;
that is, the following 2 × 2 determinant is nonzero on Q:
J=
∂w1
∂u
∂w1
∂v
∂w2
∂u
∂w2
∂v
=
∂w1 ∂w2 ∂w1 ∂w2
−
6= 0.
∂u ∂v
∂v ∂u
Then the random variables U and V are jointly continuous with the joint probability density
function g(u, v) given by
f w1 (u, v), w2 (u, v) J
(u, v) ∈ Q
g(u, v) =
(8.29)
0
elsewhere.
Theorem 8.8 is a result of the change of a variable theorem in double integrals. To see this, let
B be a subset of Q in the uv -plane. Suppose that in the xy -plane, A ⊆ R is the set that is
transformed to B by the one-to-one transformation (8.28). Clearly, the events (U, V ) ∈ B and
(X, Y ) ∈ A are equiprobable. Therefore,
ZZ
f (x, y) dx dy.
P (U, V ) ∈ B = P (X, Y ) ∈ A =
A
Using the change of variable formula for double integrals, we have
ZZ
ZZ
f (x, y) dx dy =
f w1 (u, v), w2 (u, v) J du dv.
A
B
Hence
P (U, V ) ∈ B =
ZZ
B
f w1 (u, v), w2 (u, v) J du dv.
Since this is true for all subsets B of Q, it shows that g(u, v), the joint probability density
function of U and V, is given by (8.29).
Example 8.27 Let X and Y be positive independent random variables with the identical
probability density function e−x for x > 0. Find the joint probability density function of
U = X + Y and V = X/Y .
Solution: Let f (x, y) be the joint probability density function of X and Y . Then
(
e−x
if x > 0
fX (x) =
0
if x ≤ 0,
(
e−y
if y > 0
fY (y) =
0
if y ≤ 0.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 376 — #392
✐
✐
376
Chapter 8
Bivariate Distributions
Therefore,
f (x, y) = fX (x)fY (y) =
(
e−(x+y)
0
if x > 0 and y > 0
elsewhere.
Let h1 (x, y) = x + y and h2 (x, y) = x/y . Then the system of equations
x+ y = u
x
=v
y
has the unique solution x = (uv)/(v + 1), y = u/(v + 1), and
J=
v
v+1
1
v+1
u
(v + 1)2
u
−
(v + 1)2
=−
u
6= 0,
(v + 1)2
since x > 0 and y > 0 imply that u > 0 and v > 0; that is,
Q = (u, v) : u > 0 and v > 0 .
Hence, by Theorem 8.8, g(u, v), the joint probability density function of U and V, is given by
g(u, v) = e−u
u
,
(v + 1)2
u > 0 and v > 0. The following example proves a well-known theorem called Box–Muller’s theorem,
which, as we explain in Section 13D of Chapter 13 of the Companion Website for this book, is
used to simulate normal random variables (see Theorem 13c of the same section).
Example 8.28 Let X and Y be two independent
uniform random variables
√ over (0, 1);
√
show that the random variables U = cos(2πX) −2 ln Y and V = sin(2πX) −2 ln Y are
independent standard normal random variables.
√
√
Solution: Let h1 (x, y) = cos(2πx) −2 ln y and h2 (x, y) = sin(2πx) −2 ln y. Then
the system of equations
(
√
cos(2πx) −2 ln y = u
√
sin(2πx) −2 ln y = v
defines a one-to-one transformation of the set
R = (x, y) : 0 < x < 1, 0 < y < 1
onto
Q = (u, v) : − ∞ < u < ∞, −∞ < v < ∞ ;
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 377 — #393
✐
✐
Section 8.4 Transformations of Two Random Variables
377
hence it can be solved uniquely in terms of x and y . To see this, square both sides of these
equations and sum them up. We obtain −2 ln y = u2 + v 2 , which gives
y = exp − (u2 + v 2 )/2 .
Putting −2 ln y = u2 + v 2 back into these equations, we get
u
cos 2πx = √ 2
u + v2
and
v
sin 2πx = √ 2
,
u + v2
which enable us to determine the unique value of x. For example, if u √
> 0 and v > 0, then
2πx is uniquely determined in the first quadrant from 2πx = arccos(u/ u2 + v 2 ). Hence the
first condition of Theorem 8.8 is satisfied. To check the second condition, note that for u > 0
and v > 0,
1
u
,
arccos √ 2
2π
u + v2
w2 (u, v) = exp − (u2 + v 2 )/2 .
w1 (u, v) =
Hence
J=
=
∂w1
∂u
∂w1
∂v
∂w2
∂u
∂w2
∂v
=
−v
2π(u2 + v 2 )
−u exp − (u2 + v 2 )/2
1
exp − (u2 + v 2 )/2 6= 0.
2π
u
2π(u2 + v 2 )
−v exp − (u2 + v 2 )/2
Now X and Y being two independent uniform random variables over (0, 1) imply that f, their
joint probability density function is
(
1
if 0 < x < 1, 0 < y < 1
f (x, y) = fX (x)fY (y) =
0
elsewhere.
Hence, by Theorem 8.8, g(u, v), the joint probability density function of U and V is given by
g(u, v) =
1
exp − (u2 + v 2 )/2 ,
2π
−∞ < u < ∞, −∞ < v < ∞.
The probability density function of U is calculated as follows:
Z ∞
u2 + v 2 −u2 Z ∞
−v 2 1
1
gU (u) =
exp −
dv =
exp
exp
dv,
2
2π
2
2
−∞ 2π
−∞
Z ∞
Z ∞
√
√
2
where
(1/ 2π) exp(−v /2) dv = 1 implies that
exp(−v 2 /2) dv =
2π.
−∞
−∞
Therefore,
−u2 1
gU (u) = √ exp
,
2
2π
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 378 — #394
✐
✐
378
Chapter 8
Bivariate Distributions
which shows that U is standard normal. Similarly,
−v 2 1
.
gV (v) = √ exp
2
2π
Since g(u, v) = gU (u)gV (v), U and V are independent standard normal random variables.
As an application of Theorem 8.8, we now prove the following theorem, an excellent
resource for calculation of probability density and distribution functions of sums of continuous
independent random variables.
Theorem 8.9 (Convolution Theorem)
Let X and Y be continuous independent random variables with probability density functions f1 and f2 and distribution functions F1 and
F2 , respectively. Then g and G, the probability density and distribution functions of X + Y,
respectively, are given by
Z ∞
g(t) =
f1 (x)f2 (t − x) dx,
−∞
Z ∞
G(t) =
f1 (x)F2 (t − x) dx.
−∞
Proof: Let f (x, y) be the joint probability density function of X and Y . Then f (x, y) =
f1 (x)f2 (y). Let U = X + Y, V = X, h1 (x, y) = x + y, and h2 (x, y) = x. Then the system
of equations
x+y =u
x=v
has the unique solution x = v, y = u − v, and
J=
∂x
∂u
∂x
∂v
∂y
∂u
∂y
∂v
0
1
=
1 −1
= −1 6= 0.
Hence, by Theorem 8.8, the joint probability density function of U and V, ψ(u, v), is given by
ψ(u, v) = f1 (v)f2 (u − v)|J | = f1 (v)f2 (u − v).
Therefore, the marginal probability density function of U = X + Y is
Z ∞
Z ∞
g(u) =
ψ(u, v) dv =
f1 (v)f2 (u − v) dv,
−∞
−∞
which is the same as
g(t) =
Z ∞
−∞
f1 (x)f2 (t − x) dx.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 379 — #395
✐
✐
Section 8.4 Transformations of Two Random Variables
379
To find G(t), the distribution function of X + Y, note that
G(t) =
Z t
g(u) du =
−∞
=
Z ∞ Z t
−∞
=
Z t Z ∞
Z ∞
−∞
−∞
−∞
−∞
f1 (x)f2 (u − x) dx du
f2 (u − x) du f1 (x) dx
F2 (t − x)f1 (x) dx,
where, letting s = u − x, the last equality follows from
Z t
−∞
f2 (u − x) du =
Z t−x
−∞
f2 (s) ds = F2 (t − x). Note that by symmetry we can also write
Z ∞
g(t) =
f2 (y)f1 (t − y) dy,
−∞
G(t) =
Z ∞
−∞
Definition 8.6
defined by
f2 (y)F1 (t − y) dy.
Let f1 and f2 be two probability density functions. Then the function g(t),
g(t) =
Z ∞
−∞
f1 (x)f2(t − x) dx,
is called the convolution of f1 and f2 . Theorem 8.9 shows that
If X and Y are independent continuous random variables, the probability density function of X + Y is the convolution of the probability density
functions of X and Y .
Theorem 8.9 is also valid for discrete random variables. Let pX and pY be probability mass
functions of two discrete random variables X and Y . Then the function
p(z) =
X
x
pX (x)pY (z − x)
is called the convolution of pX and pY . It is readily seen that if X and Y are independent, the
probability mass function of X + Y is the convolution of the probability mass functions of X
and Y :
P (X + Y = z) =
=
X
x
P (X = x, Y = z − x) =
x
pX (x)pY (z − x).
X
X
x
P (X = x)P (Y = z − x)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 380 — #396
✐
✐
380
Chapter 8
Bivariate Distributions
Exercise 6, a famous example given by W. J. Hall, shows that the converse of Theorem 8.9 is
not valid. That is, it may happen that the probability mass function of two dependent random
variables X and Y is the convolution of the probability mass functions of X and Y .
Example 8.29 Let X and Y be independent exponential random variables, each with
parameter λ. Find the distribution function of X + Y .
Solution: Let f1 and f2 be the probability density functions of X and Y, respectively. Then
λe−λx if x ≥ 0
f1 (x) = f2 (x) =
0
otherwise.
Hence
f2 (t − x) =
λe−λ(t−x)
0
if 0 ≤ x ≤ t
otherwise.
By convolution theorem, h, the probability density function of X + Y, is given by
Z ∞
Z t
h(t) =
f2 (t − x)f1 (x) dx =
λe−λ(t−x) · λe−λx dx = λ2 te−λt , t ≥ 0.
−∞
0
This is the probability density function of a gamma random variable with parameters 2 and
λ. Hence X + Y is gamma with parameters 2 and λ. We will study a generalization of this
important theorem in Section 11.2. EXERCISES
A
1.
Let X and Y be independent random numbers from the interval (0, 1). Find the joint
probability density function of U = −2 ln X and V = −2 ln Y .
2.
Let X and Y be two positive independent continuous random variables with the probability density functions f1 (x) and f2 (y), respectively. Find the probability density function of U = X/Y.
Hint: Let V = X; find the joint probability density function of U and V . Then calculate the marginal probability density function of U .
3.
Let X ∼ N (0, 1) and Y ∼ N (0, 1)√be independent random variables. Find the joint
probability density function of R = X 2 + Y 2 and Θ = arctan(Y /X). Show that
R and Θ are independent. Note that (R, Θ) is the polar coordinate representation of
(X, Y ).
4.
Let X and Y be continuous random variables with the joint probability density function
given by
8xy
0<y≤x<1
f (x, y) =
0
otherwise.
Find the probability density function of U = XY .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 381 — #397
✐
✐
Section 8.4 Transformations of Two Random Variables
381
g(t)
1
0
Figure 8.8
5.
1
2
t
The shape of a triangular probability density function.
From the interval (0, 1), two random numbers X and Y are selected independently.
Show that the probability density function of their sum is given by
if 0 ≤ t < 1
t
g(t) = 2 − t
if 1 ≤ t < 2
0
otherwise.
Because of the shape of the probability density function of X + Y, the random variable
X + Y is said to have a triangular distribution. (See Figure 8.8.)
6.
Let −1/9 < c < 1/9 be a constant. Let p(x, y), the joint probability mass function of
the random variables X and Y, be given by the following table:
y
7.
x
−1
0
1
−1
0
1
1/9
1/9 + c
1/9 − c
1/9 − c
1/9
1/9 + c
1/9 + c
1/9 − c
1/9
(a)
Show that the probability mass function of X + Y is the convolution function of
the probability mass functions of X and Y for all c.
(b)
Show that X and Y are independent if and only if c = 0.
All international passengers arriving at a U.S. airport must go through the first stage
of an immigration process, which takes an exponential length of time with parameter
λ. However, experience shows that for some p, 0 < p < 1, 100p% of the passengers
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 382 — #398
✐
✐
382
Chapter 8
Bivariate Distributions
have discrepancies in their travel related documents and must go through a second stage,
which is exponential with parameter µ independently of stage 1. Determine the probability density function of the time a randomly selected international passenger spends
in the Airport Immigration Office upon arrival to the United States. Such a probability
density function is called Coxian.
B
8.
Let X and Y be independent random variables with common probability density function
1
2
if x ≥ 1
x
f (x) =
0
elsewhere.
Calculate the joint probability density function of U = X/Y and V = XY.
9.
Let X and Y be independent random variables with common probability density function
(
e−x
if x > 0
f (x) =
0
elsewhere.
Find the joint probability density function of U = X + Y and V = eX .
10.
Prove that if X and Y are independent standard normal random variables, then X + Y
and X − Y are independent random variables. This is a special case of the following
important theorem.
Let X and Y be independent random variables with a common distribution F . The random variables X + Y and X − Y are independent
if and only if F is a normal distribution function.
11.
12.
13.
Let X and Y be independent (strictly positive) gamma random variables with parameters
(r1 , λ) and (r2 , λ), respectively. Define U = X + Y and V = X/(X + Y ).
(a)
Find the joint probability density function of U and V .
(b)
Prove that U and V are independent.
(c)
Show that U is gamma and V is beta.
Let X and Y be independent gamma random variables with parameters (n/2, 1/2) and
(m/2, 1/2),respectively, where n and m are positive integers. The random variable
U = (X/n) (Y /m) is said to have F -distribution with n and m degrees of freedom,
and has significant applications in statistics. Find the probability density function of U .
Let X and Y be independent (strictly positive) exponential random variables each with
parameter λ. Are the random variables X + Y and X/Y independent?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 383 — #399
✐
✐
Chapter 8
Summary
383
Self-Quiz on Section 8.4
Time allotted: 25 Minutes
1.
2.
Each problem is worth 5 points.
Let X and Y be two independently selected random numbers from the interval (0, 1).
Find the joint probability density function of U = X + Y and V = X − Y .
The joint probability density function of the continuous random variables X and Y is
given by
2e−(x+2y)
x ≥ 0, y ≥ 0
f (x, y) =
0
otherwise.
Find h, the probability density function of X + Y .
CHAPTER 8 SUMMARY
◮ Joint Probability Mass Functions Let X and Y be two discrete random variables
defined on the same sample space. Let the sets of possible values of X and Y be A and B,
respectively. The function
p(x, y) = P (X = x, Y = y)
is called the joint probability
P mass function of X and Y . The functions pX (x) =
P
x∈A p(x, y) are called, respectively, the marginal probability
y∈B p(x, y) and pY (y) =
mass functions of X and Y .
◮ If h is an ordinary function of two variables from R2 to R , then h(X, Y ) is a discrete
random variable with the expected value given by
X X
E h(X, Y ) =
h(x, y)p(x, y),
x∈A y∈B
provided that the sum is absolutely convergent. By this relation, E(X + Y ) = E(X) + E(Y ).
◮ Joint Probability Density Functions Two random variables X and Y, defined on the
same sample space, have a continuous joint distribution if there exists a nonnegative function
of two variables, f (x, y), on R × R , such that for any region R in the xy -plane that can be
formed from rectangles by a countable number of set operations,
ZZ
P (X, Y ) ∈ R =
f (x, y) dx dy.
R
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 384 — #400
✐
✐
384
Chapter 8
Bivariate Distributions
The function f (x, y) is called the joint probability
density function of XR and Y . The funcR∞
∞
tions fX and fY , given by fX (x) = −∞ f (x, y) dy and fY (y) = −∞ f (x, y) dx are
called,
density functions of X and Y . We have E(X) =
R ∞ respectively, the marginalR probability
∞
xfX (x) dx and E(Y ) = −∞ yfY (y) dy . Furthermore, if h is an ordinary function of
−∞
two variables from R2 to R , then h(X, Y ) is a random variable with the expected value
given by
Z ∞Z ∞
h(x, y)f (x, y) dx dy,
E h(X, Y ) =
−∞
−∞
provided that the integral is absolutely convergent. By this relation,
E(X + Y ) = E(X) + E(Y ).
◮ Geometric Probability Let S be a subset of the plane with area A(S). A point is said
to be randomly selected from S if for any subset R of S with area A(R), the probability that
R contains the point is A(R)/A(S).
◮ Independence of Random Variables Two random variables X and Y are called
independent if, for arbitrary subsets A and B of real numbers, the events {X ∈ A} and
{Y ∈ B} are independent, that is, if P (X ∈ A, Y ∈ B) = P (X ∈ A)P (Y ∈ B). Two
random variables are independent if and only if their joint distribution function is the product
of their marginal distribution functions. Two discrete random variables are independent if and
only if their joint probability mass function is the product of their marginal probability mass
functions. Two continuous random variables are independent if and only if their joint probability density function is the product of their marginal probability density functions.
Let X and Y be independent random variables and g : R → R and h : R → R be realvalued functions; then g(X) and h(Y ) are also independent random variables. Furthermore,
E g(X)h(Y ) = E g(X) E h(Y ) ,
In particular, E(XY ) = E(X)E(Y ). However, two random variables X and Y might be
dependent while E(XY ) = E(X)E(Y ).
◮ Conditional Distributions: Discrete Case Let X be a discrete random variable with
set of possible values A, and let Y be a discrete random variable with set of possible values
B . Let p(x, y) be the joint probability mass function of X and Y, and let pX and pY be the
marginal probability mass functions of X and Y . The conditional probability mass function of
X given that Y = y , denoted by pX|Y (x|y), is defined as follows:
pX|Y (x|y) = P (X = x | Y = y) =
p(x, y)
,
pY (y)
x ∈ A, y ∈ B,
where pY (y) > 0. The conditional expectation of the random variable X given that Y = y is
as follows:
E(X | Y = y) =
X
x∈A
xP (X = x | Y = y) =
X
xpX|Y (x|y),
x∈A
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 385 — #401
✐
✐
Chapter 8
Summary
385
where pY (y) > 0. Similarly, for pX (x) > 0,
E(Y | X = x) =
X
y∈B
ypY |X (y|x),
and if h is an ordinary function from R to R , then
X
E h(X) | Y = y =
h(x)pX|Y (x|y).
x∈A
◮ Conditional Distributions: Continuous Case Let X and Y be two continuous random variables with the joint probability density function f (x, y). The conditional probability
density function of X given that Y = y , fX|Y (x|y), is defined as follows:
fX|Y (x|y) =
f (x, y)
,
fY (y)
provided that fY (y) > 0. Similarly, The conditional probability density function of Y given
that X = x, fY |X (y|x), is defined as follows:
fY |X (y|x) =
f (x, y)
,
fX (x)
provided that fX (x) > 0. The conditional expectation of X given that Y = y is as follows:
Z ∞
xfX|Y (x|y) dx,
E(X | Y = y) =
−∞
where fY (y) > 0. Similarly, the conditional expectation of Y given that X = x is given by
Z ∞
yfY |X (y|x) dy,
E(Y | X = x) =
−∞
where fX (x) > 0. If h is an ordinary function from R to R , then,
Z ∞
E h(X) | Y = y =
h(x)fX|Y (x|y) dx.
−∞
◮ Transformations of Two Random Variables Let X and Y be continuous random variables with joint probability density function f (x, y). Let h1 and h2 be real-valued functions of
two variables, U = h1 (X, Y ) and V = h2 (X, Y ). Suppose that
(a)
u = h1 (x, y) and v = h2 (x, y) defines a one-to-one transformation of a set R in the
xy -plane onto a set Q in the uv -plane. That is, for (u, v) ∈ Q, the system of two
equations in two unknowns,
(
h1 (x, y) = u
h2 (x, y) = v,
has a unique solution x = w1 (u, v) and y = w2 (u, v) for x and y, in terms of u and v;
and
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 386 — #402
✐
✐
386
Chapter 8
(b)
the functions w1 and w2 have continuous partial derivatives, and the Jacobian of the
transformation x = w1 (u, v) and y = w2 (u, v) is nonzero at all points (u, v) ∈ Q;
that is, the following 2 × 2 determinant is nonzero on Q:
Bivariate Distributions
J=
∂w1
∂u
∂w1
∂v
∂w2
∂u
∂w2
∂v
=
∂w1 ∂w2 ∂w1 ∂w2
−
6= 0.
∂u ∂v
∂v ∂u
Then the random variables U and V are jointly continuous with the joint probability density
function g(u, v) given by
f w1 (u, v), w2 (u, v) J
(u, v) ∈ Q
g(u, v) =
0
elsewhere.
◮ Convolution Theorem Let X and Y be continuous independent random variables with
probability density functions f1 and f2 and distribution functions F1 and F2 , respectively. Then
g and G, the probability density and distribution functions of X + Y, respectively, are given by
Z ∞
g(t) =
f1 (x)f2 (t − x) dx,
−∞
G(t) =
Z ∞
−∞
f1 (x)F2 (t − x) dx.
REVIEW PROBLEMS
1.
The joint probability mass function of X and Y is given by the following table.
x
(a)
(b)
2.
y
1
2
3
2
4
6
0.05
0.14
0.10
0.25
0.10
0.02
0.15
0.17
0.02
Find P (XY ≤ 6).
Find E(X) and E(Y ).
A fair die is tossed twice. The sum of the outcomes is denoted by X and the largest
value by Y . (a) Calculate the joint probability mass function of X and Y ; (b) find the
marginal probability mass functions of X and Y ; (c) find E(X) and E(Y ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 387 — #403
✐
✐
Chapter 8
Review Problems
387
3.
Calculate the probability mass function of the number of spades in a random bridge hand
that includes exactly four hearts.
4.
Suppose that three cards are drawn at random from an ordinary deck of 52 cards. If X
and Y are the numbers of diamonds and clubs, respectively, calculate the joint probability mass function of X and Y .
5.
Calculate the probability mass function of the number of spades in a random bridge hand
that includes exactly four hearts and three clubs.
6.
Let the joint probability density function of X and Y be given by
c
if 0 < y < x, 0 < x < 2
f (x, y) = x
0
elsewhere.
7.
(a)
Determine the value of c.
(b)
Find the marginal probability density functions of X and Y .
Let X and Y have the joint probability density function below. Determine if E(XY ) =
E(X)E(Y ).
3
1
x2 y + y
if 0 < x < 1 and 0 < y < 2
4
4
f (x, y) =
0
elsewhere.
8.
Prove that the following cannot be the joint distribution function of two random variables
X and Y .
(
1
if x + y ≥ 1
F (x, y) =
0
if x + y < 1.
9.
Three concentric circles of radii r1 , r2 , and r3 , r1 > r2 > r3 , are the boundaries of the
regions that form a circular target. If a person fires a shot at random at the target, what
is the probability that it lands in the middle region?
10.
A fair coin is flipped 20 times. If the total number of heads is 12, what is the expected
number of heads in the first 10 flips?
11.
Let the joint distribution function of the lifetimes of two brands of lightbulb be given by
(
2
2
(1 − e−x )(1 − e−y )
if x > 0, y > 0
F (x, y) =
0
otherwise.
12.
Find the probability that one lightbulb lasts more than twice as long as the other.
For Ω = (x, y) : 0 < x + y < 1, 0 < x < 1, 0 < y < 1 , a region in the plane, let
(
3(x + y)
if (x, y) ∈ Ω
f (x, y) =
0
otherwise
be the joint probability density function of the random variables X and Y . Find the
marginal probability density functions of X and Y, and P (X + Y > 1/2).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 388 — #404
✐
✐
388
Chapter 8
13.
Let X and Y be continuous random variables with the joint probability density function
(
e−y
if y > 0, 0 < x < 1
f (x, y) =
0
elsewhere.
Bivariate Distributions
Find E(X n | Y = y), n ≥ 1.
14.
From an ordinary deck of 52 cards, cards are drawn successively and with replacement.
Let X and Y denote the number of spades in the first 10 cards and in the second 15
cards, respectively. Calculate the joint probability mass function of X and Y .
15.
Let the joint probability density function of X and Y be given by
(
cx(1 − x)
if 0 ≤ x ≤ y ≤ 1
f (x, y) =
0
otherwise.
(a)
Determine the value of c.
(b)
Determine if X and Y are independent.
16.
A point is selected at random from the bounded region between the curves y = x2 − 1
and y = 1 − x2 . Let X be the x-coordinate, and let Y be the y -coordinate of the point
selected. Determine if X and Y are independent.
17.
Let X and Y be two independent uniformly distributed random variables over the intervals (0, 1) and (0, 2), respectively. Find the probability density function of X/Y .
18.
If F is the distribution function of a random variable X, is G(x, y) = F (x) + F (y) a
joint distribution function?
19.
A bar of length ℓ is broken into three pieces at two random spots. What is the probability
that the length of at least one piece is less than ℓ/20?
20.
There are prizes in 10% of the boxes of a certain type of cereal. Let X be the number
of boxes of such cereal that Kim should buy to find a prize. Let Y be the number of
additional boxes of such cereal that she should purchase to find another prize. Calculate
the joint probability mass function of X and Y .
21.
Let the joint probability density function of random variables X and Y be given by
(
1
if |y| < x, 0 < x < 1
f (x, y) =
0
otherwise.
Show that E(Y | X = x) is a linear function of x while E(X | Y = y) is not a linear
function of y .
22.
(The Wallet Paradox) Consider the following “paradox” given by Martin Gardner
in his book Aha! Gotcha (W. H. Freeman and Company, New York, 1981).
Each of two persons places his wallet on the table. Whoever has the
smallest amount of money in his wallet, wins all the money in the other
wallet. Each of the players reason as follows: “I may lose what I have but
I may also win more than I have. So the game is to my advantage.”
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 389 — #405
✐
✐
Chapter 8
Self-Test Problems
389
As Kent G. Merryfield, Ngo Viet, and Saleem Watson have observed in their paper “The
Wallet Paradox” in the August–September 1997 issue of the American Mathematical
Monthly,
Paradoxically, it seems that the game is to the advantage of both players. . . . However, the inference that “the game is to my advantage” is the
source of the apparent paradox, because it does not take into account
the probabilities of winning or losing. In other words, if the game is played
many times, how often does a player win? How often does he lose? And
by how much?
Following the analysis of Kent G. Merryfield, Ngo Viet, and Saleem Watson, let X and
Y be the amount of money in the wallets of players A and B, respectively. Let WA and
WB be the amount of money that player A and B will win, respectively. WA (X, Y ) =
−WB (X, Y ) and
−X if X > Y
WA (X, Y ) =
Y if X < Y
0 if X = Y .
Suppose that the distribution function of the money in each player’s wallet is the same;
that is, X and Y are independent, identically distributed random variables on some
interval [a, b] or [a, ∞), 0 ≤ a < b < ∞. Show that
E(WA ) = E(WB ) = 0.
Self-Test on Chapter 8
Time allotted: 150 Minutes
1.
2.
Each problem is worth 10 points.
A point is selected at random from the set S = (x, y) : x2 + y 2 ≤ 4 .
(a) What is the probability that it falls inside the set E = (x, y) : x2 + y 2 ≤ 1 ?
(b) What is the probability that it falls inside the set F = (x, y) : |x| + |y| < 2 ?
Let
f (x, y) =
60x2 y
0
x ≥ 0, y ≥ 0, x + y ≤ 1
otherwise
be the joint probability density function of continuous random variables X and Y . Are
X and Y independent? Why or why not?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 390 — #406
✐
✐
390
Chapter 8
Bivariate Distributions
3.
The lifetime of a particular type of tablet PC is exponential with mean 8 years. The
manufacturer sells each tablet PC for $960.00 and offers a warranty to fully refund the
customer’s money if the PC fails within a year of purchase and to refund half of the
customer’s money if it fails during the second year following the purchase. There will
be no refund after the second year. If the manufacturer sells 11,000 tablet PC’s every
year, what is the expected value of the total refund amount it pays every year? Assume
that the set of the lifetimes of all such PC’s is an independent set of random variables.
4.
A certain admission test, given to prospective graduate applicants, consists of two sections, verbal reasoning and quantitative reasoning. Let 100X and 100Y be the respective
verbal and quantitative scores of a randomly selected applicant with an undergraduate
3.0 GPA. Suppose that X and Y are independent and both have the following probability
density function
8
4≤x≤8
2
f (x) = x
0
otherwise.
5.
6.
Find the probability that the total score, 100X + 100Y, of the randomly selected applicant is at least 1300?
A point is selected randomly from the unit disk D = (x, y) : x2 + y 2 ≤ 1 .
(a)
Find the conditional probability density function of X, the x-coordinate of the
point selected, given that Y, it’s y -coordinate, is y .
(b)
Find the conditional expected value and the conditional variance of X given that
Y = y.
Let the joint probability density function of the continuous random variables X and Y
be given by
xy/16
1 < x < 3, 1 < y < 3
f (x, y) =
0
otherwise.
Using the convolution theorem, find the probability density function of X + Y .
7.
A parallel system with two components is a system that functions if and only if at least
one of its components functions. Let X and Y, the lifetimes of the components, be
independent exponential random variables with parameters λ and µ, respectively. Find
the probability density function of |X − Y |, the time between the failures of the two
components.
8.
Let X be the amount of time a patient with a non-life threatening condition has to spend
in an emergency room before being seen by a doctor. Suppose that the emergency room
serves such patients with rate Y, where Y is an exponential random variable with mean
1/2. Also suppose that the conditional distribution for X given Y = y is exponential
with mean 1/y . Find the probability density function, the mean, and the standard deviation of the rate at which the emergency room serves the non-life threatening patients
when the waiting time is the constant x.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 391 — #407
✐
✐
Chapter 8
9.
Self-Test Problems
391
Let the joint probability density function of the continuous random variable X and Y be
given by
4e−2(x+y)
x > 0, y > 0
f (x, y) =
0
otherwise.
X
.
X +Y
Let X and Y be two independently selected random numbers from the interval (0, 1).
(a) Find the joint probability density function of U = X + Y and V = X/(X + Y ).
(b) Find the marginal probability density function of V .
Find the probability density function of the random variable U =
10.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 392 — #408
✐
✐
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 393 — #409
✐
✐
Chapter 9
Multivariate Distributions
9.1
JOINT DISTRIBUTIONS OF n > 2 RANDOM VARIABLES
Joint Probability Mass Functions
The following definition generalizes the concept of joint probability mass function of two
discrete random variables to n > 2 discrete random variables.
Definition 9.1 Let X1 , X2 , . . . , Xn be discrete random variables defined on the same sample space, with sets of possible values A1 , A2 , . . . , An , respectively. The function
p(x1 , x2 , . . . , xn ) = P (X1 = x1 , X2 = x2 , . . . , Xn = xn )
is called the joint probability mass function of X1 , X2 , . . . , Xn .
Note that
(a)
(b)
(c)
p(x1 , x2 , . . . , xn ) ≥ 0.
If for some i, 1 ≤ i ≤ n, xi 6∈ Ai , then p(x1 , x2 , . . . , xn ) = 0.
P
xi ∈Ai , 1≤i≤n p(x1 , x2 , . . . , xn ) = 1.
Moreover, if the joint probability mass function of random variables X1 , X2 , . . . , Xn ,
p(x1 , x2 , . . . , xn ), is given, then for 1 ≤ i ≤ n, the marginal probability mass function
of Xi , pXi , can be found from p(x1 , x2 , . . . , xn ) by
pXi (xi ) = P (Xi = xi ) = P (Xi = xi ; Xj ∈ Aj , 1 ≤ j ≤ n, j 6= i)
X
=
p(x1 , x2 , . . . , xn ).
(9.1)
xj ∈Aj , j6=i
More generally, to find the joint probability mass function marginalized over a given set
of k of these random variables, we sum up p(x1 , x2 , . . . , xn ) over all possible values of the
remaining n−k random variables. For example, if p(x, y, z) denotes the joint probability mass
function of random variables X, Y, and Z, then
pX,Y (x, y) =
X
p(x, y, z)
z
393
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 394 — #410
✐
✐
394
Chapter 9
Multivariate Distributions
is the joint probability mass function marginalized over X and Y, whereas
X
pY,Z (y, z) =
p(x, y, z)
x
is the joint probability mass function marginalized over Y and Z .
Example 9.1 Dr. Shams has 23 hypertensive patients, of whom five do not use any medicine
but try to lower their blood pressures by self-help: dieting, exercise, not smoking, relaxation,
and so on. Of the remaining 18 patients, 10 use beta blockers and 8 use diuretics. A random
sample of seven of all these patients is selected. Let X, Y, and Z be the number of the patients
in the sample trying to lower their blood pressures by self-help, beta blockers, and diuretics, respectively. Find the joint probability mass function and the marginal probability mass functions
of X, Y, and Z .
Solution: Let p(x, y, z) be the joint probability mass function of X, Y, and Z . Then for
0 ≤ x ≤ 5, 0 ≤ y ≤ 7, 0 ≤ z ≤ 7, x + y + z = 7,
!
! !
5
10
8
x
y
z
!
;
p(x, y, z) =
23
7
p(x, y, z) = 0, otherwise. The marginal probability mass function of X, pX , is obtained as
follows:
!
!
!
5
10
8
7−x
X
X
x
y
7−x−y
!
pX (x) =
p(x, y, z) =
23
y=0
x+y+z=7
0≤y≤7
7
0≤z≤7
!
5
!
!
7−x
x X
10
8
!
=
.
7−x−y
23 y=0 y
7
!
!
7−x
X
10
8
Now
is the total number of the ways we can choose 7 − x patients
y
7−x−y
y=0
from 18 (in!Exercise 58 of Section 2.4, let m = 10, n = 8, and r = 7 − x); hence it is equal
18
to
. Therefore, as expected,
7−x
!
!
5
18
x
7−x
!
pX (x) =
,
x = 0, 1, 2, 3, 4, 5.
23
7
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 395 — #411
✐
✐
Section 9.1
Joint Distributions of n > 2 Random Variables
395
Similarly,
!
13
7−y
!
,
23
7
!
!
8
15
z
7−z
!
,
23
7
10
y
pY (y) =
pZ (z) =
!
y = 0, 1, 2, . . . , 7,
z = 0, 1, 2, . . . , 7. Remark 9.1 A joint probability mass function, such as p(x, y, z) of Example 9.1, is called
multivariate hypergeometric. In general, suppose that a box contains n1 marbles of type 1,
n2 marbles of type 2, . . . , and nr marbles of type r . If n marbles are drawn at random and
Xi (i = 1, 2, . . . , r ) is the number of the marbles of type i drawn, the joint probability mass
function of X1 , X2 , . . . , Xr is called multivariate hypergeometric and is given by
!
!
!
n1
n2
nr
···
x1
x2
xr
!,
p(x1 , x2 , . . . , xr ) =
n1 + n2 + · · · + nr
n
where x1 + x2 + · · · + xr = n. We now extend the definition of the joint distribution from 2 to n > 2 random variables.
Let X1 , X2 , . . . , Xn be n random variables (discrete, continuous, or mixed). Then the joint
distribution function of X1 , X2 , . . . , Xn is defined by
F (t1 , t2 , . . . , tn ) = P (X1 ≤ t1 , X2 ≤ t2 , . . . , Xn ≤ tn )
(9.2)
for all −∞ < ti < +∞, i = 1, 2, . . . , n.
The marginal distribution function of Xi , 1 ≤ i ≤ n, can be found from F as follows:
FXi (ti ) = P (Xi ≤ ti )
= P (X1 < ∞, . . . , Xi−1 < ∞, Xi ≤ ti , Xi+1 < ∞, . . . , Xn < ∞)
=
lim
tj →∞
1≤j≤n, j6=i
F (t1 , t2 , . . . , tn ).
(9.3)
More generally, to find the joint distribution function marginalized over a given set of k of
these random variables, we calculate the limit of F (t1 , t2 , . . . , tn ) as tj → ∞, for every j that
belongs to one of the remaining n − k variables. For example, if F (x, y, z, t) denotes the joint
probability density function of random variables X, Y, Z, and T, then the joint distribution
function marginalized over Y and T is given by
FY,T (y, t) = lim F (x, y, z, t).
x,z→∞
Just as the joint distribution function of two random variables, we have that F, the joint distribution of n random variables, satisfies the following:
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 396 — #412
✐
✐
396
Chapter 9
(a)
F is nondecreasing in each argument.
(b)
F is right continuous in each argument.
(c)
F (t1 , t2 , . . . , ti−1 , −∞, ti+1 , . . . , tn ) = 0 for i = 1, 2, . . . , n.
(d)
Multivariate Distributions
F (∞, ∞, . . . , ∞) = 1.
We now generalize the concept of independence of random variables from two to any
number of random variables.
Suppose that X1 , X2 , . . . , Xn are random variables (discrete, continuous, or mixed) on a
sample space. We say that they are independent if, for arbitrary subsets A1 , A2 , . . . , An of
real numbers,
P (X1 ∈ A1 , X2 ∈ A2 , . . . , Xn ∈ An ) = P (X1 ∈ A1 )P (X2 ∈ A2 ) · · · P (Xn ∈ An ).
Similar to the case of two random variables, X1 , X2 , . . . , Xn are independent if and only if,
for any xi ∈ R , i = 1, 2, . . . , n,
P (X1 ≤ x1 , X2 ≤ x2 , . . . , Xn ≤ xn ) = P (X1 ≤ x1 )P (X2 ≤ x2 ) · · · P (Xn ≤ xn ).
That is, X1 , X2 , . . . , Xn are independent if and only if
F (x1 , x2 , . . . , xn ) = FX1 (x1 )FX2 (x2 ) · · · FXn (xn ).
If X1 , X2 , . . . , Xn are discrete, the definition of independence reduces to the following condition:
P (X1 = x1 , . . . , Xn = xn ) = P (X1 = x1 ) · · · P (Xn = xn )
(9.4)
for any set of points, xi , i = 1, 2, . . . , n.
Let X, Y, and Z be independent random variables. Then, by definition, for arbitrary subsets
A1 , A2 , and A3 of R ,
P (X ∈ A1 , Y ∈ A2 , Z ∈ A3 ) = P (X ∈ A1 )P (Y ∈ A2 )P (Z ∈ A3 ).
(9.5)
Now, if in (9.5) we let A2 = R, then since the event Y ∈ R is certain and has probability 1,
we get
P (X ∈ A1 , Z ∈ A3 ) = P (X ∈ A1 )P (Z ∈ A3 ).
This shows that X and Z are independent random variables. In the same way it can be shown
that {X, Y }, and {Y, Z} are also independent sets. Similarly, if {X1 , X2 , . . . , Xn } is a sequence of independent random variables, its subsets are also independent sets of random variables. This observation motivates the following definition for the concept of independence of
any collection of random variables.
Definition 9.2
A collection of random variables is called independent if all of its finite
subcollections are independent.
Sometimes independence of a collection of random variables is self-evident and requires
no checking. For example, let E be an experiment with the sample space S . Let X be a random
variable defined on S . If the experiment E is repeated n times independently and on the ith
experiment X is called Xi , the sequence {X1 , X2 , . . . , Xn } is an independent sequence of
random variables.
It is also important to know that the result of Theorem 8.5 is true for any number of random
variables:
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 397 — #413
✐
✐
Section 9.1
Joint Distributions of n > 2 Random Variables
397
If {X1 , X2 , . . .} is a sequence of independent random variables and for
i = 1, 2, . . . , gi : R → R is a real-valued function, then the sequence
{g1 (X1 ), g2 (X2 ), . . .} is also an independent sequence of random variables.
The following theorem is the generalization of Theorem 8.4. It follows from (9.4).
Theorem 9.1 Let X1 , X2 , . . . , Xn be jointly discrete random variables with the joint
probability mass function p(x1 , x2 , . . . , xn ). Then X1 , X2 , . . . , Xn are independent if and
only if p(x1 , x2 , . . . , xn ) is the product of their marginal probability mass functions pX1 (x1 ),
pX2 (x2 ), . . . , pXn (xn ).
The following is a generalization of Theorem 8.1 from dimension 2 to n.
Theorem 9.2 Let p(x1 , x2 , . . . , xn ) be the joint probability mass function of discrete random variables X1 , X2 , . . . , Xn . For 1 ≤ i ≤ n, let Ai be the set of possible values of Xi .
If h is a function of n variables from Rn to R, then Y = h(X1 , X2 , . . . , Xn ) is a discrete
random variable with expected value given by
X
X
E(Y ) =
···
h(x1 , x2 , . . . , xn )p(x1 , x2 , . . . , xn ),
xn ∈An
x1 ∈A1
provided that the sum is finite.
Using Theorems 9.1 and 9.2, an almost identical proof to that of Theorem 8.6 implies that
The expected value of the product of several independent discrete random
variables is equal to the product of their expected values.
Example 9.2
Let
p(x, y, z) = k(x2 + y 2 + yz),
x = 0, 1, 2; y = 2, 3; z = 3, 4.
(a)
For what value of k is p(x, y, z) a joint probability mass function?
(b)
Suppose that, for the value of k found in part (a), p(x, y, z) is the joint probability mass
function of random variables X, Y, and Z . Find PY,Z (y, z) and pZ (z).
(c)
Find E(XZ).
Solution:
(a)
We must have
k
2 X
3 X
4
X
(x2 + y 2 + yz) = 1.
x=0 y=2 z=3
This implies that k = 1/203.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 398 — #414
✐
✐
398
(b)
Chapter 9
Multivariate Distributions
2
PY,Z (y, z) =
=
1 X 2
(x + y 2 + yz)
203 x=0
1
(3y 2 + 3yz + 5),
203
2
pZ (z) =
=
y = 2, 3; z = 3, 4.
3
1 XX 2
(x + y 2 + yz)
203 x=0 y=2
7
15
z+ ,
203
29
z = 3, 4.
An alternate way to find pZ (z) is to use the result obtained in part (b):
pZ (z) =
3
X
3
PY,Z (y, z) =
y=2
=
(c)
15
7
z+ ,
203
29
1 X 2
(3y + 3yz + 5)
203 y=2
z = 3, 4.
By Theorem 9.2,
2
E(XZ) =
3
4
1 XXX
774
xz(x2 + y 2 + yz) =
≈ 3.81. 203 x=0 y=2 z=3
203
⋆ Example 9.3 (Reliability of Systems)† Consider a system consisting of n components denoted by 1, 2, . . . , n. Suppose that component i, 1 ≤ i ≤ n, is either functioning or
not functioning, with no other performance capabilities. Let Xi be a Bernoulli random variable
defined by
(
1 if the component i is functioning,
Xi =
0 if the component i is not functioning.
The random variable Xi determines the performance mode of the ith component of the system.
Suppose that
pi = P (Xi = 1) = 1 − P (Xi = 0).
The value pi , the probability that the ith component functions, is called the reliability of the ith
component. Throughout we assume that the components function independently of each other.
That is, {X1 , X2 , . . . , Xn } is an independent set of random variables.
Suppose that the system itself also has only two performance capabilities, functioning and
not functioning. Let X be a Bernoulli random variable defined by
(
1 if the system is functioning,
X=
0 if the system is not functioning.
†
If this example is skipped, then all exercises and examples in this and future chapters marked “(Reliability
of Systems)” should be skipped as well.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 399 — #415
✐
✐
Section 9.1
Joint Distributions of n > 2 Random Variables
399
The random variable X determines the performance mode of the system. Suppose that
r = P (X = 1) = 1 − P (X = 0).
The value r, the probability that the system functions, is called the reliability of the system.
We consider only systems for which it is possible to determine r from knowledge of the performance modes of the components.
As an example, a system is called series system when it functions if and only if all n of its
components function. For such a system,
X = min(X1 , X2 , . . . , Xn ) = X1 X2 · · · Xn .
This is because, if at least one component does not function, then Xi = 0 for at least one i,
which implies that X = 0, indicating that the system does not function. Similarly, if all n
components of the system function, then Xi = 1 for 1 ≤ i ≤ n, implying that X = 1, which
indicates that the system functions. For a series system, the preceding relationship, expressing
X in terms of X1 , X2 , . . . , Xn , enables us to find the reliability of the system:
r = P (X = 1) = P min(X1 , X2 , . . . Xn ) = 1
= P (X1 = 1, X2 = 1, . . . , Xn = 1)
= P (X1 = 1)P (X2 = 1) · · · P (Xn = 1) =
1
Figure 9.1
2
3
n
Y
pi .
i=1
n _1
n
Geometric representation for a series system.
Geometrically, a series system is often represented by the diagram of Figure 9.1. The reason
for this representation is that, for a signal fed in at the input to be transmitted to the output, all
components must be functioning. Otherwise, the signal cannot go all the way through.
As another example, consider a parallel system. Such a system functions if and only if at
least one of its n components functions. For a parallel system,
X = max(X1 , X2 , . . . , Xn ) = 1 − (1 − X1 )(1 − X2 ) · · · (1 − Xn ).
This is because, if at least one component functions, then for some i, Xi = 1, implying that
X = 1, whereas if none of the components functions, then for all i, Xi = 0 implying that
X = 0. The reliability of a parallel system is given by
r = P (X = 1) = P max(X1 , X2 , . . . , Xn ) = 1
= 1 − P max(X1 , X2 , . . . , Xn ) = 0
= 1 − P (X1 = 0, X2 = 0, . . . , Xn = 0)
= 1 − P (X1 = 0)P (X2 = 0) · · · P (Xn = 0) = 1 −
n
Y
i=1
(1 − pi ).
Geometrically, a parallel system is shown by the diagram of Figure 9.2 on the next page.
This representation shows that, for a signal fed in at the input, at least one component must
function to transmit it to the output.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 400 — #416
✐
✐
400
Chapter 9
Multivariate Distributions
1
2
3
n _1
n
Figure 9.2
Geometric representation for a parallel system.
Calculation of r, the reliability of a given system, is not, in general, as easy as it was for
series and parallel systems. One useful technique for evaluating r is to write X, if possible, in
terms of X1 , X2 , . . . , Xn , and then use the following simple relation:
r = P (X = 1) = 1 · P (X = 1) + 0 · P (X = 0) = E(X).
An example follows.
3
1
2
6
4
Figure 9.3
5
A combination of series and parallel systems.
Consider the system whose structure is as shown in Figure 9.3. This system is a combination of series and parallel systems. The fact that for a series system, X = X1 X2 · · · Xn , and
for a parallel system,
X = 1 − (1 − X1 )(1 − X2 ) · · · (1 − Xn ),
enable us to find out immediately that, for this system,
X = X1 X2 1 − (1 − X3 )(1 − X4 X5 ) X6 = X1 X2 X6 (X3 + X4 X5 − X3 X4 X5 ).
Since for 1 ≤ i ≤ 6,
E(Xi ) = 1 · P (Xi = 1) + 0 · P (Xi = 0) = pi ,
and the Xi ’s are independent random variables, r = E(X) yields
r = E(X) = E(X1 )E(X2 )E(X6 ) E(X3 ) + E(X4 )E(X5 ) − E(X3 )E(X4 )E(X5 )
= p1 p2 p6 (p3 + p4 p5 − p3 p4 p5 ). ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 401 — #417
✐
✐
Section 9.1
Joint Distributions of n > 2 Random Variables
401
Joint Probability Density Functions
We now define the concept of a jointly continuous probability density function for more than
two random variables.
Definition 9.3
Let X1 , X2 , . . . , Xn be continuous random variables defined on the same
sample space. We say that X1 , X2 , . . . , Xn have a continuous joint distribution if there exists
a nonnegative function of n variables, f (x1 , x2 , . . . , xn ), on R × R × · · · × R ≡ Rn such
that for any region R in Rn that can be formed from n-dimensional rectangles by a countable
number of set operations,
Z
Z
P (X1 , X2 , . . . , Xn ) ∈ R =
· · · f (x1 , x2 , . . . , xn ) dx1 dx2 · · · dxn .
(9.6)
R
The function f (x1 , x2 , . . . , xn ) is called the joint probability density function of X1 , X2 , . . . ,
Xn .
Let R = (x1 , x2 , . . . , xn ) : xi ∈ Ai , 1 ≤ i ≤ n , where Ai , 1 ≤ i ≤ n, is any subset
of real numbers that can be constructed from intervals by a countable number of set operations.
Then (9.6) gives
P (X1 ∈ A1 , X2 ∈ A2 , . . . , Xn ∈ An )
Z Z
Z
=
···
f (x1 , x2 , . . . , xn ) dx1 dx2 · · · dxn .
An
An−1
A1
Letting Ai = (−∞, +∞), 1 ≤ i ≤ n, this implies that
Z +∞
−∞
···
Z +∞
−∞
f (x1 , x2 , . . . , xn ) dx1 dx2 · · · dxn = 1.
Let fXi be the marginal probability density function of Xi , 1 ≤ i ≤ n. Then
fXi (xi ) =
Z +∞
|
−∞
···
{z
Z +∞
−∞
n−1 terms
}
f (x1 , x2 , . . . , xn ) dx1 · · · dxi−1 dxi+1 · · · dxn .
(9.7)
Therefore, for instance,
Z +∞
Z +∞
fX3 (x3 ) =
···
f (x1 , x2 , . . . , xn ) dx1 dx2 dx4 · · · dxn .
−∞
−∞
|
{z
}
n−1 terms
More generally, to find the joint probability density function marginalized over a given set
of k of these random variables, we integrate f (x1 , x2 , . . . , xn ) over all possible values of the
remaining n − k random variables. For example, if f (x, y, z, t) denotes the joint probability
density function of random variables X, Y, Z, and T, then
fY,T (y, t) =
Z +∞ Z +∞
−∞
f (x, y, z, t) dx dz
−∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 402 — #418
✐
✐
402
Chapter 9
Multivariate Distributions
is the joint probability density function marginalized over Y and T, whereas
Z +∞
fX,Z,T (x, z, t) =
f (x, y, z, t) dy
−∞
is the joint probability density function marginalized over X, Z, and T .
The following theorem is the generalization of Theorem 8.7 and the continuous analog of
Theorem 9.1. Its proof is similar to the proof of Theorem 8.7.
Theorem 9.3 Let X1 , X2 , . . . , Xn be jointly continuous random variables with the joint
probability density function f (x1 , x2 , . . . , xn ). Then X1 , X2 , . . . , Xn are independent if and
only if f (x1 , x2 , . . . , xn ) is the product of their marginal densities fX1 (x1 ), fX2 (x2 ), . . . ,
fXn (xn ).
Let F be the joint distribution function of jointly continuous random variables
X1 , X2 , . . . , Xn , with the joint probability density function f (x1 , x2 , . . . , xn ), then
Z t1
Z tn Z tn−1
f (x1 , x2 , . . . , xn ) dx1 · · · dxn ,
(9.8)
···
F (t1 , t2 , . . . , tn ) =
−∞
−∞
−∞
and
f (x1 , x2 , . . . , xn ) =
∂ n F (x1 , x2 , . . . , xn )
∂x1 ∂x2 · · · ∂xn
.
(9.9)
The following is the continuous analog of Theorem 9.2. It is also a generalization of Theorem 8.2 from dimension 2 to n.
Theorem 9.4 Let f (x1 , x2 , . . . , xn ) be the joint probability density function of random
variables X1 , X2 , . . . , Xn . If h is a function of n variables from Rn to R, then Y =
h(X1 , X2 , . . . , Xn ) is a random variable with expected value given by
Z ∞
Z ∞
E(Y ) =
···
h(x1 , x2 , . . . , xn )f (x1 , x2 , . . . , xn ) dx1 dx2 · · · dxn ,
−∞
−∞
provided that the integral is absolutely convergent.
Using Theorems 9.3 and 9.4, an almost identical proof to that of Theorem 8.6 implies that
The expected value of the product of several independent random variables is equal to the product of their expected values.
Example 9.4 A system has n components, whose lifetimes are exponential random variables with parameters λ1 , λ2 , . . . , λn , respectively. Suppose that the lifetimes of the components are independent random variables, and the system fails as soon as one of its components
fails. Find the probability density function and the expected value of the time until the system
fails.
Solution: Let X1 , X2 , . . . , Xn be the lifetimes of the n components, respectively. Then X1 ,
X2 , . . . , Xn are independent random variables and for i = 1 ,2, . . . , n,
P (Xi ≤ t) = 1 − e−λi t ,
t ≥ 0.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 403 — #419
✐
✐
Section 9.1
Joint Distributions of n > 2 Random Variables
403
Letting X be the time until the system fails, we have X = min(X1 , X2 , . . . , Xn ). Therefore,
P (X > t) = P min(X1 , X2 , . . . , Xn ) > t = P (X1 > t, X2 > t, . . . , Xn > t)
= P (X1 > t)P (X2 > t) · · · P (Xn > t)
= (e−λ1 t )(e−λ2 t ) · · · (e−λn t ) = e−(λ1 +λ2 +···+λn )t ,
t ≥ 0.
Let f be the probability density function of X; then
f (t) =
d
d
1 − e−(λ1 +λ2 +···+λn )t
P (X ≤ t) =
dt
dt
= (λ1 + λ2 + · · · + λn )e−(λ1 +λ2 +···+λn )t ,
t ≥ 0.
This shows that X is an exponential random variable with parameter λ1 +λ2 +· · ·+λn . Hence
E(X) =
1
. λ1 + λ2 + · · · + λn
Remark 9.2 In Example 9.4, we have shown the following important property of exponential
random variables.
If X1 , X2 , . . . , Xn are n independent exponential random variables with parameters λ1 , λ2 ,
. . . , and λn , respectively, then min(X1 , X2 , . . . , Xn ) is an exponential random variable with
parameter λ1 + λ2 + · · · + λn . Hence
E min(X1 , X2 , . . . , Xn ) =
1
.
λ1 + λ2 + · · · + λn
Example 9.5
(a)
(b)
Prove that the following is a joint probability density function.
1
if 0 < t ≤ z ≤ y ≤ x ≤ 1
f (x, y, z, t) = xyz
0
elsewhere.
Suppose that f is the joint probability density function of random variables X, Y, Z,
and T . Find fY,Z,T (y, z, t), fX,T (x, t), and fZ (z).
Solution:
(a)
Since f (x, y, z, t) ≥ 0 and
Z 1Z xZ y
Z 1Z xZ yZ z
1
1
dt dz dy dx =
dz dy dx
0
0
0
0 xyz
0
0
0 xy
Z 1Z x
Z 1
1
=
dy dx =
dx = 1,
0
0 x
0
f is a joint probability density function.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 404 — #420
✐
✐
404
Chapter 9
(b)
For 0 < t ≤ z ≤ y ≤ 1,
Multivariate Distributions
fY,Z,T (y, z, t) =
Z 1
y
1
1
ln y
1
dx =
ln x = −
.
xyz
yz
yz
y
Therefore,
ln y
−
yz
fY,Z,T (y, z, t) =
0
if 0 < t ≤ z ≤ y ≤ 1
elsewhere.
To find fX,T (x, t), we have that for 0 < t ≤ x ≤ 1,
Z xZ y
Z xh
1
ln z iy
fX,T (x, t) =
dz dy =
dy
xy t
t
t xyz
t
Z x
h 1
ix
ln y ln t ln t
=
−
dy =
(ln y)2 −
ln y
xy
xy
2x
x
t
t
1
1
1
1
(ln x)2 − (ln t)(ln x) +
(ln t)2 =
(ln x − ln t)2
2x
x
2x
2x
1 2x
=
ln .
2x
t
=
Therefore,
1
x
ln2
2x
t
fX,T (x, t) =
0
if 0 < t ≤ x ≤ 1
otherwise.
To find fZ (z), we have that for 0 < z ≤ 1,
Z 1Z x
1
1
dt dy dx =
dy dx
xyz
xy
z
z
0
z
z
Z 1h
Z 1
ix
1
1
1
=
ln y dx =
ln x − ln z dx
x
x
x
z
z
z
h1
i1
1
= (ln x)2 − (ln x)(ln z) = (ln z)2 .
2
2
z
fZ (z) =
Thus
Z 1Z xZ z
1
(ln z)2
fZ (z) = 2
0
if 0 < z ≤ 1
otherwise.
In Chapter 8, we showed how Definition 8.5 was justified. The following analog of that
definition in three dimensions is justified similarly.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 405 — #421
✐
✐
Joint Distributions of n > 2 Random Variables
Section 9.1
405
Definition 9.4 Let S be a subset of the three-dimensional Euclidean space with volume
Vol(S) 6= 0. A point is said to be randomly selected from S if for any subset Ω of S with
volume Vol(Ω), the probability that Ω contains the point is Vol(Ω)/Vol(S).
Let X, Y, and Z be the coordinates of the point randomly selected from S . Let f (x, y, z)
be the joint probability density function of X, Y, and Z . By an argument similar to the twodimensional case in Section 8.1, which resulted in (8.8), we have that
1
if (x, y, z) ∈ S
f (x, y, z) = Vol(S)
(9.10)
0
otherwise.
Conditional probability density functions and conditional probability mass functions in
higher dimensions are defined similarly to dimension two. For example, if X, Y, and Z are
three continuous random variables with the joint probability density function f, then
fX,Y |Z (x, y|z) =
f (x, y, z)
fZ (z)
at all points z for which fZ (z) > 0. As another example, let X, Y, Z, V, and W be five
continuous random variables with the joint probability density function f . Then
fX,Z,W |Y,V (x, z, w|y, v) =
f (x, y, z, v, w)
fY,V (y, v)
at all points (y, v) for which fY,V (y, v) > 0.
Example 9.6 Let X be a random point from the interval (0, 1), Y be a random point from
(0, X), and Z be a random point from (X, 1). Find f, the joint probability density function of
X, Y, and Z .
Solution: Note that
f (x, y, z) = fX (x) ·
fX,Y (x, y) f (x, y, z)
·
fX (x)
fX,Y (x, y)
= fX (x) · fY |X (y|x) · fZ|X,Y (z|x, y).
Clearly,
fX (x) =
fY |X (y|x) =
fZ|X,Y (z|x, y) =
1
0
0
0
1/x
0<x<1
otherwise.
0<y<x
otherwise.
1/(1 − x)
x<z<1
otherwise.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 406 — #422
✐
✐
406
Chapter 9
Multivariate Distributions
Thus
f (x, y, z) =
The relationship
1
x(1 − x)
0
0<y<x<z<1
otherwise.
f (x, y, z) = fX (x)fY |X (y|x)fZ|X,Y (z|x, y),
(9.11)
which was established in Example 9.6, can be generalized to n continuous random variables.
Let f be the joint probability density function of X1 , X2 , . . . , Xn . We have
f (x1 , x2 , . . . , xn ) = fX1 (x1 )fX2 |X1 (x2 |x1 )fX3 |X1 ,X2 (x3 |x1 , x2 )
· · · fXn |X1 ,X2 ,...,Xn−1 (xn |x1 , x2 , . . . , xn−1 ). (9.12)
Random Sample
Definition 9.5 We say that n random variables X1 , X2 , . . . , Xn form a random sample
of size n, from a (continuous or discrete) distribution function F, if they are independent and,
for 1 ≤ i ≤ n, the distribution function of Xi is F . Therefore, elements of a random sample
are independent and identically distributed.
To explain this definition, suppose that the lifetime distribution of the light bulbs manufactured by a company is exponential with parameter λ. To estimate 1/λ, the average lifetime of
a light bulb, for some positive integer n, we choose n light bulbs at random and independently
from those manufactured by the company. For 1 ≤ i ≤ n, let Xi be the lifetime of the ith
light bulb selected. Then {X1 , X2 , . . . , Xn } is a random sample of size n from the exponential distribution with parameter λ. That is, for 1 ≤ i ≤ n, Xi ’s are independent, and Xi is
exponential with parameter λ. Clearly, an estimation of 1/λ is the mean of the random sample
X1 , X2 , . . . , Xn denoted by X̄ :
X̄ =
X1 + X2 + · · · + Xn
n
.
Thus all we need to do is to measure the lifetime of each of the n light bulbs of the random
sample and find the average of the observed values. In Sections 11.3 and 11.5, we will discuss
methods to calculate n, the sample size, so that the error of estimation does not exceed a
predetermined quantity.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 407 — #423
✐
✐
Section 9.1
Joint Distributions of n > 2 Random Variables
407
EXERCISES
A
1.
From an ordinary deck of 52 cards, 13 cards are selected at random. Calculate the joint
probability mass function of the numbers of hearts, diamonds, clubs, and spades selected.
2.
A jury of 12 people is randomly selected from a group of eight Afro-American, seven
Hispanic, three Native American, and 20 white potential jurors. Let A, H, N, and W be
the number of Afro-American, Hispanic, Native American, and white jurors selected, respectively. Calculate the joint probability mass function of A, H, N, W and the marginal
probability mass function of A.
3.
Let p(x, y, z) = (xyz)/162, x = 4, 5, y = 1, 2, 3, and z = 1, 2, be the joint probability mass function of the random variables X, Y, Z .
4.
(a)
Calculate the joint marginal probability mass functions of X, Y ; Y, Z; and X,
Z.
(b)
Find E(Y Z).
Let the joint probability density function of X, Y, and Z be given by
(
6e−x−y−z
if 0 < x < y < z < ∞
f (x, y, z) =
0
elsewhere.
(a)
Find the marginal joint probability density function of X, Y ; X, Z; and Y, Z .
(b)
Find E(X).
5.
From the set of families with two children a family is selected at random. Let X1 = 1 if
the first child of the family is a girl; X2 = 1 if the second child of the family is a girl;
and X3 = 1 if the family has exactly one boy. For i = 1, 2, 3, let Xi = 0 in other cases.
Determine if X1 , X2 , and X3 are independent. Assume that in a family the probability
that a child is a girl is independent of the gender of the other children and is 1/2.
6.
Let X, Y, and Z be jointly continuous with the following joint probability density function:
(
x2 e−x(1+y+z)
if x, y, z > 0
f (x, y, z) =
0
otherwise.
Are X, Y, and Z independent? Are they pairwise independent?
7.
Let the joint distribution function of X, Y, and Z be given by
F (x, y, z) = (1 − e−λ1 x )(1 − e−λ2 y )(1 − e−λ3 z ), x, y, z > 0,
where λ1 , λ2 , λ3 > 0.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 408 — #424
✐
✐
408
Chapter 9
Multivariate Distributions
(a)
Are X, Y, and Z independent?
(b)
Find the joint probability density function of X, Y, and Z .
(c)
Find P (X < Y < Z).
8.
In a huge office building, the alarm system has n sensors, the lifetime of each being
exponential with mean 1/λ, independently of the other ones. If for the next t units of
time the alarm system is not checked, what is the probability that at time t only k sensors
are working? Assume that without inspection there is no way for the technical crew to
know that a sensor is down.
9.
(a)
(b)
Show that the following is a joint probability density function.
ln x
−
xy
f (x, y, z) =
0
if 0 < z ≤ y ≤ x ≤ 1
otherwise.
Suppose that f is the joint probability density function of X, Y, and Z . Find
fX,Y (x, y) and fY (y).
10.
Inside a circle of radius R, n points are selected at random and independently. Find the
probability that the distance of the nearest point to the center is at least r .
11.
A point is selected at random from the cube
Ω = (x, y, z) : − a ≤ x ≤ a, −a ≤ y ≤ a, −a ≤ z ≤ a .
What is the probability that it is inside the sphere inscribed in the cube?
12.
Is the following a joint probability density function?
(
e−xn
if 0 < x1 < x2 < · · · < xn
f (x1 , x2 , . . . , xn ) =
0
otherwise.
13.
Suppose that the lifetimes of radio transistors are independent exponential random variables with mean five years. Arnold buys a radio and decides to replace its transistor upon
failure two times: once when the original transistor dies and once when the replacement
dies. He stops using the radio when the second replacement of the transistor goes out
of order. Assuming that Arnold repairs the radio if it fails for any other reason, find the
probability that he uses the radio at least 15 years.
14.
Let X1 , X2 , . . . , Xn be identically distributed, independent, exponential random variables with parameters λ1 , λ2 , . . . , λn . Prove that
E min(X1 , . . . , Xn ) < min E(X1 ), . . . , E(Xn ) .
15.
(Reliability of Systems) Suppose that a system functions if and only if at least
k (1 ≤ k ≤ n) of its components function. Furthermore, suppose that pi = p for
1 ≤ i ≤ n. Find the reliability of this system. (Such a system is said to be a k-out-of-n
system.)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 409 — #425
✐
✐
Joint Distributions of n > 2 Random Variables
Section 9.1
409
B
16.
An item has n parts, each with an exponentially distributed lifetime with mean 1/λ. If
the failure of one part makes the item fail, what is the average lifetime of the item?
Hint: Use the result of Example 9.4.
17.
Suppose that the lifetimes of a certain brand of transistor are identically distributed and
independent random variables with distribution function F . These transistors are randomly selected, one at a time, and their lifetimes are measured. Let the N th be the first
transistor that will last longer than s hours. Let XN be the lifetime of this transistor. Are
N and XN independent random variables?
18.
(Reliability of Systems) Consider the system whose structure is shown in
Figure 9.4. Find the reliability of this system.
2
4
1
7
5
3
Figure 9.4
19.
6
A diagram for the system of Exercise 18.
(Reliability of Systems) To transfer water from point A to point B, a water-supply
system with five water pumps located at the points 1, 2, 3, 4, and 5 is designed as in
Figure 9.5. Suppose that whenever the system is turned on for water to flow from A to
B, pump i, i ≤ 5, functions with probability pi independent of the other pumps. What
is the probability that, at such a time, water reaches B ?
B
5
4
3
2
1
A
Figure 9.5
The water-supply system of Exercise 19.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 410 — #426
✐
✐
410
Chapter 9
20.
Let F be a distribution function. Prove that the functions F n and
1 − (1 − F )n are also distribution functions.
Hint:
Let X1 , X2 , . . . , Xn be independent random variables each with the distribution function F . Find the distribution functions of the random variables
max(X1 , X2 , . . . , Xn ) and min(X1 , X2 , . . . , Xn ).
21.
Let X1 , X2 , . . . , Xn be n independent
random numbers from the interval (0, 1). Find
E max Xi and E min Xi .
1≤i≤n
22.
Multivariate Distributions
1≤i≤n
Let X1 , X2 , . . . , Xn be n independent random numbers from (0, 1), and
Yn = n · min(X1 , X2 , . . . , Xn ).
Prove that
lim P (Yn > x) = e−x ,
n→∞
23.
x ≥ 0.
Suppose that h is the probability density function of a continuous random variable. Let
the joint probability density function of X, Y, and Z be
f (x, y, z) = h(x)h(y)h(z),
x, y, z ∈ R.
Prove that P (X < Y < Z) = 1/6.
24.
A point is selected at random from the pyramid
V = (x, y, z) : x, y, z ≥ 0, x + y + z ≤ 1 .
Letting (X, Y, Z) be its coordinates, determine if X, Y, and Z are independent.
Hint: Recall that the volume of a pyramid is Bh/3, where h is the height and B is
the area of the base.
25.
(Roots of Quadratic Equations) Three numbers A, B, and C are selected at
random and independently from the interval (0, 1). Determine the probability that the
quadratic equation Ax2 + Bx + C = 0 has real roots. In other words, what fraction of
“all possible quadratic equations” with coefficients in (0, 1) have real roots?
26.
(Roots of Cubic Equations) Solve the following exercise posed by S. A. Patil
and D. S. Hawkins, Tennessee Technological University, Cookeville, Tennessee, in The
College Mathematics Journal, September 1992.
Let A, B, and C be independent random variables uniformly distributed
on [0, 1]. What is the probability that all of the roots of the cubic equation
x3 + Ax2 + Bx + C = 0 are real?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 411 — #427
✐
✐
Section 9.2
Order Statistics
411
Self-Quiz on Section 9.1
Time allotted: 20 Minutes
Each problem is worth 5 points.
1.
In the inventory of a pharmacy, in a carton, there are 40 boxes of painkillers of which
15 are brand A, 10 are brand B, 6 are brand C, 4 are brand D, and 5 are brand E. An
assistant pharmacist chooses 20 of these boxes randomly to move them inside the store
for the over-the-counter sale. Let X, Y, Z, U, and V be the respective numbers of the
boxes of the aforementioned brands selected. Find the joint probability mass function of
X, Y, Z, U, and V, and the marginal probability mass function of U and V .
2.
Let X, Y, and Z be continuous random variables with the joint probability density function given by
2y e−(2x+y+z)
x > 0, y > 0, z > 0
f (x, y, z) =
0
otherwise.
Find P (X < Y < Z).
9.2
ORDER STATISTICS
Definition 9.6
Let X1 , X2 , . . . , Xn be an independent set of identically distributed
continuous random variables with the common probability
density and distribution functions f
and F, respectively. Let X(1) be the smallest value in X1 , X2 , . . . , Xn , X(2) be the second
smallest value, X(3) be the third smallest, and, in general, X(k) (1 ≤ k ≤ n) be the k th
smallest value in X1 , X2 , . . . , Xn . Then X(k) is called the kth order statistic, and the set
X(1) , X(2) , . . . , X(n) is said to consist of the order statistics of X1 , X2 , . . . , Xn .
By this definition, for example, if at a sample point ω of the sample space, X1 (ω) = 8,
X2 (ω) = 2, X3 (ω) = 5, and X4 (ω) = 6, then the order statistics of {X1 , X2 , X3 , X4 } is
{X(1) , X(2) , X(3) , X(4) }, where X(1) (ω) = 2, X(2) (ω) = 5, X(3) (ω) = 6, and X(4) (ω) = 8.
Continuity of Xi ’s implies that P (X(i) = X(j) ) = 0. Hence
P (X(1) < X(2) < X(3) < · · · < X(n) ) = 1.
Unlike Xi ’s, the random variables X(i) ’s are neither independent nor identically distributed.
There are many useful and practical applications of order statistics in different branches of
pure and applied probability, as well as in estimation theory. To show how it arises, we present
three examples.
Example 9.7 Suppose that customers arrive at a warehouse from n different locations. Let
Xi , 1 ≤ i ≤ n, be the time until the arrival of the next customer from location i; then X(1) is
the arrival time of the next customer to the warehouse. ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 412 — #428
✐
✐
412
Chapter 9
Multivariate Distributions
Example 9.8 Suppose that a machine consists of n components with the lifetimes X1 , X2 ,
. . . , Xn , respectively, where Xi ’s are independent and identically distributed. Suppose that the
machine remains operative unless k or more of its components fail. Then X(k) , the k th order
statistic of {X1 , X2 , . . . , Xn }, is the time when the machine fails. Also, X(1) is the failure
time of the first component. Example 9.9 Let X1 , X2 , . . . , Xn be a random sample of size n from a population with
continuous distribution F . Then the following important statistical concepts are expressed in
terms of order statistics:
(i)
(ii)
(iii)
The sample range is X(n) − X(1) .
The sample midrange is [X(n) + X(1) ] 2.
The sample median is
m=
X(i+1)
X(i) + X(i+1)
2
if n = 2i + 1
if n = 2i.
We will now determine the distribution and the probability density functions of X(k) , the k th
order statistic.
Theorem 9.5 Let {X(1) , X(2) , . . . , X(n) } be the order statistics of the independent and
identically distributed continuous random variables X1 , X2 , . . . , Xn with the common distribution and probability density functions F and f, respectively. Then Fk and fk , the distribution
and probability density functions of X(k) , respectively, are given by
!
n
X
i n−i
n Fk (x) =
F (x) 1 − F (x)
,
i
i=k
−∞ < x < ∞,
(9.13)
and
fk (x) =
k−1 n−k
n!
f (x) F (x)
1 − F (x)
,
(k − 1)! (n − k)!
−∞ < x < ∞.
(9.14)
Proof: Let −∞ < x < ∞. To calculate P (X(k) ≤ x), note that X(k) ≤ x if and only if at
least k of the random variables X1 , X2 , . . . , Xn are in (−∞, x]. Thus
Fk (x) = P (X(k) ≤ x)
n
X
=
P i of the random variables X1 , X2 , . . . , Xn are in (−∞, x]
i=k
!
n
X
i n−i
n =
F (x) 1 − F (x)
,
i
i=k
where the last equality follows because from the random variables X1 , X2 , . . . , Xn the number
of those that lie in (−∞, x] has binomial distribution with parameters (n, p), p = F (x).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 413 — #429
✐
✐
Section 9.2
Order Statistics
413
We will now obtain fk by differentiating Fk :
!
n
X
i−1 n−i
n
fk (x) =
if (x) F (x)
1 − F (x)
i
i=k
!
n
X
i
n−i−1
n F (x) (n − i)f (x) 1 − F (x)
−
i
i=k
=
n
X
i−1 n−i
n!
f (x) F (x)
1 − F (x)
(i − 1)! (n − i)!
i=k
−
=
n
X
i n−i−1
n!
f (x) F (x) 1 − F (x)
i! (n − i − 1)!
i=k
n
X
i−1 n−i
n!
f (x) F (x)
1 − F (x)
(i − 1)! (n − i)!
i=k
−
n
X
i−1 n−i
n!
f (x) F (x)
1 − F (x)
.
(i − 1)! (n − i)!
i=k+1
After cancellations, this gives (9.14).
Remark 9.3 Note that by (9.13) and (9.14), respectively, F1 and f1 , the distribution and the
probability density functions of X(1) = min(X1 , X2 , . . . , Xn ), are found to be
!
n
X
i n−i
n F1 (x) =
F (x) 1 − F (x)
i
i=1
!
n
X
i n−i n
n =
F (x) 1 − F (x)
− 1 − F (x)
i
i=0
n n
= F (x) + 1 − F (x)
− 1 − F (x)
n
= 1 − 1 − F (x) ,
−∞ < x < ∞.
and
n−1
f1 (x) = nf (x) 1 − F (x)
,
−∞ < x < ∞.
n−1
fn (x) = nf (x) F (x)
,
−∞ < x < ∞.
Also, Fn and fn , the distribution and probability density functions of X(n)
max(X1 , X2 , . . . , Xn ), respectively, are found to be
n
Fn (x) = F (x) ,
−∞ < x < ∞,
=
and
These quantities can also be calculated directly (see, for example, Example 9.4).
Example 9.10 Let X1 , X2 , . . . , X2n+1 be 2n + 1 random numbers from (0, 1). Then f and
F, the respective probability density and distribution functions of Xi ’s, are given by
(
1
if 0 < x < 1
f (x) =
0
elsewhere,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 414 — #430
✐
✐
414
Chapter 9
Multivariate Distributions
0
F (x) = x
1
if x < 0
if 0 ≤ x < 1
if x ≥ 1.
Using (9.12), the probability density function of X(n+1) , the median of these numbers, is found
to be
(2n + 1)! n
x (1 − x)n ,
0 < x < 1;
fn+1 (x) =
n! n!
0, elsewhere. Since for the beta function, B(n + 1, n + 1) = (n! n!)/(2n + 1)!, we have
fn+1 (x) =
1
xn (1 − x)n ,
B(n + 1, n + 1)
0 < x < 1;
0, elsewhere. Hence X(n+1) is beta with parameters (n + 1, n + 1). In fact, it is straightforward
to see that in this example, for 1 ≤ k ≤ 2n + 1, X(k) is beta with parameters k and 2n − k + 2.
The following theorem gives the joint probability density function of X(i) and X(j) . Its
proof is similar to that of Theorem 9.5.
Theorem 9.6 Let {X(1) , X(2) , . . . , X(n) } be the order statistics of the independent and
identically distributed continuous random variables X1 , X2 , . . . , Xn with the common probability density and distribution functions f and F, respectively. Then, for i < j and x < y,
fij (x, y), the joint probability density function of X(i) and X(j) , is given by
fij (x, y) =
i−1 j−i−1 n−j
n!
f (x)f (y) F (x)
F (y) − F (x)
1 − F (y)
.
(i − 1)! (j − i − 1)! (n − j)!
It is clear that for x ≥ y, fij (x, y) = 0.
Another important quantity that is used often is the joint probability density function of
X(1) , X(2) , . . . , X(n) . The following theorem, whose proof we skip, expresses the general form
of this function.
Theorem 9.7 Let {X(1) , X(2) , . . . , X(n) } be the order statistics of the independent and
identically distributed continuous random variables X1 , X2 , . . . , Xn with the common probability density and distribution functions f and F, respectively. Then f12···n , the joint probability
density function of X(1) , X(2) , . . . , X(n) , is given by
f12···n (x1 ,x2 , . . . , xn ) =
(
n!f (x1 )f (x2 ) · · · f (xn )
0
Justification:
−∞ < x1 < x2 < · · · < xn < ∞
otherwise.
For small ε > 0,
P (x1 − ε < X(1) < x1 + ε, . . . , xn − ε < X(n) < xn + ε)
Z xn +ε Z xn−1 +ε
Z x1 +ε
=
···
f12···n (x1 , x2 , . . . , xn )dx1 dx2 · · · dxn
xn −ε
xn−1 −ε
x1 −ε
≈ 2n εn f12···n (x1 , x2 , . . . , xn ).
(9.15)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 415 — #431
✐
✐
Section 9.2
415
Order Statistics
Now for x1 < x2 < · · · < xn , let P be the set of all permutations of {x1 , x2 , · · · , xn }; then
P has n! elements and we can write
P (x1 − ε < X(1) < x1 + ε, . . . , xn − ε < X(n) < xn + ε)
X
=
P (xi1 − ε < X1 < xi1 + ε, . . . , xin − ε < Xn < xin + ε)
{xi1 ,xi2 ,...,xin }∈P
≈
X
{xi1 ,xi2 ,...,xin }∈P
2n εn f (xi1 )f (xi2 ) · · · f (xin ).
(9.16)
This is because X1 , X2 , . . . , Xn are independent, and hence their joint probability density
function is the product of their marginal probability density functions. Putting (9.15) and (9.16)
together, we obtain
X
{xi1 ,xi2 ,...,xin }∈P
f (xi1 )f (xi2 ) · · · f (xin ) = f12···n (x1 , x2 , . . . , xn ).
(9.17)
But
f (xi1 )f (xi2 ) · · · f (xin ) = f (x1 )f (x2 ) · · · f (xn ).
Therefore,
X
{xi1 ,xi2 ,...,xin }∈P
f (xi1 )f (xi2 ) · · · f (xin ) =
X
{xi1 ,xi2 ,...,xin }∈P
f (x1 )f (x2 ) · · · f (xn ) = n!f (x1 )f (x2 ) · · · f (xn ). (9.18)
Relations (9.17) and (9.18) imply that
f12···n (x1 , x2 , . . . , xn ) = n!f (x1 )f (x2 ) · · · f (xn ). Example 9.11 The distance between two towns, A and B, is 30 miles. If three gas stations
are constructed independently at randomly selected locations between A and B, what is the
probability that the distance between any two gas stations is at least 10 miles?
Solution: Let X1 , X2 , and X3 be the locations at which the gas stations are constructed. The
probability density function of X1 , X2 , and X3 is given by
1
if 0 < x < 30
f (x) = 30
0
elsewhere.
Therefore, by Theorem 9.7, f123 , the joint probability density function of the order statistics of
X1 , X2 , and X3 , is as follows.
1 3
f123 (x1 , x2 , x3 ) = 3!
,
30
0 < x1 < x2 < x3 < 30.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 416 — #432
✐
✐
416
Chapter 9
Multivariate Distributions
Using this, we have that the desired probability is given by the following triple integral.
P (X(1) + 10 < X(2) and X(2) + 10 < X(3) )
Z 10 Z 20 Z 30
1
=
f123 (x1 , x2 , x3 ) dx3 dx2 dx1 =
27
0
x1 +10 x2 +10
EXERCISES
A
1.
Let X1 , X2 , X3 , and X4 be four independently selected random numbers from (0, 1).
Find P (1/4 < X(3) < 1/2).
2.
Two random points are selected from (0, 1) independently. Find the probability that one
of them is at least three times the other.
3.
A box contains 20 identical balls numbered 1 to 20. Seven balls are drawn randomly and
without replacement. Find the probability mass function of the median of the numbers
on the balls drawn.
4.
Let X1 , X2 , X3 , and X4 be independent exponential random variables, each with
parameter λ. Find P (X(4) ≥ 3λ).
5.
Let X1 , X2 , X3 , . . . , Xn be a sequence of nonnegative, identically distributed, and
independent random variables. Let F be the distribution function of Xi , 1 ≤ i ≤ n.
Prove that
Z ∞
n E[X(n) ] =
1 − F (x)
dx.
0
Hint:
6.
7.
Use Remark 6.4.
Let X1 , X2 , X3 , . . . , Xm be a sequence of nonnegative, independent binomial random variables, each with parameters (n, p). Find the probability mass function of X(i) ,
1 ≤ i ≤ m.
Prove that G, the distribution function of [X(1) + X(n) ] 2, the midrange of a random
sample of size n from a population with continuous distribution function F and probability density function f, is given by
G(t) = n
Z t
−∞
n−1
F (2t − x) − F (x)
f (x) dx.
Hint: Use Theorem 9.6 to find f1n ; then integrate over the region x + y ≤ 2t and
x ≤ y.
B
8.
Let X1 and X2 be two independent exponential random variables each with parameter
λ. Show that X(1) and X(2) − X(1) are independent.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 417 — #433
✐
✐
Section 9.3
Multinomial Distributions
417
9.
Let X1 and X2 be two independent N (0, σ 2 ) random variables. Find E[X(1) ].
Hint: Let f12 (x,RR
y) be the joint probability density function of X(1) and X(2) . The
desired quantity is
xf12 (x, y) dx dy, where the integration is taken over an
appropriate region.
10.
Let X1 , X2 , . . . , Xn be a random sample of size n from a population with continuous
distribution function F and probability density function f .
(a)
(b)
11.
Calculate the probability density function of the sample range, R
X(n) − X(1) .
=
Use (a) to find the probability density function of the sample range of n random
numbers from (0, 1).
Let X1 , X2 , . . . , Xn be n independently randomly selected points from the interval
(0, θ), θ > 0. Prove that
n−1
θ,
E(R) =
n+1
where R = X(n) − X(1) is the range of these points.
Hint: Use part (a) of Exercise 10. Also compare this with Exercise 21, Section 9.1.
Self-Quiz on Section 9.2
Time allotted: 20 Minutes
Each problem is worth 5 points.
1.
Find the expected value of the distance between two random points selected independently from the interval (0, 1).
2.
All that we know about a horse race that was held last week in Louisville, Kentucky, is
that all of the five horses that were competing reached the finish line, independently, at
random times after 2:10 P. M . and before 2:22 P. M .
9.3
(a)
Find the probability that the winning horse passed the finish line between
2:10 P. M . and 2:11 P. M .
(b)
Find the probability that the time that the last horse crossed the finish line was
after 2:21 P. M .
MULTINOMIAL DISTRIBUTIONS
Multinomial distribution is a generalization of a binomial. Suppose that, whenever an experiment is performed, one of the disjoint outcomes A1 , A2 , . . . , Ar will occur. Let P (Ai ) = pi ,
1 ≤ i ≤ r. Then p1 + p2 + · · · + pr = 1. If, in n independent performances of this experiment,
Xi , i = 1, 2, 3, . . . , r, denotes the number of times that Ai occurs, then p(x1 , . . . , xr ), the
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 418 — #434
✐
✐
418
Chapter 9
Multivariate Distributions
joint probability mass function of X1 , X2 , . . . , Xr , is called multinomial joint probability
mass function, and its distribution is said to be a multinomial distribution. For any set of
nonnegative integers {x1 , x2 , . . . , xr } with x1 + x2 + · · · + xr = n,
p(x1 , x2 , . . . , xr ) = P (X1 = x1 , X2 = x2 , . . . , Xr = xr )
n!
px1 px2 · · · pxr r .
=
x1 ! x2 ! · · · xr ! 1 2
(9.19)
To prove this relation, recall that, by Theorem 2.4, the number of distinguishable permutations of n objects of r different types where x1 are alike, x2 are alike, · · · , xr are alike
(n = x1 + · · · + xr ) is n!/(x1 ! x2 ! · · · xr !). Hence there are n!/(x1 ! x2 ! · · · xr !) sequences of A1 , A2 , . . . , Ar in which the number of Ai ’s is xi , i = 1, 2, . . . , r . Relation (9.19)
follows since the probability of the occurrence of any of these sequences is px1 1 px2 2 · · · pxr r .
By Theorem 2.6, p(x1 , x2 , . . . , xr )’s given by (9.19) are the terms in the expansion of
(p1 + p2 + · · · + pr )n . For this reason, the multinomial distribution sometimes is called the
polynomial distribution. The following relation guarantees that p(x1 , x2 , . . . , xr ) is a joint
probability mass function:
X
p(x1 , x2 , . . . , xr )
x1 +x2 +···+xr =n
=
X
n!
px1 1 px2 2 · · · pxr r = (p1 + p2 + · · · + pr )n = 1.
x
!
x
!
·
·
·
x
!
1
2
r
x1 +x2 +···+xr =n
Note that, for r = 2, the multinomial distribution coincides with the binomial distribution.
This is because, for r = 2, the experiment has only two possible outcomes, A1 and A2 . Hence
it is a Bernoulli trial.
Example 9.12 In a certain town, at 8:00 P.M., 30% of the TV viewing audience watch the
news, 25% watch a certain comedy, and the rest watch other programs. What is the probability
that, in a statistical survey of seven randomly selected viewers, exactly three watch the news
and at least two watch the comedy?
Solution: In the random sample, let X1 , X2 , and X3 be the numbers of viewers who watch the
news, the comedy, and other programs, respectively. Then the joint distribution of X1 , X2 , and
X3 is multinomial with p1 = 0.30, p2 = 0.25, and p3 = 0.45. Therefore, for i + j + k = 7,
P (X1 = i, X2 = j, X3 = k) =
7!
(0.30)i (0.25)j (0.45)k .
i! j! k!
The desired probability equals
P (X1 = 3, X2 ≥ 2) = P (X1 = 3, X2 = 2, X3 = 2)
+ P (X1 = 3, X2 = 3, X3 = 1) + P (X1 = 3, X2 = 4, X3 = 0)
7!
7!
=
(0.30)3 (0.25)2 (0.45)2 +
(0.30)3 (0.25)3 (0.45)1
3! 2! 2!
3! 3! 1!
7!
+
(0.30)3 (0.25)4 (0.45)0 ≈ 0.103. 3! 4! 0!
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 419 — #435
✐
✐
Section 9.3
Multinomial Distributions
419
Example 9.13 A warehouse contains 500 TV sets, of which 25 are defective, 300 are in
working condition but used, and the rest are brand new. What is the probability that, in a random
sample of five TV sets from this warehouse, there are exactly one defective and exactly two
brand new sets?
Solution: In the random sample, let X1 be the number of defective TV sets, X2 be the number
of TV sets in working condition but used, and X3 be the number of brand new sets. The desired
probability is given by
!
!
!
25
300
175
1
2
2
!
≈ 0.067.
P (X1 = 1, X2 = 2, X3 = 2) =
500
5
Note that, since the number of sets is large and the sample size is small, the joint distribution of
X1 , X2 , and X3 is approximately multinomial with p1 = 25/500 = 1/20, p2 = 300/500 =
3/5, and p3 = 175/500 = 7/20. Hence
P (X1 = 1, X2 = 2, X3 = 2) ≈
5! 1 1 3 2 7 2
≈ 0.066. 1! 2! 2! 20
5
20
Example 9.14 (Marginals of Multinomials) Let X1 , X2 , . . . , Xr (r ≥ 4) have the
joint multinomial probability mass function p(x1 , x2 , . . . , xr ) with parameters n and p1 , p2 ,
. . . , pr . Find the marginal probability mass functions pX1 and pX1 ,X2 ,X3 .
Solution: From the definition of marginal probability mass functions,
X
n!
px1 1 px2 2 · · · pxr r
x
!
x
!
·
·
·
x
!
2
r
x2 +x3 +···+xr =n−x1 1
X
(n − x1 )!
n!
=
px1 1
px2 2 px3 3 · · · pxr r
x1 ! (n − x1 )!
x
!
x
!
·
·
·
x
!
2
3
r
x2 +x3 +···+xr =n−x1
pX1 (x1 ) =
n!
px1 (p2 + p3 + · · · + pr )n−x1
x1 ! (n − x1 )! 1
n!
px1 (1 − p1 )n−x1 ,
=
x1 ! (n − x1 )! 1
=
where the next-to-last equality follows from multinomial expansion (Theorem 2.6) and the last
inequality follows from p1 + p2 + · · · + pr = 1. This shows that the marginal probability
mass function of X1 is binomial with parameters n and p1 . To find the joint probability mass
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 420 — #436
✐
✐
420
Chapter 9
Multivariate Distributions
function marginalized over X1 , X2 , and X3 , let k = n − x1 − x2 − x3 ; then
pX1 ,X2 ,X3 (x1 , x2 , x3 )
X
n!
px1 1 px2 2 px3 3 px4 4 · · · pxr r
=
x
!
x
!
x
!
x
!
·
·
·
x
!
2
3
4
r
x4 +···+xr =k 1
X
n!
k!
px 1 px 2 px 3
px4 · · · pxr r
=
x1 ! x2 ! x3 ! k! 1 2 3 x +···+x =k x4 ! · · · xr ! 4
4
r
n!
px1 px2 px3 (p4 + · · · + pr )k
=
x1 ! x2 ! x3 ! k! 1 2 3
n!
=
px1 px2 px3 (1 − p1 − p2 − p3 )n−x1 −x2 −x3 .
x1 ! x2 ! x3 ! (n − x1 − x2 − x3 )! 1 2 3
This shows that the joint probability mass function marginalized over X1 , X2 , and X3 is
multinomial with parameters n and p1 , p2 , p3 , 1 − p1 − p2 − p3 . Remark 9.4
The method of Example 9.14 can be extended to prove the following theorem:
Let the joint distribution of the random variables X1 , X2 , . . . , Xr be multinomial with parameters n and p1 , p2 , . . . , pr . The joint probability mass
function marginalized over a subset Xi1 , Xi2 , . . . , Xik of k (k > 1) of
these r random variables is multinomial with parameters n and pi1 , pi2 ,
. . . , pik , 1 − pi1 − pi2 − · · · − pik . EXERCISES
A
1.
Light bulbs manufactured by a certain factory last a random time between 400 and 1200
hours. What is the probability that, of eight such bulbs, three burn out before 550 hours,
two burn out after 800 hours, and three burn out after 550 but before 800 hours?
2.
An urn contains 100 chips of which 20 are blue, 30 are red, and 50 are green. We draw
20 chips at random and with replacement. Let B, R, and G be the number of blue, red,
and green chips, respectively. Calculate the joint probability mass function of B, R, and
G.
3.
Suppose that each day the price of a stock moves up 1/8 of a point with probability 1/4,
remains the same with probability 1/3, and moves down 1/8 of a point with probability
5/12. If the price fluctuations from one day to another are independent, what is the
probability that after six days the stock has its original price?
4.
At a certain college, 16% of the calculus students get A’s, 34% B’s, 34% C’s, 14% D’s,
and 2% F’s. What is the probability that, of 15 calculus students selected at random, five
get B’s, five C’s, two D’s, and at least two A’s?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 421 — #437
✐
✐
Section 9.3
Multinomial Distributions
421
5.
Of the drivers who are insured by a certain insurance company and get into at least one
accident during a random year, 15% are low-risk drivers, 35% are moderate-risk drivers,
and 50% are high-risk drivers. Suppose that drivers insured by this company get into
an accident independently of each other. If an actuary randomly chooses five drivers
insured by this company, who got into an accident one or more times last year, what is
the probability that there are at least two more high-risk drivers among them than the
low-risk and moderate-risk drivers combined?
6.
Suppose that 50% of the watermelons grown on a farm are classified as large, 30% as
medium, and 20% as small. Joanna buys five watermelons at random from this farm.
What is the probability that (a) at least two of them are large; (b) two of them are large,
two are medium, and one is small; (c) exactly two of them are medium if it is given that
at least two of them are large?
7.
Suppose that the ages of 30% of the teachers of a country are over 50, 20% are between
40 and 50, and 50% are below 40. In a random committee of 10 teachers from this
country, two are above 50. What is the probability mass function of those who are below
40?
8.
Customers enter a department store at the rate of three per minute, in accordance with a
Poisson process. If 30% of them buy nothing, 20% pay cash, 40% use charge cards, and
10% write personal checks, what is the probability that in five operating minutes of the
store, five customers use charge cards, two write personal checks, and three pay cash?
Self-Quiz on Section 9.3
Time allotted: 20 Minutes
Each problem is worth 5 points.
1.
Suppose that 40% of the students joining the mathematics department of a certain university major in actuarial science, 35% major in statistics and operation research, and
25% major in pure mathematics. There are 30 new students entering this department
next fall. What is the probability that exactly 12 of them will major in actuarial science
and either 7 or 8 students will major in pure mathematics?
2.
Of the emails that arrive in Samantha’s inbox, 45% are personal, 40% are work-related,
and 15% are unsolicited commercial emails. Samantha logs into her email account and
finds that she has 25 new emails. What is the probability that exactly 10 of them are
work-related and at most 4 of them are unsolicited commercial emails?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 422 — #438
✐
✐
422
Chapter 9
Multivariate Distributions
CHAPTER 9 SUMMARY
◮ Independence of n > 2 Random Variables Suppose that X1 , X2 , . . . , Xn are
random variables (discrete, continuous, or mixed) on a sample space. We say that they are
independent if, for arbitrary subsets A1 , A2 , . . . , An of real numbers,
P (X1 ∈ A1 , X2 ∈ A2 , . . . , Xn ∈ An ) = P (X1 ∈ A1 )P (X2 ∈ A2 ) · · · P (Xn ∈ An ).
◮ We define the joint distribution function of X1 , X2 , . . . , Xn by
F (t1 , t2 , . . . , tn ) = P (X1 ≤ t1 , X2 ≤ t2 , . . . , Xn ≤ tn )
for all −∞ < ti < +∞, i = 1, 2, . . . , n. The marginal distribution function of Xi , 1 ≤ i ≤
n, can be found from F as follows:
FXi (ti ) =
lim
tj →∞
1≤j≤n, j6=i
F (t1 , t2 , . . . , tn ).
It follows that X1 , X2 , . . . , Xn are independent if and only if
F (x1 , x2 , . . . , xn ) = FX1 (x1 )FX2 (x2 ) · · · FXn (xn ).
◮ If {X1 , X2 , . . .} is a sequence of independent random variables and for i = 1, 2, . . . ,
gi : R → R is a real-valued function, then the sequence {g1 (X1 ), g2 (X2 ), . . .} is also an
independent sequence of random variables.
◮ Joint Probability Mass Functions Let X1 , X2 , . . . , Xn be discrete random variables
defined on the same sample space, with sets of possible values A1 , A2 , . . . , An , respectively.
The function
p(x1 , x2 , . . . , xn ) = P (X1 = x1 , X2 = x2 , . . . , Xn = xn )
is called the joint probability mass function of X1 , X2 , . . . , Xn . Clearly,
(a)
(b)
(c)
p(x1 , x2 , . . . , xn ) ≥ 0.
If for some i, 1 ≤ i ≤ n, xi 6∈ Ai , then p(x1 , x2 , . . . , xn ) = 0.
P
xi ∈Ai , 1≤i≤n p(x1 , x2 , . . . , xn ) = 1.
Moreover,
for 1 ≤ i ≤ n, the marginal probability mass function of Xi , pXi , is pXi (xi ) =
P
xj ∈Aj , j6=i p(x1 , x2 , . . . , xn ). More generally, to find the joint probability mass function
marginalized over a given set of k of these random variables, we sum up p(x1 , x2 , . . . , xn )
over all possible values of the remaining n − k random variables. The discrete random variables X1 , X2 , . . . , Xn are independent if and only if p(x1 , x2 , . . . , xn ) is the product of their
marginal probability mass functions pX1 (x1 ), pX2 (x2 ), . . . , pXn (xn ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 423 — #439
✐
✐
Chapter 9
Summary
423
◮ For 1 ≤ i ≤ n, let Ai be the set of possible values of Xi . If h is a function of n variables
from Rn to R, then Y = h(X1 , X2 , . . . , Xn ) is a discrete random variable with expected
value given by
X
X
h(x1 , x2 , . . . , xn )p(x1 , x2 , . . . , xn ),
···
E(Y ) =
x1 ∈A1
xn ∈An
provided that the sum is finite. This fact can be used to prove that the expected value of the
product of several independent discrete random variables is equal to the product of their expected values.
◮ Joint Probability Density Functions
Let X1 , X2 , . . . , Xn be continuous random
variables defined on the same sample space. We say that X1 , X2 , . . . , Xn have a continuous
joint distribution if there exists a nonnegative function of n variables, f (x1 , x2 , . . . , xn ), on
Rn such that for any region R in Rn that can be formed from n-dimensional rectangles by a
countable number of set operations,
Z
Z
P (X1 , X2 , . . . , Xn ) ∈ R =
· · · f (x1 , x2 , . . . , xn ) dx1 dx2 · · · dxn .
R
The function f(x1 , x2 , . . . , xn ) is called the joint probability density function of X1 , X2 , . . . ,
Xn . Let R = (x1 , x2 , . . . , xn ) : xi ∈ Ai , 1 ≤ i ≤ n , where Ai , 1 ≤ i ≤ n, is any subset
of real numbers that can be constructed from intervals by a countable number of set operations.
Then
P (X1 ∈ A1 , X2 ∈ A2 , . . . , Xn ∈ An )
Z Z
Z
=
···
f (x1 , x2 , . . . , xn ) dx1 dx2 · · · dxn .
An
An−1
A1
Let fXi be the marginal probability density function of Xi , 1 ≤ i ≤ n. Then
Z +∞
Z +∞
f (x1 , x2 , . . . , xn ) dx1 · · · dxi−1 dxi+1 · · · dxn .
···
fXi (xi ) =
| −∞ {z −∞ }
n−1 terms
In general, to find the joint probability density function marginalized over a given set of k of
these random variables, we integrate f (x1 , x2 , . . . , xn ) over all possible values of the remaining n−k random variables. The continuous random variables X1 , X2 , . . . , Xn are independent
if and only if f (x1 , x2 , . . . , xn ) is the product of their marginal densities fX1 (x1 ), fX2 (x2 ),
. . . , fXn (xn ).
◮ If h is a function of n variables from Rn to R, then Y = h(X1 , X2 , . . . , Xn ) is a random
variable with expected value given by
Z ∞
Z ∞
E(Y ) =
···
h(x1 , x2 , . . . , xn )f (x1 , x2 , . . . , xn ) dx1 dx2 · · · dxn ,
−∞
−∞
provided that the integral is absolutely convergent. This fact can be used to prove that the
expected value of the product of several independent continuous random variables is equal to
the product of their expected values.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 424 — #440
✐
✐
424
Chapter 9
Multivariate Distributions
◮ Let f be the joint probability density function of X1 , X2 , . . . , Xn . We have
f (x1 , x2 , . . . , xn ) = fX1 (x1 )fX2 |X1 (x2 |x1 )fX3 |X1 ,X2 (x3 |x1 , x2 )
· · · fXn |X1 ,X2 ,...,Xn−1 (xn |x1 , x2 , . . . , xn−1 ).
In particular, for continuous random variables X, Y, and Z with joint probability density function f,
f (x, y, z) = fX (x)fY |X (y|x)fZ|X,Y (z|x, y),
◮ Order Statistics Let X1 , X2 , . . . , Xn be an independent set of identically distributed
continuous random variables with the common probability
density and distribution functions f
and F, respectively. Let X(1) be the smallest value in X1 , X2 , . . . , Xn , X(2) be the second
smallest value, X(3) be the third smallest, and, in general, X(k) (1 ≤ k ≤ n) be the k th
smallest value in X1 , X2 , . . . , Xn . Then X(k) is called the k th order statistic, and the set
X(1) , X(2) , . . . , X(n) is said to consist of the order statistics of X1 , X2 , . . . , Xn . Let Fk
and fk be the distribution and the probability density functions of X(k) , respectively, and, for
i < j and x < y, fij (x, y) be the joint probability density function of X(i) and X(j) . Then
!
n
X
i n−i
n Fk (x) =
F (x) 1 − F (x)
, −∞ < x < ∞,
i
i=k
fk (x) =
and
k−1 n−k
n!
f (x) F (x)
1 − F (x)
,
(k − 1)! (n − k)!
−∞ < x < ∞,
fij (x, y) =
i−1 j−i−1 n−j
n!
f (x)f (y) F (x)
F (y) − F (x)
1 − F (y)
.
(i − 1)! (j − i − 1)! (n − j)!
It is clear that for x ≥ y, fij (x, y) = 0. Moreover, f12···n , the joint probability density
function of X(1) , X(2) , . . . , X(n) , is given by
f12···n (x1 ,x2 , . . . , xn ) =
(
n!f (x1 )f (x2 ) · · · f (xn )
0
−∞ < x1 < x2 < · · · < xn < ∞
otherwise.
◮ Multinomial Distributions
Suppose that, whenever an experiment is performed, one
of the disjoint outcomes A1 , A2 , . . . , Ar will occur. Let P (Ai ) = pi , 1 ≤ i ≤ r.
Then p1 + p2 + · · · + pr = 1. If, in n independent performances of this experiment, Xi ,
i = 1, 2, . . . , r, denotes the number of times that Ai occurs, then p(x1 , . . . , xr ), the joint
probability mass function of X1 , X2 , . . . , Xr , is called multinomial joint probability mass
function, and its distribution is said to be a multinomial distribution. For any set of nonnegative
integers {x1 , x2 , . . . , xr } with x1 + x2 + · · · + xr = n,
p(x1 , x2 , . . . , xr ) = P (X1 = x1 , X2 = x2 , . . . , Xr = xr )
n!
px1 px2 · · · pxr r .
=
x1 ! x2 ! · · · xr ! 1 2
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 425 — #441
✐
✐
Chapter 9
Review Problems
425
Furthermore, the joint probability mass function marginalized over a subset Xi1 , Xi2 , . . . , Xik
of k (k > 1) of these r random variables is multinomial with parameters n and pi1 , pi2 , . . . ,
pik , 1 − pi1 − pi2 − · · · − pik .
REVIEW PROBLEMS
1.
An urn contains 100 chips of which 20 are blue, 30 are red, and 50 are green. Suppose
that 20 chips are drawn at random and without replacement. Let B, R, and G be the
number of blue, red, and green chips, respectively. Calculate the joint probability mass
function of B, R, and G.
2.
Let X be the smallest number obtained in rolling a balanced die n times. Calculate the
distribution function and the probability mass function of X .
3.
Suppose that n points are selected at random and independently inside the cube
Ω = (x, y, z) : − a ≤ x ≤ a, −a ≤ y ≤ a, −a ≤ z ≤ a .
Find the probability that the distance of the nearest point to the center is at least
r (r < a).
4.
The joint probability density function of random variables X, Y, and Z is given by
(
c(x + y + 2z)
if 0 ≤ x, y, z ≤ 1
f (x, y, z) =
0
otherwise.
(a)
Determine the value of c.
(b)
Find P (X < 1/3 | Y < 1/2, Z < 1/4).
5.
A fair die is tossed 18 times. What is the probability that each face appears three times?
6.
Alvie, a marksman, fires seven independent shots at a target. Suppose that the probabilities that he hits the bull’s-eye, he hits the target but not the bull’s-eye, and he misses
the target are 0.4, 0.35, and 0.25, respectively. What is the probability that he hits the
bull’s-eye three times, the target but not the bull’s-eye two times, and misses the target
two times?
7.
A system consists of n components whose lifetimes form an independent sequence of
random variables. In order for the system to function, all components must function. Let
F1 , F2 , . . . , Fn be the distribution functions of the lifetimes of the components of the
system. In terms of F1 , F2 , . . . , Fn , find the survival function of the lifetime of the
system.
8.
A system consists of n components whose lifetimes form an independent sequence of
random variables. Suppose that the system functions as long as at least one of its components functions. Let F1 , F2 , . . . , Fn be the distribution functions of the lifetimes of
the components of the system. In terms of F1 , F2 , . . . , Fn , find the survival function of
the lifetime of the system.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 426 — #442
✐
✐
426
Chapter 9
Multivariate Distributions
9.
A bar of length ℓ is broken into three pieces at two random spots. What is the probability
that the length of at least one piece is less than ℓ/20?
10.
Let X1 , X2 , and X3 be independent random variables from (0, 1). Find the probability density function and the expected value of the midrange of these random variables
[X(1) + X(3) ]/2.
Self-Test on Chapter 9
Time allotted: 90 Minutes
1.
Each problem is worth 20 points.
Let (X, Y, Z) be a point selected at random in the unit sphere
(x, y, z) : x2 + y 2 + z 2 ≤ 1 ;
that is, the sphere of radius 1 centered at the origin. [Note that the volume of a sphere
with radius R is (4/3)πR3 .]
(a)
Find f, the joint probability density function of X, Y, and Z .
(b)
Find the joint probability density function marginalized over X and Y .
(c)
Find the joint probability density function marginalized over Z .
Hint: To find fZ (z), convert the Cartesian coordinates (x, y) to polar coordinates
(r, θ).
2.
Let X1 be a random point from the interval (0, 1), X2 be a random point from the
interval (0, X1 ), X3 be a random point from the interval (0, X2 ), · · · , and Xn be a
random point from the interval (0, Xn−1 ). Find the joint probability density function of
X1 , X2 , . . . , Xn .
Hint: Use relation 9.12.
3.
For what value of c is the following a joint probability density function of four random
variables X, Y, Z, and T ? For that value of c find P (X < Y < Z < T ).
f (x, y, z, t) =
4.
c
(1 + x + y + z + t)6
0
x > 0, y > 0, z > 0, t > 0
otherwise.
A system has 7 components, and it functions if and only if at least one of its components
functions. Suppose that the lifetimes of the components are independent, identically
distributed exponential random variables with mean 1 year. Furthermore, suppose that
currently all the components are functional.
(a)
What is the probability that two years from now the system is still functional?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 427 — #443
✐
✐
Chapter 9
(b)
5.
Self-Test Problems
427
What is the probability that two years from now the system has at least three
functional components?
Suppose that 20% of the physicians working for a certain hospital retire before age 65,
30% retire at ages 65-69, and the remaining 50% retire at age 70 or later. If physicians
retire independently of each other, what is the probability that the median age of five
randomly selected retired physicians from this hospital is below 65?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 428 — #444
✐
✐
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 429 — #445
✐
✐
Chapter 10
More E xpectations
and Variances
10.1
EXPECTED VALUES OF SUMS OF RANDOM VARIABLES
As we have seen, for a discrete randomP
variable X with set of possible values A and probability
mass function p, E(X) is defined by x∈A xp(x). For a continuous random
R ∞ variable X with
probability density function f, the same quantity, E(X), is defined by −∞ xf (x) dx. Recall
that for a random variable X, the expected value, E(X), might not exist (see Examples 4.18
and 4.19 and Exercise 13, Section 6.3). In the following discussion we always assume that the
expected value of a random variable exists.
To begin, we prove the linearity property of expectation for continuous random variables.
Theorem 10.1 For random variables X1 , X2 , . . . , Xn defined on the same sample space,
and for real numbers α1 , α2 , . . . , αn ,
E
n
X
i=1
n
X
αi Xi =
αi E(Xi ).
i=1
Proof: For convenience, we prove this for continuous random variables. For discrete random
variables the proof is similar. Let f (x1 , x2 , . . . , xn ) be the joint probability density function of
X1 , X2 , . . . , Xn ; then, by Theorem 9.4,
E
n
X
i=1
Z ∞Z ∞
Z ∞ X
n
αi Xi =
···
αi xi f (x1 , x2 , . . . , xn ) dx1 dx2 · · · dxn
−∞
=
n
X
i=1
=
n
X
−∞
αi
−∞
Z ∞Z ∞
−∞
−∞
i=1
···
Z ∞
−∞
xi f (x1 , x2 , . . . , xn ) dx1 dx2 · · · dxn
αi E(Xi ). i=1
In Theorem 10.1, letting αi = 1, 1 ≤ i ≤ n, we have the following important corollary.
429
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 430 — #446
✐
✐
430
Chapter 10
Corollary
More Expectations and Variances
Let X1 , X2 , . . . , Xn be random variables on the same sample space. Then
E(X1 + X2 + · · · + Xn ) = E(X1 ) + E(X2 ) + · · · + E(Xn ).
Example 10.1
comes?
A die is rolled 15 times. What is the expected value of the sum of the out-
Solution: Let X be the sum of the outcomes, and for i = 1, 2, . . . , 15; let Xi be the outcome
of the ith roll. Then X = X1 + X2 + · · · + X15 . Thus
E(X) = E(X1 ) + E(X2 ) + · · · + E(X15 ).
Now since for all values of i, Xi is 1, 2, 3, 4, 5, and 6, with probability of 1/6 for all six values,
E(Xi ) = 1 ·
1
1
1
1
1
1
7
+2· +3· +4· +5· +6· = .
6
6
6
6
6
6
2
Hence E(X) = 15(7/2) = 52.5.
Example 10.1 and the following examples show the power of Theorem 10.1 and its corollary.
Note that in Example 10.1, while the formula
E(X1 + X2 + · · · + Xn ) = E(X1 ) + E(X2 ) + · · · + E(Xn )
enables us to compute E(X) readily, computing E(X) directly from probability mass function
of X is not easy. This is because calculation of the probability mass function of X is timeconsuming and cumbersome.
Example 10.2 A well-shuffled ordinary deck of 52 cards is divided randomly into four piles
of 13 each. Counting jack, queen, and king as 11, 12, and 13, respectively, we say that a match
occurs in a pile if the j th card is j . What is the expected value of the total number of matches
in all four piles?
Solution: Let Xi , i = 1, 2, 3, 4 be the number of matches in the ith pile. X = X1 + X2 +
X3 + X4 is the total number of matches, and E(X) = E(X1 ) + E(X2 ) + E(X3 ) + E(X4 ).
To calculate E(Xi ), 1 ≤ i ≤ 4, let Aij be the event that the j th card in the ith pile is j
(1 ≤ i ≤ 4, 1 ≤ j ≤ 13). Then by defining
(
1
if Aij occurs
Xij =
0
otherwise,
we have that Xi =
P13
j=1 Xij . Now P (Aij ) = 4/52 = 1/13 implies that
E(Xij ) = 1 · P (Aij ) + 0 · P (Acij ) = P (Aij ) =
1
.
13
Hence
E(Xi ) = E
13
X
13
13
X
X
1
Xij =
E(Xij ) =
= 1.
13
j=1
j=1
j=1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 431 — #447
✐
✐
Section 10.1
Expected Values of Sums of Random Variables
431
Thus on average there is one match in every pile. From this we get
E(X) = E(X1 ) + E(X2 ) + E(X3 ) + E(X4 ) = 1 + 1 + 1 + 1 = 4,
showing that on average there are a total of four matches.
Example 10.3 Exactly n married couples are living in a small town. What is the expected
number of intact couples after m deaths occur among the couples? Assume that the deaths
occur at random, there are no divorces, and there are no new marriages.
Solution: Let X be the number of intact couples after m deaths, and for i = 1, 2, . . . , n
define
(
1
if the ith couple is left intact
Xi =
0
otherwise.
Then X = X1 + X2 + · · · + Xn , and hence
E(X) = E(X1 ) + E(X2 ) + · · · + E(Xn ),
where
E(Xi ) = 1 · P (Xi = 1) + 0 · P (Xi = 0) = P (Xi = 1).
The event {Xi = 1} occurs if the ith couple is left intact, and thus all of the m deaths are
among the remaining n − 1 couples. Knowing that the deaths occur at random among these
individuals, we can write
!
2n − 2
(2n − 2)!
m
m! (2n − m − 2)!
(2n − m)(2n − m − 1)
! =
.
=
P (Xi = 1) =
2n(2n − 1)
(2n)!
2n
(2n − m)! m!
m
Thus the desired quantity is equal to
E(X) = E(X1 ) + E(X2 ) + · · · + E(Xn ) = n
=
(2n − m)(2n − m − 1)
.
2(2n − 1)
(2n − m)(2n − m − 1)
2n(2n − 1)
To have a numerical feeling for this interesting example, let n = 1000. Then the following
table shows the effect of the number of deaths on the expected value of the number of intact
couples.
m
100
300
600
900
1200
1500
1800
E(X)
902.48
722.44
489.89
302.38
159.88
62.41
9.95
This example was posed by Daniel Bernoulli (1700–1782).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 432 — #448
✐
✐
432
Chapter 10
More Expectations and Variances
Example 10.4 Dr. Windler’s secretary accidentally threw a patient’s file into the wastebasket. A few minutes later, the janitor cleaned the entire clinic, dumped the wastebasket containing the patient’s file randomly into one of the seven garbage cans outside the clinic, and left.
Determine the expected number of cans that Dr. Windler should empty to find the file.
Solution: Let X be the number of garbage cans that Dr. Windler should empty to find the
patient’s file. For i = 1, 2, . . . , 7, let Xi = 1 if the patient’s file is in the ith garbage can that
Dr. Windler will empty, and Xi = 0, otherwise. Then
X = 1 · X1 + 2 · X2 + · · · + 7 · X7 ,
and, therefore,
E(X) = 1 · E(X1 ) + 2 · E(X2 ) + · · · + 7 · E(X7 )
=1·
1
1
1
+ 2 · + · · · + 7 · = 4. 7
7
7
Example 10.5 A box contains nine light bulbs, of which two are defective. What is the
expected value of the number of light bulbs that one will have to test (at random and without
replacement) to find both defective bulbs?
Solution: For i = 1, 2, · · · , 8 and j > i, let Xij = j if the ith and j th light bulbs to be
P8 P9
examined are defective, and Xij = 0 otherwise. Then i=1 j=i+1 Xij is the number of
light bulbs to be examined. Therefore,
E(X) =
8
9
X
X
E(Xij ) =
i=1 j=i+1
8
9
X
X
i=1 j=i+1
j
1
!
9
2
8
8
9
1 X X
1 X 90 − i2 − i
≈ 6.67,
=
j=
36 i=1 j=i+1
36 i=1
2
where the next-to-last equality follows from
9
X
j=i+1
j=
9
X
j=1
j−
i
X
j=1
j=
90 − i2 − i
9 × 10 i(i + 1)
−
=
. 2
2
2
An elegant application of the corollary of Theorem 10.1 is that it can be used to calculate the
expected values of random variables, such as binomial, negative binomial, and hypergeometric.
The following examples demonstrate some applications.
Example 10.6 Let X be a binomial random variable with parameters (n, p). Recall that X
is the number of successes in n independent Bernoulli trials. Thus, for i = 1, 2, . . . , n, letting
(
1
if the ith trial is a success
Xi =
0
otherwise,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 433 — #449
✐
✐
Section 10.1
433
Expected Values of Sums of Random Variables
we get
X = X1 + X2 + · · · + Xn ,
(10.1)
where Xi is a Bernoulli random variable for i = 1, 2, . . . , n. Now, since ∀i, 1 ≤ i ≤ n,
E(Xi ) = 1 · p + 0 · (1 − p) = p,
(10.1) implies that
E(X) = E(X1 ) + E(X2 ) + · · · + E(Xn ) = np. Example 10.7 Let X be a negative binomial random variable with parameters (r, p). Then
in a sequence of independent Bernoulli trials each with success probability p, X is the number
of trials until the r th success. Let X1 be the number of trials until the first success, X2 be
the number of additional trials to get the second success, X3 the number of additional ones to
obtain the third success, and so on. Then clearly
X = X1 + X2 + · · · + Xr ,
where for i = 1, 2, . . . , n, the random variable Xi is geometric with parameter p. This is
because P (Xi = n) = (1 − p)n−1 p by the independence of the trials. Since E(Xi ) = 1/p
(i = 1, 2, . . . , r ),
E(X) = E(X1 ) + E(X2 ) + · · · + E(Xr ) =
r
.
p
This formula shows that, for example, in the experiment of throwing a fair die successively, on
the average, it takes 5/(1/6) = 30 trials to get five 6’s. Example 10.8
Let X be a hypergeometric random variable with probability mass function
!
!
D
N −D
x
n−x
!
p(x) = P (X = x) =
,
N
n
n ≤ min(D, N − D),
x = 0, 1, 2, . . . , n.
Then X is the number of defective items among n items drawn at random and without replacement from an urn containing D defective and N − D nondefective items. To calculate E(X),
let Ai be the event that the ith item drawn is defective. Also, for i = 1, 2, . . . , n, let
(
1
if Ai occurs
Xi =
0
otherwise;
then X = X1 + X2 + · · · + Xn . Hence
E(X) = E(X1 ) + E(X2 ) + · · · + E(Xn ),
where for i = 1, 2, . . . , n,
E(Xi ) = 1 · P (Xi = 1) + 0 · P (Xi = 0) = P (Xi = 1) = P (Ai ) =
D
.
N
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 434 — #450
✐
✐
434
Chapter 10
More Expectations and Variances
This follows since the ith item can be any of the N items with equal probabilities. Therefore,
E(X) =
nD
.
N
This shows that, for example, the expected number of spades in a random bridge hand is
(13 × 13)/52 = 13/4 ≈ 3.25. P∞
Remark 10.1 For n P
= ∞, Theorem 10.1 is not necessarily true. That is, E
i=1 Xi
∞
might not be equal to
i=1 E(Xi ). To show this, we give a counterexample. For i =
1, 2, 3, . . . , let
(
i
with probability 1/i
Yi =
0
otherwise,
and Xi =P
Yi+1 − Yi . Then since E(Yi ) = 1 for
) − E(Yi) =
P∞i = 1, 2, . . . , E(Xi ) = E(Yi+1 P
∞
X
=
−Y
,
we
have
that
E
E(X
)
=
0.
However,
since
0. Hence ∞
i
1
i
i=1 Xi =
i=1
Pi=1
P
∞
∞
6
E(−Y1 ) = −1. Therefore, E
i=1 E(Xi ).
i=1 Xi =
It can be shown that, in general,
P∞
If i=1 E |Xi | < ∞ or if, for all i, the random variables X1 , X1 , . . . are
nonnegative that is, P (Xi ≥ 0) = 1 for i ≥ 1 , then
E
∞
X
i=1
∞
X
Xi =
E(Xi ).
(10.2)
i=1
As an application of (10.2), we prove the following important theorem.
Theorem 10.2
{1, 2, 3, . . .}. Then
Let N be a discrete random variable with set of possible values
E(N ) =
∞
X
i=1
Proof:
For i ≥ 1, let
Xi =
P (N ≥ i).
(
1
if N ≥ i
∞
X
N
X
0
otherwise;
then
∞
X
Xi =
i=1
N
X
Xi +
i=1
Xi =
i=N +1
1+
i=1
∞
X
0 = N.
i=N +1
Since P (Xi ≥ 0) = 1, for i ≥ 1,
E(N ) = E
∞
X
i=1
Xi =
∞
X
i=1
E(Xi ) =
∞
X
i=1
P (N ≥ i). ✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 435 — #451
✐
✐
Section 10.1
Expected Values of Sums of Random Variables
435
Note that Theorem 10.2 is the analog of the fact that if X is a continuous nonnegative
random variable, then
Z ∞
E(X) =
P (X > t) dt.
0
This is explained in Remark 6.4.
We now prove an important inequality called the Cauchy–Schwarz inequality.
Theorem 10.3
finite variances,
(Cauchy–Schwarz Inequality)
E(XY ) ≤
Proof:
q
For random variables X and Y with
E(X 2 )E(Y 2 ).
For any real number t, (X − tY )2 ≥ 0. Hence for all values of t,
X 2 − 2XY t + t2 Y 2 ≥ 0.
Since nonnegative random variables have nonnegative expected values,
E(X 2 − 2XY t + t2 Y 2 ) ≥ 0,
which implies that
E(Y 2 )t2 − 2E(XY )t + E(X 2 ) ≥ 0.
This inequality is valid for all real numbers t including the minimum of the function g(t) =
E(Y 2 )t2 − 2E(XY )t + E(X 2 ), which is obtained by solving g ′ (t) = 0 for t. Now g ′ (t) =
0 implies that 2E(Y 2 )t − 2E(XY ) = 0 or t = E(XY )/E(Y 2 ). Plugging this into the
inequality above gives
2
E(XY )
E(XY )
+ E(X 2 ) ≥ 0,
E(Y ) · 2 − 2E(XY ) ·
2)
2
E(Y
E(Y )
2
or
2
E(XY ) ≤ E(X 2 )E(Y 2 ).
Taking square root of both sides gives the Cauchy–Schwarz inequality:
q
E(XY ) ≤ E(X 2 )E(Y 2 ). Corollary
Proof:
2
For a random variable X, E(X) ≤ E(X 2 ).
In Cauchy–Schwarz’s inequality, let Y = 1; then
q
q
E(X) = E(XY ) ≤ E(X 2 )E(1) = E(X 2 );
2
thus E(X) ≤ E(X 2 ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 436 — #452
✐
✐
436
Chapter 10
More Expectations and Variances
EXERCISES
A
1.
Let the probability density function of a random variable X be given by
(
|x − 1|
if 0 ≤ x ≤ 2
f (x) =
0
otherwise.
Find E(X 2 + X).
2.
A calculator is able to generate random numbers from the interval (0, 1). We need five
random numbers from (0, 2/5). Using this calculator, how many independent random
numbers should we generate, on average, to find the five numbers needed?
3.
Let X, Y, and Z be three independent random variables such that E(X) = E(Y
)=
2
2
E(Z) = 0, and Var(X) =Var(Y ) =Var(Z) = 1. Calculate E X (Y + 5Z) .
4.
Let the joint probability density function of random variables X and Y be
(
2e−(x+2y)
if x ≥ 0, y ≥ 0
f (x, y) =
0
otherwise.
Find E(X), E(Y ), and E(X 2 + Y 2 ).
5.
A company puts five different types of prizes into their cereal boxes, one in each box and
in equal proportions. If a customer decides to collect all five prizes, what is the expected
number of the boxes of cereals that he or she should buy?
6.
An absentminded professor wrote n letters and sealed them in envelopes without writing the addresses on the envelopes. Having forgotten which letter he had put in which
envelope, he wrote the n addresses on the envelopes at random. What is the expected
number of the letters addressed correctly?
Hint: For i = 1, 2, . . . , n, let
(
1
if the ith letter is addressed correctly
Xi =
0
otherwise.
Calculate E(X1 + X2 + · · · + Xn ).
7.
A cultural society is arranging a party for its members. The cost of a band to play music,
the amount that the caterer will charge, the rent of a hall to give the party, and other
expenses (in dollars) are uniform random variables over the intervals (1300, 1800),
(1800, 2000), (800, 1200), and (400, 700), respectively. If the number of party guests
is a random integer from (150, 200], what is the least amount that the society should
charge each participant to have no loss, on average?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 437 — #453
✐
✐
Section 10.1
8.
Expected Values of Sums of Random Variables
437
Let X1 , X2 , . . . , Xn be positive, identically distributed random variables. For
1 ≤ i ≤ n, let
Xi
.
Zi =
X1 + X2 + · · · + Xn
Show that Z1 , Z2 , . . . , Zn are also identically distributed and, for 1 ≤ i ≤ n, find
E(Zi ).
B
9.
Solve the following problem posed by Michael Khoury, U.S. Mathematics Olympiad
Member, in “The Problem Solving Competition,” Oklahoma Publishing Company and
the American Society for Communication of Mathematics, February 1999.
Bob is teaching a class with n students. There are n desks in the classroom, numbered from 1 to n. Bob has prepared a seating chart, but
the students have already seated themselves randomly. Bob calls off the
name of the person who belongs in seat 1. This person vacates the seat
he or she is currently occupying and takes his or her rightful seat. If this
displaces a person already in the seat, that person stands at the front of
the room until he or she is assigned a seat. Bob does this for each seat
in turn. After k (1 ≤ k < n) names have been called, what is the expected
number of students standing at the front of the room?
10.
Let {X1 , X2 , . . . , Xn } be a set of independent
P∞ random variables with
P (Xj = i) = pi (1 ≤ j ≤ n and i ≥ 1). Let hk = i=k pi . Using Theorem 10.2,
prove that
∞
X
E min(X1 , X2 , . . . , Xn ) =
hnk .
k=1
11.
A coin is tossed n times (n > 4). What is the expected number of exactly three consecutive heads?
Hint: Let E1 be the event that the first three outcomes are heads and the fourth outcome is tails. For 2 ≤ i ≤ n − 3, let Ei be the event that the outcome (i − 1) is tails, the
outcomes i, (i + 1), and (i + 2) are heads, and the outcome (i + 3) is tails. Let En−2
be the event that the outcome (n − 3) is tails, and the last three outcomes are heads. Let
(
1
if Ei occurs
Xi =
0
otherwise.
Then calculate the expected value of an appropriate sum of Xi ’s.
12.
Suppose that 80 balls are placed into 40 boxes at random and independently. What is the
expected number of the empty boxes?
13.
There are 25 students in a probability class. What is the expected number of birthdays
that belong only to one student? Assume that the birthrates are constant throughout the
year and that each year has 365 days.
Hint: Let Xi = 1 if the birthday of the ith student is not the birthday of any other
student, and Xi = 0, otherwise. Find E(X1 + X2 + · · · + X25 ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 438 — #454
✐
✐
438
Chapter 10
14.
There are 25 students in a probability class. What is the expected number of the days
of the year that are birthdays of at least two students? Assume that the birthrates are
constant throughout the year and that each year has 365 days.
15.
From an ordinary deck of 52 cards, cards are drawn at random, one by one, and without
replacement until a heart is drawn. What is the expected value of the number of cards
drawn?
Hint: See Exercise 13, Section 3.2.
16.
Let X and Y be nonnegative random variables with an arbitrary joint distribution function. Let
(
1
if X > x, Y > y
I(x, y) =
0
otherwise.
(a)
More Expectations and Variances
Show that
Z ∞Z ∞
0
(b)
I(x, y) dx dy = XY.
0
By calculating expected values of both sides of part (a), prove that
Z ∞Z ∞
E(XY ) =
P (X > x, Y > y) dx dy.
0
0
Note that this is a generalization of the result explained in Remark 6.4.
17.
Let {X1 , X2 , . . .} be a sequence of continuous, independent, and identically distributed
random variables. Let
N = min{n : X1 ≥ X2 ≥ X3 ≥ · · · ≥ Xn−1 , Xn−1 < Xn }.
Find E(N ).
18.
From an urn that contains a large number of red and blue chips, mixed in equal proportions, 10 chips are removed one by one and at random. The chips that are removed before
the first red chip are returned to the urn. The first red chip, together with all those that
follow, is placed in another urn that is initially empty. Calculate the expected number of
the chips in the second urn.
19.
Under what condition does Cauchy–Schwarz’s inequality become equality?
Self-Quiz on Section 10.1
Time allotted: 20 Minutes
1.
Each problem is worth 5 points.
Shiante does not remember in which one of her 11 disk storage wallets she stored the last
DVD that she purchased. Determine the expected number of the wallets that she should
search to find the DVD. Assume that, in her search, she chooses the wallets randomly.
Hint: For 1 ≤ i ≤ 11, let Xi = 1 if the DVD is stored in the ith wallet that Shiante
searches, and Xi = 0, otherwise. The number of wallets that she should search to find
the DVD can be written in terms of the Xi ’s.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 439 — #455
✐
✐
Section 10.2
2.
Covariance
439
Let X be the number of students in Dr. Brown-Rose’s English 101 who will fail the
course next semester. Show that
Var(X)
P (X = 0) ≤
.
E(X 2 )
Hint: Let
Y =
1
0
if X > 0
if X = 0
and apply the Cauchy–Schwarz inequality.
10.2
COVARIANCE
In Sections 4.5 and 6.3 we studied
the notion of the variance of a random variable X . We
2 showed that E X − E(X) , the variance of X, measures the average magnitude of the
fluctuations of the random variable X from its expected value, E(X). We mentioned that
this quantity measures the dispersion, or spread, of the distribution of X about its expected
value. Now suppose that X and Y are two jointly distributed random variables. Then Var(X)
and Var(Y ) determine the dispersions of X and Y independently rather than jointly. In fact,
Var(X) measures the spread, or dispersion, along the x-direction, and Var(Y ) measures the
spread, or dispersion, along the y -direction in the plane. We now calculate Var(aX + bY ), the
joint spread, or dispersion, of X and Y along the (ax+by)-direction for arbitrary real numbers
a and b:
Var(aX + bY )
2
= E (aX + bY ) − E(aX + bY )
2
= E (aX + bY ) − aE(X) − bE(Y )
2
= E a X − E(X) + b Y − E(Y )
2
2
= E a2 X − E(X) + b2 Y − E(Y ) + 2ab X − E(X) Y − E(Y )
= a2 Var(X) + b2 Var(Y ) + 2abE X − E(X) Y − E(Y ) .
(10.3)
This formula shows that the joint spread, or dispersion, of X
and Y can be
measured in
any
direction (ax + by) if the quantities Var(X), Var(Y ), and E X − E(X) Y − E(Y ) are
known. On the other hand, the joint spread, or dispersion, of X and Y depends on these three
quantities. However,
Var(X)
) determine the dispersions of X and Y independently;
and Var(Y therefore, E X − E(X) Y − E(Y ) is the quantity that gives information about the joint
spread, or dispersion, of X and Y . It is called the covariance of X and Y, is denoted by
Cov(X, Y ), and determines how X and Y covary jointly. For example, by relation (10.3), if
for random variables X, Y, and Z, Var(Y ) = Var(Z) and ab > 0, then the joint dispersion of
X and Y along the (ax + by)-direction is greater than the joint dispersion of X and Z along
the (ax + bz)-direction if and only if Cov(X, Y ) > Cov(X, Z).
Definition 10.1 Let X and Y be jointly distributed random variables; then the covariance
of X and Y is defined by
Cov(X, Y ) = E X − E(X) Y − E(Y ) .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 440 — #456
✐
✐
440
Chapter 10
More Expectations and Variances
Note that
Cov(X, X) = Var(X).
Also, by the Cauchy–Schwarz inequality (Theorem 10.3),
Cov(X, Y ) = E X − E(X) Y − E(Y )
q 2 2
≤ E X − E(X) E Y − E(Y )
q
2 2
= σX
σY = σX σY ,
which shows that if σX < ∞ and σY < ∞, then Cov(X, Y ) < ∞.
Rewriting relation (10.3) in terms of Cov(X, Y ), we obtain the following important theorem:
Theorem 10.4
Let a and b be real numbers; for random variables X and Y,
Var(aX + bY ) = a2 Var(X) + b2 Var(Y ) + 2ab Cov(X, Y ).
In particular, if a = 1 and b = 1, this gives
Var(X + Y ) = Var(X) + Var(Y ) + 2 Cov(X, Y ).
(10.4)
Similarly, if a = 1 and b = −1, it gives
Var(X − Y ) = Var(X) + Var(Y ) − 2 Cov(X, Y ).
(10.5)
Letting µX = E(X), and µY = E(Y ), an alternative formula for
Cov(X, Y ) = E X − E(X) Y − E(Y )
is calculated by the expansion of E X − E(X) Y − E(Y ) :
Cov(X, Y ) = E (X − µX )(Y − µY )
= E(XY − µX Y − µY X + µX µY )
= E(XY ) − µX E(Y ) − µY E(X) + µX µY
= E(XY ) − µX µY − µY µX + µX µY
= E(XY ) − µX µY = E(XY ) − E(X)E(Y ).
Therefore,
Cov(X, Y ) = E(XY ) − E(X)E(Y ).
(10.6)
Using this relation, we get
Cov(aX + b, cY + d) = E (aX + b)(cY + d) − E(aX + b)E(cY + d)
= E(acXY + bcY + adX + bd) − aE(X) + b cE(Y ) + d
= ac E(XY ) − E(X)E(Y ) = ac Cov(X, Y ).
Hence, for arbitrary real numbers a, b, c, d and random variables X and Y,
Cov(aX + b, cY + d) = ac Cov(X, Y ),
(10.7)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 441 — #457
✐
✐
Section 10.2
Covariance
441
which can be generalized as follows: Let ai ’s and bj ’s be constants. For random variables X1 ,
X2 , . . . , Xn and Y1 , Y2 , . . . , Ym ,
n X
m
n
m
X
X
X
ai bj Cov(Xi , Yj ).
Cov
ai Xi ,
bj Yj =
i=1
j=1
(10.8)
i=1 j=1
(See Exercise 22.)
For random variables Xand Y, Cov(X,
Y ) might be positive, negative, or zero. It is positive if the expected value of X − E(X) Y − E(Y ) is positive, that is, if X and Y decrease
together or increase together. It is negative if X increases while Y decreases, or vice versa.
If Cov(X, Y ) > 0, we say that X and Y are positively correlated. If Cov(X, Y ) < 0, we
say that they are negatively correlated. If Cov(X, Y ) = 0, we say that X and Y are uncorrelated. For example, the blood cholesterol level of a person is positively correlated with the
amount of saturated fat consumed by that person, whereas the amount of alcohol in the blood
is negatively correlated with motor coordination. Generally, the more saturated fat a person
ingests, the higher his or her blood cholesterol level will be. The more alcohol a person drinks,
the poorer his or her level of motor coordination becomes. As another example, let X be the
weight of a person before starting a health fitness program and Y be his or her weight afterward. Then X and Y are negatively correlated because the effect of fitness programs is that,
usually, heavier persons lose weight, whereas lighter persons gain weight. The best examples
for uncorrelated random variables are independent ones. If X and Y are independent, then
Cov(X, Y ) = E(XY ) − E(X)E(Y ) = 0.
However, as the following example shows, the converse of this is not true; that is,
Two dependent random variables might be uncorrelated.
Example 10.9
Let X be uniformly distributed over (−1, 1) and Y = X 2 . Then
Cov(X, Y ) = E(X 3 ) − E(X)E(X 2 ) = 0,
since E(X) = 0 and E(X 3 ) = 0. Thus the perfectly related random variables X and Y are
uncorrelated. Example 10.10 There are 300 cards in a box numbered 1 through 300. Therefore, the number on each card has one, two, or three digits. A card is drawn at random from the box. Suppose
that the number on the card has X digits of which Y are 0. Determine whether X and Y are
positively correlated, negatively correlated, or uncorrelated.
Solution: Note that, between 1 and 300, there are 9 one-digit numbers none of which is 0;
there are 90 two-digit numbers of which 81 have no 0’s, 9 have one 0, and none has two 0’s;
and there are 201 three-digit numbers of which 162 have no 0’s, 36 have one 0, and 3 have
two 0’s. These facts show that as X increases so does Y . Therefore, X and Y are positively
correlated. To show this mathematically, let p(x, y) be the joint probability mass function of
X and Y . Simple calculations will yield the following table for p(x, y).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 442 — #458
✐
✐
442
Chapter 10
More Expectations and Variances
y
x
0
1
2
pX (x)
1
2
3
9/300
81/300
162/300
0
9/300
36/300
0
0
3/300
9/300
90/300
201/300
pY (y)
252/300
45/300
3/300
To see how we calculated the entries of the table, as an example, consider p(3, 0). This quantity
is 162/300 because there are 162 three-digit numbers with no 0’s. Now from this table we have
that
90
201
9
+2·
+3·
= 2.64.
300
300
300
252
45
3
E(Y ) = 0 ·
+1·
+2·
= 0.17.
300
300
300
E(X) = 1 ·
E(XY ) =
3 X
2
X
x=1 y=0
xyp(x, y) = 2 ·
36
3
9
+3·
+6·
= 0.48.
300
300
300
Therefore,
Cov(X, Y ) = E(XY ) − E(X)E(Y ) = 0.48 − (0.17)(2.64) = 0.0312 > 0,
which shows that X and Y are positively correlated.
Example 10.11 Ann cuts an ordinary deck of 52 cards and displays the exposed card. Andy
cuts the remaining stack of cards and displays his exposed card. Counting jack, queen, and
king as 11, 12, and 13, let X and Y be the numbers on the cards that Ann and Andy expose,
respectively. Find Cov(X, Y ) and interpret the result.
Solution: Observe that the number of cards in Ann’s stack after she cuts the deck, and the
number of cards in Andy’s stack after he cuts the remaining cards will not change the probabilities we are interested in. The problem is equivalent to choosing two cards at random and
without replacement from an ordinary deck of 52 cards, and letting X be the number on one
card and Y be the number on the other card. Let p(x, y) be the joint probability mass function
of X and Y . For 1 ≤ x, y ≤ 13,
Therefore,
p(x, y) = P (X = x, Y = y) = P (Y = y | X = x)P (X = x)
1 · 1 = 1
x 6= y
156
= 12 13
0
x = y.
pX (x) =
13
X
y=1
y6=x
p(x, y) =
1
12
= ,
156
13
x = 1, 2, . . . , 13;
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 443 — #459
✐
✐
Section 10.2
pY (y) =
13
X
p(x, y) =
x=1
x6=y
1
12
=
,
156
13
Covariance
443
x = 1, 2, . . . , 13.
By these relations,
E(X) =
13
X
x
x=1
E(Y ) =
13
13
X
y
y=1
13
=
1 13 × 14
·
= 7;
13
2
=
1 13 × 14
·
= 7.
13
2
By Theorem 8.1,
13 X
13
X
xy
13 13
13
1 XX
1 X 2
E(XY ) =
=
x
xy −
156
156 x=1 y=1
156 x=1
x=1 y=1
y6=x
=
1
1 X X y −
· 819
x
156 x=1
156
y=1
=
287
1 13 × 14 13 × 14 819
·
·
−
=
.
156
2
2
156
6
13
13
Therefore,
7
287
− 49 = − .
6
6
This shows that X and Y are negatively correlated. That is, if X increses, then Y decreases; if
X decreases, then Y increases. These facts should make sense intuitively. Cov(X, Y ) = E(XY ) − E(X)E(Y ) =
Example 10.12 Let X be the lifetime of an electronic system and Y be the lifetime of
one of its components. Suppose that the electronic system fails if the component does (but not
necessarily vice versa). Furthermore, suppose that the joint probability density function of X
and Y (in years) is given by
1
e−y/7
if 0 ≤ x ≤ y < ∞
f (x, y) = 49
0
elsewhere.
(a)
Determine the expected value of the remaining lifetime of the component when the
system dies.
(b)
Find the covariance of X and Y .
Solution:
(a)
The remaining lifetime of the component when the system dies is Y − X . So the desired
quantity is
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 444 — #460
✐
✐
444
Chapter 10
More Expectations and Variances
Z ∞Z y
1 −y/7
e
dx dy
49
0
0
Z
Z
1 ∞ −y/7 2 y 2 1 ∞ 2 −y/7
=
e
y −
dy =
y e
dy = 7,
49 0
2
98 0
E(Y − X) =
(y − x)
where the last integral is calculated using integration by parts twice.
(b)
To find Cov(X, Y ) = E(XY ) − E(X)E(Y ), note that
Z ∞Z y
1
E(XY ) =
(xy) e−y/7 dx dy
49
0
0
Z ∞
Z y
1
−y/7
=
ye
x dx dy
49 0
0
Z
1 ∞ 3 −y/7
14, 406
=
y e
dy =
= 147,
98 0
98
where the last integral is calculated, using integration by parts three times. We also have
Z ∞Z y
1
E(X) =
x e−y/7 dx dy = 7,
49
0
0
Z ∞Z y
1
E(Y ) =
y e−y/7 dx dy = 14.
49
0
0
Therefore, Cov(X, Y ) = 147 − 7(14) = 49. Note that Cov(X, Y ) > 0 is expected
because X and Y are positively correlated. As Theorem 10.4 shows, one important application of the covariance of two random variables X and Y is that it enables us to find Var(aX + bY ) for constants a and b. By direct
calculations similar to (10.3), that theorem is generalized as follows: Let a1 , a2 , . . . , an be
real numbers; for random variables X1 , X2 , . . . , Xn ,
n
n
X
X
XX
Var
ai Xi =
a2i Var(Xi ) + 2
ai aj Cov(Xi , Xj ).
i=1
i=1
(10.9)
i<j
In particular, for ai = 1, 1 ≤ i ≤ n,
n
n
X
X
XX
Var
Xi =
Var(Xi ) + 2
Cov(Xi , Xj ).
i=1
(10.10)
i<j
i=1
By (10.9) and (10.10),
If X1 , X2 , . . . , Xn are pairwise independent or, more generally, pairwise
uncorrelated, then
Var
n
X
i=1
n
X
ai Xi =
a2i Var(Xi )
n
n
X
X
Var
Xi =
Var(Xi ),
i=1
(10.11)
i=1
(10.12)
i=1
the reason being that Cov(Xi , Xj ) = 0, for all i, j, i =
6 j.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 445 — #461
✐
✐
Section 10.2
Example 10.13
Covariance
445
Let X be the number of 6’s in n rolls of a fair die. Find Var(X).
Solution: Let Xi = 1 if on the ith roll the die lands 6, and Xi = 0, otherwise. Then X =
X1 + X2 + · · · + Xn . Since X1 , X2 , . . . , Xn are independent,
Var(X) = Var(X1 ) + Var(X2 ) + · · · + Var(Xn ).
But for i = 1, 2, . . . , n,
5
1
1
+0· = ,
6
6
6
5
1
1
E(Xi2 ) = (1)2 · + 02 · = ,
6
6
6
E(Xi ) = 1 ·
and hence
Var(Xi ) =
1
1
5
−
=
.
6 36
36
Therefore, Var(X) = n(5/36) = (5n)/36.
Example 10.14 Using relation (10.10), calculate the variance of a binomial random variable
X with parameters (n, p).
Solution: Recall that X is the number of successes in n independent Bernoulli trials. Thus
for i = 1, 2, . . . , n, letting
(
1
if the ith trial is a success
Xi =
0
otherwise,
we obtain
X = X1 + X2 + · · · + Xn ,
where Xi is a Bernoulli random variable for i = 1, 2, . . . , n. Note that E(Xi ) = p and
Var(Xi ) = p(1 − p). Since {X1 , X2 , . . . , Xn } is an independent set of random variables, by
(10.12),
Var(X) = Var(X1 ) + Var(X2 ) + · · · + Var(Xn ) = np(1 − p).
Example 10.15 Using relation (10.10), calculate the variance of a negative binomial random
variable X, with parameter (r, p).
Solution: Recall that in a sequence of independent Bernoulli trials, X is the number of trials
until the r th success. Let X1 be the number of trials until the first success, X2 the number of
additional trials to get the second success, X3 the number of additional ones to obtain the third
success, and so on. Then X = X1 + X2 + · · · + Xr , where for i = 1, 2, . . . , r, the random
variable Xi is geometric with parameter p. Since X1 , X2 , . . . , Xr are independent,
Var(X) = Var(X1 ) + Var(X2 ) + · · · + Var(Xr ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 446 — #462
✐
✐
446
Chapter 10
More Expectations and Variances
But Var(Xi ) = (1 − p)/p2 , for i = 1, 2, . . . , r (see Section 5.3). Thus Var(X) =
r(1 − p)/p2 . Let F be the distribution function of a certain characteristic of the elements of some population. For example, let F be the distribution function of lifetimes of light bulbs manufactured
by a company, a Scottish soldier’s chest size, or the score on an achievement test of a random
student. Let X1 , X2 , . . . , Xn be a random sample from the distribution F . For example, let
X1 , X2 , . . . , Xn be the chest sizes of n Scottish soldiers selected at random and independently. Then the following lemma enables us to find the expected value and variance of the
sample mean in terms of the expected value and variance of the distribution function F (i.e.,
the population mean and variance).
Lemma 10.1 Let X1 , X2 , . . . , Xn be a random sample of size n from a distribution F with
mean µ and variance σ 2 . Let X̄ be the mean of the random sample. Then
E(X̄) = µ
Var(X̄) =
and
σ2
n
.
Proof: By Theorem 10.1,
E(X̄) = E
X + X + · · · + X 1
2
n
n
=
1
1
E(X1 )+E(X2 )+· · · +E(Xn ) = ·nµ = µ.
n
n
Since X1 , X2 , . . . , Xn are independent random variables, by (10.11),
Var(X̄) = Var
X + X + · · · + X 1
2
n
2
=
n
1
σ
· nσ 2 =
. 2
n
n
=
1
Var(X1 ) + Var(X2 ) + · · · + Var(Xn )
2
n
EXERCISES
A
1.
Ann cuts an ordinary deck of 52 cards and displays the exposed card. After Ann places
her stack back on the deck, Andy cuts the same deck and displays the exposed card.
Counting jack, queen, and king as 11, 12, and 13, let X and Y be the numbers on the
cards that Ann and Andy expose, respectively. Find Cov(X, Y ).
2.
Let the joint probability mass function of random variables X and Y be given by
1
x(x + y) if x = 1, 2, 3, y = 3, 4
70
p(x, y) =
0
elsewhere.
Find Cov(X, Y ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 447 — #463
✐
✐
Section 10.2
Covariance
447
3.
Roll a balanced die and let the outcome be X . Then toss a fair coin X times and let Y
denote the number of tails. Find Cov(X, Y ) and interpret the result.
Hint: Let p(x, y) be the joint probability mass function of X and Y . To save time, use
the table for p(x, y) constructed in Example 8.2.
4.
Thieves stole four animals at random from a farm that had seven sheep, eight goats, and
five burros. Calculate the covariance of the number of sheep and goats stolen.
5.
In n independent Bernoulli trials, each with probability of success p, let X be the number of successes and Y the number of failures. Calculate E(XY ) and Cov(X, Y ).
6.
For random variables X, Y, and Z, prove that
(a)
Cov(X + Y, Z) = Cov(X, Z)+Cov(Y, Z).
(b)
Cov(X, Y + Z) = Cov(X, Y )+Cov(X, Z).
7.
Show that if X and Y are independent random variables, then for all random variables
Z,
Cov(X, Y + Z) = Cov(X, Z).
8.
For random variables X and Y, show that
Cov(X + Y, X − Y ) = Var(X) − Var(Y ).
9.
Prove that
Var(X − Y ) = Var(X) + Var(Y ) − 2 Cov(X, Y ).
10.
Let X and Y be two independent random variables.
(a)
(b)
Show that X − Y and X + Y are uncorrelated if and only if Var(X) =Var(Y ).
Show that Cov(X, XY ) = E(Y )Var(X).
11.
Prove that if Θ is a random number from the interval [0, 2π], then the dependent random
variables X = sin Θ and Y = cos Θ are uncorrelated.
12.
Let
coordinates of a random point selected uniformly from the unit disk
X and 2Y be the
(x, y) : x + y 2 ≤ 1 . Are X and Y independent? Are they uncorrelated? Why or
why not?
13.
Mr. Jones has two jobs. Next year, he will get a salary raise of X thousand dollars from
one employer and a salary raise of Y thousand dollars from his second. Suppose that
X and Y are independent random variables with probability density functions f and g,
respectively, where
8x/15
if 1/2 < x < 2
f (x) =
0
elsewhere,
√
6 y/13
if 1/4 < y < 9/4
g(y) =
0
elsewhere.
What are the expected value and variance of the total raise that Mr. Jones will get next
year?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 448 — #464
✐
✐
448
Chapter 10
14.
Let X and Y be independent random variables with expected values µ1 and µ2 , and
variances σ12 and σ22 , respectively. Show that
More Expectations and Variances
Var(XY ) = σ12 σ22 + µ21 σ22 + µ22 σ12 .
15.
A voltmeter is used to measure the voltage of voltage sources, such as batteries. Every
time this device is used, a random error is made, independent of other measurements,
with mean 0 and standard deviation σ . Suppose that we want to measure the voltages,
V1 and V2 , of two batteries. If a measurement with smaller error variance is preferable,
determine which of the following methods should be used:
(a)
To measure V1 and V2 separately.
(b)
To measure V = V1 + V2 and W = V1 − V2 , and then find V1 and V2 from
V1 = (V + W )/2 and V2 = (V − W )/2. There are methods available for
engineers to do these. For example, to measure V1 + V2 , they attach batteries
in series so that each battery pushes current in the same direction. To measure
V1 − V2 , they attach the batteries in series so that they push current in opposite
directions.
Assume that the internal resistances of the batteries are negligible.
16.
Let X and Y have the following joint probability density function
(
8xy
if 0 < x ≤ y < 1
f (x, y) =
0
otherwise.
(a) Calculate Var(X + Y ).
(b)
Show that X and Y are not independent. Explain why this does not contradict
Exercise 27 of Section 8.2.
17.
Find the variance of a sum of n randomly and independently selected points from the
interval (0, 1).
18.
Let X and Y be jointly distributed with joint probability density function
1
x3 e−xy−x
if x > 0, y > 0
f (x, y) = 2
0
otherwise.
19.
Determine if X and Y are positively
negatively correlated, or uncorrelated.
R ∞ correlated,
n −ax
Hint: Note that for all a > 0, 0 x e
dx = n!/an+1 .
Let X be a random variable. Prove that Var(X) = min t E (X − t)2 .
Hint: Let µ = E(X) and look at the expansion of
E (X − t)2 = E (X − µ + µ − t)2 .
B
20.
Let S be the sample space of an experiment. Let A and B be two events of S . Let IA
and IB be the indicator variables for A and B . That is,
(
1
if ω ∈ A
IA (ω) =
0
if ω 6∈ A,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 449 — #465
✐
✐
Section 10.2
IB (ω) =
(
1
Covariance
449
if ω ∈ B
if ω 6∈ B.
0
Show that IA and IB are positively correlated if and only if P (A | B) > P (A), and if
and only if P (B | A) > P (B).
21.
Show that for random variables X, Y, Z, and W and constants a, b, c, and d,
Cov(aX + bY, cZ + dW )
= ac Cov(X, Z) + bc Cov(Y, Z) + ad Cov(X, W ) + bd Cov(Y, W ).
Hint:
22.
For a simpler proof, use the results of Exercise 6.
Prove the following generalization of Exercise 21:
n
m
n X
m
X
X
X
Cov
ai Xi ,
bj Y j =
ai bj Cov(Xi , Yj ).
i=1
j=1
i=1 j=1
23.
A fair die is thrown n times. What is the covariance of the number of 1’s and the number
of 6’s obtained?
Hint: Use the result of Exercise 22.
24.
Show that if X1 , X2 , . . . , Xn are random variables and a1 , a2 , . . . , an are constants,
then
Var
n
X
i=1
25.
n
X
XX
ai Xi =
a2i Var(Xi ) + 2
ai aj Cov(Xi , Xj ).
i=1
i<j
Let X be a hypergeometric random variable with probability mass function
!
!
D
N −D
x
n−x
!
,
p(x) = P (X = x) =
N
n
n ≤ min(D, N − D),
x = 0, 1, 2, . . . , n.
Recall that X is the number of defective items among n items drawn randomly and
without replacement from a box containing D defective and N − D nondefective items.
Show that
nD(N − D) n−1
Var(X) =
1
−
.
N2
N −1
Hint:
let
Let Ai be the event that the ith item drawn is defective. Also for i = 1, 2, . . . , n,
Xi =
(
1
0
if Ai occurs
otherwise.
Then X = X1 + X2 + · · · + Xn .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 450 — #466
✐
✐
450
Chapter 10
26.
Exactly n married couples are living in a small town. What is the variance of the surviving couples after m deaths occur among them? Assume that the deaths occur at random,
there are no divorces, and there are no new marriages.
Note: This situation involves the Daniel Bernoulli problem discussed in Example 10.3.
More Expectations and Variances
Self-Quiz on Section 10.2
Time allotted: 20 Minutes
1.
Each problem is worth 5 points.
We draw 8 cards, one at a time, randomly, and without replacement from an ordinary
deck of 52 cards. Let
1
if the ith card drawn is a heart
Xi =
0
otherwise.
For 1 ≤ i < j ≤ 8, find Cov(Xi , Xj ).
2.
Let X and Y be independent and identically distributed exponential random variables
with parameter λ. Let U = max(X, Y ) and V = min(X, Y ). Using the relation
Cov(U, V ) = Cov(U − V + V, V ) = Cov(U − V, V ) + Cov(V, V ),
Calculate Cov(U, V ) in terms of Var(V ).
10.3
CORRELATION
Although for random variables X and Y, Cov(X, Y ) provides information about how X and
Y vary jointly, it has a major shortcoming: It is not independent of the units in which X and
Y are measured. For example, suppose that for random variables X and Y, when measured
in (say) centimeters, Cov(X, Y ) = 0.15. For the same random variables, if we change the
measurements to millimeters, then X1 = 10X and Y1 = 10Y will be the new observed
values, and by relation (10.7) we get
Cov(X1 , Y1 ) = Cov(10X, 10Y ) = 100 Cov(X, Y ) = 15,
showing that Cov(X, Y ) is sensitive to the units of measurement.
From
Section 4.6 we know
that for a random variable X, the standardized X, X ∗ = X − E(X) /σX , is independent of
the units in which X is measured. Thus, to define a measure of association between X and Y,
independent of the scales of measurements, it is appropriate to consider Cov(X ∗ , Y ∗ ) rather
than Cov(X, Y ). Using relation (10.7), we obtain
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 451 — #467
✐
✐
Section 10.3
Correlation
451
X − E(X) Y − E(Y ) Cov(X ∗ , Y ∗ ) = Cov
,
σX
σY
1
1
1
1
X−
E(X),
Y −
E(Y )
= Cov
σX
σX
σY
σY
=
1
1
Cov(X, Y )
·
Cov(X, Y ) =
.
σX σY
σX σY
2
Definition 10.2
Let X and Y be two random variables with 0 < σX
< ∞ and
2
0 < σY < ∞. The covariance between the standardized X and the standardized Y is called
the correlation coefficient between X and Y and is denoted by ρ = ρ(X, Y ). Therefore,
ρ=
Cov(X, Y )
σX σY
.
The quantity Cov(X, Y )/(σX σY ) gives all the important information that Cov(X, Y ) provides about how X and Y covary and, at the same time, it is not sensitive to the scales of measurement. Clearly, ρ(X, Y ) > 0 if and only if X and Y are positively correlated; ρ(X, Y ) < 0
if and only if X and Y are negatively correlated; and ρ(X, Y ) = 0 if and only if X and Y are
uncorrelated. Moreover, ρ(X, Y ) roughly measures the amount and the sign of linear relationship between X and Y . It is −1 if Y = aX + b, a < 0, and +1 if Y = aX + b, a > 0. Thus
ρ(X, Y ) = ±1 in the case of perfect linear relationship and 0 in the case of independence of
X and Y . The key to the proof of “ρ(X, Y ) = ±1 if and only if Y = aX + b” is the following
lemma.
Lemma 10.2
For random variables X and Y with correlation coefficient ρ(X, Y ),
Y = 2 + 2ρ(X, Y );
σX
σY
X
Y Var
−
= 2 − 2ρ(X, Y ).
σX
σY
Var
X
+
We prove the first relation; the second can be shown similarly. By Theorem 10.4,
Proof:
+
Y 1
1
Cov(X, Y )
= 2 Var(X) + 2 Var(Y ) + 2
= 2 + 2ρ(X, Y ). σY
σX
σY
σX σY
Theorem 10.5
For random variables X and Y with correlation coefficient ρ(X, Y ),
Var
(a)
X
σX
−1 ≤ ρ(X, Y ) ≤ 1.
(b)
With probability 1, ρ(X, Y ) = 1 if and only if Y = aX + b for some constants a, b,
a > 0.
(c)
With probability 1, ρ(X, Y ) = −1 if and only if Y = aX + b for some constants a, b,
a < 0.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 452 — #468
✐
✐
452
Chapter 10
More Expectations and Variances
Proof:
(a)
Since the variance of a random variable is nonnegative,
X
X
Y Y +
≥ 0 and Var
−
≥ 0.
Var
σX
σY
σX
σY
Therefore, by Lemma 10.2, 2 + 2ρ(X, Y ) ≥ 0 and 2 − 2ρ(X, Y ) ≥ 0. That is,
ρ(X, Y ) ≥ −1 and ρ(X, Y ) ≤ 1.
(b)
First, suppose that ρ(X, Y ) = 1. In this case, by Lemma 10.2,
X
Y Var
−
= 2 1 − ρ(X, Y ) = 0.
σX
σY
Therefore, with probability 1,
X
Y
−
= c,
σX
σY
for some constant c. Hence, with probability 1,
σY
Y =
X − c σY ≡ aX + b,
σX
with a = σY /σX > 0 and b = −c σY . Next, assume that Y = aX + b, a > 0. We
have that
Cov(X, aX + b)
ρ(X, Y ) = ρ(X, aX + b) =
σX σaX+b
=
(c)
a Cov(X, X)
a Var(X)
=
= 1.
σX (aσX )
a Var(X)
The proof of this statement is similar to that of part (b).
Example 10.16 Show that if X and Y are continuous random variables with the joint probability density function
(
x+y
if 0 < x < 1, 0 < y < 1
f (x, y) =
0
otherwise,
then X and Y are not linearly related.
Solution: Since X and Y are linearly related if and only if ρ(X, Y ) = ±1 with probability
1, it suffices to prove that ρ(X, Y ) 6= ±1. To do so, note that
Z 1Z 1
7
E(X) =
x(x + y) dx dy =
,
12
0
0
E(XY ) =
Z 1Z 1
0
1
xy(x + y) dx dy = .
3
0
Also, by symmetry, E(Y ) = 7/12; therefore,
Cov(X, Y ) = E(XY ) − E(X)E(Y ) =
1
7 7
1
−
=−
.
3 12 12
144
Similarly,
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 453 — #469
✐
✐
Section 10.3
E(X 2 ) =
Z 1Z 1
0
x2 (x + y) dx dy =
0
Correlation
453
5
,
12
r
√
q
7 2
2
5
11
−
.
σX = E(X 2 ) − E(X) =
=
12
12
12
√
Again, by symmetry, σY = 11/12. Thus
ρ(X, Y ) =
Cov(X, Y )
−1/144
1
√
=√
= − 6= ±1.
σX σY
11
11/12 · 11/12
The following example shows that even if X and Y are dependent through a nonlinear
relationship such as Y = X 2 , still, statistically, there might be a strong linear association
between X and Y . That is, ρ(X, Y ) might be very close to 1 or −1, indicating that the points
(x, y) are tightly clustered around a line.
Example 10.17 Let X be a random number from the interval (0, 1) and Y = X 2 . The
probability density function of X is
(
1
if 0 < x < 1
f (x) =
0
elsewhere,
and for n ≥ 1,
E(X n ) =
Z 1
xn dx =
0
1
.
n+1
Thus E(X) = 1/2, E(Y ) = E(X 2 ) = 1/3,
and, finally,
2
1 1
1
2
σX
= Var(X) = E(X 2 ) − E(X) = − = ,
3 4
12
2
1 2
4
1
= ,
σY2 = Var(Y ) = E(X 4 ) − E(X 2 ) = −
5
3
45
Cov(X, Y ) = E(X 3 ) − E(X)E(X 2 ) =
1
1 11
−
=
.
4 23
12
Therefore,
Cov(X, Y )
1/12
√
√ =
ρ(X, Y ) =
=
σX σY
1/2 3 · 2/3 5
√
15
= 0.968. 4
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 454 — #470
✐
✐
454
Chapter 10
More Expectations and Variances
EXERCISES
A
1.
2.
Let X and Y be jointly distributed, with ρ(X, Y ) = 1/2, σX = 2, σY = 3. Find
Var(2X − 4Y + 3).
Let the joint probability density function of X and Y be given by
(
sin x sin y
if 0 ≤ x ≤ π/2, 0 ≤ y ≤ π/2
f (x, y) =
0
otherwise.
Calculate the correlation coefficient of X and Y .
3.
A stick of length 1 is broken into two pieces at a random point. Find the correlation
coefficient and the covariance of these pieces.
4.
For real numbers α and β, let
1
sgn(αβ) = 0
−1
if αβ > 0
if αβ = 0
if αβ < 0.
Prove that for random variables X and Y,
ρ(α1 X + α2 , β1 Y + β2 ) = ρ(X, Y ) sgn(α1 β1 ).
5.
Is it possible that for some random variables X and Y, ρ(X, Y ) = 3, σX = 2, and
σY = 3?
6.
Prove that if Cov(X, Y ) = 0, then
ρ(X + Y, X − Y ) =
Var(X) − Var(Y )
.
Var(X) + Var(Y )
B
7.
Show that if the joint probability density function of X and Y is
1
sin(x + y)
f (x, y) = 2
0
if 0 ≤ x ≤
π
,
2
0≤y≤
π
2
elsewhere,
then there exists no linear relation between X and Y .
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 455 — #471
✐
✐
Section 10.4
Conditioning on Random Variables
455
Self-Quiz on Section 10.3
Time allotted: 20 Minutes
Each problem is worth 5 points.
1.
In a probability exam, for two random variables X and Y with a given joint probability
density function, Fiona’s calculations resulted in E(X) = E(Y ) = 2, E(X 2 ) = 13,
E(Y 2 ) = 40, and E(XY ) = 23. However, Fiona got no points for these answers.
Under her solution, the professor wrote, impossible results. Show that, in fact, for no
joint probability density function such results are possible.
2.
For random variables X1 , X2 , . . . , Xn , we have that Var(Xi ) = 4, 1 ≤ i ≤ n, and for
1 ≤ i 6= j ≤ n, ρ(Xi , Xj ) = −1/16. Find Var(X1 + X2 + · · · + Xn ).
Pn−1 Pn
Hint: Note that i=1 j=i+1 Cov(Xi , Xj ) has (n − 1) + (n − 2) + · · · + 1 =
[(n − 1)n]/2 terms.
10.4
CONDITIONING ON RANDOM VARIABLES
An important application of conditional expectations is that ordinary expectations and probabilities can be calculated by conditioning on appropriate random variables. First, to explain this
procedure, we state a definition.
Definition 10.3 Let X and Y be two random variables. By E(X|Y ) we mean a function
of Y that is defined to be E(X | Y = y) when Y = y .
Recall that a function of a random variable Y, say h(Y ), is defined to be h(a) at all sample
points at which Y = a. For example, if Z = log Y, then at a sample point ω, where Y (ω) = a,
we have that Z(ω) = log a. In this definition E(X | Y ) is a function of Y, which at a sample
point ω is defined to be E(X | Y = y), where y = Y (ω). Since E(X|Y ) is defined only
when pY (y) > 0, it is defined at all sample points ω, where pY (y) > 0, y = Y (ω). E(X|Y ),
being a function of Y, is a random variable. Its expected value, whenever finite, is equal to
the expected value of X . This extremely important property sometimes enables us to calculate
expectations which are otherwise, if possible, very difficult to find. To see why the expected
value of E(X|Y ) is E(X), let X and Y be discrete random variables with sets of possible
values A and B, respectively. Let p(x, y) be the joint probability mass function of X and Y .
On the one hand,
E(X) =
X
xpX (x) =
x∈A
=
XX
y∈B x∈A
X
x
x∈A
xp(x, y) =
X
p(x, y)
y∈B
XX
xpX|Y (x|y)pY (y)
y∈B x∈A
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 456 — #472
✐
✐
456
Chapter 10
More Expectations and Variances
=
XX
y∈B
=
X
y∈B
x∈A
xpX|Y (x|y) pY (y)
E(X | Y = y)P (Y = y),
(10.13)
showing that E(X) is the weighted average of the conditional expectations of X (given Y = y )
over all possible values of Y . On the other hand, we know that, for a real-valued function h,
X
E h(Y ) =
h(y)P (Y = y).
y∈B
Applying this formula to h(Y ) = E(X|Y ) yields
X
E E(X|Y ) =
E(X | Y = y)P (Y = y).
(10.14)
y∈B
Comparing (10.13) and (10.14), we obtain
X
E(X) = E E(X|Y ) =
E(X | Y = y)P (Y = y).
(10.15)
y∈B
Therefore, E E(X|Y ) = E(X) is a condensed way to say that E(X) is the weighted average of the conditional expectations of X (given Y = y ) over all possible values of Y . We have
proved the following theorem for the discrete case.
Theorem 10.6
Let X and Y be two random variables. Then
E E(X | Y ) = E(X).
Proof: We have already proven this for the discrete case. Now we will prove it for the
case where X and Y are continuous random variables with joint probability density function,
f (x, y).
Z ∞
E E(X|Y ) =
E(X|Y = y)fY (y) dy
−∞
=
Z ∞ Z ∞
−∞
−∞
xfX|Y (x|y) dx fY (y) dy
Z ∞ Z ∞
=
x
fX|Y (x|y)fY (y) dy dx
−∞
−∞
−∞
−∞
Z ∞ Z ∞
f (x, y)
fY (y) dy dx
=
x
−∞
−∞ fY (y)
Z ∞ Z ∞
=
x
f (x, y) dy dx
=
Z ∞
xfX (x) dx = E(X). −∞
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 457 — #473
✐
✐
Section 10.4
Conditioning on Random Variables
457
Example 10.18 Suppose that N (t), the number of people who pass by a museum at or
prior to t, is a Poisson process having rate λ. If a person passing by enters the museum
with probability p, what is the expected number of people who enter the museum at or prior
to t?
Solution: Let M (t) denote the number of people who enter the museum at or prior to t; then
∞
X
E M (t) = E E M (t) | N (t) =
E M (t) | N (t) = n P N (t) = n .
n=0
Given that N (t) = n, the number of people who enter
the museum at or
prior to t, is a binomial
random variable with parameters n and p. Thus E M (t) | N (t) = n = np. Therefore,
∞
∞
X
X
(λt)n−1
e−λt (λt)n
−λt
= pe λt
= pe−λt λteλt = pλt. E M (t) =
np
n!
(n
−
1)!
n=1
n=1
Example 10.19
function
Find E(X|Y ).
Let X and Y be continuous random variables with joint probability density
3
(x2 + y 2 )
2
f (x, y) =
0
if 0 < x < 1,
0<y<1
otherwise.
Solution: E(X|Y ) is a random variable that is defined to be E(X | Y = y) when Y = y .
First, we calculate E(X | Y = y):
E(X | Y = y) =
Z 1
xfX|Y (x|y) dx =
0
Z 1
0
x
f (x, y)
dx,
fY (y)
where
fY (y) =
Z 1
0
3
1
3 2
(x + y 2 ) dx = y 2 + .
2
2
2
Hence
Z 1
Z 1
(3/2)(x2 + y 2 )
3(x2 + y 2 )
x
dx
=
dx
(3/2)y 2 + 1/2
3y 2 + 1
0
0
Z 1
3
3(2y 2 + 1)
.
= 2
(x3 + xy 2 ) dx =
3y + 1 0
4(3y 2 + 1)
E(X | Y = y) =
x
Thus, if Y = y, then
E(X | Y = y) =
3(2y 2 + 1)
.
4(3y 2 + 1)
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 458 — #474
✐
✐
458
Chapter 10
More Expectations and Variances
Now since the random variable E(X|Y ) coincides with E(X | Y = y) if Y = y, we have
E(X|Y ) =
3(2Y 2 + 1)
. 4(3Y 2 + 1)
Example 10.20 What is the expected number of random digits that should be generated to
obtain three consecutive zeros?
Solution: Let X be the number of random digits to be generated until three consecutive zeros
are obtained. Let Y be the number of random digits to be generated until the first nonzero digit
is obtained. Then
∞
X
E(X) = E E(X|Y ) =
E(X | Y = i)P (Y = i)
i=1
=
3
X
i=1
=
E(X | Y = i)P (Y = i) +
∞
X
i=4
E(X | Y = i)P (Y = i)
3
∞
1 i−1 9 X
1 i−1 9 X
i + E(X)
+
,
3
10
10
10
10
i=1
i=4
which gives
E(X) = 1.107 + 0.999 E(X) + 0.003.
Solving this for E(X), we find that E(X) = 1110.
Example 10.21 Let X and Y be two random variables and f be a real-valued function from
R to R. Prove that
E f (Y )X | Y = f (Y )E(X|Y ).
Proof: If Y = y, then f (Y )E(X|Y ) is f (y)E(X | Y = y). We show that E f (Y )X|Y
is also equal to this quantity. Let fX|Y (x|y) be the conditional probability density function of
X given that Y = y; then
Z ∞
f (y)xfX|Y (x|y) dx
E f (Y )X | Y = y = E f (y)X|Y = y =
= f (y)
Z ∞
−∞
−∞
xfX|Y (x|y) dx = f (y)E(X | Y = y). Suppose that a certain airplane breaks down N times a year, where N is a random variable.
If the repair time for the ith breakdown is Xi , then the total repair time for this airplane is
the random variable X1 + X2 + · · · + XN . To find the expected length
of time that, due to
PN
breakdowns, the plane cannot fly, we need to calculate E
X
.
What
is different about
i
i=1
this sum is that, not only is each of its terms a random variable, but the number of its terms is
also a random variable. The expected values of such sums are calculated using the following
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 459 — #475
✐
✐
Section 10.4
Conditioning on Random Variables
459
theorem, discovered by Abraham Wald, a statistician who is best known for developing the
theory of sequential statistical procedures during World War II.
Theorem 10.7 (Wald’s Equation) Let X1 , X2 , . . . be independent and identically
distributed random variables with the finite mean E(X). Let N > 0 be an integer-valued
random variable, independent of {X1 , X2 , . . .}, with E(N ) < ∞. Then
E
N
X
i=1
By (10.15),
Proof:
E
Xi = E(N )E(X).
N
X
i=1
N
N
∞
X
h X
i X
Xi N = n P (N = n),
E
Xi = E E
Xi N =
i=1
n=1
(10.16)
i=1
where
E
N
X
Xi N = n = E
i=1
n
X
i=1
Xi N = n = E
n
X
Xi =
i=1
n
X
E(Xi ) = nE(X),
i=1
since N is independent of {X1 , X2 , . . .}. Hence, by (10.16),
E
N
X
i=1
∞
∞
X
X
Xi =
nE(X)P (N = n) = E(X)
nP (N = n) = E(X)E(N ). n=1
n=1
Example 10.22 Suppose that the average number of breakdowns for a certain airplane is
12.5 times a year. If the expected value of repair time is 7 days for each breakdown, and if the
repair times are identically distributed, independent random variables, find the expected total
repair time. Assume that repair times are independent of the number of breakdowns.
Solution: Let N be the number of breakdowns in a year and Xi be the repair time for the ith
breakdown. Then, by Wald’s equation, the expected total repair time is
E
N
X
i=1
Xi = E(N )E(Xi ) = (12.5)(7) = 87.5. The following theorem gives a formula, analogous to Wald’s equation, for variance. We leave
its proof as an exercise.
Theorem 10.8
Let {X1 , X2 , . . .} be an independent and identically distributed sequence
of random variables with finite mean E(X) and finite variance Var(X). Let N > 0 be
an integer-valued random variable independent of {X1 , X2 , . . .} with E(N ) < ∞ and
Var(N ) < ∞. Then
Var
N
X
i=1
2
Xi = E(N )Var(X) + E(X) Var(N ).
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 460 — #476
✐
✐
460
Chapter 10
More Expectations and Variances
We now explain a procedure for calculation of probabilities by conditioning on random
variables. Let B be an event associated with an experiment and X be a discrete random variable
with possible set of values A. Let
(
1
if B occurs
Y =
0
if B does not occur.
Then
E(Y ) = E E(Y |X) .
But
E(Y ) = 1 · P (B) + 0 · P (B c ) = P (B)
(10.17)
(10.18)
and
X
E E(Y |X) =
E(Y | X = x)P (X = x)
x∈A
=
X
x∈A
P (B | X = x)P (X = x),
(10.19)
where the last equality follows since
E(Y | X = x) = 1 · P (Y = 1 | X = x) + 0 · P (Y = 0 | X = x)
=
P (B and X = x)
P (Y = 1, X = x)
=
= P (B | X = x).
P (X = x)
P (X = x)
Relations (10.17), (10.18), and (10.19) imply the following theorem.
Theorem 10.9 Let B be an arbitrary event and X be a discrete random variable with
possible set of values A; then
P (B) =
X
x∈A
P (B | X = x)P (X = x).
If X is a continuous random variable, the relation analogous to (10.20) is
Z ∞
P (B) =
P (B | X = x)f (x) dx,
(10.20)
(10.21)
−∞
where f is the probability density function of X .
Theorem 10.9 shows that the probability of an event B is the weighted average of the
conditional probabilities of B (given X = x) over all possible values of X . For the discrete
case, this is a generalization of Theorem 3.4 and is the most general version of the law of total
probability.
Example 10.23 The time between consecutive earthquakes in Los Angeles and the time between consecutive earthquakes in San Francisco are independent and exponentially distributed
with means 1/λ and 1/µ, respectively. What is the probability that the next earthquake occurs
in Los Angeles?
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 461 — #477
✐
✐
Section 10.4
Conditioning on Random Variables
461
Solution: Let X and Y denote the times between now and the next earthquake in
Los Angeles and San Francisco, respectively. Because of the memoryless property of exponential distribution, X and Y are exponentially distributed with means 1/λ and 1/µ, respectively.
To calculate P (X < Y ), the desired probability, we will condition on X :
Z ∞
P (X < Y ) =
P (X < Y | X = x)λe−λx dx
0
=
Z ∞
−λx
P (Y > x)λe
Z ∞
dx =
0
=λ
Z ∞
e−(λ+µ)x dx =
0
where
Z ∞
e−µx λe−λx dx
0
λ
λ+µ
Z ∞
(λ + µ) e−(λ+µ)x dx =
0
λ
,
λ+µ
(λ + µ) e−(λ+µ)x dx = 1 since its integrand is the probability density function of
0
an exponential random variable with parameter λ + µ, and P (Y > x) is calculated from
P (Y > x) = 1 − P (Y ≤ x) = 1 − (1 − e−µx ) = e−µx . Remark 10.2 In Example 10.23, we have shown the following important property of exponential random variables.
Let X and Y be independent exponential random variables with parameters λ and µ, respectively. Then
λ
. P (X < Y ) =
λ+µ
Example 10.24 Suppose that Z1 and Z2 are independent standard normal random variables. Show that the ratio Z1 /|Z2 | is a Cauchy random variable. That is, Z1 /|Z2 | is a random
variable with the probability density function
f (t) =
1
,
π(1 + t2 )
−∞ < t < ∞.
Solution: Let g(x) be the probability density function of |Z2 |. To find g(x), note that, for
x ≥ 0,
Z x
Z x
2
1 −u2 /2
1
√ e
√ e−u /2 du.
du = 2
P |Z2 | ≤ x = P (−x ≤ Z2 ≤ x) =
2π
2π
0
−x
Hence
g(x) =
2
d
2
P |Z2 | ≤ x = √ e−x /2 ,
dx
2π
x ≥ 0.
To find the probability density function of Z1 /|Z2 |, note that, by Theorem 10.9,
P
Z
Z
1
≤t =1−P
> t = 1 − P Z1 > t|Z2 |
|Z2 |
|Z2 |
Z ∞
2
2
=1−
P Z1 > t|Z2 | |Z2 | = x √ e−x /2 dx
2π
0
1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 462 — #478
✐
✐
462
Chapter 10
More Expectations and Variances
Z ∞
2
2
P (Z1 > tx) √ e−x /2 dx
2π
0
Z ∞ Z ∞
2
2
2
1
√ e−u /2 du √ e−x /2 dx.
=1−
2π
2π
0
tx
=1−
Now, by the fundamental theorem of calculus,
d
dt
Z ∞
tx
2
1
d
√ e−u /2 du =
1−
dt
2π
Z tx
2
2 2
1
1
√ e−u /2 du = −x √ e−t x /2 .
2π
2π
−∞
Therefore,
d Z1
P
≤t =
dt
|Z2 |
Z ∞
0
2 2
2
x
2
2
√ e−t x /2 · √ e−x /2 dx =
2π
2π
2π
Making the change of variable y = (1 + t2 )x2 /2 yields
Z ∞
d Z1
1
1
P
≤t =
e−y dy =
,
2
dt
|Z2 |
π(1 + t ) 0
π(1 + t2 )
Z ∞
2
2
xe−(t +1)x /2 dx.
0
−∞ < t < ∞. Example 10.25 Let X be the natural lifetime of a device that also fails if a catastrophe such
as a shock occurs. Let Y be the time until the next catastrophe. Suppose that X and Y are
independent exponential random variables with parameters λ and µ, respectively. Given that
X < Y, find the expected lifetime of the device.
Warning: Since X < Y, we might fallaciously think that the expected lifetime of the device
is 1/λ.
Solution: Let F be the distribution function of the period in which the device remains operative. We have
P (X ≤ t, X < Y )
,
F (t) = P (X ≤ t | X < Y ) =
P (X < Y )
where,
P (X ≤ t, X < Y ) = P X < min{t, Y }
Z ∞
P X < min{t, Y } | X = x fX (x) dx
=
0
=
Z ∞
0
=
Z t
0
=
Z t
P min{t, Y } > x | X = x fX (x) dx
P (Y > x | X = x)fX (x) dx
P (Y > x)fX (x) dx =
0
=
Z t
e−µx λe−λx dx
0
λ 1 − e−(λ+µ)t ,
λ+µ
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 463 — #479
✐
✐
Section 10.4
Conditioning on Random Variables
463
and, by Remark 10.2,
P (X < Y ) =
Thus
λ
.
λ+µ
λ 1 − e−(λ+µ)t
λ+µ
= 1 − e−(λ+µ)t .
λ
λ+µ
This shows that, given X < Y, the lifetime of the device is exponential with parameter λ + µ.
Therefore,
1
. E(X | X < Y ) =
λ+µ
P (X ≤ t | X < Y ) =
Remark 10.3 By Example 10.25, we have established the following important property of
exponential random variables.
Let X and Y be independent exponential random variables with parameters λ and µ, respectively. Given that X < Y, the distribution function of X is exponential with parameter λ + µ.
Hence
1
E(X | X < Y ) =
. λ+µ
Let B be an event of a sample space with P (B) > 0, and X be a random variable defined
d
on the same sample space. Let F (t) = P (X ≤ t | B), and f (t) =
F (t). Then, by
dt
Z ∞
definition, E(X | B) =
xf (t) dt. The following is an important tool for calculating
−∞
E(X) in terms of E(X|B) and E(X|B c ).
Theorem 10.10 Let B be an event of a sample space with P (B) > 0 and P (B c ) > 0. Let
X be a random variable defined on the same sample space. Then
E(X) = E(X|B)P (B) + E(X|B c )P (B c ).
Proof: Let N be a discrete random variable with the set of possible values {0, 1} defined
as follows: N = 0 if and only if B occurs, and N = 1 if and only if B c occurs. Then, by
Theorem 10.6,
E(X) = E E(X|N )
= E(X | N = 0)P (N = 0) + E(X | N = 1)P (N = 1)
= E(X|B)P (B) + E(X|B c )P (B c ). The following is a natural generalization of Theorem 10.10.
Theorem 10.11 Let {B1 , B2 , . . . , Bn } be a partition of the sample space of an experiment
with P (Bi ) > 0 for i = 0, 1, . . . , n. Let X be a random variable defined on the same sample
space. Then
n
X
E(X) =
E(X|Bi )P (Bi ).
i=1
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 464 — #480
✐
✐
464
Chapter 10
More Expectations and Variances
Let X and Y be two given random variables. Define the new random variable
Var(X|Y ) by
2
Var(X|Y ) = E X − E(X|Y ) | Y .
2
Then the formula analogous to Var(X) = E(X 2 ) − E(X) is given by
Var(X|Y ) = E(X 2 |Y ) − E(X|Y )2 .
(10.22)
(See Exercise 26.) The following theorem shows that Var(X) is the sum of the expected value
of Var(X|Y ) and the variance of E(X|Y ).
Theorem 10.12
Proof:
Var(X) = E Var(X|Y ) + Var E[X|Y ] .
By (10.22),
E Var(X|Y ) = E E(X 2 |Y ) − E E(X|Y )2
= E(X 2 ) − E E(X|Y )2 .
By the definition of variance,
2
Var E[X|Y ] = E E(X|Y )2 − E E(X|Y )
2
= E E(X|Y )2 − E(X) .
Adding these two equations, we have the theorem.
Example 10.26 A fisherman catches fish in a large lake with lots of fish, at a Poisson rate
of two per hour. If, on a given day, the fisherman spends randomly anywhere between 3 and 8
hours fishing, find the expected value and the variance of the number of fish he catches.
Solution: Let X be the number of hours the fisherman spends fishing. Then X is a uniform
random variable over the interval (3, 8). Label the time the fisherman begins fishing on the
given day at t = 0. Let N (t) denote the total number of fish caught at or prior to t. Then
N (t) : t ≥ 0 is a Poisson process with parameter λ = 2. Assuming that X is independent
of N (t) : t ≥ 0 , we have
This implies that
Therefore,
E N (X) | X = t = E N (t) = 2t.
E N (X) | X = 2X.
8+3
= 11.
E N (X) = E E N (X) | X = E(2X) = 2E(X) = 2 ·
2
Similarly,
Thus
Var N (X) | X = t = Var N (t) = 2t.
Var N (X) | X = 2X.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 465 — #481
✐
✐
Section 10.4
Conditioning on Random Variables
465
By Theorem 10.12,
Var N (X) = E Var N (X) | X + Var E N (X) | X
= E(2X) + Var(2X) = 2E(X) + 4 Var(X)
=2·
8+3
(8 − 3)2
+4·
= 19.33. 2
12
EXERCISES
A
1.
A fair coin is tossed until two tails occur successively. Find the expected number of the
tosses required.
Hint: Let
(
1
if the first toss results in tails
X=
0
if the first toss results in heads,
and condition on X .
2.
3.
The orders received for grain by a farmer add up to X tons, where X is a continuous
random variable uniformly distributed over the interval (4, 7). Every ton of grain sold
brings a profit of a, and every ton that is not sold is destroyed at a loss of a/3. How
many tons of grain should the farmer produce to maximize his expected profit?
Hint: Let Y (t) be the profit if the farmer produces t tons of grain. Then
i
h
a
E Y (t) = E aX − (t − X) P (X < t) + E(at)P (X ≥ t).
3
In a box, Lynn has b batteries of which d are dead. She tests them randomly and one by
one. Every time that a good battery is drawn, she will return it to the box; every time
that a dead battery is drawn, she will replace it by a good one.
(a)
Determine the expected value of the number of good batteries in the box after n
of them are checked.
(b)
Determine the probability that on the nth draw Lynn draws a good battery.
Hint: Let Xn be the number of good batteries in the box after n of them are checked.
Show that
1
Xn−1 .
E(Xn | Xn−1 ) = 1 + 1 −
b
Then, by computing the expected value of this random variable, find a recursive relation
between E(Xn ) and E(Xn−1 ). Use this relation and induction to prove that
1 n
E(Xn ) = b − d 1 −
.
b
Note that n should approach ∞ to get E(Xn ) = b. For part (b), let En be the event that
on the nth draw she gets a good battery. By conditioning on Xn−1 prove that P (En ) =
E(Xn−1 )/b.
✐
✐
✐
✐
✐
✐
“K27443” — 2018/7/13 — 14:41 — page 466 — #482
✐
✐
466
4.
Chapter 10
More Expectations and Variances
For given random variables Y and Z, let
(
Y with probability p
X=
Z with probability 1 − p.
Find E(X) in terms of E(Y ) and E(Z).
5.
If a car is under one year old and is totaled, an insurance company replaces that car
with a new one. An actuary has calculated that when a $60,000-car gets involved in an
accident, the probability of total loss is 0.015 and the probability of partial loss is 0.038.
She has also approximated the probability density function of X, the cost to repair partial
damage, in thousands of dollars, to be
1
e−x/8
0 < x < 60
f (x) = 8
0
otherwise,
where f is only approximately a probability density function. Find the expected value
of claim payment per accident for cars that are not yet one year old, were purchased
for $60,000, and are insured by this insurance company under policies that have no
deductibles.
6.
The lifetime of a machine, in years, is a uniform random variable over the interval (0, 7).
The machine will be replaced at age 5 or when it fails, if that occurs before age 5. Find
the expected value of the age of the machine at the time of replacement.
7.
A typist, on average, makes three typing errors in every two pages. If pages with more
than two errors must be retyped, on average how many pages must she type to prepare
a report of 200 pages? Assume that the number of errors in a page is a Poisson random
variable. Note that some of the retyped pages should be retyped, and so on.
Hint: Find p, the probability that a page should be retyped. Let Xn be the number
of pages that should be typed at least n times. Show that E(X
E(X2 ) =
P1∞) = 200p,
200p2 , . . . , E(Xn ) = 200pn . The desired quantity is E
X
,
which
can be
i
i=1
calculated using relation (10.2).
8.
In data communication, usually messages sent are combinations of characters, and each
character consists of a number of bits. A bit is the smallest unit of information and is
either 1 or 0. Suppose that the length of a character (in bits) is a geometric random
variable with parameter p. Suppose that a message is combined of K characters, where
K is a random variable with mean µ and variance σ 2 . If the lengths of characters of a
message are independent of each other and of K, and if it takes
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )