See discussions, stats, and author profiles for this publication at: https://www.researchgate.net/publication/280311472
The Ubiquitous Ewens Sampling Formula
Article in Statistical Science · February 2016
DOI: 10.1214/15-STS529
CITATIONS
READS
10
400
1 author:
Harry Crane
Rutgers, The State University of New Jersey
32 PUBLICATIONS 102 CITATIONS
SEE PROFILE
All content following this page was uploaded by Harry Crane on 02 April 2016.
The user has requested enhancement of the downloaded file. All in-text references underlined in blue are added to the original document
and are linked to publications on ResearchGate, letting you access and read them immediately.
Statistical Science
2016, Vol. 31, No. 1, 1–19
DOI: 10.1214/15-STS529
© Institute of Mathematical Statistics, 2016
The Ubiquitous Ewens Sampling Formula1
Harry Crane
Abstract. Ewens’s sampling formula exemplifies the harmony of mathematical theory, statistical application, and scientific discovery. The formula
not only contributes to the foundations of evolutionary molecular genetics,
the neutral theory of biodiversity, Bayesian nonparametrics, combinatorial
stochastic processes, and inductive inference but also emerges from fundamental concepts in probability theory, algebra, and number theory. With an
emphasis on its far-reaching influence throughout statistics and probability,
we highlight these and many other consequences of Ewens’s seminal discovery.
Key words and phrases: Ewens’s sampling formula, Poisson–Dirichlet
distribution, random partition, coalescent process, inductive inference, exchangeability, logarithmic combinatorial structures, Chinese restaurant process, Ewens–Pitman distribution, Dirichlet process, stick breaking, permanental partition model, cyclic product distribution, clustering, Bayesian nonparametrics, α-permanent.
to
each allelic partition (m1 , . . . , mn ) for which
n
j =1 j · mj = n.
Derived under the assumption of selective neutrality,
equation (1) is the null hypothesis distribution needed
to test the controversial neutral theory of evolution. Of
the formula and its consequences, Ewens begins his
abstract matter-of-factly, “In this paper a beginning is
made on the sampling theory of neutral alleles” [36],
page 87. Though obviously aware of its significance to
the statistical theory of neutral sampling, Ewens could
not have foreseen far-reaching contributions to the unified neutral theory of biodiversity [50], nonparametric Bayesian inference [2, 39], combinatorial stochastic
processes [60, 73], and the philosophy of inductive inference [93], not to mention fundamental connections
to the determinant function [19], Macdonald polynomials [27], and prime factorization [10, 29] in algebra
and number theory. In the coming pages, we present
Ewens’s sampling formula in all its glory, highlighting
each of these connections in further detail.
1. INTRODUCTION
In 1972, Warren Ewens [36] derived a remarkable
formula for the sampling distribution of allele frequencies in a population undergoing neutral selection.
An allele is a type of gene, for example, the alleles A, B, and O in the ABO blood group, so that
each gene has a particular allelic type and a sample
of n = 1, 2, . . . genes can be summarized by its allelic
partition (m1 , . . . , mn ), where m1 is the number of alleles appearing exactly once, m2 is the number of alleles appearing exactly twice, and in general mj is the
number of alleles appearing exactly j times, for each
j = 1, 2, . . . , n. Ewens’s sampling formula (ESF) with
parameter θ > 0 assigns probability
(1)
p(m1 , . . . , mn ; θ )
=
n
θ mj
n!
θ (θ + 1) · · · (θ + n − 1) j =1 j mj mj !
1.1 Outline
Harry Crane is Assistant Professor of Statistics &
Biostatistics, Rutgers, the State University of New Jersey,
110 Frelinghuysen Road, Room 501, Piscataway, New
Jersey 08854, USA (e-mail: hcrane@stat.rutgers.edu).
We retrace the development of Ewens’s sampling
formula, from neutral allele sampling and Kingman’s
mathematical theory of genetic diversity [60–62], to
modern nonparametric Bayesian [2, 39] and frequentist
[21] statistical methods, and backward in time to the
1 Discussed in 10.1214/15-STS535, 10.1214/15-STS536,
10.1214/15-STS537, 10.1214/15-STS538 and
10.1214/15-STS540; rejoinder at 10.1214/15-STS544.
1
2
H. CRANE
roots of probabilistic reasoning and inductive inference
[8, 24, 54]. In between, Pitman’s [73, 76] investigation
of exchangeable random partitions and combinatorial
stochastic processes unveils many more surprising connections between Ewens’s sampling formula and classical stochastic process theory, while other curious appearances in algebra [19, 27] and number theory [10,
29] only add to its mystique.
1.2 Historical Context
Elements of Ewens’s sampling formula appeared in
Yule’s prior work [91] on the evolution of species,
which Champernowne [16] and Simon [79] later recast
in the context of income distribution in large populations and word frequencies in large pieces of text, respectively. Shortly after Ewens, Antoniak [2] independently rediscovered (1) while studying Dirichlet process priors in Bayesian statistics. We recount various
other historical aspects of Ewens’s sampling formula
in the coming pages.
1.3 Relation to Prior Work
Ewens and Tavaré [35, 82] have previously reviewed
various structural properties of and statistical applications involving the Ewens sampling formula. The
present survey provides updated and more detailed
coverage of a wider range of topics, which we hope
will serve as a handy reference for experts and an eyeopening introduction for beginners.
2. NEUTRAL ALLELE SAMPLING AND SYNOPSIS
OF EWENS’S 1972 PAPER
2.1 The Neutral Wright–Fisher Evolutionary Model
Population genetics theory studies the evolution of a
population through changes in allele frequencies. Because many random events contribute to these changes,
the theory relies on stochastic models. The simplest of
these is the Wright–Fisher model, following its independent introduction by Fisher [40] and Wright [90].
The Wright–Fisher model concerns a diploid population, that is, a bi-parental population in which every
individual has two genes, one derived from each parent, at any gene locus. The population is assumed to be
of fixed size N , so that there are 2N genes at each gene
locus in any generation. The generic Wright–Fisher
model assumes that the 2N genes in any offspring generation are obtained by sampling uniformly with replacement from the 2N genes in the parental generation.
A gene comprises a sequence of DNA nucleotides,
where is typically on the order of several thousand,
and thus each gene has exactly one of the possible 4
allelic types. The large number of allelic types motivates the infinitely many alleles assumption, by which
each transmitted gene is of the same type as its parental
gene with probability 1 − u and mutates independently
with probability u to a new allelic type “not currently
existing (nor previously existing) in the population”
[36], page 88.
Under these assumptions, Ewens [36] derives equation (1) by a partly heuristic argument made precise by
Karlin and McGregor [55]. The parameter θ in Ewens’s
sampling formula is equal to 4Nu and, therefore, admits an interpretation in terms of the mutation rate. In
more general applications, θ is best regarded as an arbitrary parameter.
Ewens goes on to discuss both deductive and inductive questions about the sample and population and, in
the latter half of [36], addresses issues surrounding statistical inference and hypothesis testing in population
genetics. Among all these contributions, Ewens’s discussion in the opening pages about the mean number
of alleles and the distribution of allele frequencies has
had the most lasting impact.
2.2 The Mean Number of Alleles
Equation (1) leads directly to the probability distribution of the number of different alleles K in the sample as
P{K = k} = Snk θk
,
θ (θ + 1) · · · (θ + n − 1)
where Snk is the (n, k)-Stirling number of the first kind
[80]:A008275. Together with equation (1), the above
expression implies that K is a sufficient statistic for θ ,
a point discussed later, and also leads to the expression
θ
θ
θ
+
+ ··· +
∼ θ log(n)
θ θ +1
θ +n−1
for the mean of K, where a ∼ b indicates that a/b → 1
as n → ∞. Contrasting this with the mean number of
distinct alleles in a population of size N ,
θ
θ
θ
+
+ ··· +
∼ θ log(2N).
θ θ +1
θ + 2N − 1
Ewens approximates the mean number of different alleles in the population that do not appear in the sample
by θ log(2N/n).
3
THE UBIQUITOUS EWENS’S SAMPLING FORMULA
2.3 Predictive Probabilities
Invoking “a variant of the ‘coupon collector’s problem’ (or the ‘law of succession’)” [36], page 94, Ewens
calculates the probability that “the (j + 1)th gene
drawn is of an allelic type not observed on the first j
draws” as θ/(θ + j ). Conversely, the probability that
the (j + 1)st gene is of a previously observed allelic
type is j/(θ + j ). These predictive probabilities precede Dubins and Pitman’s Chinese restaurant seating
rule (Section 4) and posterior probabilities based on
Ferguson’s Dirichlet process prior (Section 6) and are
also closely associated with De Morgan’s rule of succession for inductive questions (Section 7).
More than describing the allelic configuration of
genes at any given time, the neutral Wright–Fisher
model describes the evolution of a population of size
N with nonoverlapping generations. Under this evolution, equation (1) acts as the stationary distribution
for a sample of n genes from a large population. In an
effort to better understand how (1) arises from these
dynamics, Griffiths [45] and Kingman [63] considered
the behavior of allelic frequencies under the infinite
population diffusion approximation. Perhaps the most
major advance in evolutionary population genetics in
the last three decades, Kingman’s coalescent has resulted in the widely adopted coalescent theory within
population genetics [87] as well as the mathematical
study of partition-valued and more general combinatorial stochastic processes [9, 76].
Under the dynamics of Section 2.1, each of the 2N
genes in the current generation has a parent gene in the
previous generation. Tracing these parental relationships backward in time produces the ancestral lineages
of the current 2N genes. Thus, the number of children
genes X of a typical gene is a Binomial random variable with success probability 1/(2N), that is,
(2)
N−k
2N
P{X = k} =
(2N)−k 1 − (2N)−1
,
k
k = 0, 1, . . . , 2N,
and so the probability that two genes in the current
generation have the same parent is 1/(2N). More generally, the number of generations Y for which the
ancestral lines of two genes have been distinct follows the Geometric distribution with success probability 1/(2N), that is,
(3)
k 1 − 1/(2N)
k
k
1 − 2/(2N) · · · 1 − ( − 1)/(2N) .
The coalescent process arises as a natural infinite
population diffusion approximation to the Wright–
Fisher model by taking N → ∞ and putting k =
2Nt for t ≥ 0. Under this regime, we obtain the limiting probability that the lineages of genes remain
distinct for t ≥ 0 time units as
lim
N→∞,k/(2N)→t
−1
1 − j/(2N)
k
j =1
= exp −t( − 1)/2 .
2.4 The Coalescent
From (2) and (3), the probability that the ancestral lines
of genes have been distinct for k generations is
k
P{Y ≥ k} = 1 − (2N)−1 ,
k = 0, 1, . . . .
Realizing that this equals the distribution of the mini mum of 2 independent standard Exponential random
variables, Kingman [63] arrived at his description of
the coalescent, according to which distinct lineages
merge independently at the times of standard Exponential random variables.
Although we have focused mainly on its original
derivation in the context of the Wright–Fisher model,
Ewens’s sampling formula applies for a wide range
of neutral models [59]. In their treatment of Macdonald polynomials (Section 8.3), Diaconis and Ram [27]
attribute these “myriad practical appearances [to] its
connection with Kingman’s coalescent process [. . . ].”
Ethier and Griffiths’s [31] formula for the transition
function of the Fleming–Viot process [42] provides
yet another link between the coalescent, Dirichlet processes, the Poisson–Dirichlet distribution, and Ewens’s
sampling formula.
2.5 Legacy in Theoretical Population Genetics
The sufficiency of K for θ led Ewens [36] and
Watterson [89, 88] to objective tests of the controversial neutral theory of evolution [58]. Plainly, sufficiency implies that the conditional distribution of
(m1 , . . . , mn ) given K is independent of θ , so that
the unknown parameter θ need not be estimated and
is not involved in any test based on the conditional distribution of (m1 , . . . , mn ) given K. Beyond the realm
of Ewens’s seminal work on testing, Christiansen [17]
states that Ewens [36] “laid the foundations for modern molecular population genetics.” Nevertheless, in
the early 1980s Ewens shifted his focus to mapping
genes associated to human diseases. He is partly responsible for the transmission-disequilibrium test [81],
which has been used to locate at least fifty such
4
H. CRANE
genes. The paper [81] that introduced the transmissiondisequilibrium test was chosen as one of the top ten
classic papers in the American Journal of Human Genetics, 1949–2014.
In general biology, Ewens’s paper marks a seminal contribution to Hubbell’s neutral theory of biodiversity [50], and the foundations it laid have been refined by novel sampling formulas [32, 34] and statistical tests for neutrality [33] in ecology. For the rest
of the paper, we focus on applications of equation (1)
in other areas, mentioning other biological applications
only briefly.
3. CHARACTERISTIC PROPERTIES OF EWENS’S
SAMPLING FORMULA
3.1 Partition Structures
3.2 Random Set Partitions
In many respects, partition structures are more naturally developed as distributions on partitions of the set
[n] = {1, . . . , n}. Instead of summarizing the sample of
genes by the allelic partition (m1 , . . . , mn ), we can label genes distinctly 1, . . . , n and assign each an allelic
type. At each generation, the genes segregate into subsets B1 , . . . , Bk , called blocks, such that i and j in the
same block indicates that genes i and j have the same
allelic type. The resulting collection π = {B1 , . . . , Bk }
of nonempty, disjoint subsets with B1 ∪ · · · ∪ Bk = [n]
is a (set) partition of [n].
The ordering of B1 , . . . , Bk in π is inconsequential,
so we follow convention and list blocks in ascending
order of their smallest element. For example, there are
five partitions of the set {1, 2, 3},
From an allelic partition (m1 , . . . , mn ) of n ≥ 1, we
obtain an allelic partition (m1 , . . . , mn−1 ) of n − 1 by
choosing J randomly with probability
(4)
P J = j |(m1 , . . . , mn ) = j · mj /n,
and putting
(5)
j = 1, . . . , n,
⎧
⎪
⎨ mj − 1,
mj = mj + 1,
⎪
⎩
mj ,
J = j,
J = j + 1,
otherwise.
{1, 2, 3} ,
{1, 3}, {2} ,
{1, 2}, {3} ,
{1}, {2}, {3} ,
but only three allelic partitions of size 3,
(3, 0, 0),
(1, 1, 0),
(0, 0, 1).
Each set partition π = {B1 , . . . , Bk } corresponds
uniquely to an allelic partition n(π) = (m1 , . . . , mn ),
where mj counts the number of blocks of size j in π ,
for example,
n {1, 2, 3}
Alternatively, (5) is the allelic partition obtained by
choosing a gene uniformly at random and removing it
from the sample.
n {1, 2}, {3}
T HEOREM 3.1 (Kingman [60, 61]). Let (m1 , . . . ,
mn ) be a random allelic partition from Ewens’s sampling formula (1) with parameter θ > 0. Then (m1 , . . . ,
mn−1 ) obtained as in (5) is also distributed according
to Ewens’s sampling formula with parameter θ > 0.
n {1}, {2}, {3}
We call a family of distributions (pn )n≥1 a partition structure if (m1 , . . . , mn ) ∼ pn implies (m1 , . . . ,
mn−1 ) ∼ pn−1 for all n ≥ 2. This definition implies
that (pn )n≥1 is consistent under uniform deletion of
any number of genes and, thus, agrees with Kingman’s
original definition [60], which generalizes the outcome
in Theorem 3.1.
Kingman’s study of partition structures anticipates
his paintbox process correspondence, by which he
proves a de Finetti-type representation for all infinite
exchangeable random set partitions. Kingman’s correspondence establishes a link between Ewens’s sampling formula and the Poisson–Dirichlet distribution,
which in turn broadens the scope of equation (1); see
Section 4.1 below.
{1}, {2, 3} ,
= (0, 0, 1),
= n {1}, {2, 3}
= n {1, 3}, {2} = (1, 1, 0) and
= (3, 0, 0).
Conversely, every allelic partition (m1 , . . . , mn ) corresponds to the set of all partitions π for which n(π) =
(m1 , . . . , mn ).
Because every individual draws its parent genes
independently and uniformly in the Wright–Fisher
model, set partitions corresponding to the same allelic partition occur with the same probability. Consequently, we can generate a random set partition
n by first drawing an allelic partition (m1 , . . . , mn )
from Ewens’s sampling formula (1) and then selecting uniformly among partitions π for which n(π) =
(m1 , . . . , mn ). The resulting random set partition n
follows the so-called Ewens distribution with parameter θ > 0,
(6)
P n = {B1 , . . . , Bk }
k
θk
=
(#Bj − 1)!,
θ (θ + 1) · · · (θ + n − 1) j =1
5
THE UBIQUITOUS EWENS’S SAMPLING FORMULA
where #Bj is the cardinality of block Bj for each j =
1, . . . , k.
Simple enumeration and the law of total probability connects (1) and (6). As an exercise, the reader can
verify that each allelic partition (m1 , . . . , mn ) corre
sponds to n!/ nj=1 j !mj mj ! set partitions through n
and, therefore, the conditional distribution of a partition drawn uniformly among all π whose block sizes
form allelic partition (m1 , . . . , mn ) is
P n = π|n(n ) = (m1 , . . . , mn )
=
n
1 j !mj mj !,
n! j =1
3.4 Consistency Under Subsampling
Above all, Kingman’s definition of a partition structure emphasizes the “need for consistency between different sample sizes” [61], page 374. By subsampling
[m] ⊂ [n], a partition π = {B1 , . . . , Bk } of [n] restricts
to a partition of [m] by
π|[m] = B1 ∩ [m], . . . , Bk ∩ [m] \ {∅}.
For example, the restrictions of π = {{1, 4, 7, 8},
{2, 3, 5}, {6}} to samples of size m = 7, 6, 5, respectively, are
π|[7] = {1, 4, 7}, {2, 3, 5}, {6} ,
π|[6] = {1, 4}, {2, 3, 5}, {6} ,
n(π) = (m1 , . . . , mn ).
3.3 Exchangeability
The distribution in (6) depends only on the allelic
partition induced by the block sizes of n . Therefore,
for any permutation σ : [n] → [n], the relabeling σn
obtained by first taking n distributed as in (6) and
then putting σ (i) and σ (j ) in the same block of σn
if and only if i and j are in the same block of n is
also distributed as in (6). As is natural in the genetics
setting, the labels 1, . . . , n distinguish between genes
but otherwise can be assigned arbitrarily.
More generally, a random partition n is exchangeable if its distribution is invariant under relabeling by
any permutation σ : [n] → [n], that is,
P{n = π} = P n = π σ
for all permutations σ : [n] → [n].
Pitman’s exchangeable partition probability function
(EPPF) [73] captures the notion of exchangeability
through a function Pn (m1 , . . . , mn ) on allelic partitions. In particular, n is exchangeable if and only if
there exists an EPPF Pn such that
P{n = π } = Pn n(π)
for all partitions of [n].
Exchangeability of the Ewens distribution is clear from
the closed-form expression in (6), which depends on
n only through its block sizes. In general, every probability distribution pn (·) on allelic partitions of n determines an EPPF Pn by
Pn n(π) = pn (m1 , . . . , mn ) ×
(7)
n
1 j !mj mj !,
n! j =1
n(π) = (m1 , . . . , mn ).
π|[5] = {1, 4}, {2, 3, 5} .
To satisfy the partition structure requirement, the
sampling distribution of m must coincide with the
marginal distribution of n|[m] , the partition induced
by subsampling m genes from a sample of size n ≥ m.
A family of random set partitions (n )n≥1 is consistent
under subsampling, or sampling consistent, if n|[m]
is distributed the same as m for all n ≥ m ≥ 1. Similarly, a family of distributions is sampling consistent
if it governs a consistent family of random partitions.
Any partition structure (pn )n≥1 determines a family of
exchangeable, consistent EPPFs through (7). In particular, Ewens’s sampling formula (1) determines a partition structure and the Ewens distribution in (6) is derived from (1) via (7); thus, the family of Ewens distributions is consistent under subsampling.
Sampling consistency elicits an interpretation of the
distributions of (n )n≥1 as the sampling distributions
induced by a data-generating process for the whole
population. Inductively, the sampling distributions of
(n )n≥1 permit a sequential construction: given n =
π , we generate n+1 from the conditional distribution
among all partitions of [n + 1] that restrict to π under
subsampling, that is,
P n+1 = π |n = π
=
P n+1 = π /P{n = π},
0,
= π,
π|[n]
otherwise.
Thus, a consistent family of partitions implies the existence of conditional distributions for predictive inference and the sequence (n )n≥1 determines a unique
infinite random partition ∞ of the positive integers
N = {1, 2, . . .}. In the special case of Ewens distribution, these predictive probabilities determine the Chinese restaurant process (Section 4.5).
6
H. CRANE
3.5 Self-Similarity
In the Wright–Fisher model, subpopulations exhibit
the same behavior as the population at large and different genes do not interfere with each other. Together, these comprise the statistical property of selfsimilarity.
Formally, set partitions are partially ordered by the
refinement relation: we write π ≤ π if every block
of π is a subset of some block of π . For example,
π = {{1, 2}, {3, 4}, {5}} refines π = {{1, 2, 5}, {3, 4}}
but not π = {{1, 3}, {2, 4, 5}}, because {1, 2} is not a
subset of {1, 3} or {2, 4, 5}. Let P (·) be an EPPF for
an infinite exchangeable random partition, so that P
determines an exchangeable probability distribution on
partitions of any finite size by P{n = π} = P (n(π)).
The family of random set partitions (n )n≥1 is selfsimilar if for all n = 1, 2, . . . and all set partitions π
of [n],
P n = π|n ≤ π =
b∈π (8)
P n(π|b ) ,
3.6 Noninterference
A longstanding question in ecology concerns interactions between different species. In this context, we
regard the elements 1, 2, . . . as labels for different specimens, instead of genes, so that the allelic partition
(m1 , . . . , mn ) counts the number of species that appear
once, twice, and so on in a sample. Given an arbitrary
partition structure (pn )n≥1 and (m1 , . . . , mn ) ∼ pn , we
choose an index J = r as in (4) and, instead of reducing to an allelic partition of n − 1 as in (5), we define
m∗ = (m∗1 , . . . , m∗n−r ) by
mj − 1,
mj ,
P{n = π}
(9)
= exp #π log θ −
×
n−1
log(θ + j )
j =0
(#b − 1)!,
b∈π
where #π denotes the number of blocks of π . Within
statistical physics, (9) can be rewritten in canonical
Gibbs form as
P{n = π} = Zn−1
ψ(#b),
b∈π
In other words, given that n is a refinement of π ,
the further breakdown of elements within each block
of π occurs independently of other blocks and with the
same distribution determined by P . The reader can verify that (6) satisfies condition (8).
The family of Ewens distributions on set partitions
can be expressed as an exponential family with a natural parameter log(θ ) and canonical sufficient statistic
for the number of blocks of n :
(10)
π ≤ π .
m∗j =
3.7 Exponential Families, Gibbs Partitions, and
Product Partition Models
J = j,
otherwise,
to obtain an allelic partition of n − r. In effect, we
remove all specimens with the same species as one
chosen uniformly at random among 1, . . . , n. If for
every n = 1, 2, . . . the conditional distribution of m∗
given J = r is distributed according to pn−r , the family (pn )n≥1 satisfies noninterference. Kingman [60]
showed that Ewens’s sampling formula (1) is the only
partition structure with the noninterference property.
for nonnegative constants ψ(k) = θ · (k − 1)! and
normalizing constant Zn = θ (θ + 1) · · · (θ + n − 1).
The following theorem distinguishes the Ewens family among this class of Gibbs distributions.
T HEOREM 3.2 (Kerov [56]). A family (Pn )n≥1 of
Gibbs distributions (10) with common weight sequence
{ψ(k)}k≥1 is exchangeable and consistent under subsampling if and only if there exists θ > 0 such that
ψ(k) = θ · (k − 1)! for all k ≥ 1.
Without regard for statistical or physical properties
of (10), Hartigan [46] proposed the class of product
partition models for certain statistical clustering applications. For a collection of cohesion functions c(b),
b ⊆ [n], the product partition model assigns probability
P{n = π} ∝
c(b)
b∈π
to each partition π of [n]. Clearly, the product partition model is exchangeable only if c(b) depends only
on the cardinality of b ⊆ [n]. Kerov’s work (Theorem 3.2) establishes that the only nondegenerate, exchangeable, consistent product partition model is the
family of Ewens distributions in (6).
3.8 Logarithmic Combinatorial Structures
For θ > 0, let Y1 , Y2 , . . . be independent random
variables for which Yj has the Poisson distribution with
parameter θ/j for each j = 1, 2, . . . . Given nj=1 j ·
7
THE UBIQUITOUS EWENS’S SAMPLING FORMULA
Yj = n, (Y1 , . . . , Yn ) determines a random allelic partition of n with distribution
n
j · Yj = n
P (Y1 , . . . , Yn ) = (m1 , . . . , mn )
j =1
(11) ∝
∝
n
θ mj −θ/j
e
j mj mj !
j =1
n
θ mj
,
j mj mj !
j =1
that is, Ewens’s sampling formula with parameter
θ > 0.
From the above description, the collection (n )n≥1
of Ewens partitions is a special case of a logarithmic combinatorial structure, which Arratia, Barbour
and Tavaré [4] define for set partitions as follows.
Let (n )n≥1 be a collection of random set partitions
and, for each n ≥ 1, let (Nn,1 , . . . , Nn,n ) be the random allelic partition determined by the block sizes
of n . Then (n )n≥1 is a logarithmic combinatorial
structure if for every n = 1, 2, . . . the allelic partition
(Nn,1 , . . . , Nn,n ) satisfies the conditioning relation
P{Nn,1 = m1 , . . . , Nn,n = mn }
n
= P Y1 = m1 , . . . , Yn = mn j · Yj = n
j =1
for some sequence Y1 , Y2 , . . . of independent random
variables on {0, 1, . . .} that satisfies the logarithmic
condition
lim nP{Yn = 1} = lim nEYn > 0.
n→∞
n→∞
Both conditions are plainly satisfied by the independent Poisson sequence above.
Arratia, Barbour and Tavaré [3] further established
the following stronger result according to which block
sizes of a random Ewens partition can be well approximated by independent Poisson random variables as the
sample size grows.
T HEOREM 3.3 (Arratia, Barbour and Tavaré [3]).
For n ≥ 1, let Nn,j be the number of blocks of size
j in a Ewens(θ ) partition of [n]. Then (Nn,j )j ≥1 →D
(Y1 , Y2 , . . .) as n → ∞, where Y1 , Y2 , . . . are independent Poisson random variables with E(Yj ) = θ/j and
→D denotes convergence in distribution.
In the above theorem, Nn,j counts the number of
blocks of size j in a random partition of {1, . . . , n}.
More generally, Nn,j may count the number of components of size j in an arbitrary structure of size n, for
example, the components of size j in a random graph
or a random mapping. Arratia, Barbour and Tavaré’s
monograph [5] relates the component sizes of various structures to Ewens’s sampling formula, for example, 2-regular graphs (with θ = 1/2) and properties of monic polynomials (with θ = 1). Aldous [1],
Lemma 11.23, previously showed that the component
sizes of the directed graph induced by a uniform random mapping [n] → [n] converge in distribution to the
component sizes from Ewens’s sampling formula with
θ = 1/2.
4. SEQUENTIAL CONSTRUCTIONS AND URN
SCHEMES
4.1 Poisson–Dirichlet Distribution
Let
S ↓ = (s1 , s2 , . . .) : s1 ≥ s2 ≥ · · · ≥ 0,
sk ≤ 1
k≥1
be the ranked-simplex and, for s ∈ S ↓ , write ∞∼ s
to signify an infinite random partition generated as follows. Let X1 , X2 , . . . be independent random variables
with distributions
⎧
j ≥ 1,
⎨ sj , (12) P{Xi = j |s} = 1 − k≥1 sk , j = −i,
⎩
0,
otherwise.
From (X1 , X2 , . . .), we define the s-paintbox ∞ ∼ s
by putting
(13)
i and j in the same block of ∞
Xi = Xj .
if and only if
Notice that s0 = 1 − k≥1 sk is the probability that
Xi = −i for each i = 1, 2, . . . and, therefore, corresponds to the probability that element i appears as a
singleton in ∞ . By the law of large numbers and
Kingman’s correspondence, each block of an infinite
exchangeable partition is either a singleton or is infinite; there can be no blocks of size two, three, etc., or
with zero limiting frequency.
T HEOREM 4.1 (Kingman’s correspondence [61]).
Let ∞ be an infinite exchangeable partition. Then
there exists a unique probability measure ν on S ↓ such
that ∞ ∼ ν , where
(14)
ν (·) =
S↓
s (·)ν(ds)
is the mixture of s-paintbox measures with respect to ν.
8
H. CRANE
By Kingman’s correspondence, every exchangeable
random partition of N can be constructed by first sampling S ∼ ν and then “painting” elements 1, 2, . . . according to (12). In the special case of Ewens’s sampling formula, the mixing measure ν is called the
Poisson–Dirichlet distribution with parameter θ > 0.
Kingman further showed that if a sequence of populations of growing size is such that for each population
the allelic partition from a sample of n genes obeys
Ewens’s sampling formula, then the limiting distribution of allele frequencies is the Poisson–Dirichlet distribution with parameter θ . In genetics, the Poisson–
Dirichlet distribution can be viewed as an infinite
population result for selectively neutral alleles in the
infinitely many alleles setting. Within statistics, the
Poisson–Dirichlet distribution is the prior distribution
over the set of all paintboxes whose colors occur according to the proportions of s ∈ S ↓ . Below, we lay
bare several instances of the Poisson–Dirichlet distribution throughout mathematics. The limit of the
Dirichlet-Multinomial process offers perhaps the most
tangible interpretation of the Poisson–Dirichlet distribution.
to comply with the parameterization of the forthcoming Ewens–Pitman distribution (Section 5.1).
In determining a random partition n based on
X1 , . . . , Xn , we disregard the specific values of X1 , . . . ,
Xn and only retain the equivalence classes. If j distinct values appear among X1 , . . . , Xn , there are k ↓j =
k(k − 1) · · · (k − j + 1) possible assignments that induce the same partition of [n]; hence,
P{n = π} =
(16)
=
f (s1 , . . . , sk ; α1 , . . . , αk )
=
(15)
(α1 + · · · + αk ) α1 −1
· · · skαk −1 ,
s1
k
i=1 (αi )
s1 + · · · + sk = 1, s1 , . . . , sk ≥ 0,
(s) = 0∞ x s−1 e−x dx is the gamma function.
where
Given S = (S1 , . . . , Sk ) from the above density, we
draw X1 , X2 , . . . conditionally independently from
P X1 = j |S = (s1 , . . . , sk ) = sj ,
j = 1, . . . , k,
as in (13). With α <
and define a random partition ∞ 0, α1 = · · · = αk = −α, and nj = ni=1 1{Xi = j }, the
count vector (n1 , . . . , nk ) for a sample of size n has
unconditional probability
[0,1]k
=
(−kα) −α+n1 −1
s1
· · · sk−α+nk −1 ds1 · · · dsk
k
(−α)
k
1
(−α)↑nj ,
(−kα)↑n j =1
where α ↑j = α(α +1) · · · (α +j −1) is the rising factorial function. Here we specify α to be negative in order
(−kα/α)↑#π −(−α)↑#b .
(−kα)↑n b∈π
The Poisson–Dirichlet(θ ) distribution corresponds
to Ewens’s one-parameter family (6) and can be viewed
as a limiting case of the above Dirichlet-Multinomial
construction in the following sense. Let θ > 0 and,
for each m = 1, 2, . . . , let n,m have the DirichletMultinomial distribution in (16) with parameters α =
−θ/m and k = m. The distributions of n,m satisfy
P{n,m = π} =
4.2 Dirichlet-Multinomial Process
For α1 , . . . , αk > 0, the (k − 1)-dimensional Dirichlet distribution with parameter (α1 , . . . , αk ) has density
k ↓#π (−α)↑#b
(−kα)↑n b∈π
m↓#π (θ/m)↑#b
θ ↑n b∈π
and, therefore, n,m converges in distribution to a
Ewens(θ ) partition of [n] as m → ∞.
In Kingman’s paintbox process, n,m is the mixture with respect to the distribution of decreasing order statistics of the (m − 1)-dimensional Dirichlet
distribution with parameter (θ/m, . . . , θ/m), whereas
a Ewens(θ ) partition is the mixture with respect to
the Poisson–Dirichlet(θ ) distribution. By the bounded
convergence theorem, Poisson–Dirichlet(θ ) is the limiting distribution of the decreasing order statistics of
(m − 1)-dimensional Dirichlet(θ/m, . . . , θ/m) distributions as m → ∞. This connection between Ewens’s
sampling formula and the Dirichlet-Multinomial construction partially explains the utility of the Chinese
restaurant process in Bayesian inference (Section 6.1).
Consult Feng [38] for more details on the Poisson–
Dirichlet distribution and its connections to diffusion
processes.
4.3 Hoppe’s Urn
In the paintbox process, we construct a random partition by sampling X1 , X2 , . . . conditionally independently and defining its blocks as in (13). Alternatively,
Hoppe [48] devised a Pólya-type urn scheme by which
(6) arises by sampling with reinforcement from an urn.
Initiate an urn with a single black ball with label 0 and weight θ > 0. Sequentially for each n =
THE UBIQUITOUS EWENS’S SAMPLING FORMULA
1, 2, . . . , choose a ball with probability proportional to
its weight and replace it along with a new ball labeled
n, weighted 1, and colored as follows:
• the same as the chosen ball, if not black, or
• differently from all other balls in the urn, if the chosen ball is black.
The above scheme determines a partition of {1, 2, . . .}
for which i and j are in the same block if and only
if the balls labeled i and j have the same color. Hoppe
showed that the color composition (m1 , . . . , mn ) is distributed as in equation (1), where mj is the number of
colors represented by exactly j of the first n nonblack
balls.
From this point on, we leave behind the interpretation of 1, 2, . . . as labels for genes, as in Ewens’s original context, in favor of a more generic setting in which
1, 2, . . . are themselves abstract items or elements, for
examples, labels of balls in Hoppe’s urn.
4.4 De Morgan’s Process
Some 150 years before Hoppe, De Morgan [24]
posited a similar sequential scheme for explaining how
to update conditional probabilities for events that have
not yet occurred and are not even known to exist:
“When it is known beforehand that either A
or B must happen, and out of m + n times A
has happened m times, and B n times, then
[. . . ] it is m + 1 to n + 1 that A will happen the next time. But suppose we have no
reason, except that we gather from the observed event, to know that A or B must happen; that is, suppose C or D, or E, etc. might
have happened: then the next event may be
either A or B, or a new species, of which it
can be found that the respective probabilities are proportional to m + 1, n + 1, and 1
[. . . ].” (De Morgan, [24], page 66)
Thus, De Morgan considers situations for which we do
not even know all the possible outcomes in advance, as
opposed to binary outcomes such as whether the sun
will rise or not or whether a coin will land on heads or
tails. On the (n + 1)st trial, De Morgan assigns probability 1/(n + t + 1) to the event that a new type is observed and nj /(n + t + 1) to the event that a type with
nj prior occurrences is observed. With θ = 1 and t = 0,
De Morgan’s and Hoppe’s update probabilities coincide. This framework is sometimes called the species
sampling problem—before encountering an animal of
9
a new species we are not aware that the species exists—
and requires the use of exchangeable random partitions
instead of exchangeable random variables [93].
In deriving (1), Ewens invokes “a variant of the
‘coupon collector’s problem’ (or the ‘law of succession’)” [36], page 94, and must have been aware of
the sequential construction of (1) via Hoppe’s urn or,
equivalently, the Chinese restaurant process from the
coming section. In fact, the update probabilities of
Ewens’s sampling formula are distinguished within the
broader context of rules of succession: if the conditional probability that item n + 1 is of a new type depends only on n, then it must have the form θ/(θ + n)
for some θ > 0 [28]; moreover, if the conditional probability that the (n + 1)st item is of a type seen m times
previously depends only on m and n, then the underlying sampling distribution must be Ewens’s sampling
formula. If, in addition to depending on m, the conditional probability at stage n + 1 also depends on the
number of species observed so far, then the underlying
distribution has the two-parameter Ewens–Pitman distribution, which we discuss throughout Section 5. This
latter point relies on Johnson’s [54] sufficientness postulate, which we discuss further in Section 7.
4.5 Chinese Restaurant Process
Dubins and Pitman (see, e.g., [1], Section 11.19),
proposed the Chinese restaurant process, a sampling
scheme equivalent to Hoppe’s urn above. Imagine a
restaurant in which customers are labeled according to
the order in which they arrive: the first customer is labeled 1, the second is labeled 2, and so on. If the first
n customers are seated at m ≥ 1 different tables, the
(n + 1)st customer sits
• at a table occupied by t ≥ 1 customers with probability t/(θ + n) or
• at an unoccupied table with probability θ/(θ + n),
where θ > 0. By regarding the balls in Hoppe’s urn
as customers in a restaurant, it is clear that the Chinese restaurant process and Hoppe’s urn scheme are
identical and, thus, the allelic partition (m1 , . . . , mn ),
for which mj counts the number of tables with j customers, is distributed according to equation (1).
5. THE TWO-PARAMETER EWENS–PITMAN
DISTRIBUTION
5.1 Ewens–Pitman Two-Parameter Family
In some precise sense, the distribution in (16) can
be viewed as a specialization of (6) to the case of a
10
H. CRANE
partition with a bounded number of blocks. Both (6)
and (16) are special cases of Pitman’s two-parameter
extension to Ewens’s sampling formula; see [76] for
an extensive list of references.
With (α, θ ) satisfying either
Ewens. Gnedin [43] has recently studied a different
two-parameter model which allows for the possibility
of a finite, but random, number of blocks.
• α < 0 and θ = −kα, for some k = 1, 2, . . . , or
• 0 ≤ α ≤ 1 and θ > −α,
The extended parameter range of the two-parameter
model leads to different asymptotic regimes for various critical statistics of Ewens–Pitman partitions. For
(α, θ ) in the parameter space of (17), let (n )n≥1 be a
collection of random partitions generated by the above
two-parameter Chinese restaurant process. For each
n ≥ 1, let Kn denote the number of blocks of n . When
α < 0 or α = 0, the asymptotic behavior of Kn is clear
from prior discussion: the α < 0 case corresponds to
Dirichlet-Multinomial sampling (Section 4.2), so that
Kn → −θ/α a.s. as n → ∞, while the α = 0 case
corresponds to Ewens’s sampling formula, whose description as a logarithmic combinatorial structure (Section 3.8) immediately gives Kn ∼ θ log(n) a.s. as n →
∞. When α > 0, Pitman [76] obtains the following
limit law.
the Ewens–Pitman distribution with parameter (α, θ )
assigns probability
(17)
P{n = π} =
(θ/α)↑#π −(−α)↑#b .
θ ↑n b∈π
When α = 0, (17) coincides with equation (6); and
when α < 0 and θ = −kα, (17) simplifies to (16). In
terms of the paintbox process (14), the mixing measure of an infinite partition drawn from (17) is the
(two-parameter) Poisson–Dirichlet distribution with
parameter (α, θ ) [72, 77]. The one-parameter Poisson–
Dirichlet(θ ) distribution mentioned previously is the
specialization of the two-parameter case with α = 0
and may also be called the Poisson–Dirichlet(0, θ ) distribution.
5.2 Role of Parameters
A more general Chinese restaurant-type construction, as in Section 4.5, elucidates the meaning of parameters α and θ in (17). If the first n customers are
seated at m ≥ 1 different tables, the (n + 1)st customer
sits
• at a table occupied by t ≥ 1 customers with probability (t − α)/(n + θ ) or
• at an unoccupied table with probability (mα +
θ )/(n + θ ).
From Ewens’s original derivation, θ is related to the
mutation rate in the Wright–Fisher model and is therefore tied to the prior probability of observing a new
species, or sitting at an unoccupied table, at the next
stage. On the other hand, α reinforces the probability of
observing new species in the future, given that we have
observed a certain number of species so far. Thus, α affects the probability of observing new species and the
probability of observing a specific species in opposite
ways: when α < 0, the number of blocks is bounded,
so observing a new species decreases the number of
unseen species and the future probability of observing new species but increases the probability of seeing the newly observed species again in the future;
when α > 0, the opposite is true; and when α = 0,
we are in the neutral sampling scheme considered by
5.3 Asymptotic Properties
T HEOREM 5.1 (Pitman [76]). For 0 < α < 1 and
θ > −α, n−α Kn → Sα a.s., where Sα is a strictly positive random variable with continuous density
(θ + 1) θ/α
x gα (x) dx, x > 0
P{Sα ∈ dx} =
(θ/α + 1)
and
∞
1 (−1)k+1
gα (x) =
(αk + 1) sin(παk)x k−1 ,
πα k=0
k!
x > 0,
is the Mittag–Leffler density.
The random variable Sα in Theorem 5.1 is called
the α-diversity of (n )n≥1 . Pitman goes on to characterize random partitions with a certain α-diversity in
terms of the power-law behavior of their relative block
sizes [76], Lemma 3.11. Of the many fascinating properties of the Ewens–Pitman family, the next description
in terms of the jumps of subordinators provides some
of the deepest connections to classical stochastic process theory.
5.4 Gamma and Stable Subordinators
For 0 ≤ α ≤ 1 and θ > 0, the theory of Poisson–
Kingman partitions [75] brings forth a remarkable
connection between Ewens’s sampling formula, the
Poisson–Dirichlet distribution, and the jumps of subordinators. A subordinator (τ (s))s≥0 is an increasing
stochastic process with stationary, independent increments whose distribution is determined by a Lévy mea-
11
THE UBIQUITOUS EWENS’S SAMPLING FORMULA
sure
such that for each s ≥ 0
E exp −λτ (s)
= exp −s
∞
0
1 − exp(−λx)
(dx) .
From a random variable T > 0 and the closure Z of the
range of (τ (s))s≥0 , we define
V1 (T ) ≥ V2 (T ) ≥ · · · ≥ 0
as the ranked lengths of the subintervals within [0, T ] \
Z. Thus, the ranked, normalized vector
(18)
5.5 Size-Biased Sampling and the Random
Allocation Model
When studying mass partitions (S1 , S2 , . . .) ∈ S ↓ , it
is sometimes more convenient to work with a sizebiased reordering, denoted S̃ = (S̃1 , S̃2 , . . .), in the infinite simplex
V1 (T ) V2 (T )
,
,...
T
T
(dx) = θ x
−1 −bx
e
dx,
S = (s1 , s2 , . . .) : si ≥ 0,
is a random element of S ↓ and its distribution must
determine the law of an infinite exchangeable partition
by Kingman’s correspondence (Theorem 4.1).
Pitman and Yor [77] found several deep and explicit connections between the Poisson–Dirichlet distribution and the ranked vector in (18). For θ > 0 and
b > 0, let (τ (s))s≥0 be a gamma subordinator with
Lévy measure
x > 0.
Then for all θ > 0, the vector (18) with T = τ (θ ) has
the Poisson–Dirichlet(0, θ ) distribution. On the other
hand, if (τ (s))s≥0 is an α-stable subordinator for 0 <
α < 1, that is, for some C > 0 the Lévy measure satisfies
Poisson–Dirichlet distribution with parameter (α, θ ).
A deep result within the theory of Lévy processes
and exchangeable random partitions, Proposition 21
has also spurred recent progress in the applied field of
Bayesian nonparametrics [53].
E exp −λτ (s) = exp −sC (1 − α)λα ,
then, for all s > 0, (18) with T = τ (s) has the Poisson–
Dirichlet(α, 0) distribution. Other interesting special
cases include the Poisson–Dirichlet(α, α) distribution, which arises as the ranked excursion lengths
of semistable Markov bridges derived from α-stable
surbordinators. In particular, the excursion lengths
of Brownian motion on [0, 1] give rise to Poisson–
Dirichlet(1/2, 0) and the excursion lengths of Brownian bridge on [0, 1] lead to Poisson–Dirichlet(1/2,
1/2). The Poisson–Dirichlet(α, 0) distribution also
arises in the low-temperature asymptotics of Derrida’s
random energy model; see [25, 26] for further details.
In complete generality, Pitman and Yor [77], Proposition 21, derive the family of Poisson–Dirichlet measures for 0 < α < 1 and θ > 0 by taking (τ (s))s≥0
to be a subordinator with Lévy measure (dx) =
αCx −α−1 e−x dx and (γ (t))t≥0 to be a gamma subordinator independent of (τ (s))s≥0 . For θ > 0, Sα,θ =
C −1 γ (θ/α)/ (1 − α), and T = τ (Sα,θ ), (18) has the
sk ≤ 1 .
k≥1
Assuming S ∈ S ↓ satisfies k≥1 Sk = 1, we obtain
S̃ from S = (S1 , S2 , . . .) by putting S̃j = SIj , where
I1 , I2 , . . . are drawn randomly with distribution
P{I1 = i|S} = Si
and
P{Ij = i|S, S̃1 , . . . , S̃j −1 }
=
Si
1 − S̃1 − · · · − S̃j −1
j −1
1{S̃1 = i, . . . , S̃j −1 = i},
j −1
as long as k=1 S̃k < 1. If k=1 S̃k = 1, we put S̃k = 0
for all k ≥ j . In words, we sample I1 , I2 , . . . without replacement from an urn with balls labeled 1, 2, . . . and
weighted S1 , S2 , . . . , respectively. The distribution of a
size-biased reordering of S ∼ Poisson–Dirichlet(α, θ )
is called the Griffiths–Engen–McCloskey distribution
with parameter (α, θ ).
By size-biasing the frequencies of a Poisson–
Dirichlet sequence, we arrive at yet another nifty construction in terms of the random allocation model, or
stick-breaking process. For example, let U1 , U2 , . . . be
independent, identically distributed Uniform random
variables on [0, 1] and define V1 , V2 , . . . in S by
V1 = U1 ,
V2 = U2 (1 − U1 )
and
Vk = Uk (1 − Uk−1 ) · · · (1 − U1 ).
Then V = (V1 , V2 , . . .) has the Griffiths–Engen–
McCloskey distribution with parameter (0, 1), that is,
V is distributed as a size-biased reordering of the block
frequencies of a Ewens set partition with parameter
θ = 1. We can visualize the above procedure as a recursive breaking of a stick with unit length: we first break
the stick U1 units from the bottom, we then break the
remaining piece U2 (1 − U1 ) units from the bottom, and
so on.
12
H. CRANE
T HEOREM 5.2 (Pitman [74]). Let V = (V1 , V2 ,
. . .) be a size-biased reordering of components from the
Poisson–Dirichlet distribution with parameter (α, θ ).
Then V =D V ∗ = (V1∗ , V2∗ , . . .), where
V1∗ = W1 ,
V2∗ = W2 (1 − W1 )
and
Vk∗ = Wk (1 − Wk−1 ) · · · (1 − W1 ),
for W1 , W2 , . . . independent and Wj ∼ Beta(1 − α, θ +
j α) for each j = 1, 2, . . . .
Note that W1 , W2 , . . . are identically distributed only
if α = 0, and so the stick-breaking description further
explains the self-similarity property of Ewens distribution (Section 3.5). In particular, let V = (V1 , V2 , . . .)
be distributed as a size-biased sample from a Poisson–
Dirichlet mass partition with parameter (0, θ ). With
V \ {V1 } = (V2 , V3 , . . .), Theorem 5.2 implies that
(19)
1
V1 ,
V \ {V1 } =D (W, V ),
1 − V1
where W, V1 , V2 , . . . are independent Beta(1, θ ) random variables.
6. BAYESIAN NONPARAMETRICS AND
CLUSTERING METHODS
6.1 Dirichlet Process and Stick-Breaking Priors
Because of its many nice properties and convenient
descriptions in terms of stick breaking, subordinators,
etc., the Ewens–Pitman distribution (17) is widely applicable throughout statistical practice, particularly in
the fast-developing field of Bayesian nonparametrics.
In nonparametric problems, the parameter space is the
collection of all probability distributions. As a practical matter, Bayesians often neglect subjective prior beliefs in exchange for a prior distribution whose “posterior distributions [. . . are] manageable analytically”
[39], page 209. Conjugacy between the prior and posterior distributions in the Dirichlet-Multinomial process
(Section 4.2) hints at a similar relationship in Ferguson’s Dirichlet process prior [39] for nonparametric
Bayesian problems. The construction of the Dirichlet
process and Poisson–Dirichlet(0, θ ) distributions from
a gamma subordinator (Section 5.4) nails down the
connection to Ewens’s sampling formula.
A Dirichlet process with finite, non-null concentration measure β on X is a stochastic process S for
which the random vector (S(A1 ), . . . , S(Ak )) has the
(k − 1)-dimensional Dirichlet distribution with parameter (β(A1 ), . . . , β(Ak )) for every measurable partition
A1 , . . . , Ak of X . Thus, a Dirichlet process S determines a random probability measure on X for which
the conditional distribution of S, given X1 , . . . , Xn , is
againa Dirichlet process with concentration measure
β + ni=1 δXi , where δXi is a point mass at Xi . In particular, for a measurable partition A1 , . . . , Ak , let Nj
denote the number of points among X1 , . . . , Xn that
fall in Aj for each j = 1, . . . , k. Then the conditional
distribution of (S(A1 ), . . . , S(Ak )) given (N1 , . . . , Nk )
is Dirichlet with parameter (β(A1 ) + N1 , . . . , β(Ak ) +
Nk ), just as in Section 4.2.
From a realization X1 , . . . , Xn of the Dirichlet process, we can construct a partition nas in (13). The
posterior concentration measure β + ni=1 δXi and basic properties of the Dirichlet distribution yield the update rule
P{Xn+1 ∈ ·|X1 , . . . , Xn }
=
n
β(·)
δXi (·)
+
,
n + β(X ) i=1 n + β(X )
from which the relationship to the Chinese restaurant
rule with θ = β(X ) < ∞ and Blackwell and MacQueen’s [11] urn scheme is apparent.
More recently, the Ewens–Pitman distribution has
been applied to sampling applications in fish trawling [78], analysis of rare variants [15], species richness in multiple populations [6], and Bayesian clustering methods [22]. In fact, ever since Ishwaran and
James [51, 52] brought stick-breaking priors and the
Chinese restaurant process to the forefront of Bayesian
methodology, the field of Bayesian nonparametrics has
become one of the most active areas of statistical research. Others [37, 65] have further contributed to the
foundations of Bayesian nonparametrics laid down by
Ishawaran and James. This overwhelming activity forbids any possibility of a satisfactory survey of the topic
and promises to quickly outdate the contents of the
present section. For a more thorough accounting of this
rich area, we recommend other related work by the
cited authors.
6.2 Clustering and Classification
In classical and modern problems alike, statistical
units often segregate into nonoverlapping classes B =
{B1 , B2 , . . .}. These classes may represent the grouping of animals according to species, as in Fisher, Corbet and Williams’s [41] and Good and Toulmin’s [44]
consideration of the number of unseen species in a finite sample, or the grouping of literary works according to author, as in Efron and Thisted’s [30, 83] tex-
THE UBIQUITOUS EWENS’S SAMPLING FORMULA
tual analysis of an unattributed poem from the Shakespearean era. While these past analyses employ parametric empirical Bayes [41] and nonparametric [44]
models, modern approaches to machine learning and
classification problems often involve partition models
and Dirichlet process priors, for example, [12, 69]. In
many cases, the Ewens–Pitman family is a natural prior
distribution for the true clustering B = {B1 , B2 , . . .}.
For clustering based on categorical data sequences,
for example, DNA sequences, roll call data, and item
response data, Crane [21] enlarges the parameter space
of the Ewens–Pitman family to include an underlying clustering B. Given α1 , . . . , αk > 0, each block in
B partitions independently according to the DirichletMultinomial process with parameter (α1 , . . . , αk ). By
ignoring the class labels as in (13), we obtain a new
partition by aggregating across the blocks of B. By
first splitting within blocks and then aggregating across
blocks the above procedure warrants the description
as a cut-and-paste process [18, 20]. When α1 = · · · =
αk = α > 0, the cut-and-paste distribution amounts to
(20)
P{n = π } = k ↓#π
b ∈π α
b∈B
with the convention that α ↑0 = 1.
↑#(b∩b )
(kα)↑#b
,
When B = {{1, . . . , n}}, (20) coincides with the
Dirichlet-Multinomial distribution in (16), equivalently the Ewens–Pitman(−α, kα) distribution in (17);
and when B = {{1}, . . . , {n}}, every individual chooses
its block independently with probability 1/k. In both
cases, n is exchangeable, but otherwise the distribution in (20) is only invariant under relabeling by permutations that fix B, a statistical property called relative exchangeability. In general, B is a partition of the
population N, but the marginal distribution of n depends on B only through its restriction to [n]. Thus,
the family is also sampling consistent and it enjoys the
noninterference property (Section 3.6).
In the three-parameter model (20), the Ewens–
Pitman prior for B exhibits nice properties and produces reasonable inferences [22]. For continuous response data, McCullagh and Yang [68] employ the
Gauss–Ewens cluster process, whereby B obeys the
Ewens distribution with parameter θ > 0 and, given
B, (Y1 , Y2 , . . .) is a sequence of multivariate normal
random vectors with mean and covariance depending
on B. Nice properties of the Gaussian and Ewens distributions combine to permit tractable calculation of
posterior predictive probabilities and other quantities
of interest for classification and machine learning applications.
13
7. INDUCTIVE INFERENCE
7.1 Rules of Succession and the Sufficientness
Postulate
Unbeknownst to Ewens, Ferguson, or Antoniak, the
circle of ideas surrounding Ewens’s sampling formula
and the Dirichlet process prior lies at the heart of fundamental questions in inductive inference. Two centuries before Ewens’s discovery, Bayes, Laplace, and
De Morgan pondered epistemological questions about
how past information can be used to update beliefs
about the future [92]. For example, what is the probability the sun will rise tomorrow given that it has risen
each of the previous N days? Laplace’s famed rule of
succession, which attributes probability (N + 1)/(N +
2) to this event, follows from Bayes’s [8] paradigm for
“[events] concerning the probability of which we absolutely know nothing antecedently to any trials made
concerning it.” Under these circumstances, Bayes argued that he has “no reason” to assume anything other
than a uniform prior on the possible outcomes, that is,
the number of successes Sn in n trials satisfies P{Sn =
k} = 1/(n + 1) for each k = 0, 1, . . . , n. A straightforward mathematical argument [92] reveals that Bayes’s
principle of indifference implies a uniform prior distribution on the success probability of each outcome.
Johnson [54] later expanded upon Bayes’s analysis
by allowing for an event with possibly k ≥ 2 different
outcomes. In this case, the result of n trials is summarized by a vector (n1 , . . . , nk ), with ni counting the
number of outcomes of type i = 1, . . . , k. Under Johnson’s sufficientness postulate, by which the conditional
probability that the (n + 1)st observation is type i depends only on ni and n, either all outcomes are independent or the conditional probabilities have the form
of the Ewens–Pitman family with α < 0 and θ = −kα;
see Section 5.1 above. Just as Bayes’s postulate implies the uniform prior, Johnson’s sufficientness postulate implies the (k − 1)-dimensional symmetric Dirichlet prior (15). Thus, Johnson, whose work predates Ferguson and Antoniak by forty years, provides a logical
justification for the Dirichlet–Multinomial and Dirichlet process priors in Bayesian analysis.
At its core, Ewens’s sampling formula is concerned
with rules of succession or, more cavalierly, predicting
the unpredictable [93]: Given an observed allelic partition (m1 , m2 , . . .), what is the probability that the next
sampled individual is of a previously observed type or
of a new type entirely? De Morgan [24] pondered this
question more than a century before Ewens and arrived
at a similar answer, but without any formal justification; see Section 4.4 above.
14
H. CRANE
7.2 Zabell’s Universal Continuum
A primary consideration of induction surrounds universal generalizations, for example, what is the probability that the sun will rise tomorrow and every day in
the future given that it has risen today and every day
in the past? More generically, given that we have so
far observed a partition n = 1n = {{1, . . . , n}} with
all elements in the same block, what is the probability
that we are actually sampling from the universal oneblock partition 1∞ = {{1, 2, . . .}}? The Chinese restaurant process update probabilities implicitly entail the
Poisson–Dirichlet process prior, which is absolutely
continuous and assigns zero prior mass to the universal
partition 1∞ . Simply put, no amount of data is enough
to nudge the posterior probability of 1∞ above zero.
Following the path of least resistance, Zabell [94] refines the Ewens–Pitman two-parameter family by taking ν in the paintbox process (14) to be the two-point
mixture
να,θ,ε = (1 − ε)να,θ + εδ1∞ ,
where 0 ≤ ε ≤ 1 is the prior probability assigned to the
point mass δ1∞ at the universal partition and να,θ is
the Poisson–Dirichlet measure with parameter (α, θ ).
Thus, ε = 0 corresponds to the usual two-parameter
family, and the conditional distribution of n+1 , given
n , coincides with the Chinese restaurant probabilities (Section 5.2) on the event n = 1n . On the event
n = 1n , the above formulation assigns posterior probability
(1 − εn )
n−α
+ εn
n+θ
to the event that the (n + 1)st observed species is the
same type as all prior species and
(1 − εn )
α+θ
n+θ
to the event that the next species is new, where
εn =
ε
ε + (1 − ε)
n−1
j =1 (j − α)/(j + θ )
is the posterior probability of the event ∞ = 1∞
given n = 1n . Zabell’s original derivation expresses
these probabilities in terms of (α, θ, γ ), with α and θ
as before and γ = (α + θ )ε.
8. COMBINATORICS, ALGEBRA, AND NUMBER
THEORY
As if all the above instances were not enough,
Ewens’s sampling formula is also linked to two of
the most fundamental ideas in mathematics, the determinant function in algebra and prime factorization
in number theory.
8.1 The Determinant and α -Permanent
The determinant of an n × n matrix M =
(Mi,j )1≤i,j ≤n is defined as
det(M)
(21)
=
sign(σ )M1,σ (1) M2,σ (2) · · · Mn,σ (n) ,
σ :[n]→[n]
where the sum is over all n! permutations of {1, . . . , n}
and sign(σ ) is the parity of σ , which equals +1 if σ is
a product of an even number of cycles and equals −1
otherwise.
In the early 1800s, Cauchy [14] initiated the study of
“fonctions symétriques permanentes,” that is, permanent symmetric functions, which ignore the parity of σ
in (21). Cauchy’s permanent
(22) per(M) =
M1,σ (1) M2,σ (2) · · · Mn,σ (n)
σ :[n]→[n]
resembles the determinant in appearance but little else:
the determinant is easy to compute (e.g., as a product
of eigenvalues), but the permanent is #P-complete [84];
the determinant has a geometric interpretation in terms
of volume, whereas the permanent’s best interpretation
is graph-theoretic.
Though determinant and permanent appear to occupy different mathematical territory, they come together in Vere-Jones’s α-permanent [85]. For a
complex-valued parameter α, the α-permanent of M
is
per(M)
α
(23)
=
α cyc(σ ) M1,σ (1) M2,σ (2) · · · Mn,σ (n) ,
σ :[n]→[n]
where cyc(σ ) is the number of cycles of σ . Heuristically, (23) interpolates between (21) and (22), Cauchy’s
permanent (22) is the α-permanent with α = 1, and the
determinant (21) equals (−1)n per−1 (M), but its role
is most prominent in modeling bosons and fermions
in statistical physics [49, 66]. Amazingly, the αpermanent also incorporates Ewens distribution in two
different ways. [Note that the parameter α in (23)
does not correspond directly to the parameter α in the
Ewens–Pitman distribution (17).]
15
THE UBIQUITOUS EWENS’S SAMPLING FORMULA
8.1.1 Random permutations. As long as α > 0 and
Mi,j > 0 for all i, j = 1, . . . , n, the α-permanent is the
normalizing constant for the cyclic product distribution
on permutations of [n]:
(24)
P{n = σ } = α
cyc(σ ) M1,σ (1) · · · Mn,σ (n)
perα (M)
α cyc(σ )
α(α + 1) · · · (α + n − 1)
= exp cyc(σ ) log(α) −
n−1
log(α + j ) ,
j =0
with natural parameter log(α) and canonical sufficient
statistic cyc(σ ). In the context of Section 4.5, (25) results by refining the Dubins–Pitman Chinese restaurant
construction: the (n + 1)st customer
• sits to the left of customer j = 1, . . . , n with probability 1/(α + n) and
• sits alone at a table with probability α/(α + n).
Occupied tables correspond to cycles in a random permutation and the left-to-right ordering of customers at
each table determines the order of elements within each
cycle. Just as in the Chinese restaurant construction in
Section 4.5, we recover Ewens distribution (6) from
(25) by ignoring the order in which individuals are
seated at each table. In particular, n induces a random
partition n whose distribution is the sum of (25) over
all permutations with unordered cycles corresponding
to the blocks of a specific partition of [n].
Developed in this way, Ewens distribution is a subfamily of the cyclic product distribution on partitions
of [n]. For any subset b ⊆ [n], the sum of cyclic products
cyp(M)[b] =
Mi,σ (i)
σ :b→b s.t. cyc(σ )=1 i∈b
is the sum over all permutations of b with a single cycle, and so the α-permanent decomposes into a sum
over partitions of [n] by
per(M) =
α
π
α #π
b∈π
cyp(M)[b].
P{n = π} = α
∝
b∈π cyp(M)[b]
#π
perα (M)
α · cyp(M)[b];
b∈π
P{n = σ }
=
.
Computational complexity of the α-permanent makes
(24) intractable in general; however, when Mi,j = 1 for
all i, j , we recover an exponential family of distributions on permutations,
(25)
The (α, M)-cyclic product distribution has the form of
a product partition model,
see Section 3.7.
8.1.2 The two-parameter model. The permanental
decomposition theorem [19] expresses the αpermanent as a sum over partitions of [n] of permanents of related matrices. In particular, for real constants α and β,
(26)
per(M) =
αβ
β ↓#π
π
per M[b] ,
α
b∈π
where M[b] = (Mi,j )i,j ∈b is the submatrix of M
whose rows and columns are indexed by b ⊆ [n] and
the sum is over all partitions of [n]. As long as every
term in (26) is nonnegative, we obtain the permanental
partition model
(27)
P{n = π} = β
↓#π
b∈π perα (M[b])
perαβ (M)
,
where Mi,j is a measure of similarity between elements i and j and β ↓n = β(β − 1) · · · (β − n + 1). In a
homogeneous environment, that is, Mi,j ≡ 1, (27) becomes
P{n = π} = β
↓#π
↑#b
b∈π α
,
(αβ)↑n
which equals the Ewens–Pitman distribution (17) under the substitution α → −α and β → θ/α. In this
way, the permanental partition model (27) extends
the Ewens–Pitman two-parameter family to a threeparameter distribution, but (27) is neither exchangeable
nor consistent in general.
8.2 Random Numbers and Large Prime Factors
For n = 1, 2, . . . , let Nn be uniformly distributed in
{1, . . . , n}, that is,
P{Nn = i} = 1/n,
i = 1, . . . , n.
By the fundamental theorem of arithmetic, Nn can be
factored uniquely into a product of prime numbers, that
is, there exists a unique sequence Pn,1 ≥ · · · ≥ Pn,k ≥ 2
of primes such that
Nn = Pn,1 × · · · × Pn,k .
16
H. CRANE
↓
Since Nn is random, so is Pn = (Pn,1 , . . . , Pn,k ).
Moreover, we can express
log(Nn ) = log(Pn,1 ) + · · · + log(Pn,k )
so that the normalized vector
log(Pn,k )
log(Pn,1 )
(28) Sn↓ =
,...,
, 0, 0, . . .
log(n)
log(n)
is a random element of the ranked-simplex S ↓ .
Billingsley [10] first studied the distribution of the
↓
ranked, normalized prime factors Sn . Donnelly and
Grimmett [29] followed twenty years later with a simpler proof based on the size-biased reordering S̃n of
↓
Sn . In light of all previous discussion, their conclusion
is astonishing: the relative sizes of the prime factors of
a uniform random integer converge in distribution to
the asymptotic block sizes of a Ewens partition with
θ = 1.
T HEOREM 8.1 (Billingsley [10]; Donnelly and
↓
Grimmett [29]). For each n = 1, 2, . . . , let Pn =
↓
(Pn,1 , . . . , Pn,k ) be the prime factorization of a uni↓
form random integer in {1, . . . , n}, let Sn be the normalized, ranked vector in (28), and let S̃n be its sizebiased reordering. Then
(29)
Sn↓ −→D Poisson–Dirichlet(0, 1)
or, equivalently,
(30) S̃n −→D Griffiths–Engen–McCloskey(0, 1).
Knuth and Trabb Pardo [64] and Vershik [86] show
the same distributional convergence for Nn drawn uniformly in {n, n + 1, . . . , 2n} and Nn drawn from the
Riemann zeta distribution,
P{Nn = x} = x −sn /ζ (sn ),
where ζ (s) =
∞
x=1 x
−s and s
x = 1, 2, . . . ,
n = 1 + 1/ log(n).
8.3 Macdonald Polynomials
Recent work at the interface of representation theory, algebraic combinatorics, and probability has led to
interesting connections between symmetric polynomials and fundamental notions in combinatorial stochastic process theory, particularly Kingman’s theorem [57,
71]. Macdonald processes [13] are a particularly interesting family of probability distributions on sequences
of integer partitions that arise in certain models for interacting particle systems and directed random polymers in statistical physics. The Macdonald process
is named after its description in terms of Macdonald
polynomials [67], a family of orthogonal polynomials with two parameters (q, t). Macdonald polynomials
generalize various other families of symmetric polynomials, for example, Hall–Littlewood and Jack polynomials, and thus arise throughout representation theory
and algebraic combinatorics. Within probability theory, Diaconis and Ram [27] have recently provided an
interpretation in terms of the stationary distribution of
a special Markov chain on spaces of integer partitions.
Ewens’s sampling formula (1) arises as the limit of this
stationary distribution under the regime q = t 1/θ and
t → 1, for θ > 0. See [13] and [27], Section 2.4.2, for
further details.
9. CONCLUDING REMARKS
A confluence of mathematical, statistical, and scientific facts contributes to the ubiquity of Ewens’s sampling formula: Ewens’s initial assumptions of neutral
mutation and independence between individuals suited
a need for a tractable mathematical theory of allele
sampling; Laplace’s rule of succession, De Morgan’s
urn scheme, and Johnson’s sufficientness postulate all
arise from principles of indifference at the heart of
Bayesian epistemology [23]; Ferguson [39] and Antoniak [2] stumbled upon Ewens’s sampling formula
without regard for the above logical properties or principles of inductive inference; and the same mathematical properties that drive Ferguson’s and Antoniak’s approach underlie the deep connections between Ewens’s
sampling formula and classical stochastic process theory via the Poisson–Dirichlet distribution [38, 76,
77]. These discoveries along with the occurrence of
Ewens’s sampling formula in the realm of matrix permanents [19] and prime divisors [10, 29] hint at deep
roots in the foundations of mathematics. Still new uses
of the sampling formula in clustering problems [21,
22, 68, 69] manifest its utility in modern applications.
Altogether, Ewens’s sampling formula envelops a rich
history of important contributions within classical and
modern scientific, mathematical, and statistical study.
As much as space permits, the foregoing survey provides a comprehensive modern overview of Ewens’s
sampling formula. Less prominent but equally intriguing connections to rumor spreading [7], physics [47],
computation [70], and many other areas are scattered
throughout the literature. If recent trends in Bayesian
statistics and stochastic process theory are any indication, a book length monograph will soon be necessary to adequately summarize the varied occurrences
of Ewens’s sampling formula.
THE UBIQUITOUS EWENS’S SAMPLING FORMULA
REFERENCES
[1] A LDOUS , D. J. (1985). Exchangeability and related topics. In École d’été de Probabilités de Saint-Flour, XIII—
1983. Lecture Notes in Math. 1117 1–198. Springer, Berlin.
MR0883646
[2] A NTONIAK , C. E. (1974). Mixtures of Dirichlet processes
with applications to Bayesian nonparametric problems. Ann.
Statist. 2 1152–1174. MR0365969
[3] A RRATIA , R., BARBOUR , A. D. and TAVARÉ , S. (1992).
Poisson process approximations for the Ewens sampling formula. Ann. Appl. Probab. 2 519–535. MR1177897
[4] A RRATIA , R., BARBOUR , A. D. and TAVARÉ , S. (2000).
Limits of logarithmic combinatorial structures. Ann. Probab.
28 1620–1644. MR1813836
[5] A RRATIA , R., BARBOUR , A. D. and TAVARÉ , S. (2003).
Logarithmic Combinatorial Structures: A Probabilistic Approach. Eur. Math. Soc., Zürich. MR2032426
[6] BACALLADO , S., FAVARO , S. and T RIPPA , L. (2015).
Bayesian nonparametric inference for shared species richness
in multiple populations. J. Statist. Plann. Inference 166 14–
23. MR3390130
[7] BARTHOLOMEW, D. J. (1973). Stochastic Models for Social
Processes, 2nd ed. Wiley, London. MR0408041
[8] BAYES , T. (1764). An essay toward solving a problem in the
doctrine of chances. Philos. Trans. R. Soc. Lond. Ser. A Math.
Phys. Eng. Sci. 53 370–418.
[9] B ERTOIN , J. (2006). Random Fragmentation and Coagulation Processes. Cambridge Studies in Advanced Mathematics
102. Cambridge Univ. Press, Cambridge. MR2253162
[10] B ILLINGSLEY, P. (1972). On the distribution of large prime
divisors. Period. Math. Hungar. 2 283–289. MR0335462
[11] B LACKWELL , D. and M AC Q UEEN , J. B. (1973). Ferguson
distributions via Pólya urn schemes. Ann. Statist. 1 353–355.
MR0362614
[12] B LEI , D., N G , A. and J ORDAN , M. (2003). Latent Dirichlet
allocation. J. Mach. Learn. Res. 3 993–1022.
[13] B ORODIN , A. and C ORWIN , I. (2014). Macdonald processes.
Probab. Theory Related Fields 158 225–400. MR3152785
[14] C AUCHY, A. (1815). Mémoire sur les fonctions qui ne
peuvent obtenir que deux valeurs égales et de signes contraires par suite des transpositions opérées entre les variables
qu’elles renferment. Journal de l’École Polytechnique 10 91–
169.
[15] C ESARI , O., FAVARO , S. and N IPOTI , B. (2014). Posterior
analysis of rare variants in Gibbs-type species sampling models. J. Multivariate Anal. 131 79–98. MR3252637
[16] C HAMPERNOWNE , D. (1953). A model of income distribution. Econom. J. 63.
[17] C HRISTIANSEN , F. B. (2008). Theories of Population Variation in Genes and Genomes. Princeton Univ. Press, Princeton,
NJ. MR2426517
[18] C RANE , H. (2011). A consistent Markov partition process
generated from the paintbox process. J. Appl. Probab. 48
778–791. MR2884815
[19] C RANE , H. (2013). Some algebraic identities for the
α-permanent. Linear Algebra Appl. 439 3445–3459.
MR3119862
[20] C RANE , H. (2014). The cut-and-paste process. Ann. Probab.
42 1952–1979. MR3262496
17
[21] C RANE , H. (2015a). Clustering from categorical data sequences. J. Amer. Statist. Assoc. 110 810–823. MR3367266
[22] C RANE , H. (2015b). Generalized Ewens–Pitman model for
Bayesian clustering. Biometrika 102 231–238. MR3335108
[23] DE F INETTI , B. (1937). La prévision: Ses lois logiques,
ses sources subjectives. Ann. Inst. H. Poincaré 7 1–68.
MR1508036
[24] DE M ORGAN , A. (1838). An Essay on Probabilities, and on
Their Application to Life Contingencies and Insurance Offices. Longman et al., London.
[25] D ERRIDA , B. (1981). Random-energy model: An exactly
solvable model of disordered systems. Phys. Rev. B (3) 24
2613–2626. MR0627810
[26] D ERRIDA , B. (1997). From random walks to spin glasses.
Phys. D 107 186–198. MR1491962
[27] D IACONIS , P. and R AM , A. (2012). A probabilistic interpretation of the Macdonald polynomials. Ann. Probab. 40 1861–
1896. MR3025704
[28] D ONNELLY, P. (1986). Partition structures, Pólya urns, the
Ewens sampling formula, and the ages of alleles. Theoret.
Population Biol. 30 271–288. MR0865115
[29] D ONNELLY, P. and G RIMMETT, G. (1993). On the asymptotic distribution of large prime factors. J. Lond. Math. Soc.
(2) 47 395–404. MR1214904
[30] E FRON , B. and T HISTED , R. (1976). Estimating the number
of unseen species: How many words did Shakespeare know?
Biometrika 63 435–447.
[31] E THIER , S. N. and G RIFFITHS , R. C. (1993). The transition
function of a Fleming–Viot process. Ann. Probab. 21 1571–
1590. MR1235429
[32] E TIENNE , R. (2005). A new sampling formula for neutral
biodiversity. Ecology Letters 8 253–260.
[33] E TIENNE , R. (2007). A neutral sampling formula for multiple
samples and an “exact” text of neutrality. Ecology Letters 10
608–618.
[34] E TIENNE , R. and A LONSO , D. (2005). A dispersal-limited
sampling theory for species and alleles. Ecology Letters 8
1147–1156.
[35] E WENS , W. and TAVARÉ , S. (1998). The Ewens sampling
formula. In Encyclopedia of Statistical Science (S. Kotz,
C. B. Read and D. L. Banks, eds.) Wiley, New York.
[36] E WENS , W. J. (1972). The sampling theory of selectively
neutral alleles. Theoret. Population Biology 3 87–112; erratum, ibid. 3 (1972), 240; erratum, ibid. 3 (1972), 376.
MR0325177
[37] FAVARO , S., L IJOI , A. and P RÜNSTER , I. (2013). Conditional formulae for Gibbs-type exchangeable random partitions. Ann. Appl. Probab. 23 1721–1754. MR3114915
[38] F ENG , S. (2010). The Poisson–Dirichlet Distribution and Related Topics: Models and Asymptotic Behaviors. Springer,
Heidelberg. MR2663265
[39] F ERGUSON , T. S. (1973). A Bayesian analysis of some nonparametric problems. Ann. Statist. 1 209–230. MR0350949
[40] F ISHER , R. (1922). On the dominance ratio. Proceedings of
the Royal Society of Edinburgh 42 321–341.
[41] F ISHER , R., C ORBET, A. and W ILLIAMS , C. (1943). The
relation between the number of species and the number of
individuals in a random sample of an animal population. The
Journal of Animal Ecology 12 42–58.
18
H. CRANE
[42] F LEMING , W. H. and V IOT, M. (1979). Some measurevalued Markov processes in population genetics theory. Indiana Univ. Math. J. 28 817–843. MR0542340
[43] G NEDIN , A. (2010). A species sampling model with
finitely many types. Electron. Commun. Probab. 15 79–88.
MR2606505
[44] G OOD , I. J. and T OULMIN , G. H. (1956). The number of
new species, and the increase in population coverage, when a
sample is increased. Biometrika 43 45–63. MR0077039
[45] G RIFFITHS , R. C. (1979). Exact sampling distributions from
the infinite neutral alleles model. Adv. in Appl. Probab. 11
326–354. MR0526416
[46] H ARTIGAN , J. A. (1990). Partition models. Comm. Statist.
Theory Methods 19 2745–2756. MR1088047
[47] H IGGS , P. (1995). Frequency distributions in population genetics parallel those in statistical physics. Phys. Rev. E (3) 51
1–7.
[48] H OPPE , F. M. (1984). Pólya-like urns and the Ewens’ sampling formula. J. Math. Biol. 20 91–94. MR0758915
[49] H OUGH , J. B., K RISHNAPUR , M., P ERES , Y. and
V IRÁG , B. (2006). Determinantal processes and independence. Probab. Surv. 3 206–229. MR2216966
[50] H UBBELL , S., (2001). The Unified Neutral Theory of Biodiversity and Biogeography. Princeton Univ. Press, Princeton,
NJ.
[51] I SHWARAN , H. and JAMES , L. F. (2001). Gibbs sampling
methods for stick-breaking priors. J. Amer. Statist. Assoc. 96
161–173. MR1952729
[52] I SHWARAN , H. and JAMES , L. F. (2003). Generalized
weighted Chinese restaurant processes for species sampling
mixture models. Statist. Sinica 13 1211–1235. MR2026070
[53] JAMES , L. F. (2013). Stick-breaking PG(α, ζ )-generalized
Gamma processes. Available at arXiv:1308.6570v3.
[54] J OHNSON , W. (1932). Probability: The deductive and inductive problems. Mind 41 409–423.
[55] K ARLIN , S. and M C G REGOR , J. (1972). Addendum to a paper of W. Ewens. Theoret. Population Biology 3 113–116.
MR0325178
[56] K EROV, S. (2005). Coherent random allocations, and the
Ewens–Pitman formula. Zap. Nauchn. Sem. S.-Peterburg. Otdel. Mat. Inst. Steklov. (POMI) 325 127–145. MR2160323
[57] K EROV, S., O KOUNKOV, A. and O LSHANSKI , G. (1998).
The boundary of the Young graph with Jack edge multiplicities. Int. Math. Res. Not. IMRN 4 173–199. MR1609628
[58] K IMURA , M. (1968). Evolutionary rate at the molecular
level. Nature 217 624–626.
[59] K INGMAN , J. F. C. (1977). The population structure associated with the Ewens sampling formula. Theoret. Population
Biology 11 274–283. MR0682238
[60] K INGMAN , J. F. C. (1978a). Random partitions in population
genetics. Proc. R. Soc. Lond. Ser. A Math. Phys. Eng. Sci. 361
1–20. MR0526801
[61] K INGMAN , J. F. C. (1978b). The representation of partition
structures. J. Lond. Math. Soc. (2) 18 374–380. MR0509954
[62] K INGMAN , J. F. C. (1980). Mathematics of Genetic Diversity. CBMS-NSF Regional Conference Series in Applied
Mathematics 34. SIAM, Philadelphia, PA. MR0591166
[63] K INGMAN , J. F. C. (1982). The coalescent. Stochastic Process. Appl. 13 235–248. MR0671034
[64] K NUTH , D. E. and T RABB PARDO , L. (1976/77). Analysis
of a simple factorization algorithm. Theoret. Comput. Sci. 3
321–348. MR0498355
[65] L IJOI , A., M ENA , R. H. and P RÜNSTER , I. (2007). Bayesian
nonparametric estimation of the probability of discovering
new species. Biometrika 94 769–786. MR2416792
[66] M ACCHI , O. (1975). The coincidence approach to stochastic
point processes. Adv. in Appl. Probab. 7 83–122. MR0380979
[67] M ACDONALD , I. G. (1995). Symmetric Functions and
Hall Polynomials, 2nd ed. Clarendon Press, New York.
MR1354144
[68] M C C ULLAGH , P. and YANG , J. (2006). Stochastic classification models. In International Congress of Mathematicians.
Vol. III 669–686. Eur. Math. Soc., Zürich. MR2275702
[69] M C C ULLAGH , P. and YANG , J. (2008). How many clusters?
Bayesian Anal. 3 101–120. MR2383253
[70] N EAL , R. M. (2000). Markov chain sampling methods for
Dirichlet process mixture models. J. Comput. Graph. Statist.
9 249–265. MR1823804
[71] O LSHANSKI , G. (2011). Random permutations and related
topics. In The Oxford Handbook of Random Matrix Theory
510–533. Oxford Univ. Press, Oxford. MR2932645
[72] P ERMAN , M., P ITMAN , J. and YOR , M. (1992). Size-biased
sampling of Poisson point processes and excursions. Probab.
Theory Related Fields 92 21–39. MR1156448
[73] P ITMAN , J. (1995). Exchangeable and partially exchangeable
random partitions. Probab. Theory Related Fields 102 145–
158. MR1337249
[74] P ITMAN , J. (1996). Random discrete distributions invariant
under size-biased permutation. Adv. in Appl. Probab. 28 525–
539. MR1387889
[75] P ITMAN , J. (2003). Poisson–Kingman partitions. In Statistics
and Science: A Festschrift for Terry Speed. Institute of Mathematical Statistics Lecture Notes—Monograph Series 40 1–34.
IMS, Beachwood, OH. MR2004330
[76] P ITMAN , J. (2006). Combinatorial Stochastic Processes.
Lecture Notes in Math. 1875. Springer, Berlin. MR2245368
[77] P ITMAN , J. and YOR , M. (1997). The two-parameter
Poisson–Dirichlet distribution derived from a stable subordinator. Ann. Probab. 25 855–900. MR1434129
[78] S IBUYA , M. (2014). Prediction in Ewens–Pitman sampling
formula and random samples from number partitions. Ann.
Inst. Statist. Math. 66 833–864. MR3250819
[79] S IMON , H. A. (1955). On a class of skew distribution functions. Biometrika 42 425–440. MR0073085
[80] S LOANE , N. Online Encyclopedia of Integer Sequences. Published electronically at http://www.oeis.org/.
[81] S PIELMAN , R., M C G INNIS , R. and E WENS , W. (1993).
Transmission test for linkage disequilibrium: The insulin
gene region and insulin-dependent diabetes mellitus (IDDM).
American Journal of Human Genetics 52 506–516.
[82] TAVARÉ , S. and E WENS , W. (1997). The Multivariate
Ewens Distribution. In Discrete Multivariate Distributions
(N. L. Johnson, S. Kotz and N. Balakrishnan, eds.) Wiley,
New York.
[83] T HISTED , R. and E FRON , B. (1987). Did Shakespeare
write a newly-discovered poem? Biometrika 74 445–455.
MR0909350
[84] VALIANT, L. G. (1979). The complexity of computing the
permanent. Theoret. Comput. Sci. 8 189–201. MR0526203
THE UBIQUITOUS EWENS’S SAMPLING FORMULA
[85] V ERE -J ONES , D. (1988). A generalization of permanents
and determinants. Linear Algebra Appl. 111 119–124.
MR0974048
[86] V ERSHIK , A. M. (1986). Asymptotic distribution of decompositions of natural numbers into prime divisors. Dokl. Akad.
Nauk SSSR 289 269–272. MR0856456
[87] WAKELEY, J. (2008). Coalescent Theory: An Introduction.
Roberts and Company Publishers, Greenwood Village, CO.
[88] WATTERSON , G. (1978). The homozygosity test of neutrality. Genetics 88 405–417.
[89] WATTERSON , G. A. (1977). Heterosis or neutrality? Genetics
85 789–814. MR0504021
[90] W RIGHT, S. Evolution in Mendelian populations. Genetics
16 97–159.
19
[91] Y ULE , G. (1925). A mathematical theory of evolution, based
on the conclusions of Dr. J. C. Willis. F. R. S. Phil. Trans. Roy.
Soc. London, B 213 21–87.
[92] Z ABELL , S. (1988). Symmetry and its discontents. In Causation, Chance, and Credence 1 155–190. Kluwer Academic,
Norwell.
[93] Z ABELL , S. (1992). Predicting the unpredictable. Synthese
90 205–232. MR1148566
[94] Z ABELL , S. (1997). The continuum of inductive methods
revisited. In The Cosmos of Science: Essays of Exploration
Univ. Pittsburg Press, Pittsburgh, PA.
Statistical Science
2016, Vol. 31, No. 1, 20–22
DOI: 10.1214/15-STS535
Main article DOI: 10.1214/15-STS529
© Institute of Mathematical Statistics, 2016
Diffusion Processes and the Ewens
Sampling Formula
Shui Feng
Abstract. Crane [The ubiquitous Ewens sampling formula (2016) Preprint]
provides an excellent review of Ewens’ sampling formula (henceforth, ESF),
its applications in and connections with various subjects. This note intends
to extend the discussion a little bit. The focus will be on nonequilibrium ESF
involving diffusion processes, ESF with symmetric selection and asymptotics
of ESF. The references listed are by no means exhaustive.
Key words and phrases:
Asymptotics, selection, nonequilibrium.
Let S ↓ be the ranked simplex and νθ be Kingman’s Poisson–Dirichlet distribution on S ↓ with parameter θ > 0. For notational convenience, we denote the generic element of S ↓ by x = (x1 , x2 , . . .),
y = (y1 , y2 , . . .), etc. Consider a population of individuals of various types. If the random proportions of
types follow the law νθ , then ESF gives the distribution of the allelic partitions of random samples from
the population. Given the sample size n and an allelic
partition m = (m1 , . . . , mn ), set fm (x) = 1 for n = 1
and for n > 1,
random proportions of a population under the influence
of parent independent mutation with mutation rate θ
and random sampling. The generator of the process on
an appropriate domain has the form
δij is the Kronecker delta.
Starting from any point x, the distribution νθ (t) of the
process at each fixed time t > 0 is shown in Ethier
(1992) to be absolutely continuous with respect to νθ
and the density function is
n!
mj
j =1 (j !) mj !
fm (x) = n
·
∞
∞
∂
1 ∂2
Aθ =
xi (δij − xj )
−θ
xi
,
2 i,j =1
∂xi ∂xj
∂xi
i=1
q(t, x, y) = 1 +
xk11 · · · xk1m1 xk221 · · ·
∞
e−λk t ϕk (x, y),
k=2
distinct kij
)
where λk = k(k−1+θ
, and
2
· xk22m · · · xknn1 · · · xknnmn .
2
The
ESF can be written as ESF n (m; θ ) =
f
S ↓ m (x)νθ (d x).
The distribution νθ can be constructed from a sequence of Dirichlet distributions through a Poisson
type limiting procedure. Ethier and Kurtz (1981) generalized this construction to a dynamical setting and
constructed an infinite dimensional diffusion process
with reversible measure νθ . The process is an infinite
dimensional limit of a sequence of finite-dimensional
Wright–Fisher diffusions. It describes the evolution of
ϕk (x, y) =
k
2k − 1 + θ k
(−1)k−n
n
k!
n=0
·
(n + k − 1 + θ )
pn (x, y).
(n + θ )
Here (·) denotes the gamma function, and the function pn (x, y) has the form
pn (x, y) =
fm (x)fm (y)
m
ESF n (m; θ )
,
and the summation is over all allelic partitions of the
size n sample. It is clear that q(t, x, y) converges to 1
as t tends to infinity. The summation starting from 2 is
a result of the ordering procedure.
Shui Feng is Professor, Department of Mathematics and
Statistics, McMaster University, Hamilton, Ontario,
Canada L8S 4K1 (e-mail: shuifeng@mcmaster.ca).
20
21
DISCUSSION
Rearranging the terms, one obtains the following
representation for q(t, x, y):
q(t, x, y) = d0 (t) +
∞
dn (t)pn (x, y),
n=1
where
d0 (t) = 1−
∞
e−λk t
k=1
2k + θ − 1
(θ + k − 1)
(−1)k−1
k!
(θ )
and for n ≥ 1,
Griffiths (1979a) is the first to obtain the nonequilibrium or transient ESF. The integral on the right-hand
of the above equation can be calculated explicitly.
For 0 < α < 1, θ + α > 0, let να,θ denote the
two-parameter Poisson–Dirichlet distribution. Petrov
(2009) constructed an infinite dimensional diffusion
process that has να,θ as the reversible measure. Alternate constructions were obtained later in Feng and Sun
(2010), Ruggiero and Walker (2009). The generator of
the diffusion process has the form
∞
∞
1 ∂2
Aα,θ =
xi (δij − xj )
2 i,j =1
∂xi ∂xj
2k + θ − 1
k
(−1)k−n
dn (t) =
e−λk t
n
k!
k=n
(n + θ + k − 1)
.
·
(n + θ )
Here dn (t) is the probability of having n ancestors at
time t in Kingman’s coalescent and the representation
of dn (t) is derived in Tavaré (1984).
This representation of q(t, x, y) gives a clear picture
about the population structure at each positive time t.
Initially individuals in the population have types (old
types) with proportions x. The population evolves under the influence of random sampling and mutation.
Random sampling changes proportions of each type
while each mutation results in a new type not seen before. At each positive time, the population is a mixture
of individuals of old types and new types. The number
of old types is always finite and the distribution is given
by Kingman’s coalescent. For each n ≥ 1, the function
pn (x, y) reflects the details of the mixture when the
number of old types is n and the type of proportions
in the population is y. An unordered model, a particular Fleming–Viot process, is studied in Ethier and Griffiths (1993) where the distribution at each positive time
is represented as a mixture of posteriors of the Dirichlet process. The mixing factor is given by Kingman’s
coalescent.
For each fixed t > 0, the nonequilibrium ESF gives
the distribution of the allelic partitions of random samples from the population when the random proportions
follow the law νθ (t). For any n ≥ 1 and allelic partition
m, the nonequilibrium ESF is
∞
∂
−θ
(xi + α)
∂xi
i=1
on an appropriate domain and the transition density
function is obtained in Feng et al. (2011). Given the
sample size n and the allelic partition m, the Pitman’s sampling formula or the two-parameter ESF,
PSF n (m; α, θ ), also has a nonequilibrium version:
PSF n (m; t, α, θ ) = PSF n (m; α, θ ) + Gn (t),
where Gn (t) is obtained by replacing νθ with να,θ in
the expression of Fn (t). More details are found in Xu
(2011), Zhou (2015).
ESF with selection.
For any real numbers s and r ≥ 1, set
hr (x) =
Fn (t) =
k=2
xir ,
φr (x) = exp shr (x) .
The function h2 is the homozygosity and the parameter
s is the selection intensity. The probability
φ
νθ r (d x) = φr (x)νθ (d x)
is called the Poisson–Dirichlet distribution with symmetric selection. Grote and Speed (2002) studied the
φ
sampling formula under νθ 2 when s < 0 and obtained
a useful approximation. Handa (2005) studied the general case. For given sample size n and allelic partition
m, the sampling formula is
φ
S↓
∞
∞
i=1
ESF n (m; t, θ ) = ESF n (m; θ ) + Fn (t),
where
fm (x)νθ r (d x)
∞ l
θ mi
θ
Il (m),
mi m !
(j
!)
l!
i
i=1
l=1
n
e
−λk t
S↓
ϕk (x, y)fm (y)νθ (d y).
The nonequilibrium factor Fn (t) describes the impact
of finite time and diminishes as t tends to infinity.
= ESF n (m; θ ) + n!
where Il (·) has explicit integral form depending on s
and r. The structure of this formula is very similar
22
S. FENG
to the nonequilibrium ESF. An alternate derivation is
found in Huillet (2007).
Asymptotics.
In the neutral evolution model, the parameter θ is
the scaled population mutation rate and is equal to
4Nu with u being the individual mutation rate and N
the effective population size. The Poisson type limit
is to let N tend to infinity and u tend to zero while
the product Nu is held constant. If N goes to infinity faster or slower than 1/u, one would be dealing
with limiting procedures of θ tending to infinity or
zero. Given sample size n, let m0 = (0, 0, . . . , 1) and
m∞ = (n, 0, . . . , 0) be two allelic partitions. Then the
ESFs corresponding to θ = 0 and θ = ∞ are Dirac
measures at m0 and m∞ , respectively. Asymptotic results for ESF such as central limit theorems and large
deviations can be found in Griffiths (1979b), Joyce,
Krone and Kurtz (2002) and Feng (2007).
ACKNOWLEDGMENTS
Supported in part by the Natural Science and Engineering Research Council of Canada.
REFERENCES
C RANE , H. (2016). The ubiquitous Ewens sampling formula.
Statist. Sci. 31 1–19.
E THIER , S. N. (1992). Eigenstructure of the infinitely-manyneutral-alleles diffusion model. J. Appl. Probab. 29 487–498.
MR1174426
E THIER , S. N. and G RIFFITHS , R. C. (1993). The transition function of a Fleming–Viot process. Ann. Probab. 21 1571–1590.
MR1235429
E THIER , S. N. and K URTZ , T. G. (1981). The infinitely-manyneutral-alleles diffusion model. Adv. in Appl. Probab. 13 429–
452. MR0615945
F ENG , S. (2007). Large deviations associated with Poisson–
Dirichlet distribution and Ewens sampling formula. Ann. Appl.
Probab. 17 1570–1595. MR2358634
F ENG , S. and S UN , W. (2010). Some diffusion processes associated with two parameter Poisson–Dirichlet distribution and
Dirichlet process. Probab. Theory Related Fields 148 501–525.
MR2678897
F ENG , S., S UN , W., WANG , F.-Y. and X U , F. (2011). Functional
inequalities for the two-parameter extension of the infinitelymany-neutral-alleles diffusion. J. Funct. Anal. 260 399–413.
MR2737405
G RIFFITHS , R. C. (1979a). Exact sampling distributions from the
infinite neutral alleles model. Adv. in Appl. Probab. 11 326–354.
MR0526416
G RIFFITHS , R. C. (1979b). On the distribution of allele frequencies in a diffusion model. Theoret. Population Biol. 15 140–158.
MR0528914
G ROTE , M. N. and S PEED , T. P. (2002). Approximate Ewens
formulae for symmetric overdominance selection. Ann. Appl.
Probab. 12 637–663. MR1910643
H ANDA , K. (2005). Sampling formulae for symmetric selection.
Electron. Commun. Probab. 10 223–234. MR2182606
H UILLET, T. (2007). Ewens sampling formulae with and without
selection. J. Comput. Appl. Math. 206 755–773. MR2333711
J OYCE , P., K RONE , S. M. and K URTZ , T. G. (2002). Gaussian
limits associated with the Poisson–Dirichlet distribution and
the Ewens sampling formula. Ann. Appl. Probab. 12 101–124.
MR1890058
P ETROV, L. A. (2009). A two-parameter family of infinitedimensional diffusions on the Kingman simplex. Funct. Anal.
Appl. 43 279–296.
RUGGIERO , M. and WALKER , S. G. (2009). Countable representation for infinite dimensional diffusions derived from the
two-parameter Poisson–Dirichlet process. Electron. Commun.
Probab. 14 501–517. MR2564485
TAVARÉ , S. (1984). Line-of-descent and genealogical processes,
and their applications in population genetics models. Theoret.
Population Biol. 26 119–164. MR0770050
X U , F. (2011). The sampling formula and Laplace transform associated with the two-parameter Poisson–Dirichlet distribution.
Adv. in Appl. Probab. 43 1066–1085. MR2867946
Z HOU , Y. (2015). Ergodic inequality of a two-parameter infinitelymany-alleles diffusion model. J. Appl. Probab. 52 238–246.
MR3336858
Statistical Science
2016, Vol. 31, No. 1, 23–26
DOI: 10.1214/15-STS536
Main article DOI: 10.1214/15-STS529
© Institute of Mathematical Statistics, 2016
Two Early Contributions to the
Ewens Saga
Peter McCullagh
Abstract. The mixture model devised by Fisher, Corbet and Williams [Journal of Animal Ecology 12 (1943) 42–58] for species sampling and the sequential prediction approach pioneered by Good [Biometrika 40 (1953) 237–264]
and Good and Toulmin [Biometrika 43 (1956) 45–63] are both closely related
to the Ewens sampling formula. Fisher’s two-parameter joint distribution for
the species counts includes the Ewens distribution as the conditional distribution given the sample size. The log-series model, as it is known in the
ecological literature, is closely related to a Poisson process model devised by
Arratia, Barbour and Tavaré [Ann. Appl. Probab. 2 (1992) 519–535]. Oddly,
despite its advantages for statistical inference, Fisher does not mention the
conditional distribution. Likewise, athough Good (1953) pioneered the sequential prediction approach, neither he nor Toulmin discovered the Ewens
process in a form equivalent to the modern-day Chinese restaurant process.
Key words and phrases: Chinese restaurant process, Poisson process,
species richness, species sampling.
of species for which
exactly r specimens occur in the
sample, so
m
=
r>0 mr is the total species count,
·
and N = r rmr is the specimen count.
Although Fisher’s paper is a citation classic in the
ecological literature, and the approach appears to be
simple and well understood, some parts of his argument are not straightforward and other parts are not
correct. Fisher begins with the assumption that the
specimen counts for a single species are distributed according to the Poisson distribution with parameter λ,
and the species-specific λ-values are distributed according to a Poisson process with mean measure G
on (0, ∞). It follows that the joint intensity-count distribution is that of a Poisson process Z ∼ PP(μ) with
mean measure
Crane is to be commended for his survey of the
diverse areas of scientific work in which the Ewens
sampling formula has arisen. It is an impressive list
stretching from literary studies to population genetics and probabilistic number theory. The fundamental
mathematical object in all of this work is a partition—
originally an integer partition but preferably a set
partition—which splits the population units into disjoint subsets called blocks or clusters and does the
same thing to the sample units. Crane makes a strong
case that the Ewens process is to random partitions or
clusterings as the Gaussian process is to a random sequence of real numbers, or the Poisson process is to a
random series of events in time or space. I agree. It is
one of a small number of processes that deserves to be
a central part of the statistical curriculum.
Fisher, Corbet and Williams (1943) appears to be one
of the first studies of the statistical relation between the
number of specimens and the number of species in typical ecological samples. In this setting, the multiplicity vector m with components mr records the number
e−λ λr
G(dλ)
r!
at (λ, r) in the product space. The projected marginal
process on counts, the frequency of frequencies, mr =
#{λ : (λ, r) ∈ Z} is also Poisson, and if G is proportional to the gamma distribution, the marginal measure
at r ≥ 0 is
(r + ν)ηr
(1)
,
μr = E(mr ) = θ
r!
μ(dλ, r) =
Peter McCullagh is Professor, Department of Statistics,
University of Chicago, Chicago, Illinois 60637, USA
(e-mail: pmcc@galton.uchicago.edu).
23
24
P. MCCULLAGH
proportional to the negative binomial series for some
0 < η < 1. In ecological work, the observation is not
the marginal process, but its restriction to r > 0, excluding m0 .
To understand better the interpretation of these parameters, it is helpful to consider the effect of increasing the sampling effort by the factor t > 0, for example, by increasing the number of traps or the observation time. All things being equal, the effect on each
species intensity is λ → tλ, so the transformed measure is Gt (A) = G(t −1 A) for subsets A ⊂ R+ . The total species mass is unaltered. The effect on the marginal
mean measure (1) is to transform η to ηt multiplicatively on the odds scale:
ηt
η
(2)
,
=t
1 − ηt
1−η
leaving θ fixed. My guess is that Fisher was aware of
the implication, ηt = t/(t +γ ) for some constant γ > 0
independent of t, but this equation does not appear in
his paper.
On the basis of empirical evidence derived from Corbet’s series on Malayan butterflies, Fisher concluded
that ν must be close to zero; the limit value ν = 0
implies that μ0 = ∞ and μr = θ ηr /r proportional to
the coefficients in the expansion of − log(1 − η). Conveniently enough, Fisher’s limiting log-series model,
mr ∼ Po(μr ) with infinitely many independent components, is log-linear
log μr = log θ + r log η − log r.
It is a generalized linear model with canonical parameter (log θ, log η), offset − log r, and minimal sufficient
statistic (m· , N) with expected value
(3)
E(N) = θ η/(1 − η),
E(m· ) = −θ log(1 − η) = θ log 1 + E(N)/θ .
Fisher computed the maximum-likelihood estimate by
solving the simultaneous nonlinear equation
N = θ̂ η̂/(1 − η̂),
m· = −θ̂ log(1 − η̂).
In the pre-computer era, he also provided tables to assist in its solution.
For each fixed θ , the statistic N is complete and
sufficient for η, so Fisher’s two-parameter model has
a Neyman structure (Lehmann, 1986, Section 4.3).
Given N = n, the multiplicity vector m is a random
partition of the integer n. By sufficiency, the conditional distribution depends on θ alone; it is the Ewens
sampling formula with parameter θ . Given his earlier
writings on 2 × 2 tables in Statistical Methods for Research Workers (1935, Section 4) and his 1934 paper
on location-scale models, it is strange that Fisher does
not mention the conditional distribution. Presumably
it did not occur to him then or subsequently. But one
Fisherian passage is worth quoting: The quantity θ is
independent of the size of sample and is proportional
to the number of species of the group considered, at
any chosen level of abundance, relative to the means
of capture employed. Values of θ from different samples [. . .] may be compared as a measure of richness
in species. Fisher’s expressions (3) for the moments do
not justify the first part of his statement, so presumably what he had in mind was the argument leading to
(2), undoubtedly obvious to Fisher if not to most readers, that increased sampling effort leaves θ fixed but
increases η in an inverse-linear manner.
Fisher’s paper concludes with a discussion of
maximum-likelihood estimation and the computation
of standard errors, both of which would have been simpler using the conditional likelihood. For Williams’s
Macrolepidoptera series at Harpenden, the counts
N = 15,609, m· = 240 yield θ̂ = 40.248 and η̂ =
0.9974281; the conditional mle of θ based on the
Ewens model is 40.146. The unconditional and conditional standard errors for θ̂ are 2.85 and 2.84 respectively. The observed frequencies are in remarkably good agreement with the fitted series, and the fit
is not improved by taking ν = 0. The proximity of η̂ to
the upper boundary is consistent with the asymptotic
behavior of the Ewens process (Arratia, Barbour and
Tavaré, 1992), so η = 1 is the only correct value in the
limit.
For Fisher’s two-parameter model, the asymptotic
variance of θ̂ , as given by the eponymous inverse information matrix, is
θ̂ 2
θ̂ 2 (N + θ̂ )
=
;
m· − N(1 − η̂) m· (N + θ̂ ) − N θ̂
the numerical value is 8.105. Using a line of argument
that is flawed in places, Fisher deduced incorrectly that
var(m· ) ≃ θ log 2 instead of θ log(1 + N/θ ). As a result, the formula given for the variance of θ̂ is incorrect, and the reported value (1.1251) is too small by
the approximate factor log(1 + N/θ )/ log(2) = 8.60.
Ten years later, Good looked at the problem of
estimating the population relative frequency qr of a
species having r representatives in the sample, whose
size n is fixed by design. The two authors have very
different styles. Fisher’s four-page contribution is terse
to the point of obscurity; Good’s 25-page tracts are
25
DISCUSSION
discursive to the point of distraction. Fisher embraces
parametric assumptions; Good recognizes the need for
smoothing, but he avoids parametric assumptions—
even when they might be helpful.
Avoiding all assumptions about the behavior of the
expected multiplicities, μn,r = E(mr |N = n), Good
(1953) concluded that the posterior expected frequency
is
r + 1 μr+1,n+1
E(qr |data) =
(4)
.
n + 1 μr,n
Subsequently, Good and Toulmin (1956) looked at the
problem of sample extension, in an effort to determine the conditional distribution of the number of new
species that occur in an extended sample. Their approach puts the emphasis on prediction, where it properly belongs, rather than on distribution fitting and parameter estimation. It may be viewed as the first attempt to express an infinitely exchangeable partition
process dynamically using the conditional distribution
given the current configuration. If they had adapted
Fisher’s two-parameter model, they might have succeeded in developing a Chinese-restaurant description.
As it stands, their analysis is unavoidably complicated
because of the deliberate avoidance of parametric assumptions.
For a point-process model in which n is not fixed,
Good’s predictive ratio is replaced by the Papangelou
conditional intensity ratio with μr ∝ ηr /r, which
yields
rη,
r ≥ 1;
E(qr |data) ∝
θ η,
r = 0,
for Fisher’s model. The value for r = 0 is the combined intensity for all unrepresented species. Had Good
or Good and Toulmin taken a more positive view of
parametric models, they could easily have arrived at
the simpler expression
r
E(qr |data) =
,
n+θ
leading to q0 = θ/(n + θ ), which is exact for Fisher’s
model. But their determination to develop the story
without the benefit of parametric smoothing led them
elsewhere. Recognizing that the posterior expected
value of qr is the same as the conditional probability
that the next specimen belongs to that block, we can
speculate on how the subject might have developed differently.
It is fair to say that Fisher almost discovered the
Ewens sampling formula. He had it in his grasp. After all, he had only to compute the conditional distribution given the sample size, a task that was both statistically natural and, for him, mathematically trivial. But
he did not. And even had he done so, the conditional
distribution would not necessarily have led quickly to
the Ewens process as we know it today. It is also fair
to say that Good should have discovered the process.
He also had it in his grasp, in the sense that he was
aware of Fisher’s work, he was asking the right questions, and his analysis was correct. But his refusal to
consider Fisherian parametric smoothing was a critical
blinker. In the end, Fisher did not discover the sampling
formula, and Good discovered neither the formula nor
the process.
In closing, it is of some interest to examine
Williams’s Macrolepidoptera data for the years 1933–
36 from the perspective of the Ewens process. For the
year 1933, 178 species were observed in a sample of
3540 specimens, yielding θ̂1 = 39.355 with standard
error 3.34. The data given by Williams in Table 4 is
not sufficient to determine the additional species numbers for each of the following three years, but, for the
three years combined, a further set of 62 species was
recorded among 12,069 specimens. In the Ewens process, this additional species count is a sum of independent Bernoulli variables with mean 58.06 and variance
57.73, so the observed value is only 0.52 standard deviations above what is predicted under the model of
temporal homogeneity using the fitted 1933 value θ̂1 .
Given the observed data for the year 1933, the Ewens
conditional likelihood yields an estimated richness parameter θ̂2 = 42.042 with standard error 5.36 for the
period 1934–36. Since these estimators are statistically
independent, at least asymptotically, we may compute
a standardized difference in the usual way, giving the
value T = 2.687/6.314 = 0.425 as a test for temporal homogeneity. The overall maximum likelihood estimate, assuming temporal homogeneity for the combined sample, is θ̂ = 40.146, and the likelihood ratio statistic is 0.184, in good agreement with the Wald
statistic T 2 = 0.181. Although the specimen count for
1935 was more than twice the count in any other year,
there is no evidence for a change in the relative composition of the Macrolepidoptera population at Harpenden over these four years.
REFERENCES
A RRATIA , R., BARBOUR , A. D. and TAVARÉ , S. (1992). Poisson
process approximations for the Ewens sampling formula. Ann.
Appl. Probab. 2 519–535. MR1177897
F ISHER , R. A. (1934). Two new properties of maximum likelihood. Proc. R. Soc. Lond. Ser. A Math. Phys. Eng. Sci. 144 285–
307.
26
P. MCCULLAGH
F ISHER , R. A. (1935). Statistical Methods for Research Workers.
Oliver & Boyd, Edinburgh.
F ISHER , R. A., C ORBET, A. S. and W ILLIAMS , C. B. (1943). The
relation between the number of species and the number of individuals in a random sample of an animal population. Journal of
Animal Ecology 12 42–58.
G OOD , I. J. (1953). The population frequencies of species and the
estimation of population parameters. Biometrika 40 237–264.
MR0061330
G OOD , I. J. and T OULMIN , G. H. (1956). The number of new
species, and the increase in population coverage, when a sample
is increased. Biometrika 43 45–63. MR0077039
L EHMANN , E. L. (1986). Testing Statistical Hypotheses, 2nd ed.
Wiley, New York. MR0852406
Statistical Science
2016, Vol. 31, No. 1, 27–29
DOI: 10.1214/15-STS537
Main article DOI: 10.1214/15-STS529
© Institute of Mathematical Statistics, 2016
Exploiting the Feller Coupling for the
Ewens Sampling Formula
Richard Arratia, A. D. Barbour and Simon Tavaré
We congratulate Harry Crane on a masterful survey,
showing the universal character of the Ewens sampling
formula.
There are two grand ways to get a simple handle
on the Ewens sampling formula; one is the Chinese
restaurant coupling, and the other is the Feller coupling. Since Crane has discussed the Chinese Restaurant process, but not the Feller coupling, we will give a
brief survey of the latter.
The Ewens sampling formula, given in Crane’s (1),
has an interpretation in terms of the cycle type of a
random permutation of n objects. For θ = 1, it is just
Cauchy’s formula, expressed in terms of the fraction of
permutations of n objects that have exactly mi cycles
of order i, 1 ≤ i ≤ n. For general θ , the power
The Feller coupling, motivated by the example in Feller
([6], page 815) is defined as follows. Take independent Bernoulli random variables ξi , i = 1, 2, 3, . . . ,
with the simple odds ratios P(ξi = 0)/P(ξi = 1) =
(i − 1)/θ . Thus, Eξi = P(ξi = 1) = θ/(θ + i − 1),
and P(ξi = 0) = (i − 1)/(θ + i − 1). Say that an spacing occurs in a sequence a1 , a2 , . . . , of zeros and
ones, starting at position i − and ending at position i,
if ai− ai−+1 · · · ai−1 ai = 10−1 1, a one followed by
− 1 zeros followed by another one. Then if, for each
≥ 1, we define
C (n) := the number of -spacings in
ξ1 , ξ2 , . . . , ξn−1 , ξn , 1, 0, 0, . . . ,
the joint distribution of C1 (n), . . . , Cn (n) is the Ewens
sampling formula, as per Crane’s (1) and our (1). This
can be seen directly, for the case θ = 1: consider a
random permutation of 1 to n, write the canonical
cycle notation one symbol at a time, and let ξi indicate the decision to complete a cycle, when there is
an i-way choice of which element to assign next. The
general case θ > 0 follows by biasing, with respect
to θ K : since K = ξ1 + · · · + ξn , and the ξ1 , . . . , ξn
are independent, biasing their joint distribution by
θ ξ1 +···+ξn = θ ξ1 · · · θ ξn preserves their independence
and Bernoulli distributions, while changing the odds
P(ξi = 0)/P(ξi = 1) from (i − 1)/1 to (i − 1)/θ .
Now, the wonderful thing that happens is that, with
Y defined to be the number of -spacings in the infinite sequence ξ1 , ξ2 , . . . , it turns out that Y1 , Y2 , . . .
are mutually independent, and that Y is Poisson distributed, with EY = θ/, as in formula (11) in Section 3.8. This shows that the Ewens sampling formula
is closely related to the simpler independent process
Y1 , Y2 , . . . , Yn . Explicitly, let Rn be the position of
the rightmost one in ξ1 , ξ2 , . . . , ξn−1 , ξn —noting that
always ξ1 = 1 so Rn is well-defined—and let Jn :=
(n + 1) − Rn . We have
θ m1 +m2 +···+mn = θ K
appearing in the formula, where K denotes the number
of cycles, biases the uniform random choice of a permutation by weighting with the factor θ K , the remaining factors involving θ merely reflecting the new normalization constant required to specify a probability
distribution. We use the notation (C1 (n), . . . , Cn (n))
to denote a random object distributed according to the
Ewens sampling formula, suppressing the parameter
θ but making explicit the parameter n, so that, with
Crane’s notation (1),
(1)
P C1 (n) = m1 , . . . , Cn (n) = mn
= p(m1 , . . . , mn ; θ ).
Richard Arratia is Professor, Department of Mathematics,
University of Southern California, 3620 S. Vermont Ave,
KAP 104, Los Angeles, California 90089-2532, USA
(e-mail: rarratia@usc.edu). A. D. Barbour is Professor
Emeritus, Institut für Mathematik, Universität Zürich,
Winterthurerstrasse 190, 8057 Zürich, Switzerland (e-mail:
a.d.barbour@math.uzh.ch). Simon Tavaré is Professor,
Department of Applied Mathematics and Theoretical
Physics, University of Cambridge, Centre for Mathematical
Sciences, Wilberforce Road, Cambridge CB3 0WA, United
Kingdom (e-mail: st321@cam.ac.uk).
(2)
C (n) ≤ Y + 1(Jn = ),
1 ≤ ≤ n,
with contributions to strict inequality whenever, for
some 1 ≤ ≤ n, an -spacing occurred in ξ1 , ξ2 , . . .
starting at i − and ending at i > n.
27
28
R. ARRATIA, A. D. BARBOUR AND S. TAVARÉ
We view (2) as saying that the Ewens sampling
formula distributed (C1 (n), . . . , Cn (n)) can be constructed from the independent Poisson Y ’s using at
most one insertion, together with a random number of
deletions. The expected number of deletions is Oθ (1),
that is, bounded over all n, with the upper bound depending on the value of θ . A concrete upper bound
is given in [3], but the limit value, call it c(θ ), is
cleaner. This limit is the expected number of spacings
of length at most 1, with right end greater than 1, in the
scale invariant Poisson process on (0, ∞) with intensity θ/x dx; see [1]. We have
θ P at least one arrival in (x − 1, x) dx
x>1 x
x
θ
θ
dy dx
1 − exp −
=
x−1 y
x>1 x
c(θ ) =
θ
1
1− 1−
=
x
x>1 x
θ dx
and, using the substitution v = 1 − 1/x, we get
c(θ ) = θ
=θ
1
0
n≥0
=θ
(1 − v)−1 1 − v θ dv
1
1
−
n+1 n+1+θ
1
1 1
+
−
θ n≥0 n + 1 n + θ
= 1 + θ γ + ψ(θ ) ,
where γ is Euler’s constant and ψ is the digamma
function.
The simple fact that one can transform the Ewens
sampling formula into the highly tractable Poisson process Y1 , Y2 , . . . , Yn using a bounded (in expectation)
number of insertions and deletions is, in itself, quite
powerful, since there are interesting aspects of the joint
distribution which are insensitive to a bounded number
of insertions and deletions. For example, consider the
Erdős–Turán law for the order of a random permutation. The order of a permutation is the least common
multiple of the lengths of its cycles, and the Erdős–
Turán law is the statement of convergence to the standard normal distribution, for the log of the order, centered by subtracting an asymptotic mean log2 n/2, and
scaling by dividing by an asymptotic standard deviation, log3/2 n/3. The effect of a finite number of cycle
lengths is washed away by the scaling; see [5] for details.
In a similar spirit, and modeled after the Feller coupling for the Ewens sampling formula, [2] shows that
for a random integer chosen uniformly from 1 to n,
the counts Cp (n) of prime factors, including multiplicity, can be coupled to independent Z2 , Z3 , Z5 , . . .
with P(Zp ≥ k) = p−k for
prime p and k = 0, 1, 2, . . .
in such a way that E p≤n |Cp (n) − Zp | ≤ 2 +
O((log log n)2 / log n); informally, the prime factorization can be converted into the process of independent
geometric random variables, using on average no more
than 2 + εn insertions and deletions. The fact of being
able to convert with o(log log n) insertions and deletions already easily implies the Hardy–Ramanujan theorem for the normal order of the number of prime
divisors,
and the fact of being able to convert with
√
o( log log n) insertions and deletions readily implies
the Erdős–Kac central limit theorem for the number of
prime divisors.
The Feller coupling expresses the Ewens sampling
formula in terms of the spacings of the independent
Bernoulli sequence ξ1 , ξ2 , . . . , ξn . The conditioning
relation, described in Crane’s article at the start of
Section 3.8, expresses the Ewens sampling formula
in terms of the independent Poisson Y1 , Y2 , . . . , Yn .
Both these independent processes have the same limit
upon rescaling, namely, the scale invariant Poisson
process on (0, ∞) with intensity θ/x dx. This leads
to a property of the scale invariant Poisson process:
the set of its spacings has the same distribution as
the set of its arrivals. This property can be exploited
to bound the distance to the Poisson–Dirichlet limit,
which is mentioned in Crane’s Section 4.2. Write
(X1 , X2 , . . .) for the random vector distributed according to the Poisson–Dirichlet(θ = 1). For random
permutations, writing Li (n) for the size of the ith
largest cycle,
[4] shows that there are couplings which
achieve E i≥1 |Li (n) − nXi | ∼ 14 log n, and that no
coupling can achieve a constant smaller than 1/4.
For prime factorizations, writing Pi (n) for the ith
largest prime factor of a random integer distributed
uniformly from 1 to n, [2] shows that there is a coupling
of random integers to Poisson–Dirichlet such that
E i≥1 | log Pi (n)− (log n)Xi | = O(log log n), and the
conjecture that O(1) can be achieved remains open.
REFERENCES
[1] A RRATIA , R. (1998). On the central role of scale invariant Poisson processes on (0, ∞). In Microsurveys in Discrete Probability (Princeton, NJ, 1997). DIMACS Ser. Discrete Math. Theoret. Comput. Sci. 41 21–41. Amer. Math. Soc.,
Providence, RI. MR1630407
DISCUSSION
[2] A RRATIA , R. (2002). On the amount of dependence in the
prime factorization of a uniform random integer. In Contemporary Combinatorics. Bolyai Soc. Math. Stud. 10 29–91. János
Bolyai Math. Soc., Budapest. MR1919568
[3] A RRATIA , R., BARBOUR , A. D. and TAVARÉ , S. (1992).
Poisson process approximations for the Ewens sampling formula. Ann. Appl. Probab. 2 519–535. MR1177897
[4] A RRATIA , R., BARBOUR , A. D. and TAVARÉ , S. (2006).
A tale of three couplings: Poisson–Dirichlet and GEM approx-
29
imations for random permutations. Combin. Probab. Comput.
15 31–62. MR2195574
[5] BARBOUR , A. D. and TAVARÉ , S. (1994). A rate for the
Erdős–Turán law. Combin. Probab. Comput. 3 167–176.
MR1288438
[6] F ELLER , W. (1945). The fundamental limit theorems in probability. Bull. Amer. Math. Soc. (N.S.) 51 800–832. MR0013252
Statistical Science
2016, Vol. 31, No. 1, 30–33
DOI: 10.1214/15-STS538
Main article DOI: 10.1214/15-STS529
© Institute of Mathematical Statistics, 2016
Relatives of the Ewens Sampling Formula
in Bayesian Nonparametrics
Stefano Favaro and Lancelot F. James
Abstract. We commend Harry Crane on his review paper which serves to
not only point out the ubiquity of the Ewens sampling formula (ESF) but also
highlights some connections to more recent developments. As pointed out by
Harry Crane, it is impossible to cover all aspects of the ESF and its relatives
in the pages generously provided by this journal. Our task is to present additional commentary in regards to some, perhaps not so well-known, related
developments in Bayesian noparametrics.
Key words and phrases: Age ordered ESF, Bayesian nonparametrics, posterior ESF, spatial neutral to the right process, species sampling problem.
where it was assumed that X = (X1 , . . . , Xn ) comes
from a true distribution P̃ . Furthermore, it was often
assumed that P̃ was a nonatomic distribution, meaning there are no ties among the Xi ’s. So while it was
of interest to calculate the Bayes estimate, which is the
prediction rule Qn (dx) = E[P (dx)|X], its behavior is
evaluated relative to P̃ . That is to say, from the point of
view of such statistical problems, the combinatorial aspects
of the joint exchangeable distribution Mn (dx) =
E[ 1≤i≤n P (dxi )] was not of primary interest.
In the case of the DP, with prior law P0,θ (dP |H ),
Mn (dx) = Mn (dx|θ ) is the Blackwell–MacQueen
Pólya urn distribution in Blackwell and MacQueen [2]
which produces the Chinese restaurant process (CRP)
and the Ewens’ sampling formula (ESF). The result
by Blackwell and MacQueen [2] shows that if the
empirical measure Pn is based on X sampled from
Mn (dx|θ ), then as n → ∞ it converges to a DP with
law P0,θ (dP |H ). The problem considered in Antoniak [1] involves placing a prior distribution π(dθ )
on θ , which necessitates the involvement of Mn (dx|θ )
and gives impetus to Antoniak’s independent derivation of the ESF. The reader is referred to Proposition 3
in Antoniak [1] for a detailed account. The model in
Lo [12] represents the prototype for the modern usage of Bayesian hierarchical mixture models where X
are viewed as latent variables drawn from Mn (dx|θ ).
The approach proposed by Lo [12] incorporates Fubini’s theorem which expresses the joint distribution
of (P , X) in terms of P |X, and X ∼ Mn (dx|θ ). Furthermore, Lemma 2 in Lo [12] establishes characteri-
1. INTRODUCTION
Let X1 , . . . , Xn |P be independent and identically
distributed as P , where P is an unknown probability
measure P (A) = Pr(X1 ∈ A) with A being a subset of
a measurable space X . Despite the complexity of X ,
the classic nonparametric
estimator of P is given by
−1
Pn (x) = n
1≤i≤n δXi (x). Bayesian nonparametric
statistics has its origins in David Blackwell’s desire in
the late 1960s to find appropriate Bayesian solutions
to this problem, that is, to find priors P (dP ) over
the space of distributions that lead to posterior distributions P (dP |X1 , . . . , Xn ) that are tractable and ideally mimic nice properties of their frequentist counterparts. The Dirichlet process (DP) by Ferguson [6] was
the first and still most prominent solution to this problem, producing a posterior distribution that is again a
DP. Doksum [3] soon followed with neutral to the right
(NTR) priors defined as random distribution functions
over the real line R, and showed that the posterior was
also a NTR process. Doksum also showed that when
X = R the DP arises as a special case of NTR priors.
These early works focused on the statistical problem
Stefano Favaro is Associate Professor, Department of
Economics and Statistics, Corso Unione Sovietica 218/bis,
10134 Torino, Italy (e-mail: stefano.favaro@unito.it).
Lancelot F. James is Professor, Department of Information
Systems, Business Statistics and Operations Management,
Clear Water Bay, Kowloon, Hong Kong (e-mail:
lancelot@ust.hk).
30
31
DISCUSSION
zations via the partition of [n] = {1, . . . , n} induced by
X and hence the CRP. Nonetheless, the combinatorial
complexity involved in such characterizations limited
their practical usage at that time.
The purpose of our exposition so far was to note that
in the 1970s and 1980s the primary focus in Bayesian
nonparametrics was on priors over spaces of measures and the somewhat limited employment of combinatorial structures such as the CRP/ESF. The field
has changed in a dramatic fashion since Ferguson [6].
Due in large part to advances in computer technology, Bayesian nonparametric ideas are now being applied to tackle a wide range of statistical problems
where parametric assumptions are often infeasible. In
particular, many applications exploit the flexible modeling features exhibited by the ESF/CRP, and generalized Mn , as applied to intricate missing data problems. The availability of explicit probability distributions for these mechanisms provides formal generative
processes that allow one to learn model structure in the
presence of incoming data via the standard Bayesian
updating mechanism. For instance, generalized notions
of the CRP provide natural priors on the space of partitions/clusters/groups that grow with the sample size.
Of note are some recent developments related to problems arising out of Bayesian machine learning; see, for
example, Teh and Jordan [14] and references therein.
Rather than recount points that are generally familiar to
a more specialized Bayesian nonparametric audience,
in the next two sections we focus on less well-known,
but highly pertinent, connections between the ESF and
its relatives and models arising in Bayesian nonparametrics.
2. AGE-ORDERED ESF AND SPATIAL NTR
PROCESSES
It is now well recognized (see, for instance, Pitman [13]) that, via Kingman’s correspondence, there
are bijective relations between the DP, the Blackwell–
McQueen Pólya urn and the ESF. Here we point
out a correspondence that is not so well known.
Let T1 , . . . , Tn |F be independent and identically distributed survival times with unknown survival distriS(t) = 1 − F (t) and cumulative hazard (t) =
bution
t
F
(ds)/S(s−).
If there are Kn = k ≤ n distinct or0
dered values T(1) > T(2) > · · · > T(k) , then, using the
language of survival analysis, one can form death sets
Dj = {i : Ti = T(j ) } and risk sets Rj = {i : Ti ≥ T(j ) },
with sizes dj= |Dj | and rj = |Rj |. Hence, rj =
rj −1 + dj = j ≤≤k d with rk = n and r0 = 0. It is
evident that Dn = {D1 , . . . , Dk } constitutes one of k!
orderings of a partition π = (B1 , . . . , Bk ) of [n]. Let
d = (d1 , . . . , dk ), and consider the probabilities
p0,θ (d) =
(2.1)
k
θk dj !
↑n
(θ ) j =1 rj
and
p̃0,θ (d) = k
n!
j =1 dj !
p0,θ (d).
These probabilities agree with variants of the ageordered ESF, as derived in Donelley and Tavaré [4],
that is, formula for k ≤ n distinct allelic types ordered
by their ages. As the name suggests, there should be
a connection between (2.1) and the DP, however, it is
not an immediately obvious one. Formula (2.1) arises
as a special case of a general formula in James [10]
and Gnedin and Pitman [7]. Specifically, (2.1) is derived by choosing F to be a NTR process, as defined in Doksum [3], which involves a representation
of
the marginal distribution Mn (dt) = E[ 1≤i≤n F (dti )]
in terms of a joint distribution of (D1 , . . . , Dk ) and
(T(1) , . . . , T(k) ). James [10] showed that if one integrates Mn (dt), with respect to (T(1) , . . . , T(k) ), then this
yields generalized distributions on Dn where p0,θ is a
special case. With the exception of the DP, Mn (dt) does
not produce independent and identically distributed
unique values, when F is selected to be a NTR process.
For specifics, Ŝ(t) = {j :T(j ) ≤t} (1 − dj /rj ) and
ˆ
(t)
= {j :T(j ) ≤t} dj /rj are the Kaplan–Meier and
Nelson–Aalen estimators for S and , respectively.
Recall that a NTR survival
process can always be rep
resented as S(t) = {j :τj ≤t} (1 − j ), where (j , τj )
are the points of a Poisson random measure N , with
mean E[N(du, ds)] = ρ(u|s)0 (ds) du, on ([0, 1],
(0, ∞)) where 0 (ds) = F0 (ds)/S0 (s−) is a hazard
rate. In particular,
it follows that F (ds) = S(s−)(ds),
where (t) = {j :τj ≤t} j , is a random cumulative
hazard corresponding to processes described in Hjort
[9]. It is known that F has the DP law P0,θ (dF |F0 ) if
−
ρ(u|s)
=
θ S0 (s−)u−1 (1
θ
S
(s−)−1
0
u)
. The choice of F that produces (2.1) is
ρ(u) = θ (θ + 1)(1 − u)θ −1 , which is the term (θ + 1)
multiplied by a Beta density function with parameter
(1, θ ). This can be seen as a special case of equation
(42) in Gnedin and Pitman [7]. The prediction rule for
this case can be deduced from James [10],
E F (dt)|T =
k+1
=1
q∗ F̃:n (dt) +
k
j =1
pj∗ δT(j ) (dt).
32
S. FAVARO AND L.F. JAMES
Here, setting 0 (s) = s, F̃:n denotes a truncated exponential distribution with parameter (θ (θ + 1)/(θ +
r−1 )(θ + r−1 + 1)) and support T() < t < T(−1) .
Furthermore, the transition probabilities for generating
Dn according to p0,θ in (2.1) are given by
pj = E pj∗ |Dn =
k
n rj (dj + 1) r
(θ + n) n rj + 1 =j +1 r + 1
and
qj = E qj∗ |Dn =
k
1
r
θ
,
(θ + n) rj −1 + 1 =j r + 1
which can be used to generate a special case of a generalized ordered CRP. See Section 5.2.1 in James [10].
F (t) is clearly not a DP. However, as in James [10], if
one adds a spatial component (j , τj , xj ) with mean
θ (θ + 1)(1 − u)θ −1 0 (ds)H (dx), then this creates a
spatial NTR process F (dt, dx), where for (t, x) ∈
R+ × X , P (dx) = F (∞, dx) has the Dirchlet law
P0,θ (dP |H ). Equation (42) in Gnedin and Pitman [7]
yields the ordered partition distribution of the twoparameter family for the range 0 ≤ α < 1 and θ ≥ 0.
Proposition 6.1 in James [10] used that result to show
that the Pitman–Yor process with law Pα,θ (dP |H ) can
be represented as a spatial NTR process F (∞, dx).
The analogous pj and qj can also be worked out
explicitly. A remarkable aspect of this is that while
Doksum’s NTR specification of F (t) does not contain the Pitman–Yor processes, there is a NTR process
F (t) = F (t, X ) which produces a random partition of
[n], (B1 , . . . , Bk ) whose law for θ ≥ 0 follows the two
parameter (α, θ ) CRP scheme.
ecology, and their importance has grown considerably in recent years, driven by challenging applications arising from bioinformatics, genetics, linguistics,
networking and data confidentiality, design of experiments, machine learning, etc. The ESF, together with
Pitman’s two-parameter generalization, represents the
cornerstone of the Bayesian nonparametric approach
proposed by Lijoi et al. [11] for making inference
measures of species variety. Indeed, it takes on the
natural interpretation of a prior “sampling” model induced by assuming a DP prior on the unknown species
composition of the population. Given that, the estimation of global and local measures of species variety relies on the study of distributional properties of
a posterior ESF, namely, distributional properties of
(M1,n , . . . , Mn+m,n+m ) given (M1,n , . . . , Mn,n ).
Under the assumption that P has the two-parameter
Poisson–Dirchlet law (i.e., Pitman–Yor process)
Pα,θ (dP |H ), the distribution of Ml,n+m given (M1,n ,
. . . , Mn,n ) arises by a direct application of Theorem 3 in Favaro et al. [5]. Along similar arguments,
one may obtain the joint conditional distribution of
(M1,n+m , . . . , Mn+m,n+m ) given (M1,n , . . . , Mn,n ). Of
particular interest in Bayesian nonparametric inference
for species sampling problems is the expectation and
the large m asymptotic behavior of the distribution of
Ml,n+m given (M1,n , . . . , Mn,n ). In particular, one has
=
l
i=1
·
(3.1)
3. POSTERIOR ESF AND SPECIES SAMPLING
PROBLEMS
Species sampling problems refer to a broad class
of statistical problems where samples are drawn from
a population of individuals belonging to an (ideally) infinite number of species with unknown proportions. Given a sample of size n featuring Kn = k
species with frequency counts (M1,n , . . . , Mn,n ) =
(m1,n , . . . , mn,n ), interest lies in estimating global and
local measures of species variety induced by considering an additional unobservable sample of size m: the
former refer to the species variety of the whole additional sample, whereas the latter refer to the discovery
probability at the (n + m + 1)th step of the sampling
process. These problems have originally appeared in
E Ml,n+m |Mn = (m1 , . . . , mn )
(θ + n − i + α)↑(m−l+i)
(θ + n)↑m
+
·
m
mi (i − α)↑(l−i)
l−i
m
(1 − α)↑(l−1) (θ + kα)
l
(θ + n + α)↑(m−l)
(θ + n)↑m
with k = 1≤i≤n mi . Furthermore, let Ba,b be a Beta
random variable with parameter (a, b), and let Sq be
a nonnegative random variable with density function
proportional to s q−1−1/α fα (s −1/α ), where fα is the
positive α-stable density function. As m → +∞,
(3.2)
Ml,n+m
Mn = (m1 , . . . , mn )
mα
→
α(1 − α)↑(l−1)
Bk+θ/α,n/α−k S(θ +n)/α
l!
33
DISCUSSION
almost surely, where Bk+θ/α,n/α−k and S(θ +n)/α are
independent random variables. The limiting scale mixture Bk+θ/α,n/α−k S(θ +n)/α provides the posterior counterpart of the well-known α-diversity of the twoparameter Ewens–Pitman sampling formula, which is
recovered by setting n = k = 0. In particular, if α = 0,
then Ml,n+m → Pl weakly m → +∞, where Pl is a
Poisson random variable with parameter θ/ l.
Let Ml,m = E[Ml,n+m |Mn = (m1 , . . . , mn )] be the
posterior expectation displayed in (3.1). Intuitively,
Ml,m takes on the interpretation of the Bayesian nonparametric estimator, with respect to a squared loss
function, of the number of species with frequency l
in the enlarged sample of size (n + m). Accordingly,
(3.2) provides a tool for deriving corresponding large
m asymptotic credible intervals. The estimator Ml,m
is the prototypical example of a measure of global
species variety, and other measures of global species
variety may be introduced as suitable functions of it.
For
instance, for any integer 1 ≤ τ ≤ n + m, Mm (τ ) =
1≤l≤τ Ml,m is the estimator of the so-called rare
species variety, namely, the number of species with
frequency less or the equal of a threshold τ . Then an
estimator of the overall species variety is obtained by
setting τ = n + m. Besides these measures of global
species variety, Ml,m also leads to measures of local
species variety. Indeed, it is easy to show that Dl,m =
(l − α)Ml,m /(θ + n + m) is the estimator of the probability that the observation at the (n + m + 1)th draw
coincideswith a species with frequency l. In particular, 1 − 1≤l≤n+m Dl,m is the estimator of the probability of discovering a new species at the (n + m + 1)th
draw. For any m ≥ 1, these estimators provide the natural Bayesian nonparametric counterparts of the celebrated Good–Toulmin estimator proposed by Good and
Toulmin [8].
ACKNOWLEDGMENTS
Stefano Favaro is also affiliated with Collegio Carlo
Alberto, Moncalieri, Italy and is supported in part by
the European Research Council through StG N-BNP
306406. Lancelot F. James is supported in part by the
Grant RGC-HKUST 601712 of the HKSAR.
REFERENCES
[1] A NTONIAK , C. E. (1974). Mixtures of Dirichlet processes
with applications to Bayesian nonparametric problems. Ann.
Statist. 2 1152–1174. MR0365969
[2] B LACKWELL , D. and M AC Q UEEN , J. B. (1973). Ferguson
distributions via Pólya urn schemes. Ann. Statist. 1 353–355.
MR0362614
[3] D OKSUM , K. (1974). Tailfree and neutral random probabilities and their posterior distributions. Ann. Probab. 2 183–201.
MR0373081
[4] D ONNELLY, P. and TAVARÉ , S. (1986). The ages of alleles
and a coalescent. Adv. in Appl. Probab. 18 1–19. MR0827330
[5] FAVARO , S., L IJOI , A. and P RÜNSTER , I. (2013). Conditional formulae for Gibbs-type exchangeable random partitions. Ann. Appl. Probab. 23 1721–1754. MR3114915
[6] F ERGUSON , T. S. (1973). A Bayesian analysis of some nonparametric problems. Ann. Statist. 1 209–230. MR0350949
[7] G NEDIN , A. and P ITMAN , J. (2005). Regenerative composition structures. Ann. Probab. 33 445–479. MR2122798
[8] G OOD , I. J. and T OULMIN , G. H. (1956). The number of
new species, and the increase in population coverage, when a
sample is increased. Biometrika 43 45–63. MR0077039
[9] H JORT, N. L. (1990). Nonparametric Bayes estimators based
on beta processes in models for life history data. Ann. Statist.
18 1259–1294. MR1062708
[10] JAMES , L. F. (2006). Poisson calculus for spatial neutral to
the right processes. Ann. Statist. 34 416–440. MR2275248
[11] L IJOI , A., M ENA , R. H. and P RÜNSTER , I. (2007). Bayesian
nonparametric estimation of the probability of discovering
new species. Biometrika 94 769–786. MR2416792
[12] L O , A. Y. (1984). On a class of Bayesian nonparametric
estimates. I. Density estimates. Ann. Statist. 12 351–357.
MR0733519
[13] P ITMAN , J. (1996). Some developments of the Blackwell–
MacQueen urn scheme. In Statistics, Probability and Game
Theory (T. S. Ferguson, L. S. Shapley and J. B. MacQueen, eds.). Institute of Mathematical Statistics Lecture
Notes—Monograph Series 30 245–267. IMS, Hayward, CA.
MR1481784
[14] T EH , Y. W. and J ORDAN , M. I. (2010). Hierarchical Bayesian nonparametric models with applications. In
Bayesian Nonparametrics (N. L. Hjort, C. C. Holmes,
P. Müller and S. G. Walker, eds.) 158–207. Cambridge Univ.
Press, Cambridge. MR2730663
Statistical Science
2016, Vol. 31, No. 1, 34–36
DOI: 10.1214/15-STS540
Main article DOI: 10.1214/15-STS529
© Institute of Mathematical Statistics, 2016
Bayesian Nonparametric Modeling and the
Ubiquitous Ewens Sampling Formula
Yee Whye Teh
I would like to thank Harry Crane for a most enlightening review of the many ways and guises in which the
Ewens sampling formula pops up throughout statistics
and mathematics. Given the simplicity and the almost
inevitability of Ewens’ sampling formula when working with distributions over partitions, one could say
that it plays a similar role for random partitions as the
normal distribution plays for random real-valued variables. And just as the normal distribution plays an important role as a core building block for more complex
models, for example, hierarchical Bayesian models or
graphical models, Ewens’ sampling formula and the
associated Chinese restaurant process distribution over
set partitions and Dirichlet process distribution over
probability measures play increasingly important roles
as building blocks of more complex Bayesian nonparametric models. Crane has noted, and I agree, that this
is “one of the most active areas of statistical research,”
whose “overwhelming activity forbids any possibility
of a satisfactory survey of the topic and promises to
quickly outdate the contents of the present section.” In
this discussion I will attempt to present an (already outdated) overview of the use of Ewens’ sampling formula
in Bayesian nonparametrics, specifically focusing on
the many creative ways the community has built more
complex models out of these simpler building blocks.
Much of the work is motivated by recent trends toward
using the analysis of “Big Data” sets to derive scientific
understanding and drive technological progress. Such
modern data sets are often not just tall, they are also
wide, and not just tall and wide, but also complex and
structured, and it is important to model the nontrivial
dependencies hidden behind the data.
Good introductions to Bayesian nonparametrics can
be found in the collection edited by Hjort et al. [14] and
the book by Ghosh and Ramamoorthi [13], while more
recent works can be found in the IEEE TPAMI special
issue [1] and a number of other forthcoming special
issues. Finally, shorter introductions and tutorials for
less mathematically inclined readers can be found in
[11, 12, 22].
1. NONPARAMETRIC MIXTURE MODELS AND
CLUSTERING
One of the most popular uses of Ewens’ sampling
formula in Bayesian nonparametrics is via the Chinese restaurant process (CRP), a distribution over set
partitions described in Section 4, for mixture modeling and clustering. Consider a data set of size n modeled as observations of exchangeable random variables
Y1 , . . . , Yn . Assuming that these come from a number
of heterogenous sources or clusters, we can model the
assignment of the observations to different sources using a partition of the index set. If the number of
sources is unknown and taking a Bayesian formalism,
a sensible prior should then place positive mass over all
possible partitions. A simple example of such a prior is
given by the Chinese restaurant process CRP([n], θ ),
leading to the following model:
∼ CRP [n], θ ,
ind.
Yi | ∼ F Xc∗ ,
i.i.d.
Xc∗ ∼ H,
where i ∈ c ∈ , Xc∗ is the unknown parameters describing cluster c in , H is its prior, and CRP([n], θ )
denotes the CRP distribution over partitions of the set
[n] = {1, . . . , n} with parameter θ . Such a model was
first proposed by Lo [17] for density estimation problems, and rediscovered for clustering in machine learning [19, 23]. It is now commonly known as the Dirichlet process mixture model, so named as the de Finetti
measure underlying the CRP mixture is the Dirichlet
process DP(θ, H ).
2. NESTED PARTITIONS AND TREES
In certain applications, for example, phylogenetics
and unsupervised categorization learning, it is of interest to model data as arising from a nested collection of
clusters. For example, a beagle is a dog is an animal is a
living organism. These can be modeled as nested partitions, for example, {{{1, 4}, {5}}, {{2, 6}, {3}}, {{7}}} is
Yee Whye Teh is Professor of Statistical Machine Learning,
Department of Statistics, University of Oxford, 1 South
Parks Road, Oxford OX1 3TG, United Kingdom (e-mail:
y.w.teh@stats.ox.ac.uk).
34
35
DISCUSSION
a two-level nested partition of {1, . . . , 7}. Distributions
over nested partitions can be constructed from CRPs
in two different ways: by subpartitioning or fragmenting the clusters of a partition recursively in a top-down
fashion or by coagulating the clusters of a partition recursively in a bottom-up fashion.
A fragmentation process starts with the trivial partition with just one cluster and recursively fragments
clusters using independent CRPs L times to construct
an L level nested partition. Such a fragmentation process was called a nested CRP in [5] who explored it as
a model for unsupervised learning of topic hierarchies
in text analysis. Conversely, a coagulation process can
start with another trivial partition with all items in their
own clusters, and recursively coagulate clusters in the
following way: If π is a partition, let κ be a partition
of the clusters in π , say, drawn from CRP(π, θ ). The
coagulation of π by κ is then π = { c∈γ c : γ ∈ κ},
where the clusters in π belonging to the same cluster in
κ are merged. This coagulation process can be shown
to be the dual genealogical process of the hierarchical
Dirichlet process [25] (described later), where it is a
simple case of the Chinese restaurant franchise.
Fragmentation and coagulation processes are more
conveniently represented mathematically as Markov
chains on set partitions, with fragmentations being sequences of partition refinements (see Section 3.5 of
main article), while coagulations are coarsenings [4].
Viewed in this way, one can also ask for continuous time limits of the Markov chains associated with
the nested CRPs and the Chinese restaurant franchise,
leading to Dirichlet diffusion trees [20] and Kingman’s
coalescents (Section 2.4) respectively. It is also possible to construct partition-valued Markov chains with
both fragmentations and coagulations in operation. The
mathematical properties of such processes were studied in [3], and they were applied to haplotype modeling
and genetic imputation in [10, 24].
3. HIERARCHICAL BAYESIAN NONPARAMETRIC
MODELS
A common theme across both frequentist and
Bayesian statistics is when data are separated into
groups and it is important to model groups individually
while sharing statistical strength across groups to provide more fine-grained control over model flexibility.
In Bayesian statistics this is achieved using hierarchical
Bayesian models where each group has an associated
random parameter with a common prior distribution
across groups parameterized by a random hyperparameter. The randomness of the hyperparameter induces
the sharing of statistical strength across groups.
In a Bayesian nonparametric setting, where the random parameter is typically an infinite-dimensional
stochastic process, control over model flexibility is arguably even more important than in parametric models.
For example, if each group is modeled with a Dirichlet process mixture, with Gj ∼ DP(θ, G0 ) for group j ,
one can place a hierarchical DP prior on the base distribution, G0 ∼ DP(θ0 , H ) [25], which induces sharing
of the mixture components across groups. Such hierarchical constructions also arise naturally elsewhere in
Bayesian nonparametrics, for example, Gaussian processes for regression [26] and beta processes/Indian
buffet processes for feature allocations [9, 27].
4. DEPENDENT AND RELATIONAL MODELS
Hierarchical models effectively assume exchangeability among groups and induce relatively simple
forms of statistical strength sharing across groups. This
can be relaxed to various forms of partial exchangeability. For example, if there are group level covariates,
or spatial or temporal structure, then dependent models reflecting this structure may be appropriate. There
are two levels at which general dependencies can be
induced. At the random measures level, for each covariate value t we introduce a random measure Gt
and work with the measure-valued stochastic process
(Gt ). When each Gt is a DP, such dependent DPs were
first explored by MacEachern [18]. At the random partitions level, one instead works with partition-valued
stochastic processes, for example, [6, 7]. A significant
number of constructions have been provided in the literature and reviewed in [8].
The random set partitions associated with the Ewens
sampling formula have also been used in modeling relational data such as social networks and collaborative
filtering. These are data where observations (e.g., of
links or friendships) are associated with relations between two or more objects, rather than with objects
themselves (although there can be object-level covariates). In the infinite relational model [16, 28], objects
are partitioned into clusters via the CRP, and observed
relations between objects are mediated by the clusters
that they belong to. For relational data, de Finetti’s theory of exchangeability is generalized to relational exchangeability by Aldous and Hoover [2, 15]; see [21]
for an introduction.
36
Y. W. TEH
5. SUMMARY
The Ewens sampling formula and the associated distributions over partitions, set partitions and probability measures have very many mathematically elegant
properties, which have been well explored in the literature and well reviewed in the present paper. With a
good understanding of such distributions and in a datarich world, the Bayesian nonparametrics community
is now engaged in the practical uses of Ewens’ sampling formula for modeling more complex phenomena. Important approaches have included covariatedependence, hierarchical Bayesian models, constructions of nested partitions and trees, and applications to
non-i.i.d. settings like relational and network data.
ACKNOWLEDGMENTS
Supported in part by the European Research Council
under the European Union’s Seventh Framework Programme (FP7/2007-2013) ERC Grant agreement no.
617071.
REFERENCES
[1] A DAMS , R. P., F OX , E. B., S UDDERTH , E. B. and
T EH , Y. W. (2015). Guest editors’ introduction to the special issue on Bayesian nonparametrics. IEEE Transactions on
Pattern Analysis and Machine Intelligence.
[2] A LDOUS , D. J. (1985). Exchangeability and related topics. In École d’été de Probabilités de Saint-Flour, XIII—
1983. Lecture Notes in Math. 1117 1–198. Springer, Berlin.
MR0883646
[3] B ERESTYCKI , J. (2004). Exchangeable fragmentationcoalescence processes and their equilibrium measures. Electron. J. Probab. 9 770–824 (electronic). MR2110018
[4] B ERTOIN , J. (2006). Random Fragmentation and Coagulation Processes. Cambridge Studies in Advanced Mathematics
102. Cambridge Univ. Press, Cambridge. MR2253162
[5] B LEI , D. M., G RIFFITHS , T. L. and J ORDAN , M. I. (2010).
The nested Chinese restaurant process and Bayesian nonparametric inference of topic hierarchies. J. ACM 57 Art. 7, 30.
MR2606082
[6] C ARON , F., DAVY, M. and D OUCET, A. (2007). Generalized Polya urn for time-varying Dirichlet process mixtures.
In Proceedings of the Conference on Uncertainty in Artificial
Intelligence 23.
[7] D UAN , J. A., G UINDANI , M. and G ELFAND , A. E. (2007).
Generalized spatial Dirichlet process models. Biometrika 94
809–825. MR2416794
[8] D UNSON , D. B. (2010). Nonparametric Bayes applications
to biostatistics. In Bayesian Nonparametrics 223–273. Cambridge Univ. Press, Cambridge. MR2730665
[9] E CK , D., B ENGIO , Y. and C OURVILLE , A. C. (2009). An infinite factor model hierarchy via a noisy-or mechanism. In Advances in Neural Information Processing Systems 405–413.
[10] E LLIOTT, L. and T EH , Y. W. (2012). Scalable imputation of
genetic data with a discrete fragmentation–coagulation process. In Advances in Neural Information Processing Systems.
[11] G ERSHMAN , S. J. and B LEI , D. M. (2012). A tutorial on
Bayesian nonparametric models. J. Math. Psych. 56 1–12.
MR2903470
[12] G HAHRAMANI , Z. (2013). Bayesian non-parametrics and
the probabilistic approach to modelling. Philos. Trans. R.
Soc. Lond. Ser. A Math. Phys. Eng. Sci. 371 20110553, 20.
MR3005667
[13] G HOSH , J. K. and R AMAMOORTHI , R. V. (2003). Bayesian
Nonparametrics. Springer, New York. MR1992245
[14] H JORT, N., H OLMES , C., M ÜLLER , P. and WALKER , S.,
eds. (2010). Bayesian Nonparametrics. Cambridge Series
in Statistical and Probabilistic Mathematics 28. Cambridge
Univ. Press, Cambridge. MR2722987
[15] H OOVER , D. (1979). Relations on probability spaces and arrays of random variables. Technical report, Princeton, NJ.
[16] K EMP, C., T ENENBAUM , J. B., G RIFFITHS , T. L., YA MADA , T. and U EDA , N. (2006). Learning systems of concepts with an infinite relational model. In Proceedings of the
AAAI Conference on Artificial Intelligence 21.
[17] L O , A. Y. (1984). On a class of Bayesian nonparametric
estimates. I. Density estimates. Ann. Statist. 12 351–357.
MR0733519
[18] M AC E ACHERN , S. (1999). Dependent nonparametric processes. In Proceedings of the Section on Bayesian Statistical
Science. Amer. Statist. Assoc., Alexandria, VA.
[19] N EAL , R. M. (1992). Bayesian mixture modeling. In Proceedings of the Workshop on Maximum Entropy and Bayesian
Methods of Statistical Analysis 11 197–211.
[20] N EAL , R. M. (2001). Defining priors for distributions using
Dirichlet diffusion trees. Technical Report 0104, Dept. Statistics, Univ. Toronto.
[21] O RBANZ , P. and ROY, D. M. (2015). Bayesian models of
graphs, arrays and other exchangeable random structures.
IEEE Transactions on Pattern Analysis and Machine Intelligence Special Issue on Bayesian Nonparametrics.
[22] O RBANZ , P. and T EH , Y. W. (2010). Bayesian nonparametric models. In Encyclopedia of Machine Learning. Springer,
Berlin.
[23] R ASMUSSEN , C. E. (2000). The infinite Gaussian mixture
model. In Advances in Neural Information Processing Systems 12.
[24] T EH , Y. W., B LUNDELL , C. and E LLIOTT, L. T. (2011).
Modelling genetic variations with fragmentation–coagulation
processes. In Advances in Neural Information Processing Systems.
[25] T EH , Y. W., J ORDAN , M. I., B EAL , M. J. and B LEI , D. M.
(2006). Hierarchical Dirichlet processes. J. Amer. Statist. Assoc. 101 1566–1581. MR2279480
[26] T EH , Y. W., S EEGER , M. and J ORDAN , M. I. (2005). Semiparametric latent factor models. In Proceedings of the International Workshop on Artificial Intelligence and Statistics 10.
[27] T HIBAUX , R. and J ORDAN , M. I. (2007). Hierarchical beta
processes and the Indian buffet process. In Proceedings of the
International Workshop on Artificial Intelligence and Statistics 11 564–571.
[28] X U , Z., T RESP, V., Y U , K. and K RIEGEL , H.-P. (2006). Infinite hidden relational models. In Proceedings of the Conference on Uncertainty in Artificial Intelligence 22.
Statistical Science
2016, Vol. 31, No. 1, 37–39
DOI: 10.1214/15-STS544
Main article DOI: 10.1214/15-STS529
© Institute of Mathematical Statistics, 2016
Rejoinder: The Ubiquitous Ewens
Sampling Formula
Harry Crane
The main article and extended discussion point to
Ewens’s sampling formula (ESF) as one of a few essential probability distributions. Arratia, Barbour and
Tavaré explain the emergence of ESF by the Feller coupling and also touch on number theoretic considerations; Feng provides deeper background on diffusion
processes and nonequilibrium versions of ESF; and
McCullagh regales us with a story from the works of
Fisher and Good, putting historical context around the
more specialized topics covered by Favaro and James
and Teh. The breadth of these comments exemplifies
the expansive sphere of influence of Ewens’s sampling
formula on integer partitions, Ewens’s distribution on
set partitions, and the Ewens process. I thank all of the
discussants for their participation in this important survey.
For the most part, these contributions bolster my
main thesis which, in the words of Arratia, Barbour and Tavaré, emphasizes the universal character of the Ewens sampling formula. As McCullagh
notes, the contents and subsequent discussion comprise an impressive list stretching from literary studies to population genetics and probabilistic number
theory. Both comments accord with my opening remark that Ewens’s sampling formula exemplifies the
harmony of mathematical theory, statistical application, and scientific discovery. As a whole, however, the
discussion skews disproportionately toward Bayesian
nonparametrics in a way that works against the theme
of ubiquity. I attempt to rebalance the conversation in
these final pages.
a further analogy between the Ewens process and the
Poisson process for events in time or space. Its tangible connections to population genetics, inductive inference, stochastic process theory, prime factorization,
and statistical applications earn Ewens’s sampling formula and the Poisson–Dirichlet distribution a place
alongside the Bernoulli, Gaussian, and Poisson in the
pantheon of probability distributions.
The applicability of Ewens’s sampling formula is
neither limited to specific methods nor tied to ongoing
trends: Teh centers his commentary around contemporary topics in machine learning and big data, Favaro
and James deal with problems in survival modeling
and species sampling, and McCullagh showcases the
adaptability of ESF with an enlightening application to
a problem considered by Fisher three decades before
Ewens’s discovery. As McCullagh details, Ewens’s
sampling formula and its derivatives, the Ewens distribution and Ewens process, could have—indeed, should
have—been first discovered in a purely parametric context, when data sets were small and computers were in
their infancy.
McCullagh rightly identifies Ewens’s process as one
of a small number of processes that deserves to be a
central part of the statistical curriculum. Indeed, there
are compelling reasons to teach ESF at every level of
statistics, and yet it is often reserved for special topics
or not covered at all. Its most salient features, namely,
exchangeability, sampling consistency, and noninterference, highlight subtleties that do not arise in i.i.d.
sampling models and which can be covered without
any need to delve into population genetics, stochastic
processes, or Bayesian nonparametrics.
1. EWENS’S SAMPLING FORMULA IN MODERN
STATISTICS
Wherever random partitions appear, with few exceptions, so does Ewens’s sampling formula. Teh compares its inevitability to that of the Gaussian distribution for real-valued sequences, and McCullagh makes
2. EWENS’S SAMPLING FORMULA AND
BAYESIAN NONPARAMETRICS
Of the three commentaries covering statistical elements of ESF, two (Favaro and James, Teh) focus
on recent work in Bayesian nonparametrics while the
other (McCullagh) presents an application from seventy years ago. Together these comments fit into a
Harry Crane is Assistant Professor of Statistics &
Biostatistics, Rutgers, the State University of New Jersey,
110 Frelinghuysen Road, Room 501, Piscataway, New
Jersey 08854, USA (e-mail: hcrane@stat.rutgers.edu).
37
38
H. CRANE
broader, but misleading, narrative that Bayesian nonparametrics is the lifeblood of ESF in present-day statistical research. Though several authors do build substantially on the prior work of Ewens, Kingman, and
Pitman, for example, Ishwaran and James’s [10] work
on the generalized Chinese restauarant process, Favaro,
et al.’s [8] analysis of conditional sampling formulas,
and Ruggiero and Walker’s [12] study of the Fleming–
Viot process, the Dirichlet process prior remains the
primary mechanism by which Ewens’s sampling formula arises in Bayesian nonparametrics. I have two
major comments regarding how this connection is covered in the larger literature.
First, of all the recent surveys cited by Teh ([8, 11,
12, 13, 14, 21, 22] in Teh’s numbering), only one [14],
page 108, acknowledges Ewens’s 1972 article or refers
to Ewens’s sampling formula by name. This tendency
isolates the occurrence of ESF in Bayesian nonparametrics from the rest of the literature, fostering the impression that ESF is a byproduct of purely nonparametric Bayesian concerns. Second, the Dirichlet process is primarily chosen to address practical concerns
of tractability [and] computational convenience ([9],
page 37), which sell short the ESF’s more critical statistical and inferential properties (Sections 3 and 7).
Both of these oversights undermine the significance of
the Ewens family of distributions: the first completely
ignores the larger body of work on ESF and the second
presents ESF merely as a quick fix for computational
challenges.
3. THREE VIGNETTES ON EWENS’S SAMPLING
FORMULA IN POPULATION GENETICS
While it is true that Bayesian nonparametrics is one
of the most active areas of statistical research, the field
of population genetics provides the primary context
and is the most prominent venue for ESF. Notwithstanding Feng’s account, which provides an insightful overview of how variants of ESF arise by diffusion
process approximations, the population genetics angle
warrants much more attention than it has received so
far. Below I touch on three direct consequences of ESF
in population genetics.
First, Ewens’s derivation had an immediate impact
on mutation rate estimation. Before [7], geneticists estimated θ , or functions of θ , from the empirical allele
frequencies. Ewens showed that the number K of alleles is a sufficient statistic for θ , indicating that these
early procedures used precisely the wrong part of the
data in estimation of θ . This is a rare, and perhaps
unique, example of a case where a previously unsuspected sufficient statistic changed standard inference
procedures.
Second, some geneticists, including Wright [13],
claimed that in the selectively neutral case all alleles
observed in a sample should have approximately equal
frequencies. Ewens’s sampling formula shows that this
is the least likely outcome under selective neutrality.
The two main reasons for this phenomenon are simple random sampling and history—older alleles have a
greater probability of reaching a high frequency than
alleles that have recently arisen by mutation. This observation is relevant when testing whether data from a
sample of genes supports the neutrality hypothesis.
The third and most lasting effect of Ewens’s sampling formula is that it partially influenced Kingman’s
development of the coalescent [11], now the main vehicle for research in population genetics. The coalescent leads not only to a beautiful mathematical theory, which still provides the most elegant derivation of
ESF, but also to a practical scientific framework which
has moved the field of theoretical population genetics toward largely retrospective questions like: “When
did the most recent ancestor of all humans alive today
live?” and “How can we detect the signatures of past
selective events in contemporary genomes?”
4. OTHER INSTANCES OF EWENS’S SAMPLING
FORMULA
4.1 Independent Process Approximation and the
Feller Coupling
Arratia, Barbour, and Tavaré expound a clear and
well-motivated account of how Ewens’s sampling formula emerges from the Feller coupling, which rightly
deserves a place in the main survey alongside the Chinese restaurant process (CRP). The Feller coupling is
more mathematically natural than the CRP construction, and it also illustrates the powerful technique of
approximating statistics of combinatorial structures using independent processes; see [2].
4.2 Markov Survival Processes
Favaro and James discuss a connection between
Ewens’s sampling formula and neutral to the right survival models in Bayesian nonparametrics. Dempsey
and McCullagh [6] observe the same connection but
without resorting to the Bayesian nonparametrics
framework. In the so-called pilgrim process, risk sets
evolve according to an asymmetric version of Aldous’s
beta-splitting model [1] with parameter β > −1. The
39
REJOINDER
β = 1 case yields a random partition distributed according to Ewens’s distribution which, upon extension
to recurrent events, elicits a connection to the so-called
Indian buffet process.
4.3 Scale-Free Interaction Networks
The family of Ewens distributions also comes up in
ongoing work on statistical network analysis. The degree distributions of many observed networks behave
according to a power law, that is, the proportion pk
of vertices with degree k ≥ 1 grows like pk ∼ k −γ
as k → ∞ for some γ > 1. Barabási and Albert’s [3]
preferential attachment model is the most widely cited
generating mechanism for power-law networks, but
its dynamics do not translate to a viable statistical
model for two important reasons. First, its dynamics
are too rigid to adequately reflect how most networks
form, and its lack of exchangeability often prevents inference beyond selected summary statistics. Second,
the preferential attachment dynamics can only explain
power-law behavior in the range γ > 2, but Crane and
Dempsey [4] have recently found that many networks
formed by repeated interactions within a population exhibit power-law behavior with exponent in the complementary range 1 < γ < 2. Based on the Ewens–Pitman
two-parameter family (Section 5.1), we have put forth
a new model that produces a network with the powerlaw exponent in the correct range. Our model is a precursor to the broader framework of edge exchangeable
network models [5], which is the correct notion of invariance for many network data sets.
View publication stats
REFERENCES
[1] A LDOUS , D. (1996). Probability distributions on cladograms. In Random Discrete Structures (Minneapolis, MN,
1993). IMA Vol. Math. Appl. 76 1–18. Springer, New York.
MR1395604
[2] A RRATIA , R. and TAVARÉ , S. (1994). Independent process
approximations for random combinatorial structures. Adv.
Math. 104 90–154. MR1272071
[3] BARABÁSI , A.-L. and A LBERT, R. (1999). Emergence
of scaling in random networks. Science 286 509–512.
MR2091634
[4] C RANE , H. and D EMPSEY, W. (2015). Atypical scaling behavior persists in real world interaction networks. Available
at arXiv:1509.08184.
[5] C RANE , H. and D EMPSEY, W. (2015). Edge exchangeable
network models and the power law. Unpublished manuscript.
[6] D EMPSEY, W. and M C C ULLAGH , P. (2015). The pilgrim
process. Available at arXiv:1412.1490.
[7] E WENS , W. J. (1972). The sampling theory of selectively neutral alleles. Theoret. Population Biology 3 87–112.
MR0325177
[8] FAVARO , S., L IJOI , A. and P RÜNSTER , I. (2013). Conditional formulae for Gibbs-type exchangeable random partitions. Ann. Appl. Probab. 23 1721–1754. MR3114915
[9] G HOSAL , S. (2010). The Dirichlet process, related priors and
posterior asymptotics. In Bayesian Nonparametrics (N. L.
Hjort, C. Holmes, P. Müller and S. G. Walker, eds.) 35–79.
Cambridge Univ. Press, Cambridge. MR2730660
[10] I SHWARAN , H. and JAMES , L. F. (2003). Generalized
weighted Chinese restaurant processes for species sampling
mixture models. Statist. Sinica 13 1211–1235. MR2026070
[11] K INGMAN , J. F. C. (1982). The coalescent. Stochastic Process. Appl. 13 235–248. MR0671034
[12] RUGGIERO , M. and WALKER , S. G. (2009). Bayesian nonparametric construction of the Fleming–Viot process with fertility selection. Statist. Sinica 19 707–720. MR2514183
[13] W RIGHT, S. (1978). Evolution and the Genetics of Populations 4. Univ. Chicago Press, Chicago, IL.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )