Jump to content

Lehmann–Scheffé theorem

From Wikipedia, the free encyclopedia

In statistics, the Lehmann–Scheffé theorem provides sufficient conditions for the existence of a best unbiased estimator in a statistical model. The theorem states that any unbiased estimator for a quantity that depends on the data only through a complete, sufficient statistic is the unique uniformly minimum-variance unbiased estimator (UMVUE) of that quantity. The Lehmann–Scheffé theorem is named after Erich Leo Lehmann and Henry Scheffé, given their two early papers.[1][2]

Introduction

[edit]

Given a vector {\displaystyle X=(X_{1},X_{2},\dots ,X_{n})} of random samples from a distribution {\displaystyle \mathbb {P} _{\theta }} for some parameter {\displaystyle \theta \in \Theta }, the goal is to establish sufficient conditions for the existence of an UMVU estimator {\displaystyle T^{\ast }(X)} for some quantity {\displaystyle g(\theta )}, that is, {\displaystyle \mathbb {E} _{\theta }[T^{\ast }(X)]=g(\theta )} and for any unbiased estimator {\displaystyle S(X)} it holds

{\displaystyle \mathrm {Var} _{\theta }(T^{\ast }(X))\leq \mathrm {Var} _{\theta }(S(X)),\quad \forall \theta \in \Theta .}

The Rao–Blackwell theorem already shows that, given a sufficient statistic {\displaystyle T}, the estimator {\displaystyle T^{\ast }(X)=\mathbb {E} [S(X)\mid T(X)]} has a uniformly smaller variance than {\displaystyle S(X)}, but it does not guarantee that {\displaystyle T^{\ast }} is already UMVU. This is where the Lehmann–Scheffé theorem comes in, if {\displaystyle T} is also complete, that is, for any real-valued measurable function {\displaystyle \varphi } it holds

{\displaystyle \mathbb {E} _{\theta }[\varphi (T(X))]=0\implies \varphi (T(X))=0\quad \mathbb {P} _{\theta }{\text{-a.s.}}}

Statement

[edit]

As above, let {\displaystyle X=(X_{1},X_{2},\dots ,X_{n})} be a vector of random samples from a distribution {\displaystyle \mathbb {P} _{\theta }} for some parameter {\displaystyle \theta \in \Theta } and {\displaystyle \Theta } an arbitrary set.

Assume that there exists a complete, sufficient statistic {\displaystyle T} for the family of distributions {\displaystyle (\mathbb {P} _{\theta })_{\theta \in \Theta }}. Then, the following two equivalent statements hold:

  • There exists at most one measurable function {\displaystyle h} such that {\displaystyle T^{\ast }=h(T(X))} is unbiased for {\displaystyle g(\theta )} and {\displaystyle \mathbb {E} _{\theta }[T^{\ast }(X)^{2}]<\infty } for all {\displaystyle \theta \in \Theta }, in which case {\displaystyle T^{\ast }(X)} is the unique UMVUE for {\displaystyle g(\theta )}.[3]
  • For any unbiased estimator {\displaystyle S(X)}, if it exists, with {\displaystyle \mathbb {E} _{\theta }[S(X)^{2}]<\infty } for all {\displaystyle \theta \in \Theta }, the estimator {\displaystyle T^{\ast }(X)=\mathbb {E} _{\theta }[S(X)\mid T(X)]} is the unique UMVUE for {\displaystyle g(\theta )}.[4]

In fact, the theorem does not state that unbiased estimators exist in the first place. However, if they do, then there exists a unique square-integrable UMVUE. Moreover, the estimator {\displaystyle T^{\ast }(X)=\mathbb {E} _{\theta }[S(X)\mid T(X)]} does neither depend on {\displaystyle \theta }, since {\displaystyle T} is sufficient, nor on {\displaystyle S}, since {\displaystyle T} is also complete.

Proof

[edit]

In the following, the dependence of an estimator on the data {\displaystyle X} will not be written out explicitly, i.e., we write {\displaystyle S} instead of {\displaystyle S(X)}.

First of all, if there is no unbiased estimator for {\displaystyle g(\theta )}, then there is obviously no UMVUE, and if all unbiased estimator are not square-integrable, then their variances are infinity and the statement is trivial. Thus, we focus on the case where a square-integrable unbiased estimator exists.

Uniqueness of {\displaystyle h}: Let {\displaystyle S=h_{1}(T)} and {\displaystyle R=h_{2}(T)} be unbiased estimators of {\displaystyle g(\theta )} for some measurable functions {\displaystyle h_{1}} and {\displaystyle h_{2}}. For the expectation of the difference it holds{\displaystyle \mathbb {E} _{\theta }[h_{1}(T)-h_{2}(T)]=g(\theta )-g(\theta )=0.}Since {\displaystyle T} is complete, this implies {\displaystyle h_{1}-h_{2}=0}. Thus, {\displaystyle T^{\ast }=h_{1}(T)=:h(T)} is the unique unbiased estimator that is a function of {\displaystyle T}.[3]

{\displaystyle h(T)} is the UMVUE: Let {\displaystyle S} be any square-integrable unbiased estimator and {\displaystyle S^{\ast }=\mathbb {E} _{\theta }[S\mid T]}. By the factorization lemma there exists measurable function {\displaystyle {\tilde {h}}} such that {\displaystyle S^{\ast }={\tilde {h}}(T)} and since {\displaystyle S^{\ast }} is unbiased as well, by the above, it must hold {\displaystyle S^{\ast }={\tilde {h}}(T)=h(T)=T^{\ast }}. Thus, by the Rao–Blackwell theorem, it follows{\displaystyle \mathrm {Var} _{\theta }(T^{\ast })=\mathrm {Var} _{\theta }(S^{\ast })=\mathrm {Var} _{\theta }(\mathbb {E} _{\theta }[S\mid T])\leq \mathrm {Var} _{\theta }(S),\quad \forall \theta \in \Theta .}In other words, {\displaystyle T^{\ast }=h(T)} is an UMVUE and according to the first part it is unique.[4]

Application

[edit]

The Lehmann–Scheffé theorem motivates two general methods to construct UMVU estimators for {\displaystyle g(\theta )} in models which allow for a complete sufficient statistic {\displaystyle T}.[5][6]

Method 1: Determining the function {\displaystyle h}

The UMVUE, if it exists, is the (unique) solution of the equation{\displaystyle \mathbb {E} _{\theta }[h(T)]=g(\theta )}for all {\displaystyle \theta \in \Theta }. If {\displaystyle h} is, for example, a linear function, then solving this equation is fairly easy.

Method 2: Conditioning on an unbiased estimator

First, it suffices to find any unbiased estimator {\displaystyle S} of {\displaystyle g(\theta )}, which is often easily feasible. The UMVUE can then be determined by evaluating the condition expectation {\displaystyle \mathbb {E} _{\theta }[S\mid T]}. Since the choice of {\displaystyle S} is arbitrary, it is preferable to choose it such that the conditional expectation is as simple as possible.

Examples

[edit]

Bernoulli distribution

[edit]

Let {\displaystyle X_{1},\dots ,X_{n}} be Bernoulli-distributed with probability {\displaystyle p\in (0,1)}. The joint probability mass function {\displaystyle f_{p}} is of the form{\displaystyle f_{p}(x_{1},\dots ,x_{n})=\prod _{i=1}^{n}p^{x_{i}}(1-p)^{1-x_{i}}=\left({\frac {p}{1-p}}\right)^{\sum _{i=1}^{n}x_{i}}(1-p)^{n},\quad x_{i}\in \{0,1\},}thus, by the Fisher–Neyman factorization theorem, {\displaystyle T(x)=\sum _{i=1}^{n}x_{i}} is a sufficient statistic. It is also complete: Let {\displaystyle \varphi } be any measurable function such that {\displaystyle \mathbb {E} _{\theta }[\varphi (T(X))]=0}. Since {\displaystyle T} is {\displaystyle \mathrm {Ber} (n,p)}-distributed, this means{\displaystyle 0=\mathbb {E} _{\theta }[\varphi (T(X))]=\sum _{k=0}^{n}\varphi (k){\binom {n}{k}}p^{k}(1-p)^{n-k}=(1-p)^{n}\sum _{k=0}^{n}\varphi (k){\binom {n}{k}}r^{k},}with {\displaystyle r=p/(1-p)}. Since the right hand side is a polynomial in {\displaystyle r>0} that is equal to zero, each coefficient must be zero as well, implying {\displaystyle \varphi (k)=0} for all {\displaystyle k}.

For the estimation of the parameter {\displaystyle p}, it is now easy to see that the UMVUE is the sample mean{\displaystyle {\overline {X}}={\frac {1}{n}}\sum _{i=1}^{n}X_{i}={\frac {1}{n}}T(X),}since it is unbiased and a function of {\displaystyle T(X)}.

Finding the UMVUE for the parameter {\displaystyle p(1-p)} (that is, the variance of the distribution) is less obvious. According to method 1, we seek a function {\displaystyle h} such that {\displaystyle p(1-p)=\mathbb {E} _{p}[h(T(X))]=\sum _{k=0}^{n}h(k){\binom {n}{k}}p^{k}(1-p)^{n-k}.}By defining {\displaystyle \rho =p/(1-p)}, the above rewrites as{\displaystyle \sum _{k=0}^{n}h(k){\binom {n}{k}}\rho ^{k}=\rho (1+\rho )^{n-2}=\sum _{k=1}^{n-1}{\binom {n-1}{k-1}}\rho ^{k}.}Comparing the coefficients shows that {\displaystyle h(k)=k(n-k)/(n(n-1))}, thus, the UMVUE is given by[5]{\displaystyle h(T(X))={\frac {n}{n-1}}{\overline {X}}(1-{\overline {X}}).}Noting that in the Bernoulli model {\displaystyle X_{i}=X_{i}^{2}} and that {\displaystyle \sum _{i=1}^{n}(X_{i}-{\overline {X}})^{2}=\sum _{i=1}^{n}X_{i}^{2}-n{\overline {X}}^{2}}, the UMVUE is, in fact, the unbiased sample variance {\displaystyle S_{n}^{2}}.

Uniform distribution

[edit]

Let {\displaystyle X_{1},\dots ,X_{n}} be uniformly distributed on the interval {\displaystyle (0,\theta )} for some {\displaystyle \theta } to be determined. The joint probability mass function {\displaystyle f_{p}} is of the form{\displaystyle f_{p}(x_{1},\dots ,x_{n})=\prod _{i=1}^{n}{\frac {1}{\theta }}1_{(0,\theta )}(x_{i})={\frac {1}{\theta ^{n}}}1_{[0,\infty )}\left(\min _{i=1,\dots ,n}x_{i}\right)1_{(-\infty ,\theta ]}\left(\max _{i=1,\dots ,n}x_{i}\right),\quad x\in \mathbb {R} ^{n},}thus, by the Fisher–Neyman factorization theorem, {\displaystyle T(x)=\max _{i=1,\dots ,n}x_{i}} is a sufficient statistic.

To see the completeness of {\displaystyle T}, the distribution of {\displaystyle T(X)} is

{\displaystyle \mathbb {P} _{\theta }(T(X)\leq t)=\mathbb {P} _{\theta }(X_{i}\leq t,\forall i=1,\dots ,n)=\prod _{i=1}^{n}\mathbb {P} _{\theta }(X_{i}\leq t)=\prod _{i=1}^{n}{\frac {t}{\theta }}={\frac {t^{n}}{\theta ^{n}}}.}

Thus, for any measurable function {\displaystyle \varphi } such that {\displaystyle \mathbb {E} _{\theta }[\varphi (T(X))]=0} it holds for any {\displaystyle \theta >0}

{\displaystyle 0=\mathbb {E} _{\theta }[\varphi (T(X))]=\int _{0}^{\theta }\varphi (x){\frac {n}{\theta ^{n}}}x^{n-1}\mathrm {d} x\implies \int _{0}^{\theta }\varphi (x)x^{n-1}\mathrm {d} x=0,}

and consequently for any {\displaystyle b>a>0}

{\displaystyle \int _{a}^{b}\varphi (x)x^{n-1}\mathrm {d} x=\int _{0}^{b}\varphi (x)x^{n-1}\mathrm {d} x-\int _{0}^{a}\varphi (x)x^{n-1}\mathrm {d} x=0.}

This implies that {\displaystyle \varphi =0} (almost everywhere) on {\displaystyle [0,\infty )}. So {\displaystyle T} is also complete. Its expectation is

{\displaystyle \mathbb {E} _{\theta }[T(X)]=\int _{0}^{\theta }x{\frac {n}{\theta ^{n}}}x^{n-1}\mathrm {d} x={\frac {n}{\theta ^{n}}}{\frac {\theta ^{n+1}}{n+1}}={\frac {n}{n+1}}\theta .}

Thus, the rescaled estimator

{\displaystyle T^{\ast }(X)={\frac {n+1}{n}}T(X)={\frac {n+1}{n}}\max _{i=1,\dots ,n}X_{i}}

is unbiased and since it is a function of {\displaystyle T}, it is the UMVUE for {\displaystyle \theta }.[7]

Exponential distribution

[edit]

Let {\displaystyle X_{1},\dots ,X_{n}} be exponentially distributed with parameter {\displaystyle \lambda } and suppose the parameter {\displaystyle g(\lambda )=\mathbb {P} _{\lambda }(X_{1}>t)} should be estimated for some fixed {\displaystyle t>0}. A straightforward unbiased estimator for {\displaystyle g(\lambda )} would be {\displaystyle {\widehat {g(\lambda )}}={\frac {1}{n}}\sum _{i=1}^{n}1_{\{X_{i}>t\}}}, however, it turns out to be sub-optimal.

Since the exponential distribution is an exponential family, it is known that its sufficient statistic {\displaystyle {\overline {X}}={\frac {1}{n}}\sum _{i=1}^{n}X_{i}} is also complete. Moreover, the estimator {\displaystyle T=1_{\{X_{1}>t\}}} is unbiased. Thus, according to method 2, the estimator {\displaystyle T^{\ast }=\mathbb {E} [T\mid {\overline {X}}]=\mathbb {P} (X_{1}>t\mid {\overline {X}})} is the UMVUE. It remains to evaluate the conditional expectation.

First of all, {\displaystyle X_{1}/(n{\overline {X}})\sim \mathrm {Beta} (1,n-1)}, which is independent of {\displaystyle \lambda } (also called ancillary). In fact, since {\displaystyle X_{1}} and {\displaystyle S=\sum _{i=2}^{n}X_{i}} are independent and {\displaystyle S\sim \Gamma (n-1,\lambda )}, the cumulative distribution function at {\displaystyle y\in (0,1)} reads as

{\displaystyle {\begin{aligned}\mathbb {P} \left({\frac {X_{1}}{n{\overline {X}}}}\leq y\right)&=\mathbb {P} \left({\frac {X_{1}}{X_{1}+S}}\leq y\right)\\&=\int _{0}^{\infty }\int _{0}^{\infty }1_{\{x/(x+s)\leq y\}}\lambda e^{-\lambda x}{\frac {\lambda ^{n-1}}{(n-2)!}}s^{n-2}e^{-\lambda s}\mathrm {d} x\mathrm {d} s\\&=\int _{0}^{1}\int _{0}^{\infty }1_{\{u\leq y\}}{\frac {\lambda ^{n}}{(n-2)!}}r^{n-2}(1-u)^{n-2}e^{-\lambda r}r\mathrm {d} r\mathrm {d} u\\&=\int _{0}^{1}1_{\{u\leq y\}}(1-u)^{n-2}{\frac {1}{(n-2)!}}\underbrace {\int _{0}^{\infty }\lambda ^{n}r^{n-1}e^{-\lambda r}\mathrm {d} r} _{=(n-1)!}\mathrm {d} u\\&=\int _{0}^{y}(n-1)(1-u)^{n-2}\mathrm {d} u,\end{aligned}}}

with the substitution {\displaystyle x=ru} and {\displaystyle s=r(1-u)}. Thus, by Basu's theorem, {\displaystyle X_{1}/(n{\overline {X}})} and {\displaystyle {\overline {X}}} are independent and the conditional expectation simplifies to

{\displaystyle \mathbb {P} (X_{1}>t\mid {\overline {X}}=x)=\mathbb {P} \left({\frac {X_{1}}{n{\overline {X}}}}>{\frac {t}{n{\overline {X}}}}\mid {\overline {X}}=x\right)=\mathbb {P} \left({\frac {X_{1}}{n{\overline {X}}}}>{\frac {t}{nx}}\right)=\left(1-{\frac {t}{nx}}\right)^{n-1}.}

due to the above derivation.[8] The constructed estimator

{\displaystyle T^{\ast }=\left(1-{\frac {t}{n{\overline {X}}}}\right)^{n-1}}

is unbiased and has a smaller variance than the naive estimator {\displaystyle {\widehat {g(\lambda )}}} for all possible {\displaystyle \lambda >0}.

Counterexample with incomplete statistics

[edit]

An example of an improvable Rao–Blackwell improvement, when using a minimal sufficient statistic that is not complete, was provided by Galili and Meilijson in 2016.[9] Let {\displaystyle X_{1},\ldots ,X_{n}} be a random sample from a scale-uniform distribution {\displaystyle X\sim U((1-k)\theta ,(1+k)\theta ),} with unknown mean {\displaystyle \operatorname {E} [X]=\theta } and known design parameter {\displaystyle k\in (0,1)}. In the search for "best" possible unbiased estimators for {\displaystyle \theta }, it is natural to consider {\displaystyle X_{1}} as an initial (crude) unbiased estimator for {\displaystyle \theta } and then try to improve it. Since {\displaystyle X_{1}} is not a function of {\displaystyle T=\left(X_{(1)},X_{(n)}\right)}, the minimal sufficient statistic for {\displaystyle \theta } (where {\displaystyle X_{(1)}=\min _{i}X_{i}} and {\displaystyle X_{(n)}=\max _{i}X_{i}}), it may be improved using the Rao–Blackwell theorem as follows:

{\displaystyle {\hat {\theta }}_{RB}=\operatorname {E} _{\theta }[X_{1}\mid X_{(1)},X_{(n)}]={\frac {X_{(1)}+X_{(n)}}{2}}.}

However, the following unbiased estimator can be shown to have lower variance:

{\displaystyle {\hat {\theta }}_{LV}={\frac {1}{k^{2}{\frac {n-1}{n+1}}+1}}\cdot {\frac {(1-k)X_{(1)}+(1+k)X_{(n)}}{2}}.}

And in fact, it could be even further improved when using the following estimator:

{\displaystyle {\hat {\theta }}_{\text{BAYES}}={\frac {n+1}{n}}\left[1-{\frac {{\frac {X_{(1)}(1+k)}{X_{(n)}(1-k)}}-1}{\left({\frac {X_{(1)}(1+k)}{X_{(n)}(1-k)}}\right)^{n+1}-1}}\right]{\frac {X_{(n)}}{1+k}}}

The model is a scale model. Optimal equivariant estimators can then be derived for loss functions that are invariant.[10]

See also

[edit]

Notes

[edit]
  1. Lehmann, E. L.; Scheffé, H. (1950). "Completeness, similar regions, and unbiased estimation. I." Sankhyā. 10 (4): 305–340. doi:10.1007/978-1-4614-1412-4_23. JSTOR 25048038. MR 0039201.
  2. Lehmann, E.L.; Scheffé, H. (1955). "Completeness, similar regions, and unbiased estimation. II". Sankhyā. 15 (3): 219–236. doi:10.1007/978-1-4614-1412-4_24. JSTOR 25048243. MR 0072410.
  3. 1 2 Lehmann & Casella (1998), p. 87
  4. 1 2 Czado & Schmidt (2011), p. 111
  5. 1 2 Lehmann & Casella (1998), p. 88–89
  6. Shao (2003), p. 162
  7. Lehmann & Casella (1998), p. 89
  8. Shao (2003), p. 163–164
  9. Tal Galili; Isaac Meilijson (31 Mar 2016). "An Example of an Improvable Rao–Blackwell Improvement, Inefficient Maximum Likelihood Estimator, and Unbiased Generalized Bayes Estimator". The American Statistician. 70 (1): 108–113. doi:10.1080/00031305.2015.1100683. PMC 4960505. PMID 27499547.
  10. Taraldsen, Gunnar (2020). "Micha Mandel (2020), "The Scaled Uniform Model Revisited," The American Statistician, 74:1, 98–100: Comment". The American Statistician. 74 (3): 315. doi:10.1080/00031305.2020.1769727. S2CID 219493070.

References

[edit]
  • Casella, George; Berger, Roger L. (2002). Statistical Inference (2nd ed.). Duxbury. ISBN 0-534-24312-6.
  • Czado, Claudia; Schmidt, Thorsten (2011). Mathematische Statistik [Mathematical Statistics] (in German). Springer. ISBN 978-3-642-17261-8.
  • Lehmann, Erich Leo; Casella, George (1998). Theory of Point Estimation (2nd ed.). New York: Springer. ISBN 0-387-98502-6.
  • Shao, Jun (2003). Mathematical Statistics (2nd ed.). Springer. ISBN 978-0-387-95382-3.
Lehmann–Scheffé theorem
Morty Proxy This is a proxified and sanitized view of the page, visit original site.