Generalized extreme value distribution

From Wikipedia, the free encyclopedia
(Redirected from GEV distribution)

Template:Short description

Page Module:Message box/ambox.css has no content.

Page Module:Infobox/styles.css has no content.Page Template:Infobox probability distribution/styles.css has no content.

Notation GEV(μ,σ,ξ)
Parameters μ  (location)
σ>0 (scale)
ξ (shape)
Support {x[μσξ,+) when ξ>0x(,) when ξ=0x(,μσξ] when ξ<0
PDF 1σt(x)ξ+1et(x)
where t(x)={[1+ξ(xμσ)]1/ξif ξ0exp(xμσ)if ξ=0
CDF et(x) for x  in the support (see above)
Mean {μ+σ(g11)ξif ξ0 ,ξ<1μ+σγif ξ=0if ξ1
where gk=Γ(1kξ) (see Gamma function)
and γ is Euler’s constant
Median {μ+σ(ln2)ξ1ξif ξ0μσlnln2if ξ=0
Mode {μ+σ(1+ξ)ξ1ξif ξ0μif ξ=0
Variance {σ2g2g12ξ2if ξ0 and ξ<12σ2π26if ξ=0if ξ12
Skewness {sgn(ξ)g33g2g1+2g13(g2g12)3/2if ξ0 and ξ<13126ζ(3)π3if ξ=0
where sgn(x) is the sign function
and ζ(x) is the Riemann zeta function
Excess kurtosis {g44g3g1+6g12g23g14(g2g12)2if ξ0 and ξ<14125if ξ=0
Entropy ln(σ)+γξ+γ+1
MGF see Template:Harvp[1]
CF see Template:Harvp[1]

In probability theory and statistics, the generalized extreme value (GEV) distribution[2] is a family of continuous probability distributions developed within extreme value theory to combine the Gumbel, Fréchet and Weibull families also known as type I, II and III extreme value distributions. By the extreme value theorem the GEV distribution is the only possible limit distribution of properly normalized maxima of a sequence of independent and identically distributed random variables.[3] Note that a limit distribution needs to exist, which requires regularity conditions on the tail of the distribution. Despite this, the GEV distribution is often used as an approximation to model the maxima of long (finite) sequences of random variables.

In some fields of application the generalized extreme value distribution is known as the Fisher–Tippett distribution, named after R.A. Fisher and L.H.C. Tippett who recognised three different forms outlined below. However use of this name is sometimes restricted to mean the special case of the Gumbel distribution. The origin of the common functional form for all three distributions dates back to at least Template:Harvp,[4] though allegedly[3] it could also have been given by Template:Harvp.[5]

Specification

Using the standardized variable  sxμσ , where  μ , is the location parameter and can be any real number, and  σ>0 , is the scale parameter; the cumulative distribution function of the GEV distribution is then

F(s;ξ)={exp(es)for ξ=0 ,exp((1+ξs)1/ξ)for ξ0 and ξs>1 ,0for ξ>0 and s1ξ ,1for ξ<0 and s1|ξ| ,

where  ξ , the shape parameter, can be any real number. Thus, for  ξ>0 , the expression is valid for  s>1ξ , while for  ξ<0 , it is valid for  s<1ξ. In the first case,  1ξ  is the negative, lower end-point, where  F  is 0; in the second case,  1ξ  is the positive, upper end-point, where F is 1. For  ξ=0 ,, the second expression is formally undefined and is replaced with the first expression, which is the result of taking the limit of the second, as  ξ0  in which case  s  can be any real number.

In the special case of  x=μ , we have  s=0 , so then  F(0;ξ)=e10.368  regardless of the values of  ξ  and  σ.

The probability density function of the standardized distribution is

f(s;ξ)={esexp(es)for ξ=0 ,(1+ξs)(1+1/ξ)exp((1+ξs)1/ξ)for ξ0 and ξs>1 ,0otherwise;

again valid for  s>1ξ  in the case  ξ>0 , and for  s<1ξ  in the case  ξ<0. The density is zero outside of the relevant range. In the case  ξ=0 , the density is positive on the whole real line.

Since the cumulative distribution function is invertible, the quantile function for the GEV distribution has an explicit expression, namely

Q(p;μ,σ,ξ)={μσ ln(lnp)for ξ=0 and p(0,1) ,μ+σξ( (lnp)ξ1 )for ξ>0 and p[0,1) , or for ξ<0 and p(0,1];

and therefore the quantile density function  q=dQdp  is

q(p;σ,ξ)=σ(lnp)ξ+1pfor p(0,1) ,

valid for  σ>0  and for any real  ξ.

Example of probability density functions for distributions of the GEV family. [6]

Summary statistics

Using gkΓ(1kξ) for k{1,2,3,4} , where Γ() is the gamma function, some simple statistics of the distribution are given by:[citation needed]

𝔼(X)=μ+(g11)σξ for ξ<1 ,
Var(X)=(g2g12)σ2ξ2 ,
Mode(X)=μ+((1+ξ)ξ1)σξ.

The skewness is

 skewness(X)={g33g2g1+2g13(g2g12)3/2sgn(ξ)ξ0 ,126ζ(3)π31.14ξ=0.

The excess kurtosis is:

 kurtosis excess(X)=g44g3g1+6g2g123g14(g2g12)23.


The shape parameter  ξ  governs the tail behavior of the distribution. The sub-families defined by three cases:  ξ=0 ,  ξ>0 , and  ξ<0 ; these correspond, respectively, to the Gumbel, Fréchet, and Weibull families, whose cumulative distribution functions are displayed below.

  • Type I or Gumbel extreme value distribution, case ξ=0 , for all x(  , + ) :
F( x; μ, σ, 0 )=exp(exp( xμ σ)).
  • Type II or Fréchet extreme value distribution, case ξ>0 , for all x( μσ ξ  , + ) :
Let α 1 ξ>0 and y1+ξσ(xμ) ;
F( x; μ, σ, ξ )={0y0𝗈𝗋 𝖾𝗊𝗎𝗂𝗏.xμσ ξ exp(1yα )y>0𝗈𝗋 𝖾𝗊𝗎𝗂𝗏.x>μσ ξ .
  • Type III or reversed Weibull extreme value distribution, case ξ<0 , for all x( , μ+σ | ξ |  ) :
Let α1 ξ >0 and y1 | ξ | σ(xμ) ;
F( x; μ, σ, ξ )={exp(yα)y>0𝗈𝗋 𝖾𝗊𝗎𝗂𝗏.x<μ+σ | ξ | 1y0𝗈𝗋 𝖾𝗊𝗎𝗂𝗏.xμ+σ | ξ | .

The subsections below remark on properties of these distributions.

Modification for minima rather than maxima

The theory here relates to data maxima and the distribution being discussed is an extreme value distribution for maxima. A generalised extreme value distribution for data minima can be obtained, for example by substituting  x for x in the distribution function, and subtracting the cumulative distribution from one: That is, replace  F(x)  with  1F(x)  . Doing so yields yet another family of distributions.

Alternative convention for the Weibull distribution

The ordinary Weibull distribution arises in reliability applications and is obtained from the distribution here by using the variable  t=μx , which gives a strictly positive support, in contrast to the use in the formulation of extreme value theory here. This arises because the ordinary Weibull distribution is used for cases that deal with data minima rather than data maxima. The distribution here has an addition parameter compared to the usual form of the Weibull distribution and, in addition, is reversed so that the distribution has an upper bound rather than a lower bound. Importantly, in applications of the GEV, the upper bound is unknown and so must be estimated, whereas when applying the ordinary Weibull distribution in reliability applications the lower bound is usually known to be zero.

Ranges of the distributions

Note the differences in the ranges of interest for the three extreme value distributions: Gumbel is unlimited, Fréchet has a lower limit, while the reversed Weibull has an upper limit. More precisely, univariate extreme value theory describes which of the three is the limiting law according to the initial law  X  and in particular depending on the original distribution's tail.

Distribution of log variables

One can link the type I to types II and III in the following way: If the cumulative distribution function of some random variable  X  is of type II, and with the positive numbers as support, i.e.  F( x; 0, σ, α ) , then the cumulative distribution function of lnX is of type I, namely  F( x; lnσ, 1 α , 0 ). Similarly, if the cumulative distribution function of  X  is of type III, and with the negative numbers as support, i.e.  F( x; 0, σ, α ) , then the cumulative distribution function of  ln(X)  is of type I, namely  F( x; lnσ,  1 α, 0 ).


Multinomial logit models, and certain other types of logistic regression, can be phrased as latent variable models with error variables distributed as Gumbel distributions (type I generalized extreme value distributions). This phrasing is common in the theory of discrete choice models, which include logit models, probit models, and various extensions of them, and derives from the fact that the difference of two type-I GEV-distributed variables follows a logistic distribution, of which the logit function is the quantile function. The type-I GEV distribution thus plays the same role in these logit models as the normal distribution does in the corresponding probit models.

Properties

The cumulative distribution function of the generalized extreme value distribution solves the stability postulate equation.[citation needed] The generalized extreme value distribution is a special case of a max-stable distribution, and is a transformation of a min-stable distribution.

Applications

  • The GEV distribution is widely used in the treatment of "tail risks" in fields ranging from insurance to finance. In the latter case, it has been considered as a means of assessing various financial risks via metrics such as value at risk.[7][8]
File:GEV Surinam.png
Fitted GEV probability distribution to monthly maximum one-day rainfalls in October, Surinam
  • However, the resulting shape parameters have been found to lie in the range leading to undefined means and variances, which underlines the fact that reliable data analysis is often impossible.[9][full citation needed]
  • In hydrology the GEV distribution is applied to extreme events such as annual maximum one-day rainfalls and river discharges.[10] The blue picture illustrates an example of fitting the GEV distribution to ranked annually maximum one-day rainfalls showing also the 90% confidence belt based on the binomial distribution. The rainfall data are represented by plotting positions as part of the cumulative frequency analysis.

Prediction

  • It is often of interest to predict probabilities of out-of-sample data under the assumption that both the training data and the out-of-sample data follow a GEV distribution.
  • Predictions of probabilities generated by substituting maximum likelihood estimates of the GEV parameters into the cumulative distribution function ignore parameter uncertainty. As a result, the probabilities are not well calibrated, do not reflect the frequencies of out-of-sample events, and, in particular, underestimate the probabilities of out-of-sample tail events.[11]
  • Predictions generated using the objective Bayesian approach of calibrating prior prediction have been shown to greatly reduce this underestimation, although not completely eliminate it.[11] Calibrating prior prediction is implemented in the R software package fitdistcp.[12]

Example for Normally distributed variables

Let  { Xi | 1in }  be i.i.d. normally distributed random variables with mean 0 and variance 1. The Fisher–Tippett–Gnedenko theorem[13] tells us that  max{ Xi | 1in }GEV(μn,σn,0) , where

μn=Φ1(1 1 n)σn=Φ1(11 n e )Φ1(1 1 n).

This allow us to estimate e.g. the mean of  max{ Xi | 1in }  from the mean of the GEV distribution:

𝔼{ max{ Xi | 1in } }μn+γ𝖤 σn=(1γ𝖤) Φ1(1 1 n)+γ𝖤 Φ1(11 e n )=log(n2 2π log(n22π) )  (1+γ logn +(1 logn )) ,

where  γ𝖤  is the Euler–Mascheroni constant.

  1. If  XGEV(μ,σ,ξ)  then  mX+bGEV(mμ+b, |m|σ, ξ) 
  2. If  XGumbel(μ, σ)  (Gumbel distribution) then  XGEV(μ,σ,0) 
  3. If  XWeibull(σ,μ)  (Weibull distribution) then  μ(1σlogXσ)GEV(μ,σ,0) 
  4. If  XGEV(μ,σ,0)  then  σexp(Xμμσ)Weibull(σ,μ)  (Weibull distribution)
  5. If  XExponential(1)  (Exponential distribution) then  μσlogXGEV(μ,σ,0) 
  6. If  XGumbel(αX,β)  and  YGumbel(αY,β)  then  XYLogistic(αXαY,β)  (see Logistic distribution).
  7. If  X  and  YGumbel(α,β)  then  X+YLogistic(2α,β)  (The sum is not a logistic distribution).
Note that  𝔼{ X+Y }=2α+2βγ2α=𝔼{ Logistic(2α,β) }.

Proofs

4. Let  XWeibull(σ,μ) , then the cumulative distribution of  g(x)=μ(1σlogXσ)  is:

{ μ(1σlog X σ)<x }={ logXσ>1x/μσ } 𝖲𝗂𝗇𝖼𝖾 𝗍𝗁𝖾 𝗅𝗈𝗀𝖺𝗋𝗂𝗍𝗁𝗆 𝗂𝗌 𝖺𝗅𝗐𝖺𝗒𝗌 𝗂𝗇𝖼𝗋𝖾𝖺𝗌𝗂𝗇𝗀: ={ X>σexp[1x/μσ] }=exp((σexp[1x/μσ]1σ)μ)=exp((exp[1μx/μσ])μ)=exp(exp[μxσ])=exp(exp[s]),s=xμσ ,
which is the cdf for GEV(μ,σ,0).

5. Let  XExponential(1) , then the cumulative distribution of  g(X)=μσlogX  is:

{ μσlogX<x }={ logX>μxσ } 𝖲𝗂𝗇𝖼𝖾 𝗍𝗁𝖾 𝗅𝗈𝗀𝖺𝗋𝗂𝗍𝗁𝗆 𝗂𝗌 𝖺𝗅𝗐𝖺𝗒𝗌 𝗂𝗇𝖼𝗋𝖾𝖺𝗌𝗂𝗇𝗀: ={ X>exp( μx σ) }=exp[exp( μx σ)]=exp[exp(s)] ,𝗐𝗁𝖾𝗋𝖾sxμσ ;
which is the cumulative distribution of  GEV(μ,σ,0).

See also

References

  1. ^ a b Page Module:Citation/CS1/styles.css has no content.Muraleedharan, G.; Guedes Soares, C.; Lucas, Cláudia (2011). "Characteristic and moment generating functions of generalised extreme value distribution (GEV)". In Wright, Linda L. (ed.). Sea Level Rise, Coastal Engineering, Shorelines, and Tides. Nova Science Publishers. Chapter 14, pp. 269–276. ISBN 978-1-61728-655-1.
  2. ^ Page Module:Citation/CS1/styles.css has no content.Weisstein, Eric W. "Extreme value distribution". mathworld.wolfram.com. Retrieved 2021-08-06.
  3. ^ a b Page Module:Citation/CS1/styles.css has no content.Haan, Laurens; Ferreira, Ana (2007). Extreme Value Theory: An introduction. Springer.
  4. ^ Page Module:Citation/CS1/styles.css has no content.Jenkinson, Arthur F. (1955). "The frequency distribution of the annual maximum (or minimum) values of meteorological elements". Quarterly Journal of the Royal Meteorological Society. 81 (348): 158–171. Bibcode:1955QJRMS..81..158J. doi:10.1002/qj.49708134804.
  5. ^ Page Module:Citation/CS1/styles.css has no content.von Mises, R. (1936). "La distribution de la plus grande de n valeurs". Rev. Math. Union Interbalcanique. 1: 141–160.
  6. ^ Page Module:Citation/CS1/styles.css has no content.Norton, Matthew; Khokhlov, Valentyn; Uryasev, Stan (2019). "Calculating CVaR and bPOE for common probability distributions with application to portfolio optimization and density estimation" (PDF). Annals of Operations Research. 299 (1–2). Springer: 1281–1315. arXiv:1811.11301. doi:10.1007/s10479-019-03373-1. S2CID 254231768. Archived from the original (PDF) on 2023-03-31. Retrieved 2023-02-27.
  7. ^ Page Module:Citation/CS1/styles.css has no content.Moscadelli, Marco (30 July 2004). The modelling of operational risk: Experience with the analysis of the data collected by the Basel Committee (PDF) (non-peer reviewed article). doi:10.2139/ssrn.557214. SSRN 557214. Archived from the original (PDF) on 22 September 2015. Retrieved 17 June 2015 – via Archivos curso Riesgo Operativo de N.D. Girald (unalmed.edu.co).
  8. ^ Page Module:Citation/CS1/styles.css has no content.Guégan, D.; Hassani, B.K. (2014). "A mathematical resurgence of risk management: An extreme modeling of expert opinions". Frontiers in Finance and Economics. 11 (1): 25–45. SSRN 2558747.
  9. ^ Page Module:Citation/CS1/styles.css has no content.Aas, Kjersti (23 January 2008). "[no title cited]". citeseerx.ist.psu.edu (lecture). Trondheim, NO: Norges teknisk-naturvitenskapelige universitet. CiteSeerX 10.1.1.523.6456. Archived from the original (PDF) on 17 April 2023. Retrieved 4 December 2019.
  10. ^ Page Module:Citation/CS1/styles.css has no content.Liu, Xin; Wang, Yu (September 2022). "Quantifying annual occurrence probability of rainfall-induced landslide at a specific slope". Computers and Geotechnics. 149 104877. Bibcode:2022CGeot.14904877L. doi:10.1016/j.compgeo.2022.104877. hdl:2031/1583ef0b-d776-4922-8036-42c1a0448ff5. S2CID 250232752.
  11. ^ a b Page Module:Citation/CS1/styles.css has no content.Jewson, Stephen; Sweeting, Trevor; Jewson, Lynne (2025-02-20). "Reducing reliability bias in assessments of extreme weather risk using calibrating priors". Advances in Statistical Climatology, Meteorology and Oceanography. 11 (1): 1–22. Bibcode:2025ASCMO..11....1J. doi:10.5194/ascmo-11-1-2025. ISSN 2364-3579.
  12. ^ Page Module:Citation/CS1/styles.css has no content.Jewson, Stephen (2025-04-23). fitdistcp: Distribution Fitting with Calibrating Priors for Commonly Used Distributions (Report). Comprehensive R Archive Network. doi:10.32614/cran.package.fitdistcp.
  13. ^ Page Module:Citation/CS1/styles.css has no content.David, Herbert A.; Nagaraja, Haikady N. (2004). Order Statistics. John Wiley & Sons. p. 299.

Further reading

Page Template:Refbegin/styles.css has no content.

Lua error in package.lua at line 80: module 'Module:Navbox/configuration' not found.