Abstract
Modeling excesses remains to be an important topic in insurance data modeling. Among the alternatives to modeling excess, the Peaks Over Threshold (POT) framework with Generalized Pareto distribution (GPD) is regarded as an efficient approach due to its flexibility. However, selecting an appropriate threshold for such a framework is a major difficulty. To address such difficulty, we applied several accumulation tests along with Anderson-Darling test to determine an optimal threshold. Limited simulations were conducted to assess the performance of accumulation tests. Based on the selected thresholds, the fitted GPD with the estimated quantiles can be found. We applied the procedure to the wellknown Norwegian fire insurance data and AON Re Belgian fire loss data. With the selected thresholds, confidence intervals for the Value-at-Risks (VaR) were constructed correspondingly. The accumulation test approach provides satisfactory performance in modeling the VaR of both data sets compared to the previous graphical methods.
Keywords
Introduction
The selection of a proper risk modeling approach is a major topic in the insurance industry. Among all existing approaches, parametric modeling is preferred for its flexibility and interpretability. Within parametric methods, the Peaks Over Threshold (POT) approach with Generalized Pareto Distribution (GPD) remains popular due to the support of the Pickands-Balkema-De Haan theorem (McNeil, 1997). However, the performance of such approach is mainly determined by the selection of thresholds. If the threshold is chosen too low, the excess distribution might not be well approximated by GPD; If the threshold is chosen too high, number of excesses would not be enough to establish a useful model.
To identify an appropriate threshold in GPD modeling, many methods have been proposed in the literature (Scarrott & MacDonald, 2012; Langousis et al., 2016). The Gertensgarbe-Werner (GW) plot was first developed to identify the change point that separates non-extreme and extreme parts of the data (Gerstengarbe & Werner, 1989). However, the uniform assumption on the data in this method is an essential limitation. Therefore, this method should be avoided in real applications (Langousis et al., 2016). Mean Residual Life (MRL) plot was also widely used to determine GPD thresholds (Davison & Smith, 1990; Lang et al., 1999; Coles, 2001). This method aims to identify the threshold by looking for the linear trend of the candidate thresholds in GPD parameters. However, the selection of the threshold is usually done by a visual inspection, which is subjective. Compared to the methods mentioned above, threshold selection procedures using Goodness-of-Fit (GoF) tests can be interpreted with better support from statistical theory.
The objective of GPD threshold selection procedures based on the GoF tests is to identify the lowest threshold such that the exceedances above the threshold fit a GPD well. For a single fixed threshold, the GoF test can be easily carried out (Choulakian & Stephens, 2001). However, a candidate set of possible thresholds is usually involved in finding the 'optimal' threshold. Thus, error control is necessary since the candidate thresholds are tested simultaneously, which generates a multipletesting problem. Because of the natural ordering of the candidate thresholds, the hypotheses for the GoF tests should also be ordered. Therefore, False Discovery Rate (FDR) control procedures such as the Benjamini-Hochberg (Benjamini & Hochberg, 1995) or Benjamini-Yekutieli (Benjamini & Yekutieli, 2001) cannot be used without modifications. To adapt FDR control procedures to the ordered hypotheses testing structure, ForwardStop was proposed (G'Sell et al., 2016). ForwardStop provides a way to choose a stopping point for ordered hypotheses such that the hypotheses occur before the stopping point can be rejected with adequate error control. (Bader et al., 2018) Then, this procedure was adapted to the selection of GPD thresholds and applied to rainfall data.
The accumulation tests (Li & Barber, 2017) can be seen as a generalization of the ForwardStop procedure. Any functions that satisfy certain regularity conditions can be utilized in the ordered hypotheses testing procedure with FDR control at a desired level (Li & Barber, 2017). In addition to ForwardStop, procedures like SeqStep (Barber & Candès, 2015) and HingeExp (Li & Barber, 2017) can also control FDR under different scenarios.
In this paper, we attempt to adapt the accumulation test procedures to GPD threshold selection problems and propose a generalized framework of GPD threshold selection based on accumulation tests. We aim to assess the performances of accumulation test threshold selection methods with numerical simulations. To investigate the applicability of the methods with real data and provide useful insights to decision-makers in the field of insurance, we also apply the methods to two fire insurance data sets and assess the performances of the methods in terms of estimated Value-at-Risk (VaR).
The rest of the paper is organized as follows: In section 2, an introduction to GPD model and POT method is given, with a brief introduction of accumulation tests for sequential hypotheses testing. Section 3 provides simulations to assess the performance of different accumulation tests in selecting GPD thresholds. The real data analysis of the fire insurance data is given in Section 4. The final section concludes with a discussion and our thoughts on the future directions.
Methodology
POT Modeling and Threshold Selection Procedure
A common practice in extreme value modeling is to apply the Peaks Over Threshold (POT) approach to exceedances beyond a certain threshold
The choice of
Let
Consider Let For any fixed value of Plug Equation (4) back to Equation (3). The optimization becomes a onedimensional search of Eventually, the estimates of
Grimshaw showed that the MLE for
In order to assess the performance of GPD fit, several tests were developed generally based on the GoF tests such as Anderson-Darling (AD) test. The details for the AD test are provided as follows: Given a sample z is arranged in increasing order as
The automated threshold selection for GPD was developed by utilizing an ordered hypothesis testing procedure named ForwardStop (G'Sell et al., 2016; Bader et al., 2018). The selection procedure for GPD could be summarized as follows:
Pick l candidate thresholds as For each Then, for each The largest index for the candidate thresholds to be rejected is determined by applying the ForwardStop procedure (G'Sell et al., 2016) as follows:
The above procedure could produce good estimates for the thresholds of GPD if the candidates are chosen properly. For instance, when modeling the rainfall data, (Bader et al., 2018) selected the thresholds within the
could be examined by using GoF tests. l different p -values are obtained correspondingly.
Li and Barber (Li & Barber, 2017) later generalized the results from different ordered hypothesis testing procedures (G'Sell et al., 2016; Barber & Candès, 2015) including ForwardStop. A family of accumulation tests were developed to choose a cutoff point k such that first k hypotheses are rejected, with a modified false discovery rate (FDR) being controlled. Assume n hypotheses are ordered sequentially as

The flow chart of the accumulation test threshold selection framework where the hypotheses
The estimation of VaR is important for the insurance data modeling. For a loss random variable, VaR at the level of p is defined as:
Simulation Settings
To assess the performance of sequential testing procedures, we conducted limited simulations. Three different accumulation tests were assessed for our simulations. The accumulation functions of these three accumulation tests are provided as follows:
ForwardStop: SeqStep: HingeExp:
For the value of C involved in SeqStep and HingExp, we selected
Scenarios of Composite Density with GPD Tails
To demonstrate the ability of accumulation tests in selecting the thresholds for a GPD distribution in a POT model, we generated samples from four parametric composite distributions. The composite distributions were widely used in the modeling of insurance claim sizes (Cooray & Ananda, 2005; Scollnik, 2007; Scollnik & Sun, 2012; Brazauskas & Kleefeld, 2016; Grün & Miljkovic, 2019; Liu & Ananda, 2023, 2022; Mutali & Vernic, 2022; Deng & Aminzadeh, 2022; Aminzadeh & Deng, 2019; Nadarajah, 2005; CalderínOjeda, 2018). The details of the simulation scenarios are listed in Table 1. For each scenario,
Simulation Results for Three Different Tests Under Different Simulation Scenarios
.
Simulation Results for Three Different Tests Under Different Simulation Scenarios
The mean and the RMSE were used to assess the performances of three tests under all scenarios. The formula for RMSE is provided as follows:
The simulation results are presented in Table 1
Simulation Results for Three Different Tests Under Different Simulation Scenarios
.
Simulation Results for Three Different Tests Under Different Simulation Scenarios
In this section, the well-known Norwegian fire insurance data and AON Re Belgian fire loss data are used to assess the performance of different methods of GPD threshold selection.
Selection of Accumulation Functions and Methods for Comparison
The accumulation tests that we chose in the simulations were applied to the Norwegian fire insurance data set. In addition, we selected two most commonly-used methods (GW plot and MRL plot) for comparison purposes. We utilized the R packages "tea" and "eva" to construct the plots and select GPD thresholds.
Case 1: Norwegian Fire Insurance Data
Data
Norwegian fire insurance data contains the fire insurance claims from a Norwegian company from year 1972 to 1992. The dataset is available via the R package 'ReIns’ (Reynkens & Verbelen, 2020). No information regarding inflation adjustments was provided so we chose the claims from year 1985 to 1989. Moreover, only the damages over 500,000 Norwegian Krones (NKK) are available. For analysis concern, we scaled the data by
Summary Statistics for Norwegian Fire Insurance Claims (1985–1989).
Summary Statistics for Norwegian Fire Insurance Claims (1985–1989).
Result: Threshold selection
Table 4 presents the selection of GPD thresholds
Selection of Thresholds
Notice based on Table 4, the selected thresholds with ForwardStop are significantly lower compared to the selected thresholds with HingeExp or SeqStep. This could be relevant since, as mentioned in Section 4.2.1, only claims with very high values (over
Result: Measuring tail risk using var
Table 5 provides the estimates of VaRs at the
Var of the Fitted GPD Models with Norwegian Fire Insurance Claims (1985–1992).
Text in italic indicates the CI does not cover the corresponding empirical VaR estimate.
The
Data
AON Re Belgian data contains the 1823 fire losses collected by the reinsurance company AON. The building type of the associated fire losses and the loss amount in thousands of Danish Krones (DKK) are provided in this data set. We selected building type 'C' and 'D' for our analysis. The summary statistics of the losses are presented in Table 6. Notice the fire losses for both types of the buildings are characterized with right-skewed distributions.
Summary Statistics for AON re Belgian Fire Loss Claims (Building C and D).
Summary Statistics for AON re Belgian Fire Loss Claims (Building C and D).
Result: Threshold selection
Table 7 demonstrates the selected GPD thresholds

Comparison of the CDFs using different GPD modeling methods for AON re Belgian fire insurance loss data.
Chosen Thresholds
For building Type C, all the accumulation test procedure selected the same threshold as 0.9 while GW and MRL procedure selected significantly higher thresholds in comparison. The comparison of CDF plots (Figure 2) suggest that the accumulation test procedure provided closer fit to the empirical CDF, while the estimated CDF based on GW and MRL procedure cannot provide satisfactory performance.
For building Type D, SeqStep and HingeExp made the same selection regardless of chosen alpha level, while ForwardStop selected significantly lower threshold in comparison. The selection based on MRL plot, however, indicates that the mean excess graph is characterized with linear trend from the minimum of the data set. Thus, the chosen threshold based on MRL is the minimum of the data set. The CDF plots (Figure 2) indicate that GW, SeqStep, and HingeExp failed to provide satisfactory performance, while ForwardStop and MRL procedure can produce the estimated CDFs closer to the empirical one.
Result: Measuring tail risk using Var
The estimated
Var of the Fitted GPD Models with AON Re Belgian Fire Loss Data.
Text in italic indicates the CI does not cover the corresponding empirical VaR estimate.
For building type C, the accumulation test procedures produced closer estimate to the empirical estimate, compared to GW and MRL procedures. The asymptotic 95% CI created based on threshold selected with accumulation tests covered the empirical VaR while the CIs created using thresholds selected with GW and MRL procedure failed to cover the empirical estimate.
Among all the procedures, GPD model selected based on ForwardStop procedure produced the closest estimates to both
In this article, three different accumulation tests have been used to generate GPD models for claim size distributions and VaR estimates over the Norwegian fire insurance data and AON Re Belgian fire loss data. Among the accumulation tests, the ForwardStop utilizes a smooth logarithmic function as the accumulation function while the SeqStep and the HingeExp employ a discrete step function with a pre-specified parameter C as the accumulation function. We also used the previous graphical methods including GW-Plot and MRL-Plot to generate the models for comparison purposes. Among all the models, the ForwardStop selection demonstrates the best performance as it produces the closest fits to the empirical CDFs and the closest estimates to the empirical VaRs
at the
The threshold selection based on the accumulation tests is attractive due to their good performance fitting both data sets. However, one needs to take precautions in terms of the following: the quality of the modeling for the accumulation tests are affected by the chosen FDP
It is quite interesting to see that ForwardStop demonstrated great performance for all the chosen data sets in comparison to other two accumulation test procedures. In fact, the accumulation test procedure for GPD threshold selection can be described as accumulating 'evidence' till the lowest GPD threshold is detected, which 'evidence' is quantified as GoF test p-values for GPD. Based on the accumulation functions we chose in 3.1, ForwardStop procedure collects "evidence" from all candidate thresholds, while SeqStep and HingeExp procedure only collects "evidence" when p-value is greater than
Thus, pre-specifying
Finally, we conclude that the accumulation tests have great ability to choose the thresholds in GPD models for Norwegian fire insurance data and AON Re Belgian fire loss data. More extended uses and improvements for such method in GPD threshold selection are warranted, especially designing a useful framework to select C for both SeqStep and HingeExp.
Footnotes
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
