Uncertainty Quantification Techniques in Statistics Printed Edition of the Special Issue Published in Mathematics www.mdpi.com/journal/mathematics Jong-Min Kim Edited by Uncertainty Quantification Techniques in Statistics Uncertainty Quantification Techniques in Statistics Special Issue Editor Jong-Min Kim MDPI • Basel • Beijing • Wuhan • Barcelona • Belgrade • Manchester • Tokyo • Cluj • Tianjin Special Issue Editor Jong-Min Kim University of Minnesota at Morris USA Editorial Office MDPI St. Alban-Anlage 66 4052 Basel, Switzerland This is a reprint of articles from the Special Issue published online in the open access journal Mathematics (ISSN 2227-7390) (available at: https://www.mdpi.com/journal/mathematics/special issues/uncertainty quantification techniques statistics). For citation purposes, cite each article independently as indicated on the article page online and as indicated below: LastName, A.A.; LastName, B.B.; LastName, C.C. Article Title. Journal Name Year , Article Number , Page Range. ISBN 978-3-03928-546-4 (Pbk) ISBN 978-3-03928-547-1 (PDF) c © 2020 by the authors. Articles in this book are Open Access and distributed under the Creative Commons Attribution (CC BY) license, which allows users to download, copy and build upon published articles, as long as the author and publisher are properly credited, which ensures maximum dissemination and a wider impact of our publications. The book as a whole is distributed by MDPI under the terms and conditions of the Creative Commons license CC BY-NC-ND. Contents About the Special Issue Editor . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . vii Preface to ”Uncertainty Quantification Techniques in Statistics” . . . . . . . . . . . . . . . . . . ix Javier E. Contreras-Reyes, Mohsen Maleki and Daniel Devia Cort ́ es Skew-Reflected-Gompertz Information Quantifiers with Application to Sea Surface Temperature Records Reprinted from: Mathematics 2019 , 7 , 403, doi:10.3390/math7050403 . . . . . . . . . . . . . . . . . 1 Md Showaib Rahman Sarker , Michael Pokojovy and and Sangjin Kim On the Performance of Variable Selection and Classification via Rank-Based Classifier Reprinted from: Mathematics 2019 , 7 , 457, doi:10.3390/math7050457 . . . . . . . . . . . . . . . . . 15 Sangjin Kim and Jong-Min Kim Two-Stage Classification with SIS Using a New Filter Ranking Method in High Throughput Data Reprinted from: Mathematics 2019 , 7 , 493, doi:10.3390/math7060493 . . . . . . . . . . . . . . . . . 31 Gi-Sung Lee, Ki-Hak Hong and Chang-Kyoon Son An Estimation of Sensitive Attribute Applying Geometric Distribution under Probability Proportional to Size Sampling Reprinted from: Mathematics 2019 , 7 , 1102, doi:10.3390/math7111102 . . . . . . . . . . . . . . . . 47 Abhijeet R Patil and Sangjin Kim Combination of Ensembles of Regularized Regression Models with Resampling-Based Lasso Feature Selection in High Dimensional Data Reprinted from: Mathematics 2020 , 8 , 110, doi:10.3390/math8010110 . . . . . . . . . . . . . . . . . 63 Jung Yeon Lee, Myeong-Kyu Kim, Wonkuk Kim Robust Linear Trend Test for Low-Coverage Next-Generation Sequence Data Controlling for Covariates Reprinted from: Mathematics 2020 , 8 , 217, doi:10.3390/math8020217 . . . . . . . . . . . . . . . . . 87 Hohsuk Noh and Seong J. Yang Comparing Groups of Decision-Making Units in Efficiency Based on Semiparametric Regression Reprinted from: Mathematics 2020 , 8 , 233, doi:10.3390/math8020233 . . . . . . . . . . . . . . . . . 101 v About the Special Issue Editor Jong-Min Kim (Dr.) is currently working as a Full Professor in the Statistics Division at the Science and Mathematics University of Minnesota-Morris, Minnesota, USA. He received his Ph.D. (Statistics) in 2002 from the Department of Statistics, Oklahoma State University, Stillwater, Oklahoma, USA (Minor: Mathematics). He worked as a Research Fellow at SAMSI-The Statistical and Applied Mathematical Sciences Institute (NSF, Duke, NCSU, and UNC). He received the Morris Faculty Distinguished Research Award. His research involves statistical genetics, cluster analysis, big data analytics, text mining for patent data, copula directional dependence, and cryptocurrency data analysis. He has published about 120 refereed journal research papers in statistics, biostatistics, bioinformatics, and economics. He is Guest Editor, Associate Editor, and Editorial Board member of several peer reviewed journals. Some of his important research publications in the field of copula directional dependence and big data analytics are listed at: https://academics.morris.umn.edu/ jong-min-kim. vii Preface to ”Uncertainty Quantification Techniques in Statistics” Uncertainty quantification (UQ) is a mainstream research topic in applied mathematics and statistics. To identify UQ problems, diverse modern techniques for large and complex data analyses have been developed in applied mathematics, computer science, and statistics. This Special Issue of Mathematics (ISSN 2227-7390) includes diverse modern data analysis methods such as skew-reflected-Gompertz information quantifiers with application to sea surface temperature records, the performance of variable selection and classification via a rank-based classifier, two-stage classification with SIS using a new filter ranking method in high throughput data, an estimation of sensitive attribute applying geometric distribution under probability proportional to size sampling, combination of ensembles of regularized regression models with resampling-based lasso feature selection in high dimensional data, robust linear trend test for low-coverage next-generation sequence data controlling for covariates, and comparing groups of decision-making units in efficiency based on semiparametric regression. Jong-Min Kim Special Issue Editor ix mathematics Article Skew-Reflected-Gompertz Information Quantifiers with Application to Sea Surface Temperature Records Javier E. Contreras-Reyes 1, *, Mohsen Maleki 2 and Daniel Devia Cortés 3 1 Departamento de Estadística, Facultad de Ciencias, Universidad del Bío-Bío, Concepción 4081112, Chile 2 Department of Statistics, College of Sciences, Shiraz University, Shiraz 71946 85115, Iran; m.maleki.stat@gmail.com 3 Departamento de Evaluación de Pesquerías, Instituto de Fomento Pesquero, Valparaíso 2361827, Chile; ddeviac@gmail.com * Correspondence: jcontreras@ubiobio.cl or jecontrr@uc.cl; Tel.: +56-41-311-1199 Received: 19 March 2019; Accepted: 2 May 2019; Published: 6 May 2019 Abstract: The Skew-Reflected-Gompertz (SRG) distribution, introduced by Hosseinzadeh et al. (J. Comput. Appl. Math. (2019) 349, 132–141), produces two-piece asymmetric behavior of the Gompertz (GZ) distribution, which extends the positive to a whole dominion by an extra parameter. The SRG distribution also permits a better fit than its well-known classical competitors, namely the skew-normal and epsilon-skew-normal distributions, for data with a high presence of skewness. In this paper, we study information quantifiers such as Shannon and Rényi entropies, and Kullback–Leibler divergence in terms of exact expressions of GZ information measures. We find the asymptotic test useful to compare two SRG-distributed samples. Finally, as a real-world data example, we apply these results to South Pacific sea surface temperature records. Keywords: Skew-Reflected-Gompertz distribution; Gompertz distribution; entropy; Kullback–Leibler divergence; sea surface temperature 1. Introduction The Skew-Reflected-Gompertz (SRG) distribution was recently introduced by [ 1 ] and corresponds to an extension of the Gompertz distribution [ 2 ], named after Benjamin Gompertz (1779–1865). It extends the positive dominion R + to the whole of R by an extra parameter, ε , − 1 < ε < 1, and produces two-piece asymmetric behavior of Gompertz (GZ) density. The SRG distribution has as particular cases the Reflected-GZ and GZ distributions, when ε → 1 and ε → − 1, respectively. The SRG distribution family can also represent a suitable competitor against the skew-normal (SN, [ 3 ]) and epsilon-skew-normal (ESN, [ 4 ]) distributions as a way to fit asymmetrical datasets. Indeed, refs. [ 5 , 6 ] dealt with the frequentist and Bayesian inferences of ESN distribution. Contributions by [ 1 ] provided probability density function (pdf), cumulative distribution function (cdf), quantile function, moment-generating function (MGF), stochastic representation, the Expectation-Maximization (EM) algorithm for SRG parameter estimates and the Fisher information matrix (FIM). Moreover, several recent investigations confirmed the usefulness of entropic quantifiers in the study of asymmetric distributions [ 3 , 7 , 8 ] and their applications to topics such as thermal wake [ 9 ], marine fish biology [ 3 , 8 ], sea surface temperature (SST), relative humidity measured in the Atlantic Ocean [ 10 ], and more. We build on the study of [ 3 ], which developed hypothesis testing for normality, i.e., if the shape parameter is close to zero. They considered the Kullback–Leibler (KL) divergence in terms of moments and cumulants of the modified SN distribution. Posteriorly, we consider a real-world data set of the anchovy condition factor for testing the shape parameter to decide if a food deficit produced by environmental conditions such as El Niño exists [11]. Mathematics 2019 , 7 , 403; doi:10.3390/math7050403 www.mdpi.com/journal/mathematics 1 Mathematics 2019 , 7 , 403 This work arose from a motivation to tackle the problem of determining the adequate pdf of SST [ 9 , 10 ]. Indeed, probabilistic modelling of SST is key for accurate predictions [ 9 ]. Therefore, we propose that the SRG model based on two-piece distributions could be more suitable for interpreting annual bimodal and asymmetric SST data. We also considered the existent results of Shannon and Rényi entropies, and KL divergence for GZ distributions for developed entropic quantifiers for SRG distributions. Posteriorly, we considered SST along the South Pacific and Chilean coasts from 2012 to 2014 to illustrate our results. Specifically, we introduced hypothesis testing developed by [ 12 ] for the SRG distribution, which is useful to compare two data sets with bimodal and asymmetric behavior such as SST. 2. The Skew-Reflected-Gompertz Distribution The Gompertz (GZ, [ 2 ]) distribution is a continuous probability distribution with the following pdf f ( x | σ , η ) = η σ e x σ e − η ( e x σ − 1 ) , x ≥ 0, (1) where σ > 0 and η > 0 are the scale and shape parameters, respectively, and are denoted by X ∼ GZ ( σ , η ) . The mean and variance of X are E ( X ) = σ e η Ei ( − η ) , (2) Var ( X ) = σ 2 e η τ , respectively; where Ei ( z ) = ∫ ∞ − z e − u u du , τ = − 2 η F ( − η ) + γ 2 + π 2 6 + 2 γ log η + ( log η ) 2 − e η [ Ei ( − η )] 2 , γ = 0.5772156649 is the Euler constant and F ( z ) = + ∞ ∑ k = 0 z k k ! ( k + 1 ) 3 The SRG distribution is an extension of the GZ proposed by [ 1 ]. If Y follows, the SRG distribution is denoted by Y ∼ SRG ( μ , σ , η , ε ) and has pdf g ( y | μ , σ , η , ε ) = ⎧ ⎨ ⎩ 1 2 f ( μ − y 1 + ε ∣ ∣ ∣ σ , η ) , y ≤ μ , 1 2 f ( y − μ 1 − ε ∣ ∣ ∣ σ , η ) , y > μ , (3) where μ ∈ R is the location parameter and ε ∈ ( − 1, 1 ) is the slant parameter. Note that SRG is the GZ distribution when μ = 0 and ε → − 1, GZ distribution with negative support when ε → 1, and Reflected-GZ distribution when ε = 0. Also, the Reflected-GZ distribution corresponds to a particular case of a more general class of two-piece asymmetric distributions proposed by [ 13 , 14 ]. The mean, variance and MGF of Y are E ( Y ) = μ − 2 εσ e η Ei ( − η ) , Var ( Y ) = σ 2 { τ e η + 2 ( 1 − ε 2 ) e 2 η [ Ei ( − η )] 2 } , M Y ( t ) = 1 2 η e η + μ t [( 1 − ε ) − σ t ( η ) + ( 1 + ε ) σ t ( η )] , (4) respectively; where s ( z ) = ∫ ∞ 1 v s + 1 e − vz dv . Jafari et al. [ 15 ] provide the MGF of X using expansion series. However, (4) is considered a clearer expression that depends only on integral s ( z ) . See Section 4.1 for some details of the MLE EM-based algorithm related to SRG parameters. According to [ 1 ], the SRG distribution can be re-parametrized in terms of GZ and Reflected-GZ distributions as g ( y | μ , σ + , σ − , η ) = p 1 f ( μ − y | σ + , η ) I ( − ∞ , μ ] ( y ) + p 2 f ( y − μ | σ − , η ) I ( μ , + ∞ ) ( y ) , (5) 2 Mathematics 2019 , 7 , 403 where σ ± = σ ( 1 ± ε ) , p 1 + p 2 = 1, and p 1 = σ + / ( σ + + σ − ) = ( 1 + ε ) / 2. Let Y = ( Y 1 , . . . , Y n ) be an i.i.d sample from the SRG distribution with parameters ( μ , σ ± , η ) and latent vectors Z = ( Z 1 , . . . , Z n ) , thus (5) can be equivalently represented as ( − 1 ) j ( Y i − μ ) | Z ij = 1 ∼ GZ ( σ ± , η ) , i = 1, . . . , n , j = 1, 2, where Z i = ( Z i 1 , Z i 2 ) ∼ Mult ( 1, p 1 , p 2 ) is a multinomial vector, P ( Z i 1 = z i 1 , Z i 2 = z i 2 ) = p z i 1 1 p z i 2 2 , z ij = { 0, 1 } , and z i 1 + z i 2 = 1. Given that P ( Z i 1 = 1 ) = P ( Z i 1 = 1, Z ik = 0; ∀ j = k ) , the complete log-likelihood function is ( μ , σ + , σ − , η | Y , Z ) = − n log ( 2 σ ) + n ( η + log η ) + n ∑ i = 1 [ z i 1 ( μ − y i σ + − η e μ − yi σ + ) + z i 2 ( y i − μ σ − − η e yi − μ σ − )] (6) Conditional expectations of latent variables Z i are given by ̂ z i 1 = E [ Z i 1 | ̂ μ , ̂ σ + , ̂ σ − , y i ] = ̂ p 1 f ( ̂ μ − y i | ̂ σ + , ̂ η ) g ( y i | ̂ μ , ̂ σ + , ̂ σ − , ̂ η ) I ( − ∞ , ̂ μ ] ( y i ) , (7) ̂ z i 2 = 1 − ̂ z i 1 , i = 1, . . . , n (8) The E- and M-steps on the ( k + 1 ) th iteration of the EM algorithm are E-step. From (6)–(8), we have Q ( μ , σ + , σ − , η | μ ( k ) , σ ( k ) + , σ ( k ) − , η ( k ) ) = E [ ( μ , σ + , σ − , η | Y , Z ) | μ ( k ) , σ ( k ) + , σ ( k ) − , η ( k ) ] = − n log ( 2 σ ) + n ( η + log η ) + n ∑ i = 1 [ ̂ z ( k ) i 1 ( μ − y i σ + − η e μ − yi σ + ) + ̂ z ( k ) i 2 ( y i − μ σ − − η e yi − μ σ − )] and M-step. Update σ ± , by solving the following equation n ∑ i = 1 ̂ z ( k ) ij ( η ( k ) | y i − μ ( k ) | σ 2 ± e | yi − μ ( k ) | σ ± − | y i − μ ( k ) | σ 2 ± ) = n 2 σ Update μ by solving the following equation ̂ μ ( k + 1 ) = argmax μ n ∑ i = 1 ⎧ ⎨ ⎩̂ z ( k ) i 1 ⎛ ⎝ μ − y i ̂ σ ( k + 1 ) + − η e μ − yi ̂ σ ( k + 1 ) + ⎞ ⎠ + ̂ z ( k ) i 2 ⎛ ⎝ μ − y i ̂ σ ( k + 1 ) − − η e μ − yi ̂ σ ( k + 1 ) − ⎞ ⎠ ⎫ ⎬ ⎭ Update η by ̂ η = n ⎛ ⎝ n ∑ i = 1 ⎧ ⎨ ⎩̂ z ( k ) i 1 e μ − yi ̂ σ ( k + 1 ) + + ̂ z ( k ) i 2 e μ − yi ̂ σ ( k + 1 ) − ⎫ ⎬ ⎭ ⎞ ⎠ − 1 The EM-algorithm must be iterated until the sufficient convergence rule is satisfied: ‖ ( ̂ μ ( k + 1 ) , ̂ σ ( k + 1 ) + , ̂ σ ( k + 1 ) − , ̂ η ( k + 1 ) ) − ( ̂ μ ( k ) , ̂ σ ( k ) + , ̂ σ ( k ) − , ̂ η ( k ) ) ‖ < τ , for a tolerance τ close to zero. The FIM for standard deviations of MLEs ( ̂ μ , ̂ σ , ̂ η , ̂ ε ) and additional details of the EM-algorithm are described in [1]. 3. Entropic Quantifiers In the next section, we present the main results of entropic quantifiers for SRG distribution. 3 Mathematics 2019 , 7 , 403 3.1. Shannon Entropy The Shannon entropy (SE), introduced by [ 16 ] in the context of univariate continuous distributions, quantifies the information contained in a random variable X with pdf f ( x ) through the expression H ( X ) = − ∫ + ∞ − ∞ f ( x ) log f ( x ) dx (9) The SE concept is attributed to the uncertainty of the information presented in X [ 17 ]. Propositions 1 and 2 present the SE for GZ and SRG distributions, respectively. Proposition 1. [15]. The SE of X ∼ GZ ( σ , η ) is H ( X ) = log { B ( 1, 1 ) η } − ση − E ( X ) σ + ση M X ( σ − 1 ) , where B ( · , · ) is the usual Beta function and E ( X ) is given in (2). Substituting μ = 0 and ε = − 1 into (4) (i.e., reducing SRG to its special case GZ), we obtain M X ( σ − 1 ) = η e η − 1 ( η ) = 1. Therefore, H ( X ) in Proposition 1 is reduced to H ( X ) = − log η − e η E i ( − η ) , (10) i.e., the SE of the GZ random variable only depends on shape parameter η Proposition 2. The SE of Y ∼ SRG ( μ , σ , η , ε ) is H ( Y ) = 1 + ε 2 { H ( X + ε ) − log ( 1 + ε 2 )} + 1 − ε 2 { H ( X − ε ) − log ( 1 − ε 2 )} , where X ± ε ∼ GZ ( σ ( 1 ± ε ) , η ) and H ( X ± ε ) are obtained using Proposition 1. Proof. From (3) and (9), we obtained H ( Y ) = − ∫ + ∞ − ∞ g ( y | μ , σ , η , ε ) log g ( y | μ , σ , η , ε ) dy = − 1 2 ∫ + ∞ 0 f ( x 1 + ε ∣ ∣ ∣ σ , η ) log { 1 2 f ( x 1 + ε ∣ ∣ ∣ σ , η )} dx − 1 2 ∫ + ∞ 0 f ( x 1 − ε ∣ ∣ ∣ σ , η ) log { 1 2 f ( x 1 − ε ∣ ∣ ∣ σ , η )} dx = − 1 2 ∫ + ∞ 0 ( 1 + ε ) f ( x | σ ( 1 + ε ) , η ) log { 1 + ε 2 f ( x | σ ( 1 + ε ) , η ) } dx − 1 2 ∫ + ∞ 0 ( 1 − ε ) f ( x | σ ( 1 − ε ) , η ) log { 1 − ε 2 f ( x | σ ( 1 − ε ) , η ) } dx , which concludes the proof. From (10), given that H ( X ± ε ) only depends on shape parameter η , we obtain H ( X ± ε ) = H ( X ) , and H ( Y ) only depends on η and ε parameters. Therefore, H ( Y ) = − log η − e η E i ( − η ) − 1 + ε 2 log ( 1 + ε 2 ) − 1 − ε 2 log ( 1 − ε 2 ) (11) Figure 1 illustrates SE behavior for random variable Y . We observed that SE increases when η decreases. For each η , SE is maximized and minimized at ε = 0 (Reflected-GZ) and ε → − 1 4 Mathematics 2019 , 7 , 403 (Truncated-GZ and GZ), respectively. More details appear in [ 3 , 8 ] for the SE expressions of other asymmetric distributions. Figure 1. Shannon entropy of Skew-Reflected-Gompertz (SRG) distributions for ε ∈ ( − 1, 1 ) and several values of η 3.2. Rényi Entropy The α th-order Rényi entropy (RE), introduced by [ 18 ] in the context of univariate continuous distributions, extends the concept of SE information contained in a random variable X with pdf f ( x ) through a level α , α ∈ N , α > 0, and the expression R α ( X ) = 1 1 − α log ∫ + ∞ − ∞ [ f ( x )] α dx (12) RE information can be negative and is ordered with respect to α , i.e., R α 1 ( X ) ≥ R α 2 ( X ) for any α 1 < α 2 (see, e.g., [ 7 ] and other properties of RE). From (12), the SE is obtained by the limit of H ( X ) = lim α → 1 R α ( X ) by applying l’Hôpital’s rule to R α ( X ) with respect to α (see e.g., [ 7 ]). The RE of the GZ and SRG distributions is presented in Propositions 3 and 4, respectively. Proposition 3. [15,19]. The RE of X ∼ GZ ( σ , η ) with α > 1 , α ∈ N , is R α ( X ) = − log α 1 − α + log η σ + 1 1 − α log { α − 1 ∑ j = 0 ( α − 1 j ) Γ ( j + 1 ) ( αη ) j } , where Γ ( u ) = ∫ ∞ 0 t u − 1 e − t dt is the gamma function. Proposition 4. The RE of Y ∼ SRG ( η , ε ) with α > 1 , α ∈ N , is R α ( Y ) = 1 1 − α log {( 1 + ε 2 ) α e ( 1 − α ) R α ( X + ε ) + ( 1 − ε 2 ) α e ( 1 − α ) R α ( X − ε ) } , where X ± ε ∼ GZ ( σ ( 1 ± ε ) , η ) and R α ( X ± ε ) are obtained using Proposition 3. 5 Mathematics 2019 , 7 , 403 Proof. From (3) and (12), we obtained R α ( Y ) = 1 1 − α log ∫ + ∞ − ∞ [ g ( y | μ , σ , η , ε )] α dy , = 1 1 − α log { ∫ + ∞ 0 [ 1 2 f ( x 1 + ε ∣ ∣ ∣ σ , η )] α dx + ∫ + ∞ 0 [ 1 2 f ( x 1 − ε ∣ ∣ ∣ σ , η )] α dx } , = 1 1 − α log {( 1 + ε 2 ) α ∫ + ∞ 0 [ f ( x | σ ( 1 + ε ) , η )] α dx + ( 1 − ε 2 ) α ∫ + ∞ 0 [ f ( x | σ ( 1 − ε ) , η )] α dx } , which concludes the proof. Figure 2a illustrates the behavior of RE for random variable Y when α = 2 (quadratic RE). As in the SE case, we also observed that RE increases when η decreases and reaches maximum and minimum at ε = 0 (Reflected-GZ) and ε → − 1 (Truncated-GZ and GZ), respectively. When α = 5 (or α > 2) (see Figure 2b), RE decays faster than in the quadratic RE case as ε → − 1. More details appear in [ 7 ] for the RE expressions of other asymmetric distributions. Figure 2. Rényi entropy of SRG distributions for σ = 1, − 1 < ε < 1, several values of η and ( a ) α = 2 and ( b ) α = 5 values. 3.3. Kullback–Leibler Divergence The Kullback–Leibler (KL) divergence introduced by [ 20 ] in the context of univariate continuous distributions, extends the concept of SE between two random variables X 1 and X 2 with pdfs f 1 ( x 1 ) and f 2 ( x 2 ) , respectively, through the expression K ( X 1 , X 2 ) = ∫ + ∞ − ∞ f 1 ( x ) log { f 1 ( x ) f 2 ( x ) } dx (13) The KL divergence measures the disparity between the pdfs of X 1 and X 2 , and is non-negative, non-symmetric and zero only if X 1 = X 2 in distribution. Also, the KL divergence does not satisfy the triangular inequality (see, e.g., [ 8 , 17 ] for other properties of KL and other divergences). The KL divergence for two GZ and two SRG distributions are presented in Propositions 5 and 6. Proposition 5. [21]. The KL divergence between X 1 ∼ GZ ( σ 1 , η 1 ) and X 2 ∼ GZ ( σ 2 , η 2 ) is K ( X 1 , X 2 ) = log { e η 1 σ 2 η 1 e η 2 σ 1 η 2 } + e η 1 [( σ 1 σ 2 − 1 ) E i ( − η 1 ) + η 2 η σ 1 / σ 2 1 Γ ( σ 1 σ 2 − 1, η 1 )] − ( η 1 + 1 ) , 6 Mathematics 2019 , 7 , 403 where Γ ( u , v ) = ∫ ∞ v t u − 1 e − t dt is the upper incomplete gamma function. Proposition 6. The KL divergence between Y 1 ∼ SRG ( 0, σ 1 , η 1 , ε 1 ) and Y 2 ∼ SRG ( 0, σ 2 , η 2 , ε 2 ) is K ( Y 1 , Y 2 ) = 1 + ε 1 2 [ log { 1 + ε 1 1 + ε 2 } + K ( X + ε 1 , X + ε 2 ) ] + 1 − ε 1 2 [ log { 1 − ε 1 1 − ε 2 } + K ( X − ε 1 , X − ε 2 ) ] , where X ± ε i ∼ GZ ( σ i ( 1 ± ε i ) , η i ) , i = 1, 2 , and K ( X ± ε 1 , X ± ε 2 ) are obtained using Proposition 5. Proof. From (3) and (13), we obtained K ( Y 1 , Y 2 ) = ∫ + ∞ − ∞ g ( x | 0, σ 1 , η 1 , ε 1 ) log { g ( x | 0, σ 1 , η 1 , ε 1 ) g ( x | 0, σ 2 , η 2 , ε 2 ) } dx , = 1 2 ∫ + ∞ 0 f ( x 1 + ε 1 ∣ ∣ ∣ σ 1 , η 1 ) log ⎧ ⎪ ⎨ ⎪ ⎩ f ( x 1 + ε 1 ∣ ∣ ∣ σ 1 , η 1 ) f ( x 1 + ε 2 ∣ ∣ ∣ σ 2 , η 2 ) ⎫ ⎪ ⎬ ⎪ ⎭ dx + 1 2 ∫ + ∞ 0 f ( x 1 − ε 1 ∣ ∣ ∣ σ 1 , η 1 ) log ⎧ ⎪ ⎨ ⎪ ⎩ f ( x 1 − ε 1 ∣ ∣ ∣ σ 1 , η 1 ) f ( x 1 − ε 2 ∣ ∣ ∣ σ 2 , η 2 ) ⎫ ⎪ ⎬ ⎪ ⎭ dx , = 1 + ε 1 2 [ log { 1 + ε 1 1 + ε 2 } + ∫ + ∞ 0 f ( x | σ 1 ( 1 + ε 1 ) , η 1 ) log { f ( x | σ 1 ( 1 + ε 1 ) , η 1 ) f ( x | σ 2 ( 1 + ε 2 ) , η 2 ) } dx ] + 1 − ε 1 2 [ log { 1 − ε 1 1 − ε 2 } + ∫ + ∞ 0 f ( x | σ 1 ( 1 − ε 1 ) , η 1 ) log { f ( x | σ 1 ( 1 − ε 1 ) , η 1 ) f ( x | σ 2 ( 1 − ε 2 ) , η 2 ) } dx ] , which concludes the proof. More details appear in [ 3 , 8 ] for the KL divergence expressions of other asymmetric distributions. Using Proposition 6, the asymptotic KL divergence between Y ∼ SRG ( 0, σ , η , ε ) and X ∼ GZ ( σ , η ) is K ( Y , X ) ≈ 1 + ε 2 [ lim ε 2 →− 1 log ( 1 + ε 1 + ε 2 ) + K ( X + ε , X ) ] + 1 − ε 2 [ log ( 1 − ε 2 ) + K ( X − ε , X ) ] , as ε 2 → − 1. However, we see that log ( 1 + ε 1 + ε 2 ) = + ∞ as ε 2 → − 1 and K ( Y , X ) is not finite. However, from Proposition 6 the asymptotic KL divergence between Y 1 and Y 2 is K ( Y 1 , Y 2 ) ≈ K ( X , Y ) = log ( 2 1 − ε ) + K ( X , X − ε ) , (14) as ε 1 → − 1, where X − ε ∼ GZ ( σ ( 1 − ε ) , η ) . Therefore, while K ( Y , X ) is not finite, K ( X , Y ) is finite and can be used to study the disparity of ε from − 1. Thus, hypothesis testing for H 0 : ε = − 1 can be addressed. Besides, we further study hypothesis testing for scale and shape parameters between two SRG distributions in Section 3.4. From (14), we also took that K ( Y 1 , Y 2 ) ≈ K ( X , X 1 ) as ε → − 1, with X 1 ∼ GZ ( 2 σ , η ) Figure 3 illustrates the KL divergence between two SRG distributions. We observed that for the critical points of ( ε 1 , ε 2 ) → { ( − 1, 1 ) ; ( 1, − 1 ) } , the KL divergence reaches the highest values and is close to zero in the other values [panels (a) and (b)]. For large η ’s [panel (c)], the KL divergence is zero for a concentrated region of the dominion where ε 1 = ε 2 7 Mathematics 2019 , 7 , 403 Figure 3. Plots of Kullback–Leibler (KL) divergence between Y 1 ∼ SRG ( 0, σ 1 , η 1 , ε 1 ) and Y 2 ∼ SRG ( 0, σ 2 , η 2 , ε 2 ) for values σ 1 = σ 2 = 1 and ( a ) η 1 = η 2 = 0.25; ( b ) η 1 = η 2 = 3; and ( c ) η 1 = η 2 = 10. All information quantifiers and the EM algorithm for SRG distribution were implemented in [ 22 ]. 3.4. Asymptotic Test Consider two independent samples of sizes n 1 and n 2 from Y 1 and Y 2 , respectively; where θ , θ © ∈ Θ ⊂ R p , and X 1 and X 2 have pdfs g ( y ; θ 1 ) and g ( y ; θ 2 ) , respectively; with θ i = ( σ i , η i , ε i ) , i = 1, 2. Suppose partition θ i = ( θ i 1 , θ i 2 ) , and assume θ 21 = θ 11 ∈ Θ 1 ⊂ R r , so that θ i 2 ∈ Θ ∩ Θ c 1 ⊂ R p − r Let ̂ θ i = ( ̂ θ 11 , ̂ θ i 2 ) be the MLE of θ i = ( θ 11 , θ i 2 ) for i = 1, 2, which corresponds to the MLE of the full model parameters ( θ 1 , θ 2 ) under the null hypothesis H 0 : θ 21 = θ 11 . Thus, part b) of Corollary 1 in [ 12 ] establishes that if the null hypothesis H 0 : θ 22 = θ 12 holds and n 1 n 1 + n 2 −→ n 1 , n 2 → ∞ λ , with 0 < λ < 1, then K 0 = 2 n 1 n 2 n 1 + n 2 K ( ̂ θ 1 , ̂ θ 2 ) d −→ n 1 , n 2 → ∞ χ 2 p − r , (15) where r = 3 is the number of parameters of the SRG distribution (location parameter is not considered for KL divergence). Thus, a test of level α for the above homogeneity null hypothesis consists of rejecting H 0 if K 0 > χ 2 p − r ,1 − α , where χ 2 p − r , α is the α th percentile of the χ 2 p − r -distribution. As [ 3 ] stated, the proposed asymptotic test is only valid for regular conditions of the SRG distribution, in particular for a non-singular FIM. Therefore, given that the SRG distributions’ FIM is singular at ε → ± 1 [ 1 ], the SRG model does not serve for testing the null hypothesis using (15) when ε is close to − 1 or 1. 4. Application 4.1. Sea Surface Temperature Data The spatial information and SST data analyzed in this study were recorded by a scientific observer (whose labor concerns biological sampling of fishes, incidental captures of birds, turtles and marine mammals. Biological sampling was complemented with information such as time, longline and hook features, number of buoys, baits, etc.) (SO) in the Chilean longline fleet (industrial and artisanal), which was oriented to capture swordfish ( Xiphias gladius , [ 23 ]) from 2012 to 2014 (obtaining a sampling of 83% in 2012, 55% in 2013, 90% in 2014, and 75% in 2012–2014). The covered area of the study was at 21 ◦ 31 © –36 ◦ 39 © LS and 71 ◦ 08 © –85 ◦ 52 © LW (see Figure 4). 8 Mathematics 2019 , 7 , 403 Figure 4. Spatial distribution of Sea Surface Temperature (SST) observations by year (21 ◦ 31 © –36 ◦ 39 © LS, 71 ◦ 08 © –85 ◦ 52 © LW). SST records in swordfish captures are crucial for distributional analysis and fish abundance. Specifically, variations in SST are physical factors that control productivity, growth and migration of species [ 24 ]. In addition, SST is strongly correlated with atmospheric pressure at sea level and thus climatic time scales. Therefore, changes in SST overlap with ecosystem changes [ 25 ]. However, SST influence on ecosystems is not clear because other physical processes such as superficial warming, horizontal advection of currents, upwelling, etc. [ 11 ], modify SST. Therefore, SST anomalies could be symptomatic rather than causal. 9