/
/
/
Clustering Analysis of Univariate Probability Distribution Model

Clustering Analysis of Univariate Probability Distribution Model

Original Research ArticleJul 20, 2026Online First Articles https://doi.org/10.55003/cast.2026.265686

Abstract

Accurately modeling chlorophyll distribution is essential for understanding plankton dynamics and assessing marine ecosystem health. This study evaluates the suitability of various univariate probability distribution models for chlorophyll data, including the exponential, gamma, log-normal, logistic, log-logistic, normal, Rayleigh, Generalized Extreme Value (GEV), and Weibull distributions. The model selection process incorporates Goodness-of-Fit (GoF) tests (Kolmogorov-Smirnov and Anderson-Darling) and information criteria (AIC, BIC, AICc, CAIC, and HQC) to assess both statistical fit and model complexity. Additionally, the k-means clustering algorithm is applied to classify distributions based on their GoF test performance, with the Calinski-Harabasz Index (CHI) used to determine the optimal number of clusters. The results indicate that the gamma, log-normal, logistic, log-logistic, normal, and GEV distributions consistently demonstrate superior fit across multiple sample sizes, making them the most appropriate models for chlorophyll data. Conversely, the Exponential and Rayleigh distributions exhibit poor performance, particularly for larger datasets, due to their inability to capture data heterogeneity. The clustering analysis reveals that the GEV distribution tends to group with the exponential distribution in small samples ( ) but aligns with the gamma and log-normal distributions as the sample size increases ( ). This finding suggests that sample size significantly influences distribution stability and clustering behavior. This study contributes to probability distribution modeling by integrating GoF-based clustering approaches, reducing ambiguity in model selection for heterogeneous marine datasets. The findings emphasize the importance of selecting flexible distributions, particularly in ecological and oceanographic applications, where data variability is high. Future research should explore alternative GoF tests (e.g., Cramer-von Mises, Kuiper) and advanced clustering methods (e.g., DBSCAN, Gaussian Mixture Models) to enhance distribution selection for complex marine environments, supporting Marine Life-related SDGs.

probability distribution
goodness of fit
k-means clustering
information criteria
chlorophyll data
Marine Life

How to Cite

Reba, F. ., Saifudin, T. ., & Hendradi, R. . (2026). Clustering Analysis of Univariate Probability Distribution Model. Current Applied Science and Technology, e0265686. https://doi.org/10.55003/cast.2026.265686

References

  • Ashari, I. F., Dwi Nugroho, E., Baraku, R., Novri Yanda, I., & Liwardana, R. (2023). Analysis of elbow, silhouette, Davies-Bouldin, Calinski-Harabasz, and Rand-index evaluation on K-means algorithm for classifying flood-affected areas in Jakarta. Journal of Applied Informatics and Computing, 7(1), 95-103. https://doi.org/10.30871/jaic.v7i1.4947
  • Badr, M. M. (2019). Goodness-of-fit tests for the Compound Rayleigh distribution with application to real data. Heliyon, 5(8), Article e02225. https://doi.org/10.1016/j.heliyon.2019.e02225
  • Balakrishnan, N., Chimitova, E., & Vedernikova, M. (2015). An empirical analysis of some nonparametric goodness-of-fit tests for censored data. Communications in Statistics: Simulation and Computation, 44(4), 1101-1115. https://doi.org/10.1080/03610918.2013.796982
  • Basheer, A. M. (2019). Alpha power inverse Weibull distribution with reliability application. Journal of Taibah University for Science, 13(1), 423-432. https://doi.org/10.1080/16583655.2019.1588488
  • Behrenfeld, M. J., & Boss, E. S. (2014). Resurrecting the ecological underpinnings of ocean plankton blooms. Annual Review of Marine Science, 6(1), 167-194. https://doi.org/10.1146/annurev-marine-052913-021325

Author Information

Felix Reba

Doctoral Program of Mathematics and Natural Sciences, Faculty of Science and Technology, Universitas Airlangga, Surabaya, Indonesia

Toha Saifudin

Mathematics Department, Faculty of Sciences and Technology, Universitas Airlangga, Surabaya, Indonesia

Rimuljo Hendradi

Information Systems Study Program, Faculty of Science and Technology, Universitas Airlangga, Surabaya, Indonesia

About this Article

Journal

Online First Articles

Type of Manuscript

Original Research Article

Published

20 July 2026