G5 Artikkeliväitöskirja

Developing methods for predictive data science: with applications to Bayesian inference, cross-validation, and kernel methods;




TekijätNumminen, Riikka

KustannuspaikkaTurku

Julkaisuvuosi2026

Sarjan nimiAnnales Universitatis Turkuensis F

Numero sarjassa90

ISBN978-952-02-0775-5

eISBN978-952-02-0776-2

ISSN2736-9390

eISSN2736-9684

Julkaisun avoimuus kirjaamishetkelläAvoimesti saatavilla

Julkaisukanavan avoimuus Kokonaan avoin julkaisukanava

Verkko-osoitehttps://urn.fi/URN:ISBN:978-952-02-0776-2


Tiivistelmä

Utilizing large amounts of data requires high-quality and efficient data analytics methods. However, the methods are not universal nor ready-to-use in every case, but instead, existing methods often need to be modified or developed to be suitable for the case. The quality of methods can be assessed by several criteria and the practical utility of the methods is determined by the amount of resources needed for using them. Method development is often balancing between different characteristics, i.e. some need to be weakened in order to achieve an improvement in another.This thesis consists of three publications, in which methods of predictive data science were developed. The first method was developed for predicting the profitability of free-to-play mobile games using data collected within a short time interval, and is based on Bayesian inference, survival analysis, and a mixture model. The second method is a quicksort-based acceleration of the tournament leave-pair-out cross-validation for estimating binary classification performance via Receiver Operating Characteristics (ROC) curve analysis. The third method is a learning algorithm that was developed based on the learning algorithm referred to as CGKronRLS by using nonsmooth multiobjective optimization, and its goal is to learn sparse (easier to interpret) hypotheses on pairwise data. The results obtained with simulated and available real-world data sets were consistent in every research study. The profitability prediction model was observed to be able to predict the monetization percentage, but it was noticed that badly selected prior distribution might lead to biased estimates. The cross-validation method produced equally good estimates of the ROC curves and the areas under them as the earlier, slower method, but it was noticed that the ranking of the data may not be unequivocal with the presented method. The learning algorithm was noticed to learn sparse hypotheses when the decrease in the performance measure was restricted to five percents in comparison to the prediction performance of the reference method, and even sparser when the decrease in prediction performance was not limited. All the methods presented in this thesis improve the existing methods or serve as an alternative for solving these tasks.



Last updated on