Detecting virtual concept drift of regressors without ground truth values
作者:Emilia Oikarinen, Henri Tiittanen, Andreas Henelius, Kai Puolamäki
摘要
Regression analysis is a standard supervised machine learning method used to model an outcome variable in terms of a set of predictor variables. In most real-world applications the true value of the outcome variable we want to predict is unknown outside the training data, i.e., the ground truth is unknown. Phenomena such as overfitting and concept drift make it difficult to directly observe when the estimate from a model potentially is wrong. In this paper we present an efficient framework for estimating the generalization error of regression functions, applicable to any family of regression functions when the ground truth is unknown. We present a theoretical derivation of the framework and empirically evaluate its strengths and limitations. We find that it performs robustly and is useful for detecting concept drift in datasets in several real-world domains.
论文关键词:Concept drift, Generalization error, Unknown ground truth
论文评审过程:
论文官网地址:https://doi.org/10.1007/s10618-021-00739-7