Application-independent feature construction based on almost-closedness properties

作者:Dominique Gay, Nazha Selmaoui-Folcher, Jean-François Boulicaut

摘要

Feature construction has been studied extensively, including for 0/1 data samples. Given the recent breakthroughs in closedness-related constraint-based mining, we are considering its impact on feature construction for classification tasks. We investigate the use of condensed representations of frequent itemsets based on closedness properties as new features. These itemset types have been proposed to avoid set counting in difficult association rule mining tasks, i.e. when data are noisy and/or highly correlated. However, our guess is that their intrinsic properties (say the maximality for the closed itemsets and the minimality for the δ-free itemsets) should have an impact on feature quality. Understanding this remains fairly open, and we discuss these issues thanks to itemset properties on the one hand and an experimental validation on various data sets (possibly noisy) on the other hand.

论文关键词:Feature construction, Pattern-based classification, δ-free itemsets, δ-strong rules, Closure equivalence classes, Noise-tolerance

论文评审过程:

论文官网地址:https://doi.org/10.1007/s10115-010-0369-x