A new encoding technique for peptide classification

作者:

Highlights:

摘要

Research on peptide classification problems has focused mainly on the study of different encodings and the application of several classification algorithms to achieve improved prediction accuracies. The main drawback of the literature is the lack of an extensive comparison among the available encoding methods on a wide range of classification problems. This paper addresses the fundamental issue of which peptide encoding promises the best results for machine learning classifiers. Two novel encoding methods based on physicochemical properties of the amino acids are proposed and an extensive comparison with several standard encoding methods is performed on three different classification problems (HIV-protease, recognition of T-cell epitopes and prediction of peptides that bind human leukocyte antigens). The experimental results demonstrate the effectiveness of the new encodings and show that the frequently used orthonormal encoding is inferior compared to other methods.

论文关键词:Peptide classification,Amino acid encoding,Physicochemical properties,Machine learning,HIV-protease,Human immune system

论文评审过程:Available online 7 September 2010.

论文官网地址:https://doi.org/10.1016/j.eswa.2010.09.005