SciELO - Scientific Electronic Library Online

 
vol.21 issue4Hybrid Attention Networks for Chinese Short Text ClassificationExtractive Summarization: Limits, Compression, Generalized Model and Heuristics author indexsubject indexsearch form
Home Pagealphabetic serial listing  

Services on Demand

Journal

Article

Indicators

Related links

  • Have no similar articlesSimilars in SciELO

Share


Computación y Sistemas

Print version ISSN 1405-5546

Abstract

WAHEEB, Waddah  and  GHAZALI, Rozaida. Content-based SMS Classification: Statistical Analysis for the Relationship between Number of Features and Classification Performance. Comp. y Sist. [online]. 2017, vol.21, n.4, pp.771-785. ISSN 1405-5546.  http://dx.doi.org/10.13053/cys-21-4-2593.

High dimensionality of the feature space is one of the difficulty that affect short message service (SMS) classification performance. Some studies used feature selection methods to pick up some features, while other studies used the full extracted features. In this work, we aim to analyse the relationship between features size and classification performance. For that, a classification performance comparison was carried out between ten features sizes selected by varies feature selection methods. The used methods were chi-square, Gini index and information gain (IG). Support vector machine was used as a classifier. Area Under the ROC (Receiver Operating Characteristics) Curve between true positive rate and false positive rate was used to measure the classification performance. We used the repeated measures ANOVA at p < 0.05 level to analyse the performance. Experimental results showed that IG method outperformed the other methods in all features sizes. The best result was with 50% of the extracted features. Furthermore, the results explicitly showed that using larger features size in the classification does not mean superior performance but sometimes leads to less classification performance. Therefore, feature selection step should be used. By reducing the used features for the classification, without degrading the classification performance, it means reducing memory usage and classification time.

Keywords : Short text classification; content-based SMS spam filtering; SMS classification; dimension reduction; feature selection; support vector machine; ANOVA.

        · text in English     · English ( pdf )