Volta
On the Impact of Dataset Characteristics on Arabic Document Classification
Unknown authors
· 2014
hash_id: 94aa2b9e1332038e8ebb74626d3231f726325fcdd9d49688ed52a2bb3a7fb24b
· DOI: 10.5120/17701-8680
This paper describes the impact of dataset characteristics on the results of Arabic document classification algorithms using TF-IDF representations.The experiments compared different stemmers, different categories and different training set sizes, and found that different dataset characteristics produced widely differing results, in one case attaining a remarkable 99% recall (accuracy).The use of a standard dataset would eliminate this variability and enable researchers to gain comparable knowledge from the published results.
Reference & gravity metrics
Citations
0
Citations / yr
0.00
RCR
Mass
0.00
Depth
0.00
Momentum
0.000
Burn rate
0.00 ATP/day
Start price
25.00 ATP
Secondary-market trade history
No secondary-market trades recorded for this Volta yet.
References (0)
No outbound references recorded.
Cited by (1)
| A framework for retrieving Arabic documents based on queries written in Arabic … | secondary |