Volta

On the Impact of Dataset Characteristics on Arabic Document Classification

Unknown authors · 2014
hash_id: 94aa2b9e1332038e8ebb74626d3231f726325fcdd9d49688ed52a2bb3a7fb24b · DOI: 10.5120/17701-8680

This paper describes the impact of dataset characteristics on the results of Arabic document classification algorithms using TF-IDF representations.The experiments compared different stemmers, different categories and different training set sizes, and found that different dataset characteristics produced widely differing results, in one case attaining a remarkable 99% recall (accuracy).The use of a standard dataset would eliminate this variability and enable researchers to gain comparable knowledge from the published results.

Reference & gravity metrics

Citations
0
Citations / yr
0.00
RCR
Mass
0.00
Depth
0.00
Momentum
0.000
Burn rate
0.00 ATP/day
Start price
25.00 ATP

Secondary-market trade history

No secondary-market trades recorded for this Volta yet.

References (0)

No outbound references recorded.

Cited by (1)

A framework for retrieving Arabic documents based on queries written in Arabic … secondary