PRETO: A high-performance text mining tool for preprocessing Turkish texts
13th International Conference on Computer Systems and Technologies, CompSysTech 2012, Ruse, Bulgaristan, 22 - 23 Haziran 2012, ss.134-140, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Doi Numarası: 10.1145/2383276.2383297
- Basıldığı Şehir: Ruse
- Basıldığı Ülke: Bulgaristan
- Sayfa Sayıları: ss.134-140
- Anahtar Kelimeler: Data mining, Natural language processing, Text mining, Text preprocessing
- Maltepe Üniversitesi Adresli: Evet
Özet
Text documents are usually unstructured and written in natural language. To apply conventional data mining techniques on text documents, a preprocessing operation is indispensable. In this paper, we introduce PRETO, a cross-platform, powerful and scalable preprocessing tool developed specifically for preprocessing Turkish texts, with a wide range of preprocessing options like stemming, stopword filtering, statistical term filtering, and n-gram generation. We demonstrate the performance and scalability of PRETO with some experiments on large document collections. Copyright ©2012 ACM.