J. A. Bakar, K. Omar, M. F. Nasrudin, Mohd Zamri Murah
{"title":"Tokenizer for the Malay language using pattern matching","authors":"J. A. Bakar, K. Omar, M. F. Nasrudin, Mohd Zamri Murah","doi":"10.1109/ISDA.2014.7066258","DOIUrl":null,"url":null,"abstract":"Tokenization is a fundamental task focused on text processing. Among other tasks, the segmentation process is used to identify information units, such as sentences and words. In this paper, we discuss the Natural Language ToolKit (NLTK) tokenizer as a step to manipulate patterns within text. The purpose of this work is to build up Natural Language Processing (NLP) base for Jawi corpus. A series of experiments was performed, to validate the corpus and fulfill the requirement of the Jawi script tokenizer, with the promising results. Based on these promising results, the token will be used for tagging process.","PeriodicalId":328479,"journal":{"name":"2014 14th International Conference on Intelligent Systems Design and Applications","volume":"60 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2014-11-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"4","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2014 14th International Conference on Intelligent Systems Design and Applications","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ISDA.2014.7066258","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 4
Abstract
Tokenization is a fundamental task focused on text processing. Among other tasks, the segmentation process is used to identify information units, such as sentences and words. In this paper, we discuss the Natural Language ToolKit (NLTK) tokenizer as a step to manipulate patterns within text. The purpose of this work is to build up Natural Language Processing (NLP) base for Jawi corpus. A series of experiments was performed, to validate the corpus and fulfill the requirement of the Jawi script tokenizer, with the promising results. Based on these promising results, the token will be used for tagging process.