Difference between revisions of "Language/Multiple-languages/Culture/Text-Processing-Tools"

From Polyglot Club WIKI
Jump to navigation Jump to search
Line 46: Line 46:
* Roy_VnTokenizer https://github.com/roy-a/Roy_VnTokenizer
* Roy_VnTokenizer https://github.com/roy-a/Roy_VnTokenizer
* VietSeg https://github.com/manhtai/vietseg
* VietSeg https://github.com/manhtai/vietseg
* VnCoreNLP https://github.com/vncorenlp/VnCoreNLP


Many of them are written in Python, which is an important language in machine learning. It's very easy to use: just create a blank “.py” file, write a line to import from the library, write a line to segment the text, write a line to save the result.
Many of them are written in Python, which is an important language in machine learning. It's very easy to use: just create a blank “.py” file, write a line to import from the library, write a line to segment the text, write a line to save the result.

Revision as of 14:07, 20 May 2021


In some languages, words are not separated by spaces, for example: Chinese, Japanese, Lao, Thai. In Vietnamese, spaces are used to divide syllables instead of words. This brings about difficulties for computer programs like VocabHunter, gritz and text-memorize, where words are detected only with spaces.

The solution is called “word segmentation”, which detects words and insert spaces in between. Sounds like easy, but they have to deal with ambiguities and unknown words including proper names, the time and accuracy are both to be considered. It can be based on the dictionary or machine learning.

Here are free and open-source tools to do it:

Chinese:

Japanese:

Thai:

Vietnamese:

Many of them are written in Python, which is an important language in machine learning. It's very easy to use: just create a blank “.py” file, write a line to import from the library, write a line to segment the text, write a line to save the result.

If you don't know Python, please try this: