Chinese Word Segmenter
Paste any Chinese text and see the words it is made of, counted, with pinyin and how common each one is. Finds the word boundaries Chinese does not write. Runs entirely in your browser.
0 of 5,000 characters
Try:
Words
No Chinese words found in that text.
| Pinyin |
|---|
Adds 400,000 mostly-name entries, so people and places stay whole instead of splitting into characters.
The full dictionary is kept on this device, so it loads without downloading again.
Pinyin settings
Chinese writes no spaces. 我喜欢学习中文 is seven characters and four words, and knowing which four is most of what reading a sentence takes. This tool finds them, so paste an article in and you get the words it uses rather than the characters it is spelled with.
Each word is listed once however often the text uses it, with the number of times beside it. Order the list by that number to see what the text is about, or by how common each word is in general to see which of them are worth learning first.
One writing, two readings
得 is de after a verb and děi in front of one. 行 is xíng in 旅行 and
háng in 银行. Those are different words that happen to be written the same
way, and this list keeps them apart: each row is a writing and a reading
together, read the way that occurrence in your text is read.
That is the same arrangement the Chinese dictionary uses, where a word’s meanings sit inside the reading they belong to rather than beside it.
What “how common” means
The last column is how often the word turns up in Chinese generally, which is a different question from how often your text uses it. A text about banking uses 银行 a dozen times and 的 twice, and only one of those two is a common word.
The scale runs from very common down to rare, measured over a corpus of about 60 million words. Words are grouped into bands rather than ranked, because the underlying measurement has sixteen levels and printing a number would claim a precision it does not have.
The band belongs to the written form, so where one writing is read two ways both rows carry the same band. 得 is very common however it is read, and the corpus counts the character rather than the two words it spells.
Names and places
The tool opens with a dictionary of 66,730 entries, which is every character and the fifty thousand commonest words. A name it has never seen splits into its characters, so 马克思 comes out as three words rather than one.
The full dictionary adds 400,000 mostly-name entries and fixes that. It is a 2.6 MB download and it belongs to readers with an account, so there is a button above the list once you are signed in.
Powered by pinyinjs