Chinese Word Segmenter

Paste any Chinese text and see the words it is made of, counted, with pinyin and how common each one is. Finds the word boundaries Chinese does not write. Runs entirely in your browser.

0 of 5,000 characters

Try:

Words

The words in your text. Choose a heading to reorder the list.
Pinyin

Adds 400,000 mostly-name entries, so people and places stay whole instead of splitting into characters.

Pinyin settings
Tones

Tone marks are standard. Numbers are easier to type and to search for; superscript keeps the word searchable either way.

Word spacing

Standard pinyin spaces by word rather than by syllable: 南京市 is Nánjīng Shì, not Nán jīng shì.

Capital letters

Automatic capitalises proper nouns and the first word of each sentence, which needs the Chinese to be punctuated.

Punctuation

Whether 。,、;:?! become . , ; : ? ! with the spacing English uses.

Numbers

Read says the digits aloud, and takes its cue from what follows: 1998年 is yī jiǔ jiǔ bā nián but 3个 is sān gè.

一 and 不 tone changes

The tone changes these two characters take from what follows them: 不是 is bú shì.

Third-tone sandhi

Off by default, because standard pinyin writes the underlying tone: 你好 is written nǐ hǎo even though it is said ní hǎo.

Apostrophes

The 隔音符号, which keeps a syllable boundary readable. Standard writes it only where leaving it out would be ambiguous: Xī'ān, but Tiānānmén.

Reading standard

Where the mainland and Taiwan differ: 垃圾 is lājī in 普通话 and lèsè in 國語.

Chinese writes no spaces. 我喜欢学习中文 is seven characters and four words, and knowing which four is most of what reading a sentence takes. This tool finds them, so paste an article in and you get the words it uses rather than the characters it is spelled with.

Each word is listed once however often the text uses it, with the number of times beside it. Order the list by that number to see what the text is about, or by how common each word is in general to see which of them are worth learning first.

One writing, two readings

得 is de after a verb and děi in front of one. 行 is xíng in 旅行 and háng in 银行. Those are different words that happen to be written the same way, and this list keeps them apart: each row is a writing and a reading together, read the way that occurrence in your text is read.

That is the same arrangement the Chinese dictionary uses, where a word’s meanings sit inside the reading they belong to rather than beside it.

What “how common” means

The last column is how often the word turns up in Chinese generally, which is a different question from how often your text uses it. A text about banking uses 银行 a dozen times and 的 twice, and only one of those two is a common word.

The scale runs from very common down to rare, measured over a corpus of about 60 million words. Words are grouped into bands rather than ranked, because the underlying measurement has sixteen levels and printing a number would claim a precision it does not have.

The band belongs to the written form, so where one writing is read two ways both rows carry the same band. 得 is very common however it is read, and the corpus counts the character rather than the two words it spells.

Names and places

The tool opens with a dictionary of 66,730 entries, which is every character and the fifty thousand commonest words. A name it has never seen splits into its characters, so 马克思 comes out as three words rather than one.

The full dictionary adds 400,000 mostly-name entries and fixes that. It is a 2.6 MB download and it belongs to readers with an account, so there is a button above the list once you are signed in.

Powered by pinyinjs