Chinese Word Segmenter

Paste any Chinese text and see the individual words it is made of, with counts, pinyin, HSK level and a link to each word's dictionary entry. Segments Chinese text into individual words. Runs entirely in your browser.

0 of 5,000 characters

Try:

Words

Words saved from this text use its name as a tag.

The words in your text. Choose a heading to reorder the list.
Pinyin Save

Adds 400,000 mostly-name entries, so people and places stay whole instead of splitting into characters.

Pinyin settings
Tones

Tone marks are standard. Numbers are easier to type and to search for; superscript keeps the word searchable either way.

Word spacing

Standard pinyin spaces by word rather than by syllable: 南京市 is Nánjīng Shì, not Nán jīng shì.

Capital letters

Automatic capitalises proper nouns and the first word of each sentence, which needs the Chinese to be punctuated.

Punctuation

Whether 。,、;:?! become . , ; : ? ! with the spacing English uses.

Numbers

Read says the digits aloud, and takes its cue from what follows: 1998年 is yī jiǔ jiǔ bā nián but 3个 is sān gè.

一 and 不 tone changes

The tone changes these two characters take from what follows them: 不是 is bú shì.

Third-tone sandhi

Off by default, because standard pinyin writes the underlying tone: 你好 is written nǐ hǎo even though it is said ní hǎo.

Apostrophes

The 隔音符号, which keeps a syllable boundary readable. Standard writes it only where leaving it out would be ambiguous: Xī'ān, but Tiānānmén.

Reading standard

Where the mainland and Taiwan differ: 垃圾 is lājī in 普通话 and lèsè in 國語.

Powered by pinyinjs