HSK StudioTools

    Which words in this text don't I know yet?

    Paste a chapter, an article or a chat log. You get every word in it, ranked by how often it appears, with the ones already in your review deck marked off — so the list you are left with is the list of words you actually have to learn.

    Your text

    0 Chinese characters

    Preparing HSK default… 1%

    Loading the dictionary…

    Why the words you know are part of the answer

    Every free text analyser tells you the same thing: here is your text, split into words, sorted by HSK level. That is a fact about the text. It is not the fact you need, because the useful question is not "how hard is this in general" but "how hard is this for me" — and the answer depends on the few thousand words you happen to have met.

    The tools that do answer it are paid desktop programs, and the workflow around them is a round trip: export your known words from Anki or Pleco, import the file into the analyser, analyse the text, export the unknown words, import those back into your deck. It works, and people do it, because the analyser and the deck are different programs that have to be introduced to each other.

    Here they are the same program. The words you have saved anywhere in HSK Studio — from the dictionary, the reader, a lesson, or this page — are the known-word list, so there is no import step and no export step unless you want one.

    What the coverage number means

    Coverage is measured in running words rather than distinct words. If 的 appears forty times, that is forty easy encounters, not one — and a percentage built on distinct words would badly understate how readable a text is. The thresholds the tool reports against come from research on extensive reading: about 98% of running words known for comfortable independent reading, about 95% before a text is workable with a dictionary at hand.

    Characters that no dictionary word matched — names, technical terms, typos — count against you rather than being quietly dropped. A text full of unfamiliar proper nouns really is harder, and a number that ignored them would flatter it.

    How the text is split into words

    Chinese is written without spaces, so any word list has to start by deciding where words end. This uses longest-match segmentation against an 11,643-entry HSK 3.0 dictionary: at each position it takes the longest dictionary word that fits. That is why 中国人 comes back as one word rather than three characters, and it is also the reason the pinyin is right — the reading comes from the word, so 银行 is yínháng and never yínxíng.

    Longest-match is not perfect; no unsupervised segmenter is. Where it fails it splits a rare compound into shorter real words rather than inventing anything, so the word list stays honest even when the segmentation is arguable.

    Your text stays in your browser

    The dictionary is downloaded once and the analysis runs on your device. Nothing you paste is uploaded, there is no account, and the last thing you analysed is kept in this browser's local storage so it is still here when you come back — and nowhere else.

    The other Chinese tools

    Same engine, same dictionary, different question. Whatever you pasted here works in all of them.