Word Counter & Text Cleaner
Count Chinese characters, Latin words, characters and paragraphs locally. Preview text cleanup, undo changes and continue to token planning.
Up to 200,000 UTF-16 units. Text stays in this tab; it is not saved or uploaded.
Clean up text
Preview first. Apply only the changes you want.
How these counts work
Chinese characters use the Unicode Han script. Latin words include accented letters and internal apostrophes or hyphens (don’t and well-known each count as one). Numbers and punctuation are excluded from word counts. Other writing systems are included in character counts but not the Chinese + Latin total.
Characters are Unicode code points, not bytes or tokens: a combined emoji can contain several characters. Lines include empty lines; paragraphs are separated by blank lines. Reading time estimates use 300 Chinese characters or 200 Latin words per minute and do not measure comprehension.
Cleanup normalizes line endings to LF. Duplicate lines are matched exactly after the selected whitespace cleanup; the first occurrence is kept. The invisible-character option removes only U+200B and U+FEFF, preserving emoji joiners. Review indentation in code and tables before applying.