F

中英文字数统计为什么对不上:字符、码点、Word 与编辑器的计数差异Why Chinese and English Word Counts Disagree: Characters, Code Points, Word vs Editor Counting

同一段文字,Word 说 320 字,编辑器说 412 字符,公众号后台又给了一个 356。这篇文章把"字数"背后的字符、码点、码元与分词规则讲清楚,让你再也不会被不同工具的数字搞晕The same paragraph: Word says 320 words, your editor says 412 characters, and the publishing backend reports 356. This article unpacks characters, code points, code units and tokenisation rules so you stop being confused by inconsistent counters.

一个真实的对不上案例A real counting mismatch

上个月帮同事校对一篇技术稿,他把同一段 500 字左右的中文分别贴进 Word、VS Code 和某公众号后台,得到了三个完全不同的数字:Word 显示"字数 487",VS Code 右下角显示"512 字符",公众号后台则是"463 字"。他第一反应是某个工具算错了,其实三个都没错,只是"字数"这个词在不同语境下指代的东西完全不同。Last month a colleague pasted the same 500-character Chinese paragraph into Word, VS Code and a publishing backend and got three different numbers: 487, 512 and 463. His first instinct was that one tool was broken, but all three were correct — "word count" simply means different things in different contexts.

我自己踩过更严重的坑:一个按"中文字数"计费的外包项目,甲方用 Word 统计是 12000 字,我用脚本 `str.length` 算出来是 13800,差了 15%。最后才发现是因为文中有大量 Emoji 和英文术语,Word 把一个 Emoji 算一个字,而 JavaScript 的 `length` 把代理对拆成了两个码元。这个差额差点让我白干两天。I once hit a worse version: a freelance project billed by Chinese word count came out at 12,000 in Word but 13,800 when I ran `str.length` in a script — a 15% gap. The cause was heavy use of Emoji and English terms: Word counts one Emoji as one character, while JavaScript's `length` splits surrogate pairs into two code units. That gap almost cost me two days of unpaid work.

字符、码点、码元:三个被混用的概念Characters, code points and code units

要理解计数差异,先分清三个层次。最底层是**码元(code unit)**,即编码方案里最小的存储单位:UTF-16 里一个码元是 16 位,UTF-8 里是 8 位。往上是**码点(code point)**,Unicode 给每个字符分配的唯一编号,范围 U+0000 到 U+10FFFF。一个码点在 UTF-16 里可能占 1 个或 2 个码元(Emoji 就是 2 个),在 UTF-8 里占 1 到 4 个字节。最上层才是用户感知的**字符(character)**,但即便是这个概念也不简单——"👨‍👩‍👧"是一个家庭 Emoji,由 4 个码点加零宽连接符组成,用户眼里它是一个字符。To understand the mismatch, separate three layers. At the bottom is the **code unit**, the smallest storage unit of an encoding: 16 bits in UTF-16, 8 bits in UTF-8. Above it is the **code point**, the unique number Unicode assigns to each character, ranging from U+0000 to U+10FFFF. One code point takes 1 or 2 code units in UTF-16 (Emoji take 2) and 1 to 4 bytes in UTF-8. At the top is the user-perceived **character** — but even that is fuzzy: "👨‍👩‍👧" is one family Emoji built from 4 code points plus zero-width joiners, yet users see a single character.

中文的情况更特殊。现代汉语没有空格分词,所以"字数"通常等同于"字符数",每个汉字算一个。但英文是按空格分词的,"word count"数的是词而不是字母。于是一段中英混排的文字,"字数"到底是汉字数、英文单词数,还是两者相加,不同工具各有各的定义。Word 的"字数"对中文按字符计、对英文按词计,再把标点单独处理;而很多编辑器只统计码元数,自然对不上。Chinese adds another twist. Modern Chinese has no whitespace between words, so "word count" usually equals "character count" — one Han character is one unit. English, by contrast, splits on whitespace and counts words, not letters. For mixed Chinese-English text, whether "word count" means Han characters, English words, or a sum of both depends entirely on the tool. Word counts Chinese by character and English by word, then handles punctuation separately; many editors only count code units, so the numbers never agree.

Word 与编辑器到底在数什么What Word and editors actually count

Microsoft Word 的统计逻辑最复杂。它的"字数"对中文文本按字符计(含标点),对西文按单词计,数字串算一个词,而"字符数(不计空格)"则是所有可见字符的码点级计数。Word 还会把段落标记、制表符排除在某些统计之外。这就是为什么同样一段文字,Word 的"字数"和"字符数"是两个值。Microsoft Word has the most complex logic. Its "word count" treats Chinese as characters (including punctuation), Western text as words, and numeric strings as one word, while "characters (no spaces)" is a code-point-level count of all visible glyphs. Word also excludes paragraph marks and tabs from some metrics. That is why the same paragraph yields different "words" and "characters" values inside Word itself.

VS Code 右下角的"字符"统计的是文档的 UTF-16 码元数,因为它的编辑器内核基于 UTF-16。一个 Emoji 在这里会被算成 2,而一个汉字只算 1。Vim 的 `g-` 统计的是字节数,又不一样。公众号、知乎这类平台的后台通常自己实现了一套"字数"规则,往往是"中文字符 + 英文单词"的混合口径,并且会剔除 Markdown 标记或 HTML 标签。VS Code's status bar counts UTF-16 code units because its editor core is built on UTF-16. An Emoji counts as 2 there, while a Han character counts as 1. Vim's `g-` counts bytes — yet another number. Publishing platforms like WeChat or Zhihu usually implement their own "word count", often a hybrid of "Chinese characters + English words" that also strips Markdown or HTML tags.

我在做字数统计工具时,最纠结的就是口径选择。最后采用了分栏展示:同时给出"中文字符数""英文单词数""总字符数(码点)""不含空格字符数"四个数字,让用户自己选需要的口径,而不是替用户做一个可能错的决定。When I built a word counter, the hardest decision was which metric to show. I ended up displaying four columns — Chinese characters, English words, total code points, and characters excluding whitespace — so the user picks the metric they need instead of the tool guessing wrong.

怎么选一个靠谱的字数统计How to pick a reliable counter

我的经验是三条。第一,先问清楚"字数"的定义方:是甲方、平台还是你自己?不同场景的结算口径不同,投稿前务必用对方指定的工具复核。第二,涉及 Emoji、组合字符或多语言混排时,不要用任何语言的 `string.length` 直接当字数——JavaScript、Java、C# 都是 UTF-16 码元计数,Python 3 的 `len()` 是码点计数但仍会把组合字符拆开。第三,需要精确到"用户看到的字符"时,用字形簇(grapheme cluster)分割,比如 JavaScript 的 `Intl.Segmenter` 或库 `grapheme-splitter`。My rule of thumb has three parts. First, ask who defines "word count": the client, the platform, or you? Different scenarios settle on different metrics, so always re-check with the tool the other side specifies before submitting. Second, when Emoji, combining marks or mixed scripts are involved, never trust any language's `string.length` — JavaScript, Java and C# count UTF-16 code units, and Python 3's `len()` counts code points but still splits combining sequences. Third, when you need "what the user sees", split on grapheme clusters with `Intl.Segmenter` or a library like `grapheme-splitter`.

日常写作和校对里,我更习惯用一个能同时展示多种口径的在线工具,把原文贴进去,一眼看到中文字符、英文单词、总字符和不含空格数,再对照目标平台的要求取数。本站的字数统计工具就是按这个思路做的,支持实时统计、中英文分栏和不含空格模式,写稿、投稿、对账时都能省掉反复切换工具的麻烦。For everyday writing and proofreading, I prefer an online tool that shows multiple metrics at once: paste the text and see Chinese characters, English words, total code points and no-space count side by side, then pick the number the target platform expects. Our word counter is built exactly this way — real-time counting, separate Chinese/English columns, and a no-whitespace mode — so you stop switching between tools while drafting, submitting or reconciling a bill.

← 返回教程列表← Back to all guides