F

HTML 实体编解码:  与空格的区别,以及 XSS 防护中的转义边界HTML Entity Encoding: Why   Is Not a Space, and Escaping Boundaries in XSS Defense

` ` 看起来就是个空格,但 `trim()` 去不掉、`split(' ')` 切不开,因为它根本不是空格。本文从实体编码的原理讲起,结合 XSS 转义边界的真实案例,说清什么时候该转义、转义哪些字符。` ` looks like a space, but `trim()` won't remove it and `split(' ')` won't split on it — because it is not a space at all. This article starts from how entity encoding works, then uses real XSS escaping boundary cases to explain when to escape and which characters to escape.

实体编码是什么:为什么浏览器能把 & 还原成 &What entity encoding is: why the browser turns & back into &

HTML 实体(HTML entity)是一种用可打印字符表示特殊字符的机制,格式是 `&名称;` 或 `&#数字;`。比如 `&lt;` 表示小于号 `<`,`&amp;` 表示和号 `&`,`&#x4e2d;` 用十六进制码点表示"中"。浏览器解析 HTML 时会自动把这些实体还原成对应的字符,所以你在源码里写 `&amp;`,页面上显示的是 `&`。An HTML entity is a mechanism for representing special characters with printable ones, written as `&name;` or `&#number;`. For example, `&lt;` is the less-than sign `<`, `&amp;` is the ampersand `&`, and `&#x4e2d;` represents "中" by its hexadecimal code point. The browser automatically decodes these entities while parsing HTML, so writing `&amp;` in source renders as `&` on the page.

为什么需要这套机制?因为 `<`、`>`、`&` 这些字符在 HTML 里有语法含义——`<` 开始标签,`&` 开始实体。如果正文里直接出现 `<script>`,浏览器会把它当成标签解析而不是文本显示。实体编码就是给这些"有歧义"的字符一个无歧义的写法。数字实体 `&#NNNN;` 则可以表示任意 Unicode 字符,在不支持直接输入生僻字的环境里特别有用。Why does this mechanism exist? Because `<`, `>` and `&` have syntactic meaning in HTML — `<` opens a tag and `&` opens an entity. If the body text literally contains `<script>`, the browser parses it as a tag rather than displaying it as text. Entity encoding gives these "ambiguous" characters an unambiguous representation. Numeric entities `&#NNNN;` can represent any Unicode character, which is especially useful in environments that cannot input rare characters directly.

&nbsp; 不是空格:U+00A0 的三个坑&nbsp; is not a space: three pitfalls of U+00A0

`&nbsp;` 的全称是 non-breaking space,对应 Unicode 码点 U+00A0。它在视觉上和普通空格(U+0020)一模一样,但语义完全不同:浏览器不会在 `&nbsp;` 处换行,所以它常用来把"100 km"这种数字和单位粘在一起,避免行尾拆开。`&nbsp;` stands for non-breaking space, corresponding to Unicode code point U+00A0. Visually it is identical to a regular space (U+0020), but semantically they are different: the browser will not break a line at `&nbsp;`, which is why it is often used to glue a number and its unit together — "100 km" — so they do not get separated at line end.

我踩过的第一个坑是字符串比较:从富文本编辑器导出的内容里混着 `&nbsp;`,解码后变成 U+00A0,和普通空格做 `===` 比较返回 false,去重逻辑全部失效。第二个坑是 `trim()`:JavaScript 的 `String.prototype.trim()` 能去掉 U+00A0,但很多后端语言的 trim 只去 ASCII 空白,导致"看起来已经 trim 了"的字符串其实还带着不可见字符。第三个坑是 `split(/\s+/)`:正则里的 `\s` 在大多数引擎里匹配 U+00A0,但如果你写的是 `split(' ')` 字面量空格,就完全切不开。排查这类问题的关键是把可疑字符转成码点看——U+00A0 和 U+0020 一眼就能区分。The first pitfall I hit was string comparison: content exported from a rich-text editor contained `&nbsp;`, which decoded to U+00A0, and a `===` comparison against a regular space returned false, breaking the entire deduplication logic. The second was `trim()`: JavaScript's `String.prototype.trim()` does remove U+00A0, but many backend languages only strip ASCII whitespace, so a string that "looks trimmed" may still carry invisible characters. The third was `split(/\s+/)`: the `\s` shorthand matches U+00A0 in most engines, but if you write `split(' ')` with a literal space, it will not split at all. The key to debugging these issues is converting suspicious characters to code points — U+00A0 and U+0020 are instantly distinguishable.

XSS 转义边界:不是所有上下文都用同一套规则XSS escaping boundaries: not every context uses the same rules

很多人以为"把用户输入做 HTML 转义就能防 XSS",这只对了一半。转义的规则取决于输出位置——HTML 正文、属性值、JavaScript 字符串、CSS、URL,每个上下文的危险字符都不一样。在 HTML 正文里,转义 `<`、`>`、`&`、`"`、`'` 这五个字符基本就够了;但如果输出在 `onclick="..."` 这种属性里,光转义 HTML 还不够,因为攻击者可以用 `&quot;` 闭合引号后注入事件处理器——浏览器会先解码实体再执行属性值。Many people assume "HTML-escape user input and XSS is prevented" — that is only half right. Escaping rules depend on the output context: HTML body, attribute values, JavaScript strings, CSS and URLs each have different dangerous characters. In HTML body text, escaping `<`, `>`, `&`, `"` and `'` is usually sufficient; but inside an attribute like `onclick="..."`, HTML escaping alone is not enough, because an attacker can use `&quot;` to close the quote and inject an event handler — the browser decodes entities before executing attribute values.

我在一次代码评审中发现过一个真实漏洞:模板把用户名直接塞进 `<a href="/user/{{name}}">`,开发者只做了 HTML 转义。但攻击者可以注册名为 `" onclick="alert(1)` 的用户,HTML 转义会把双引号变成 `&quot;`,看起来安全了——可一旦这个值又被某个 JS 逻辑取出来重新写入 DOM,实体被解码后引号就恢复了。正确做法是按上下文分层转义:HTML 正文用 HTML 转义,属性值额外确保引号闭合,URL 用 `encodeURIComponent`,JS 字符串用 `JSON.stringify`。永远不要用"一套转义打天下"。I found a real vulnerability in a code review: a template inserted the username directly into `<a href="/user/{{name}}">`, and the developer only applied HTML escaping. An attacker could register a user named `" onclick="alert(1)`; HTML escaping turns the double quote into `&quot;`, which looks safe — but once that value is extracted by some JavaScript logic and written back into the DOM, the entity is decoded and the quote is restored. The correct approach is context-aware layered escaping: HTML escaping for body text, extra quote-closure guarantees for attributes, `encodeURIComponent` for URLs, and `JSON.stringify` for JS strings. Never try to cover every context with one escaping function.

编解码的日常:用工具快速验证实体结果Everyday encoding: quickly verifying entity results with a tool

日常开发中我经常需要确认一段文本的实体编码结果是否正确——比如富文本导出后 `&nbsp;` 有没有被错误转成 `&amp;nbsp;`,或者接口返回的 `&#x4e2d;` 解码后是不是预期的中文。手动查实体表效率太低,用一个在线工具把文本贴进去,编码和解码结果一目了然。In day-to-day development I often need to verify whether a piece of text has been entity-encoded correctly — for example, whether `&nbsp;` was mistakenly double-encoded to `&amp;nbsp;` after rich-text export, or whether an API-returned `&#x4e2d;` decodes to the expected Chinese character. Looking up entity tables manually is slow; paste the text into an online tool and the encode/decode results are instantly clear.

本站的 HTML 实体工具支持命名实体(`&nbsp;`、`&amp;`)和数字实体(`&#x4e2d;`)的双向转换,可以帮你快速验证转义结果、排查 `&nbsp;` 混入导致的比较失败。处理用户生成内容时,先在工具里跑一遍编码结果,再决定用哪一层转义策略,比上线后被安全扫描打回来要省事得多。Our HTML entity tool supports bidirectional conversion between named entities (`&nbsp;`, `&amp;`) and numeric entities (`&#x4e2d;`), helping you quickly verify escaping results and debug comparison failures caused by stray `&nbsp;` characters. When handling user-generated content, running the encoding result through a tool first — and then deciding which escaping layer to apply — is far easier than getting flagged by a security scan after launch.

← 返回教程列表← Back to all guides