URL 编码的三个坑:空格变 + 还是 %20、中文与保留字符、双重编码Three URL Encoding Pitfalls: Space as + or %20, Chinese and Reserved Characters, Double Encoding
URL 编码看似简单,就是把不安全字符换成 `%XX`,但实际联调时空格变 `+` 还是 `%20`、中文该不该编码、参数被双重编码后服务端死活解析不出来——这些坑我一个个踩过。本文把三个最常见的问题讲透,附排查思路URL encoding looks simple — just replace unsafe characters with `%XX` — but in real integrations, whether a space becomes `+` or `%20`, whether Chinese needs encoding, and parameters that are double-encoded and can't be parsed on the server side are pitfalls I've hit one by one. This article explains the three most common problems and how to debug them.
空格的两种命运:+ 与 %20The two fates of a space: + and %20
这是最经典的坑。URL 里的空格有两种编码方式:`%20` 和 `+`。`%20` 是 RFC 3986 定义的百分号编码,适用于 URL 的所有部分(路径、查询参数、片段)。`+` 则来自 `application/x-www-form-urlencoded` 这个 MIME 类型的规范——也就是 HTML 表单提交时默认的编码方式,只在查询字符串(query string)和表单 body 里有效。This is the classic pitfall. A space in a URL has two encodings: `%20` and `+`. `%20` is the percent-encoding defined by RFC 3986, valid in all parts of a URL (path, query, fragment). `+` comes from the `application/x-www-form-urlencoded` MIME type spec — the default encoding for HTML form submissions — and is only valid in query strings and form bodies.
问题就出在这里。如果你用 JavaScript 的 `encodeURIComponent("hello world")`,得到的是 `hello%20world`;但如果你用 Java 的 `URLEncoder.encode("hello world", "UTF-8")`,得到的是 `hello+world`。两边对不上时,服务端解析就会出问题——比如一个用 `+` 编码的参数被放到了 URL 路径里(路径里 `+` 就是字面量加号,不会被解码成空格),或者一个用 `%20` 编码的查询参数被一个只认 `+` 的旧服务端解析,空格就变成了字面量 `%20`。The problem is right there. If you use JavaScript's `encodeURIComponent("hello world")`, you get `hello%20world`; but if you use Java's `URLEncoder.encode("hello world", "UTF-8")`, you get `hello+world`. When the two sides don't match, server-side parsing breaks — for example, a `+`-encoded parameter placed in the URL path (where `+` is a literal plus sign, not decoded to space), or a `%20`-encoded query parameter parsed by a legacy server that only recognizes `+`, so the space becomes the literal string `%20`.
我的经验是:查询参数里统一用 `%20`,因为 `%20` 在所有上下文里都能被正确解码为空格,而 `+` 只在表单编码上下文里有效。如果你的服务端框架把 `+` 解码成空格(大多数现代框架都会),那两种都能用,但跨语言、跨系统对接时,`%20` 更安全。My rule of thumb: use `%20` uniformly in query parameters, because `%20` is correctly decoded to a space in every context, while `+` only works in form-encoding contexts. If your server framework decodes `+` to space (most modern frameworks do), both work, but for cross-language, cross-system integration, `%20` is safer.
中文与保留字符:什么时候必须编码Chinese and reserved characters: when encoding is mandatory
中文字符在 URL 里不是"可选项",而是必须编码。RFC 3986 规定 URL 只能包含 ASCII 可打印字符中的一个子集,中文(以及任何非 ASCII 字符)必须按 UTF-8 编码后再做百分号编码。比如"中"字的 UTF-8 是 `E4 B8 AD`,编码后就是 `%E4%B8%AD`。Chinese characters in URLs are not "optional" — they must be encoded. RFC 3986 specifies that URLs can only contain a subset of printable ASCII characters; Chinese (and any non-ASCII character) must be UTF-8 encoded and then percent-encoded. For example, the UTF-8 bytes of "中" are `E4 B8 AD`, so the encoded form is `%E4%B8%AD`.
实际中的坑有两个。第一,编码用错字符集。如果服务端用 GBK 解码而客户端用 UTF-8 编码,中文就会乱码。现代框架默认都是 UTF-8,但一些老系统(尤其是国内的政务、企业系统)可能还在用 GBK,对接时一定要确认字符集。第二,浏览器地址栏的"欺骗性"。你在浏览器里看到 `https://example.com/search/中文`,浏览器显示的是中文,但实际发出去的 HTTP 请求里已经是 `%E4%B8%AD%E6%96%87` 了。所以用 Charles 或 F12 看实际请求时,不要被地址栏的显示迷惑。There are two practical pitfalls. First, using the wrong charset. If the server decodes with GBK but the client encodes with UTF-8, Chinese becomes mojibake. Modern frameworks default to UTF-8, but some legacy systems (especially domestic government and enterprise systems) may still use GBK — always confirm the charset when integrating. Second, the browser address bar is "deceptive." You see `https://example.com/search/中文` in the browser, but the actual HTTP request sent is already `%E4%B8%AD%E6%96%87`. So when inspecting actual requests with Charles or F12, don't be fooled by the address bar display.
保留字符(reserved characters)是另一类容易搞混的。`:` `/ ? # [ ] @ ! $ & ' ( ) * + , ; =` 这些字符在 URL 里有特殊含义,比如 `&` 分隔参数、`=` 分隔键值、`#` 标记片段。如果你的参数值本身包含这些字符(比如一个密码里有 `&`,或者一个搜索词是 `a+b`),就必须编码,否则会被当成 URL 语法的一部分。比如 `q=a+b` 里的 `+` 在查询参数里会被解码成空格,你实际想搜的是字面量 `a+b` 的话,应该写成 `q=a%2Bb`。Reserved characters are another source of confusion. `: / ? # [ ] @ ! $ & ' ( ) * + , ; =` have special meanings in URLs — `&` separates parameters, `=` separates keys and values, `#` marks the fragment. If your parameter value itself contains these characters (a password with `&`, or a search term `a+b`), you must encode them, otherwise they'll be treated as URL syntax. For example, in `q=a+b`, the `+` is decoded to a space in query parameters; if you actually want to search for the literal `a+b`, write `q=a%2Bb`.
双重编码:服务端死活解析不出来的元凶Double encoding: the culprit when the server can't parse anything
双重编码是最隐蔽的坑,也是我排查时间最长的一类。场景是这样的:客户端对参数做了一次编码(比如 `hello world` → `hello%20world`),然后某个中间层(网关、代理、或者另一个 SDK)又对整个 URL 做了一次编码,`%` 本身被编码成了 `%25`,结果就变成了 `hello%2520world`。服务端解码一次后得到 `hello%20world`——它以为这就是最终值,但实际上 `%20` 应该再解码一次才是空格。Double encoding is the most insidious pitfall, and the one I've spent the longest debugging. The scenario: the client encodes a parameter once (`hello world` → `hello%20world`), then some intermediate layer (gateway, proxy, or another SDK) encodes the entire URL again — the `%` itself gets encoded to `%25`, resulting in `hello%2520world`. The server decodes once and gets `hello%20world` — it thinks this is the final value, but `%20` should actually be decoded one more time to a space.
更常见的情况是:前端用 `encodeURIComponent` 编码了参数,然后又用 `fetch` 或 `axios` 发请求——这些库内部可能会对 URL 再做一次编码。或者反过来,服务端框架自动解码了一次,你的代码里又手动 `URLDecoder.decode` 了一次,导致 `%20` 被解码成空格后,空格又被当成普通字符保留——这倒不会出错,但如果原始值里有 `%` 开头的序列,第二次解码就会出问题。A more common case: the frontend encodes parameters with `encodeURIComponent`, then sends the request with `fetch` or `axios` — these libraries may internally encode the URL again. Or the reverse: the server framework automatically decodes once, and your code manually calls `URLDecoder.decode` again, causing `%20` to be decoded to a space, which is then kept as a normal character — this won't error, but if the original value contains `%`-prefixed sequences, the second decode can cause problems.
排查双重编码的方法很直接:第一步,在客户端打印编码前的原始值和编码后的值;第二步,用 Charles 或 tcpdump 抓包,看实际发出去的 URL 是什么;第三步,在服务端打印收到的原始 query string(不要用框架解析后的参数值,要看 raw query)。对比这三个值,哪一步多了一次编码就一目了然。Debugging double encoding is straightforward: Step one, print the original value before encoding and the value after encoding on the client side. Step two, capture the actual outgoing URL with Charles or tcpdump. Step three, print the raw query string received on the server side (don't use the framework-parsed parameter value — look at the raw query). Compare these three values, and which step added an extra encoding becomes obvious.
一个实用技巧:如果不确定一个字符串是不是被双重编码了,可以看 `%` 的数量。正常编码后,每个非 ASCII 字符对应 3 个字符(`%XX`);如果看到 `%25` 开头,那 `%` 本身被编码了,几乎可以确定是双重编码。本站的 URL 编解码工具支持编码和解码双向操作,也能检测是否存在双重编码的迹象,排查时可以把可疑的字符串粘进去对比。A practical trick: if you're unsure whether a string is double-encoded, count the `%` signs. After normal encoding, each non-ASCII character corresponds to 3 characters (`%XX`); if you see `%25` at the start, the `%` itself was encoded — almost certainly double encoding. Our URL encoder/decoder supports both encode and decode operations and can detect signs of double encoding — paste a suspicious string in and compare during debugging.