大小写转换不简单:土耳其语 i、Unicode 大小写与 locale 陷阱Case Conversion Is Not Trivial: Turkish i, Unicode Case Mapping and Locale Traps
`"I".toLowerCase()` 谁不会写?但在土耳其语 locale 下它返回的不是 "i" 而是 "ı"。本文从这个经典 bug 出发,讲清 Unicode 大小写映射的一对多、无条件映射与 locale 映射,以及数据库、URL、HTTP 头里的大小写坑Who can't write `"I".toLowerCase()`? But under a Turkish locale it returns "ı", not "i". Starting from this classic bug, this article explains Unicode one-to-many case mapping, unconditional vs locale-aware mapping, and the case pitfalls in databases, URLs and HTTP headers.
土耳其语 i:最著名的 locale 陷阱Turkish i: the most famous locale trap
2010 年左右,一个流传甚广的 bug 是:在土耳其语版本的 Windows 上,某些 Java 程序启动直接报错 `ClassNotFoundException`,类名明明是 `InputStream`,却被转换成了 `İnputStream`(注意那个带点的大写 I)。根因是 JDK 早期在做字符串大小写转换时,默认使用了系统 locale,而土耳其语里 `i` 的大写是 `İ`(带点),`I` 的小写是 `ı`(不带点)——两套 i 互不对应。Around 2010 a widely circulated bug was: on Turkish-language Windows, some Java programs crashed at startup with `ClassNotFoundException` — the class name was clearly `InputStream`, but it had been converted to `İnputStream` (note the dotted capital I). The root cause was that early JDK used the system locale for case conversion by default, and in Turkish the uppercase of `i` is `İ` (dotted) while the lowercase of `I` is `ı` (dotless) — two separate i's that do not map to each other.
我自己第一次踩到类似的坑是在做一个邮件系统时,用户邮箱统一转小写存库。代码写的是 `email.toLowerCase()`,测试环境一切正常,上线后有土耳其用户反馈登录失败。排查发现邮箱里的 `i` 被转成了 `ı`,和注册时存的值不一致。修复方式是 `email.toLowerCase(Locale.ROOT)` 或 `toLowerCase(Locale.ENGLISH)`,强制使用不带 locale 语义的规则。这个教训让我从此对任何"默认 locale"的 API 都保持警惕。I first hit a similar pitfall building an email system: we normalised user emails to lowercase before storing them. The code was `email.toLowerCase()`, everything passed in testing, but after launch Turkish users reported login failures. Investigation showed that the `i` in their email had become `ı`, mismatching the value stored at registration. The fix was `email.toLowerCase(Locale.ROOT)` or `toLowerCase(Locale.ENGLISH)`, forcing locale-free rules. That lesson made me wary of every API that defaults to the system locale.
Unicode 大小写映射不是一对一Unicode case mapping is not one-to-one
很多人以为大小写转换是一个字符对应一个字符的查表操作,其实远非如此。Unicode 标准里定义了三种大小写映射:简单映射(Simple_Case_Mapping)是一对一,全映射(Full_Case_Mapping)则可能是一对多。最经典的例子是德语小写字母 `ß`(eszett),它的大写在传统正字法里没有对应单字符,全映射会把 `ß` 转成 `SS` 两个字符——也就是说,`"straße".toUpperCase()` 得到 `"STRASSE"`,长度从 6 变成了 7。Many people assume case conversion is a one-to-one lookup table. It is far from it. The Unicode standard defines three kinds of case mapping: Simple_Case_Mapping is one-to-one, while Full_Case_Mapping can be one-to-many. The classic example is the German lowercase `ß` (eszett): traditional orthography has no single uppercase equivalent, so full mapping turns `ß` into two characters `SS` — `"straße".toUpperCase()` yields `"STRASSE"`, growing from 6 to 7 characters.
反过来也有:希腊语的小写 `σ` 在词尾写成 `ς`,大写都是 `Σ`,但大写转小写时需要根据位置决定用哪个。还有一些字符的大小写转换涉及组合标记,比如 `dž`(U+01C6)的大写是 `DŽ`,标题格是 `Dž`,三种形态各占一个码点。更复杂的是 Lithuanian 等语言的大小写规则,会在转换时增删重音符号。The reverse also happens: Greek lowercase `σ` becomes `ς` at the end of a word, both uppercase to `Σ`, but uppercase-to-lowercase must pick the right form based on position. Some characters involve combining marks: `dž` (U+01C6) has uppercase `DŽ` and titlecase `Dž`, each a separate code point. Languages like Lithuanian have even more complex rules that add or remove accents during conversion.
这意味着"先转大写再转小写"不一定能还原原文,`str.toUpperCase().toLowerCase()` 不是幂等操作。做大小写不敏感的比较时,正确做法通常不是转成同一种大小写再比较,而是用专门的大小写折叠(case folding),比如 Unicode 的 `CaseFolding.txt` 或 Java 的 `String.equalsIgnoreCase`。This means "uppercase then lowercase" does not always round-trip; `str.toUpperCase().toLowerCase()` is not idempotent. For case-insensitive comparison, the correct approach is usually not converting to a single case and comparing, but using dedicated case folding — such as Unicode's `CaseFolding.txt` or Java's `String.equalsIgnoreCase`.
编程里的大小写坑:数据库、URL、HeaderCase pitfalls in code: databases, URLs, headers
大小写问题在工程里无处不在。数据库层面,MySQL 的 `utf8_general_ci` 排序规则是大小写不敏感的,而 `utf8_bin` 是敏感的,同一个 `WHERE name = 'Alice'` 在两种规则下结果不同。更坑的是,`utf8_general_ci` 还会把 `ß` 和 `ss` 视为相等,把 `ä` 和 `a` 视为相等——做唯一索引时可能出现意料之外的冲突。PostgreSQL 默认大小写敏感,需要 `citext` 扩展或 `ILIKE` 才能不敏感,迁移时很容易踩。Case issues pop up everywhere in engineering. At the database level, MySQL's `utf8_general_ci` collation is case-insensitive while `utf8_bin` is case-sensitive — the same `WHERE name = 'Alice'` returns different results. Worse, `utf8_general_ci` also treats `ß` as equal to `ss` and `ä` as equal to `a`, which can cause unexpected conflicts on unique indexes. PostgreSQL is case-sensitive by default and needs the `citext` extension or `ILIKE` for insensitive matching — easy to trip over during migration.
URL 的域名部分(host)按规范是大小写不敏感的,应该用 Punycode 转写后统一比较;但路径和查询参数是大小写敏感的,`/User` 和 `/user` 是两个不同资源。很多框架的路由匹配默认大小写敏感,也有一些(如早期的 ASP.NET)不敏感,跨框架对接时要特别注意。HTTP 头字段名按 RFC 7230 是大小写不敏感的,但头字段值通常敏感——`Content-Type` 写成 `content-type` 没问题,但 `Authorization` 的值里 Bearer token 大小写必须原样保留。URL hostnames are case-insensitive by spec and should be compared after Punycode normalisation, but paths and query parameters are case-sensitive: `/User` and `/user` are different resources. Many frameworks match routes case-sensitively by default, while some (like early ASP.NET) do not — pay attention when integrating across frameworks. HTTP header field names are case-insensitive per RFC 7230, but values are usually sensitive: writing `Content-Type` as `content-type` is fine, but the Bearer token inside `Authorization` must be preserved exactly.
我还踩过一个隐蔽的坑:用 `HashMap<String, Object>` 做不区分大小写的 header 容器,直接 `put(header.toLowerCase(), value)`,在土耳其 locale 的服务器上 `Content-Type` 变成了 `content-type` 没问题,但 `Set-Cookie` 里的某些值处理出了问题。后来改用 `TreeMap` 带 `String.CASE_INSENSITIVE_ORDER` 比较器,才彻底解决。I once hit a subtle bug: using a `HashMap<String, Object>` as a case-insensitive header container with `put(header.toLowerCase(), value)`. On a Turkish-locale server, `Content-Type` became `content-type` without issue, but processing certain `Set-Cookie` values broke. Switching to `TreeMap` with `String.CASE_INSENSITIVE_ORDER` fixed it for good.
安全地做大小写转换Doing case conversion safely
总结几条我自己的实践准则。第一,凡是涉及标识符、协议字段、代码符号的大小写转换,永远显式指定 locale:Java 用 `Locale.ROOT` 或 `Locale.ENGLISH`,JavaScript 用 `toLowerCase()`(JS 不区分 locale,但若需要 locale 感知用 `toLocaleLowerCase('tr')`),C# 用 `ToLowerInvariant()`。第二,不要假设转换前后长度不变——涉及 `ß`、组合字符时长度会变,做字符串截断或固定长度字段时要注意。第三,大小写不敏感的比较优先用 case folding 或语言内置的不敏感比较,不要自己转大小写后 `equals`。Here are my practical rules. First, whenever case conversion touches identifiers, protocol fields or code symbols, always specify the locale explicitly: Java uses `Locale.ROOT` or `Locale.ENGLISH`, JavaScript uses `toLowerCase()` (JS is locale-independent by default; use `toLocaleLowerCase('tr')` when you need locale awareness), C# uses `ToLowerInvariant()`. Second, never assume length is preserved — `ß` and combining characters change length, so be careful with truncation or fixed-width fields. Third, for case-insensitive comparison, prefer case folding or the language's built-in insensitive comparison rather than converting case and calling `equals`.
第四,数据库选型时就确认排序规则,不要等上线后再改——改 collation 可能需要重建索引,代价很大。第五,用户可见的文本(如姓名、标题)不要强制转大小写存储,保留原始输入,只在显示时按需格式化。Fourth, decide the collation at database design time, not after launch — changing collation may require rebuilding indexes, which is expensive. Fifth, never force user-visible text (names, titles) to a specific case for storage; keep the original input and format only at display time.
日常开发里,我经常需要把一段文本批量转成大写、小写、驼峰或标题格,手写正则容易在边界情况出错。本站的大小写转换工具支持大写、小写、首字母大写、标题格、驼峰/下划线互转等多种模式,并且默认使用 Unicode 规则处理多语言字符,适合快速验证转换结果或批量处理配置文本。In daily work I often need to batch-convert text to uppercase, lowercase, camelCase or title case, and hand-written regexes tend to fail on edge cases. Our case converter supports uppercase, lowercase, sentence case, title case, and camelCase/snake_case conversion, using Unicode rules for multilingual characters by default — handy for quickly verifying results or batch-processing config text.