F

Base64 编码是什么?原理、用途与常见误区详解What Is Base64? Principles, Use Cases and Common Pitfalls

Base64 出现在 Data URL、接口鉴权、邮件附件等无数场景中。本文用通俗语言讲清 Base64 的工作原理、为什么体积会增大三分之一,以及处理中文与二进制时最容易踩的三个坑。Base64 shows up everywhere — Data URLs, API auth, email attachments. This article explains in plain terms how Base64 works, why it inflates size by a third, and the three pitfalls people hit most with Chinese text and binary data.

为什么需要“把二进制当文本传”Why carry binary through text channels

打开网页、发邮件、调接口,这些传输通道最初都是为“文本”设计的:JSON、XML、HTTP 头部、邮件正文只能安全承载可打印字符。但现实世界充满了二进制——图片、音频、压缩包——直接塞进文本通道会产生乱码甚至数据损坏。Web pages, emails and APIs were originally designed for text: JSON, XML, HTTP headers and mail bodies can only carry printable characters safely. Yet the real world is full of binary — images, audio, archives — and shoving raw bytes into text channels causes mojibake or corruption.

Base64 正是为此而生:它把每 3 个字节的二进制数据,映射为 4 个来自 64 个可打印字符(A–Z、a–z、0–9、+、/)的字符,不足处用 = 补齐。编码后的内容只含安全字符,因此可以穿过几乎任何文本通道而不被破坏。Base64 exists for exactly this: it maps every 3 bytes of binary onto 4 characters drawn from 64 printable symbols (A–Z, a–z, 0–9, +, /), padded with = when needed. The result contains only safe characters and can travel through almost any text channel intact.

体积增大三分之一的代价The one-third size tax

Base64 用 4 个字符表示 3 个字节,因此编码后体积约为原来的 4/3,即增大约 33%。这也是把大图直接内嵌成 Data URL 反而更慢的原因:省下的请求数抵不上膨胀的体积。Because Base64 represents 3 bytes with 4 characters, encoded data is about 4/3 of the original — a 33% inflation. That is why inlining large images as Data URLs can actually slow pages down: the saved request is outweighed by the extra bytes.

所以工程上的惯例是:小图标、短二进制片段用 Base64 内嵌很划算;大文件应作为独立资源传输并配合缓存。理解这个权衡,能帮你在性能评审中做出正确判断。The engineering rule of thumb is therefore: inline tiny icons and short binary snippets with Base64, but serve large files as separate cached resources. Understanding this trade-off leads to better performance decisions.

处理中文与二进制的三个坑Three pitfalls with Chinese and binary

第一个坑:把 Base64 当加密。Base64 可以被任何人无条件解码,不提供任何保密性——接口鉴权里的 Basic 头只是“身份声明”,不是加密。第二个坑:中文直接按字符解码会乱码,正确做法是先解成字节序列再按 UTF-8 解释。第三个坑:部分实现会按 MIME 规范每 76 个字符插入换行,拼接或比较前要先去掉空白。Pitfall one: treating Base64 as encryption. Anyone can decode it — the Basic auth header is an identity statement, not secrecy. Pitfall two: decoding Chinese text character-by-character produces mojibake; decode to bytes first, then interpret as UTF-8. Pitfall three: some implementations insert a line break every 76 characters per MIME, so strip whitespace before concatenating or comparing.

理解这三点,就能避开日常开发中绝大多数 Base64 相关 bug。本站的 Base64 工具按 UTF-8 正确处理中文与 Emoji,支持文本与文件互转,可以随时用来验证你的编码结果。With these three points in mind, most everyday Base64 bugs disappear. Our Base64 tool handles UTF-8 correctly — Chinese and Emoji included — for both text and files, so you can verify encodings any time.

← 返回教程列表← Back to all guides