MD5/SHA1/SHA256 怎么选:哈希碰撞、加盐与密码存储的正确姿势Choosing Between MD5, SHA1 and SHA256: Collisions, Salting and the Right Way to Store Passwords
哈希不是加密,MD5 也不是"过时了就不能用"——算法选型取决于场景。本文从碰撞攻击的真实数据出发,讲清通用哈希与密码哈希的本质区别,以及加盐、bcrypt、Argon2 各自的适用边界Hashing is not encryption, and MD5 is not "obsolete so never use it" — algorithm choice depends on the scenario. Starting from real collision-attack data, this article explains the fundamental difference between general-purpose hashes and password hashes, and where salting, bcrypt and Argon2 each fit.
哈希算法的本质:单向、定长、雪崩The essence of hashing: one-way, fixed-length, avalanche
哈希函数把任意长度的输入映射成固定长度的输出,三个核心特性决定了它的用途:单向性(从输出无法反推输入)、定长性(无论输入 1 字节还是 1GB,输出长度固定)、雪崩效应(输入改 1 位,输出约一半的位发生变化)。MD5 输出 128 位(32 个十六进制字符),SHA1 输出 160 位,SHA256 输出 256 位。A hash function maps input of any length to a fixed-length output, and three core properties define its use: one-wayness (you cannot reverse the output to get the input), fixed length (output is the same size whether the input is 1 byte or 1GB), and the avalanche effect (changing 1 bit of input flips roughly half the output bits). MD5 outputs 128 bits (32 hex characters), SHA1 outputs 160 bits, and SHA256 outputs 256 bits.
我在做文件校验时经常用哈希:下载一个大文件后算一下 SHA256,和官方提供的值对比,就能确认文件没有被篡改或损坏。这里选 SHA256 不是因为 MD5"不安全",而是因为 MD5 的输出空间只有 128 位,在海量文件场景下碰撞概率更高,而 SHA256 的 256 位输出空间在可预见的未来都足够安全。但如果只是做内部缓存 key 或去重指纹,MD5 完全够用——它速度快、输出短,碰撞风险在这些场景下不构成实际威胁。算法选型的第一步永远是问:我在防什么?I use hashes constantly for file integrity checks: after downloading a large file, compute its SHA256 and compare it against the official value to confirm the file was not tampered with or corrupted. Choosing SHA256 here is not because MD5 is "insecure", but because MD5's output space is only 128 bits — collision probability is higher at massive file scales — whereas SHA256's 256-bit space is secure for the foreseeable future. But for internal cache keys or deduplication fingerprints, MD5 is perfectly fine: it is fast, its output is short, and collision risk is not a practical threat in those scenarios. The first step in algorithm selection is always: what am I defending against?
碰撞攻击的真实数据:MD5 和 SHA1 是怎么被攻破的Real collision data: how MD5 and SHA1 were broken
碰撞(collision)是指找到两个不同的输入产生相同的哈希输出。理论上,输出 n 位的哈希函数,找到碰撞需要约 2^(n/2) 次运算(生日悖论)。MD5 的 128 位意味着理论上需要 2^64 次运算,但实际攻击远比这高效——2004 年王小云团队公布了 MD5 的碰撞攻击方法,2008 年研究者用碰撞构造了两个不同但哈希相同的 X.509 证书,2012 年火焰病毒(Flame)利用 MD5 碰撞伪造了微软的代码签名证书。A collision means finding two different inputs that produce the same hash output. Theoretically, an n-bit hash requires about 2^(n/2) operations to find a collision (birthday paradox). MD5's 128 bits imply 2^64 operations in theory, but practical attacks are far more efficient — in 2004 Wang Xiaoyun's team published a collision attack on MD5, in 2008 researchers used a collision to construct two different X.509 certificates with the same hash, and in 2012 the Flame malware used an MD5 collision to forge a Microsoft code-signing certificate.
SHA1 的命运类似:2017 年 Google 公布了 SHAttered 攻击,用 2^63 次 SHA1 计算(约 6500 年单 CPU 时间,但分布式计算仅需约 2 个月、成本约 1.1 万美元)构造了两个不同 PDF 的 SHA1 碰撞。2020 年进一步把 chosen-prefix 碰撞的成本降到了几千美元级别。这意味着在数字签名、证书、软件分发这些场景下,MD5 和 SHA1 已经不可信任——攻击者可以构造"看起来合法但内容不同"的文件。但注意,这些都是"碰撞攻击"(找两个不同输入相同输出),不是"原像攻击"(给定输出反推输入)。对于校验文件是否意外损坏,MD5 的原像抗性依然足够;对于防恶意篡改,必须用 SHA256 及以上。SHA1 followed a similar path: in 2017 Google published the SHAttered attack, using 2^63 SHA1 computations (about 6,500 years of single-CPU time, but roughly two months of distributed computing at a cost of about $11,000) to construct a SHA1 collision between two different PDFs. In 2020, chosen-prefix collisions were further reduced to the few-thousand-dollar range. This means that in digital signatures, certificates and software distribution, MD5 and SHA1 can no longer be trusted — an attacker can construct "legitimate-looking but different-content" files. Note, however, that these are collision attacks (finding two inputs with the same output), not preimage attacks (reversing an output to find its input). For checking whether a file was accidentally corrupted, MD5's preimage resistance is still sufficient; for defending against malicious tampering, you must use SHA256 or above.
密码存储为什么不能用 SHA256:通用哈希 vs 密码哈希Why SHA256 is wrong for passwords: general-purpose vs password hashing
很多人知道"密码不能明文存",于是用 MD5 或 SHA256 哈希后存库——这依然是错的。通用哈希函数(MD5、SHA1、SHA256)的设计目标是"快":现代 GPU 每秒可以算数十亿次 SHA256。如果数据库泄露,攻击者可以用彩虹表或 GPU 暴力破解,每秒尝试几十亿个密码,常见密码几小时内就能还原。Many people know "passwords must not be stored in plaintext", so they hash them with MD5 or SHA256 and store the result — this is still wrong. General-purpose hash functions (MD5, SHA1, SHA256) are designed to be fast: a modern GPU can compute billions of SHA256 hashes per second. If the database leaks, an attacker can use rainbow tables or GPU brute force, trying billions of passwords per second and recovering common ones within hours.
密码哈希的设计目标恰恰相反——"慢"。bcrypt、scrypt、Argon2 这些算法故意引入计算成本(迭代次数)和内存成本,让每次哈希计算需要几十到几百毫秒。这样攻击者每秒只能试几千次而不是几十亿次,暴力破解的时间从"几小时"变成"几百年"。我在一个老项目里见过用 SHA256 存密码的,数据库泄露后用户密码被批量破解,迁移到 bcrypt 时还遇到了一个坑:bcrypt 有 72 字节的输入上限,超过部分会被静默截断,所以超长密码需要先做一次 SHA256 预处理再喂给 bcrypt。Password hashes are designed for the opposite goal — slowness. bcrypt, scrypt and Argon2 deliberately introduce computational cost (iterations) and memory cost, making each hash take tens to hundreds of milliseconds. This means an attacker can try only thousands of passwords per second instead of billions, turning brute-force time from "hours" into "centuries". I once worked on a legacy project that stored passwords with SHA256; after a database leak, user passwords were cracked in bulk, and when migrating to bcrypt we hit another pitfall — bcrypt has a 72-byte input limit, and anything beyond is silently truncated, so very long passwords need a SHA256 pre-hash before being fed to bcrypt.
加盐(salt)是另一个关键:每个用户的密码哈希前拼一个随机盐值再哈希,这样即使两个用户密码相同,哈希结果也不同,彩虹表失效。盐不需要保密,存在数据库里和哈希一起就行,但必须每个用户独立生成、长度至少 16 字节。Argon2 是目前的推荐选择(2015 年密码哈希竞赛冠军),它同时支持可调的时间成本、内存成本和并行度,能有效对抗 GPU 和 ASIC 攻击。bcrypt 因为生态成熟、实现简单,在很多场景下依然是合理选择。Salting is another key measure: prepend a random salt to each user's password before hashing, so even if two users have the same password, their hashes differ and rainbow tables become useless. Salts do not need to be secret — store them in the database alongside the hash — but they must be generated independently per user and be at least 16 bytes long. Argon2 is the current recommended choice (winner of the 2015 Password Hashing Competition); it supports tunable time cost, memory cost and parallelism, effectively resisting GPU and ASIC attacks. bcrypt remains a reasonable choice in many scenarios due to its mature ecosystem and simple implementation.
选型决策树:场景决定算法A decision tree: the scenario decides the algorithm
把上面的内容整理成一个简单的决策树:文件完整性校验(防意外损坏)→ MD5 或 SHA256 都可以;防恶意篡改/数字签名 → SHA256 起步,重要场景用 SHA3;密码存储 → 绝对不要用通用哈希,用 bcrypt 或 Argon2,必须加盐;API 签名/HMAC → HMAC-SHA256(注意 HMAC 的密钥和消息分开处理,不要自己拼字符串再哈希);区块链/工作量证明 → SHA256(需要快且确定性强)。Putting it all together into a simple decision tree: file integrity check (accidental corruption) — MD5 or SHA256 both work; malicious tampering prevention / digital signatures — SHA256 at minimum, SHA3 for critical scenarios; password storage — never use a general-purpose hash, use bcrypt or Argon2 with a salt; API signing / HMAC — HMAC-SHA256 (keep the key and message separate, do not concatenate strings and hash them yourself); blockchain / proof-of-work — SHA256 (needs speed and determinism).
我在实际项目中还踩过一个坑:不同语言的哈希库默认输出格式不一样。Java 的 `MessageDigest` 返回字节数组,需要手动转十六进制;Python 的 `hashlib.hexdigest()` 直接返回字符串;Node.js 的 `crypto.createHash().digest('hex')` 也是字符串。如果前后端对同一个字符串算哈希但结果对不上,先检查是不是编码问题(UTF-8 vs GBK)和输出格式问题(字节 vs 十六进制 vs Base64),而不是怀疑算法实现。本站的 Hash 生成器支持 MD5、SHA1、SHA256、SHA512 等多种算法,输出十六进制和 Base64 两种格式,可以用来快速验证不同语言实现的哈希结果是否一致。I also hit a pitfall in real projects: hash libraries in different languages have different default output formats. Java's `MessageDigest` returns a byte array that you must manually convert to hex; Python's `hashlib.hexdigest()` returns a string directly; Node.js's `crypto.createHash().digest('hex')` also returns a string. If frontend and backend hash the same string but get different results, check encoding (UTF-8 vs GBK) and output format (bytes vs hex vs Base64) before suspecting the algorithm implementation. Our Hash generator supports MD5, SHA1, SHA256, SHA512 and more, with both hex and Base64 output — useful for quickly verifying whether hash results match across different language implementations.