F

文本清洗流水线:去重、去空行、去 BOM 与批量处理的经验总结A Text Cleaning Pipeline: Deduplication, Blank-Line Removal, BOM Stripping and Batch Processing Lessons

拿到一份"看起来没问题"的文本数据,跑进系统却乱码、重复、空行满天飞——文本清洗是数据分析和迁移中最容易被低估的环节。本文总结我在多个项目中沉淀的清洗流水线:从 BOM 到行尾,从去重到批量处理,以及每一步最容易踩的坑You receive a text dataset that "looks fine", feed it into the system, and get mojibake, duplicates and blank lines everywhere — text cleaning is the most underrated step in data analysis and migration. This article distils the cleaning pipeline I have built across multiple projects: from BOM to line endings, from deduplication to batch processing, and the pitfalls at every step.

第一步:先看字节,不要相信眼睛Step one: look at the bytes, do not trust your eyes

文本清洗的第一原则是:眼睛看到的"干净"不等于字节层面的干净。一个最典型的例子是 BOM(Byte Order Mark)——UTF-8 文件开头的 `EF BB BF` 三个字节。大多数编辑器会自动隐藏 BOM,你打开文件看到的第一行就是正常内容,但程序读取时,BOM 会被当作第一个字符的一部分,导致 JSON 解析失败、CSV 第一列表名带不可见字符、字符串比较永远不相等。我曾经花了一下午排查"为什么这个 JSON 总是 parse 失败",最后用 `xxd` 看文件头才发现是 BOM 搞的鬼。The first principle of text cleaning is: what looks "clean" to your eyes is not necessarily clean at the byte level. A classic example is the BOM (Byte Order Mark) — the three bytes `EF BB BF` at the start of a UTF-8 file. Most editors hide the BOM automatically, so when you open the file the first line looks normal, but when a program reads it, the BOM becomes part of the first character, causing JSON parse failures, CSV header columns with invisible characters, and string comparisons that never match. I once spent an afternoon debugging "why does this JSON always fail to parse", only to discover the BOM after checking the file header with `xxd`.

除了 BOM,还有几种不可见字符需要警惕:行尾的 `\r\n`(Windows)vs `\n`(Unix),全角空格(U+3000)和不间断空格(U+00A0)混在普通空格中间,零宽空格(U+200B)从富文本或网页复制时带进来。我的习惯是拿到任何文本文件,先用 `file` 命令看编码,再用 `xxd | head` 看前几个字节,确认有没有 BOM 和异常字符,然后才开始后续清洗。Besides the BOM, several other invisible characters need attention: line endings `\r\n` (Windows) vs `\n` (Unix) vs `\r` (old Mac), full-width spaces (U+3000) and non-breaking spaces (U+00A0) mixed in with regular spaces, and zero-width spaces (U+200B) and right-to-left marks (U+200E) carried in from rich text or web copy. My habit with any text file is to first run `file` to check the encoding, then `xxd | head` to inspect the first few bytes for BOM and abnormal characters, before starting any further cleaning.

去重与去空行:顺序、大小写与空白的陷阱Deduplication and blank-line removal: order, case and whitespace traps

去重看起来就是"保留唯一行",但实际操作中有三个细节决定结果。第一是是否保留原顺序:用 `Set` 去重会打乱行的顺序,如果后续逻辑依赖行序(比如配置文件、翻译文件),就必须用"有序去重"——遍历每一行,用哈希集合记录已出现的行,只保留第一次出现的。第二是是否忽略大小写和首尾空白:"Hello"和"hello"算不算重复?" data "和"data"算不算重复?这些必须在去重前明确规则。Deduplication seems like just "keep unique lines", but three details determine the result in practice. First, whether to preserve original order: using a `Set` shuffles line order, and if downstream logic depends on it (config files, translation files), you need ordered deduplication — iterate each line, use a hash set to track seen lines, and keep only the first occurrence. Second, whether to ignore case and leading/trailing whitespace: are "Hello" and "hello" duplicates? Are " data " and "data" duplicates? These rules must be decided before deduplication, or you will finish and wonder "why are there still duplicates".

第三是空行的定义:是完全空的行(长度为 0),还是只含空白字符的行(空格、Tab)也算空行?我在处理一份导出的用户评论时,一开始只去掉了完全空行,结果文件里还剩大量"只有一个空格"的行,统计行数时虚高。正确做法是先对每一行做 `trim`,如果结果为空就视为空行去掉。但要注意:如果文本本身有意义的前导空格(比如代码缩进),就不能全局 trim,需要按行类型区分处理。去空行还有一个常见需求是"合并连续空行为一个"——在 Markdown 场景下,保留段落间的单个空行是有意义的,全部删掉反而破坏格式。Third, the definition of a blank line: is it a completely empty line (length 0), or does a line containing only whitespace (spaces, tabs) also count? I once processed an exported user-comment file and initially removed only completely empty lines, leaving many lines with "just one space" that inflated the line count. The correct approach is to `trim` each line first and treat it as blank if the result is empty. But be careful: if the text has meaningful leading whitespace (code indentation, Markdown blockquotes), you cannot globally trim — you need to handle it by line type. Another common blank-line requirement is "merge consecutive blank lines into one" — in Markdown and typography, preserving a single blank line between paragraphs is meaningful, and removing them all destroys formatting.

批量处理:大文件、编码不一致与幂等性Batch processing: large files, inconsistent encoding and idempotency

当清洗对象从一个文件变成几百个文件时,问题就变了。第一个挑战是编码不一致:从不同来源收集的文本文件,可能是 UTF-8、GBK、GB2312、Big5 甚至 Latin-1 混在一起。我的做法是先用 `chardet` 或 `uchardet` 批量检测每个文件的编码,记录下来,然后统一转成 UTF-8 再处理。不要假设"都是 UTF-8"——我曾经在一个 200 个文件的批量任务里,有 3 个文件是 GBK 编码,直接按 UTF-8 读取后中文全部乱码,而且因为乱码不报错,直到最后人工抽检才发现。When the cleaning target grows from one file to hundreds, the problem changes. The first challenge is inconsistent encoding: text files collected from different sources may be UTF-8, GBK, GB2312, Big5 or even Latin-1 mixed together. My approach is to first batch-detect each file's encoding with `chardet` or `uchardet`, record the results, then convert everything to UTF-8 before processing. Never assume "it is all UTF-8" — in one batch task of 200 files, three were GBK-encoded, and reading them as UTF-8 turned all Chinese into mojibake; because mojibake does not throw errors, it was only caught during a final manual spot-check.

第二个挑战是大文件内存。一个 500MB 的日志文件,用 `read().splitlines()` 全部读进内存可能直接 OOM。正确做法是逐行流式处理:打开文件后用迭代器一行一行读,处理完立即写入输出文件,内存占用恒定。去重时哈希集合记录所有已见行,大文件下内存仍会涨——可以用布隆过滤器做预筛选,或先按行哈希排序再相邻去重(外部排序),用磁盘换内存。The second challenge is large-file memory. A 500MB log file read entirely into memory with `read().splitlines()` may cause an OOM. The correct approach is streaming line-by-line processing: open the file, use an iterator to read one line at a time, write to the output immediately, and keep memory usage constant. For deduplication, a hash set tracking all seen lines will still grow with large files — in that case consider a Bloom Filter for "possibly duplicate" pre-screening, or sort by line hash first and then remove adjacent duplicates (external sort), trading disk for memory.

第三个挑战是幂等性:清洗脚本应该可以反复运行而不产生副作用。我见过的一个坑是"去 BOM"脚本每次运行都在文件头加一个 BOM 而不是去掉——逻辑写反了,跑了三遍之后文件头有三个 BOM。正确的清洗脚本应该先检测状态,只在需要时才修改,并且最好输出到新文件而不是原地覆盖,保留原始文件作为回滚。The third challenge is idempotency: a cleaning script should be runnable repeatedly without side effects. A pitfall I saw was a "remove BOM" script that added a BOM on every run instead of removing it — the logic was inverted, and after three runs the file had three BOMs. A correct cleaning script should first detect the state (BOM present? line ending type?), modify only when needed, and ideally write to a new file rather than overwriting in place, keeping the original as a rollback.

工具链与经验总结:把清洗变成可复用的流水线Toolchain and lessons: turning cleaning into a reusable pipeline

总结我沉淀下来的文本清洗流水线,固定顺序是:1)检测编码与 BOM,统一转 UTF-8 并去 BOM;2)统一行尾为 `\n`;3)去除不可见字符(零宽空格、全角空格按需替换);4)按规则去空行;5)有序去重(保留首次出现,按需忽略大小写和空白);6)输出到新文件并校验行数变化。这个顺序不能乱——比如先去重再去 BOM,BOM 行和非 BOM 行会被当成不同内容而无法去重。The text-cleaning pipeline I have settled on follows a fixed order: 1) detect encoding and BOM, convert to UTF-8 and strip BOM; 2) normalise line endings to `\n`; 3) remove invisible characters (zero-width spaces, BOM remnants, replace full-width spaces as needed); 4) remove blank lines by rule (completely empty or empty after trim); 5) ordered deduplication (keep first occurrence, optionally ignore case and whitespace); 6) write to a new file and verify the line-count change. This order matters — for example, deduplicating before stripping BOM means BOM-prefixed lines and non-BOM lines are treated as different and will not deduplicate.

在工具选择上,简单的单行任务用 `sed`、`awk`、`sort -u` 就够了,但涉及编码检测、BOM、多规则组合时,我更倾向于用一个 Python 脚本把所有步骤串起来,参数化配置,方便复用和审计。对于临时的小文件清洗,在线工具反而更快——把文本贴进去,勾选去重、去空行、去 BOM,立刻看到结果。本站的文本去重工具支持有序去重、忽略大小写、去除空行和首尾空白,适合快速处理中小规模的文本清洗需求;遇到大文件或复杂规则,再上脚本流水线。For tooling, simple one-off tasks are fine with `sed`, `awk` and `sort -u`, but when encoding detection, BOM handling and multiple combined rules are involved, I prefer a Python script that chains all steps with parameterised config, making it reusable and auditable. For quick small-file cleaning, an online tool is faster — paste the text, check deduplicate, remove blank lines, strip BOM, and see results instantly without writing a script. Our text dedup tool supports ordered deduplication, case-insensitive mode, blank-line removal and leading/trailing whitespace stripping, ideal for quickly handling small-to-medium text cleaning needs; for large files or complex rules, move to a scripted pipeline.

← 返回教程列表← Back to all guides