F

正则表达式入门:核心概念与 8 个即用高频模式Regex for Beginners: Core Concepts and 8 Ready-to-Use Patterns

正则表达式是“学会一次、受益十年”的技能。本文用最短篇幅讲清字符类、量词、锚点三大核心,并给出邮箱、手机号、中文等 8 个高频模式与避坑建议。Regex is a "learn once, benefit for a decade" skill. This article covers the three cores — character classes, quantifiers, anchors — in minimal space, plus 8 high-frequency patterns and pitfall advice.

三分钟建立核心心智模型A three-minute mental model

正则的本质是“描述一类字符串的规则”。三个核心构件:字符类描述“这一个位置可以是什么”,如 \d(数字)、[a-z](小写字母)、[\u4e00-\u9fa5](中文);量词描述“重复几次”,如 *(0 次或多次)、+(1 次或多次)、{3,5}(3 到 5 次);锚点描述“位置”,如 ^(开头)、$(结尾)、\b(单词边界)。A regex is simply "a rule describing a class of strings". Three core building blocks: character classes say what may appear at one position — \d (digit), [a-z] (lowercase), [\u4e00-\u9fa5] (Chinese); quantifiers say how many times — * (0+), + (1+), {3,5} (3 to 5); anchors say where — ^ (start), $ (end), \b (word boundary).

把正则读成“位置 + 重复”的句子,例如 ^\d{3}-\d{4}$ 就是“开头 3 位数字、一个连字符、4 位数字、结尾”,神秘感立刻消失。Read a regex as a sentence of "position + repetition": ^\d{3}-\d{4}$ means "start, three digits, a hyphen, four digits, end" — and the mystery dissolves.

8 个高频模式直接抄Eight patterns you can copy today

① 邮箱(实用版):^[\w.+-]+@[\w-]+(\.[\w-]+)$;② 中国大陆手机号:^1[3-9]\d{9}$;③ 中文:[\u4e00-\u9fa5];④ 空白行:^\s*$;⑤ 连续空白:\s+;⑥ 16 进制颜色:^#([0-9a-fA-F]{3}|[0-9a-fA-F]{6})$;⑦ 身份证号(18 位粗校):^\d{17}[\dXx]$;⑧ 双引号字符串:"[^"]*"。① Email (practical): ^[\w.+-]+@[\w-]+(\.[\w-]+)$; ② Mainland-CN mobile: ^1[3-9]\d{9}$; ③ Chinese character: [\u4e00-\u9fa5]; ④ Blank line: ^\s*$; ⑤ Runs of whitespace: \s+; ⑥ Hex colour: ^#([0-9a-fA-F]{3}|[0-9a-fA-F]{6})$; ⑦ 18-digit ID (rough): ^\d{17}[\dXx]$; ⑧ Double-quoted string: "[^"]*".

注意:邮箱的“完全标准”正则长达数千字符且无必要,工程上“实用版 + 发送验证邮件”才是正解。所有模式建议在实时测试工具中用真实样本验证后再入库。Note: the "fully standard" email regex is thousands of characters long and unnecessary — in practice, a simple pattern plus a confirmation email is the right answer. Always verify patterns against real samples in a live tester before shipping.

避坑:贪婪、转义与双重转义Pitfalls: greediness, escaping, double escaping

贪婪量词会“能多匹配就多匹配”,提取 <div>.*</div> 时会从第一个 div 吞到最后一个;改成懒惰的 .*? 才符合直觉。特殊字符(. * + ? ( ) [ ] 等)要匹配字面意义时需反斜杠转义。Greedy quantifiers match as much as they can: <div>.*</div> swallows from the first div to the last; the lazy .*? behaves as intended. Special characters (. * + ? ( ) [ ] …) need backslash escaping to match literally.

最隐蔽的坑是代码里的双重转义:字符串中的 "\d" 要写成 "\\d",排查“代码里不生效、工具里生效”的问题时,先检查这一点。The sneakiest pitfall is double escaping in code: "\d" in a string literal must be written "\\d". When a pattern works in a tester but not in code, check this first.

← 返回教程列表← Back to all guides