登录
首页 >  文章 >  前端

HTML5中提取XML标签文本的注意事项

时间:2026-01-28 13:51:41 217浏览 收藏

一分耕耘,一分收获!既然打开了这篇文章《HTML5正则提取XML标签文本的注意事项》,就坚持看下去吧!文中内容包含等等知识点...希望你能在阅读本文后,能真真实实学到知识或者帮你解决心中的疑惑,也欢迎大佬或者新人朋友们多留言评论,多给建议!谢谢!

应使用 DOMParser 解析 XML 字符串为 XML Document 后用 querySelector 等提取;innerHTML 和正则易因 HTML 解析规则、命名空间、CDATA 等导致不可靠结果。

html5用正则匹配xml内容_快速提取特定标签文本的注意点【说明】

XML 不是 HTML,别用 innerHTML 或正则硬刚结构

HTML5 环境下直接对 XML 字符串用正则提取标签内容,本质是拿非结构化工具处理半结构化数据。浏览器原生不解析 XML 字符串为 DOM,innerHTML 会按 HTML 规则解析(比如自闭合标签被补全、命名空间丢失、 变成 ),结果不可靠。

真正安全的做法是用 DOMParser 解析 XML 字符串为 XML Document,再用 querySelectorgetElementsByTagName 提取。正则只适合极简、已知格式、无嵌套、无命名空间的临时场景。

如果非要用正则,必须避开这三类坑

  • XML 允许换行、缩进、注释、CDATA、处理指令(如 ),正则很难全覆盖;
  • 标签可能带属性、命名空间前缀(如 )、大小写混用(XML 区分大小写);
  • 内容本身含 <>(如嵌套 CDATA 或转义字符 <),会导致正则提前截断或误匹配。

例如想提取 ...></code> 文本,写成 <code>/<title>(.*?)<\/title>/s</code> 在多数简单 XML 中看似能用,但一旦遇到:</p> <pre><title>A &lt; B</title></pre> <p>就会因未处理实体而漏掉内容;若 XML 含 <code><title xmlns="http://example.com"></code>,该正则直接失效。</p> <h3><code>DOMParser</code> 解析 XML 的最小可靠写法</h3> <p>这是 HTML5 中最接近“开箱即用”的方案,兼容所有现代浏览器,且自动处理命名空间、CDATA、实体等。</p> <p>关键点:</p> <ul><li>必须指定 <code>"application/xml"</code> 或 <code>"text/xml"</code> 类型,否则解析失败;</li> <li>解析失败时 <code>parseFromString</code> 返回空文档,需检查 <code>documentElement.nodeName === "parsererror"</code>;</li> <li>用 <code>textContent</code> 而非 <code>innerHTML</code> 获取文本,避免二次 HTML 解析污染。</li> </ul><pre>const xmlStr = `<?xml version="1.0"?><root><title>Hello &amp; World</title></root>`; const parser = new DOMParser(); const xmlDoc = parser.parseFromString(xmlStr, "application/xml"); if (xmlDoc.documentElement.nodeName === "parsererror") { console.error("XML parse error:", xmlDoc.documentElement.textContent); } else { const titleEl = xmlDoc.querySelector("title"); const text = titleEl ? titleEl.textContent : null; // → "Hello & World" }</pre> <h3>命名空间和多级嵌套时的 <code>querySelector</code> 写法</h3> <p>XML 常含命名空间(如 <code>xmlns:dc="http://purl.org/dc/elements/1.1/"</code>),此时不能直接写 <code>dc\\:title</code> —— 浏览器不支持命名空间前缀的 CSS 选择器(除非用 <code>getElementsByTagNameNS</code>)。</p> <p>稳妥做法:</p> <ul><li>忽略命名空间:用通配符 <code>*|title</code>(部分浏览器支持,但非标准);</li> <li>明确指定命名空间 URI:用 <code>getElementsByTagNameNS("http://purl.org/dc/elements/1.1/", "title")</code>;</li> <li>若只需取第一个匹配项,可遍历 <code>getElementsByTagName("title")</code> 并检查 <code>namespaceURI</code> 属性。</li> </ul><p>例如提取 RSS 中的 <code><dc:creator></code>:</p> <pre>const creators = xmlDoc.getElementsByTagNameNS("http://purl.org/dc/elements/1.1/", "creator"); const firstCreator = creators.length ? creators[0].textContent : null;</pre> <p>注意:命名空间 URI 必须一字不差,包括末尾斜杠,否则匹配失败。</p> XML 解析的边界往往不在语法,而在你是否意识到那个看似普通的 <code><item></code> 标签其实裹着 <code>xmlns=""</code> 或藏在 <code><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"></code> 里。正则能省事,但省下的时间常被调试命名空间和实体转义吃掉。<p>好了,本文到此结束,带大家了解了《HTML5中提取XML标签文本的注意事项》,希望本文对你有所帮助!关注golang学习网公众号,给大家分享更多文章知识!</p> <div style="margin:16px auto;width:100%;max-width:720px;box-sizing:border-box;padding:16px;border:1px solid #e5e7eb;border-radius:12px;background:#fff;box-shadow:0 6px 24px rgba(16,24,40,0.08);text-align:center;overflow:hidden;"> <a onclick="showThirdParty('flex')" style="display:inline-flex;width:100%;max-width:100%;justify-content:center;align-items:center;gap:8px;padding:12px 18px;border-radius:10px;background:#2d8cf0;color:#fff;text-decoration:none;font-weight:600;box-sizing:border-box;overflow-wrap:anywhere;word-break:break-word;"> 前往漫画官网入口并下载 ➜ </a> </div> <div id="third-party-overlay" style="position:fixed;left:0;top:0;width:100%;height:100%;display:none;justify-content:center;align-items:center;background:rgba(0,0,0,0.4);z-index:9999;"> <div style="background:#FFF3CD;border:1px solid #FFEEBA;padding:16px;border-radius:6px;box-sizing:border-box;max-width:480px;width:90%;text-align:center;"> <div style="font-size:14px;color:#856404;margin-bottom:12px;">您即将跳转至第三方网站,请注意保护好个人信息和财产安全!</div> <a href="https://comicdow.pdlcomic.top/1273%2F%E5%9B%A7%E6%AC%A1%E5%85%83.apk" target="_blank" rel="nofollow noopener noreferrer" style="color:#2d8cf0;text-decoration:none;" onclick="showThirdParty('none');">继续访问</a> </div> </div> <script> function showThirdParty(mode){ var el = document.getElementById('third-party-overlay'); if (!el) return; el.style.display = (mode === 'none' ? 'none' : 'flex'); } </script> </div> <div class="labsList"> </div> </div> <!-- 最新阅读 --> <div class="contBoxNor"> <div class="contTit"> <div class="tit">相关阅读</div> <a href="/articlelist.html" class="more">更多></a> </div> <ul class="latestReadList"> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  8个月前  |   <a href="javascript:;" class="aLightGray" title="提升">提升</a> <a href="javascript:;" class="aLightGray" title="箭头函数">箭头函数</a> <a href="javascript:;" class="aLightGray" title="函数表达式">函数表达式</a> <a href="javascript:;" class="aLightGray" title="函数声明">函数声明</a> <a href="javascript:;" class="aLightGray" title="Function构造函数">Function构造函数</a> </div> <div class="tit lineOverflow"><a href="/article/207000.html" title="JavaScript函数定义及示例详解" class="aBlack">JavaScript函数定义及示例详解</a></div> <div class="opt"> <span><i class="view"></i>502</span> <span class="collectBtn user_collection" data-id="207000" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  2年前  |   <a href="javascript:;" class="aLightGray" title="CSS">CSS</a> <a href="javascript:;" class="aLightGray" title="优化">优化</a> <a href="javascript:;" class="aLightGray" title="体验">体验</a> </div> <div class="tit lineOverflow"><a href="/article/72840.html" title="优化用户界面体验的秘密武器:CSS开发项目经验大揭秘" class="aBlack">优化用户界面体验的秘密武器:CSS开发项目经验大揭秘</a></div> <div class="opt"> <span><i class="view"></i>501</span> <span class="collectBtn user_collection" data-id="72840" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  2年前  |   <a href="javascript:;" class="aLightGray" title="图片轮播">图片轮播</a> <a href="javascript:;" class="aLightGray" title="微信小程序">微信小程序</a> <a href="javascript:;" class="aLightGray" title="特效">特效</a> </div> <div class="tit lineOverflow"><a href="/article/76259.html" title="使用微信小程序实现图片轮播特效" class="aBlack">使用微信小程序实现图片轮播特效</a></div> <div class="opt"> <span><i class="view"></i>501</span> <span class="collectBtn user_collection" data-id="76259" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  2年前  |   <a href="javascript:;" class="aLightGray" title="sessionStorage">sessionStorage</a> <a href="javascript:;" class="aLightGray" title="存储能力">存储能力</a> <a href="javascript:;" class="aLightGray" title="限制解析">限制解析</a> </div> <div class="tit lineOverflow"><a href="/article/83771.html" title="解析sessionStorage的存储能力与限制" class="aBlack">解析sessionStorage的存储能力与限制</a></div> <div class="opt"> <span><i class="view"></i>501</span> <span class="collectBtn user_collection" data-id="83771" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  2年前  |   <a href="javascript:;" class="aLightGray" title="团队合作">团队合作</a> <a href="javascript:;" class="aLightGray" title="冒泡事件">冒泡事件</a> <a href="javascript:;" class="aLightGray" title="促进作用">促进作用</a> </div> <div class="tit lineOverflow"><a href="/article/85057.html" title="探索冒泡活动对于团队合作的推动力" class="aBlack">探索冒泡活动对于团队合作的推动力</a></div> <div class="opt"> <span><i class="view"></i>501</span> <span class="collectBtn user_collection" data-id="85057" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> </ul> </div> <!-- 最新阅读 --> <div class="contBoxNor"> <div class="contTit"> <div class="tit">最新阅读</div> <a href="/articlelist.html" class="more">更多></a> </div> <ul class="latestReadList"> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  3分钟前  |   <a href="javascript:;" class="aLightGray" title="HTML与前端技术结合">HTML与前端技术结合</a> </div> <div class="tit lineOverflow"><a href="/article/475425.html" title="HTMLCSS布局教程详解" class="aBlack">HTMLCSS布局教程详解</a></div> <div class="opt"> <span><i class="view"></i>231</span> <span class="collectBtn user_collection" data-id="475425" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  6分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/475422.html" title="动态规划解决经典背包问题全解析" class="aBlack">动态规划解决经典背包问题全解析</a></div> <div class="opt"> <span><i class="view"></i>374</span> <span class="collectBtn user_collection" data-id="475422" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  8分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/475419.html" title="CSStransition未生效?检查初始状态写法" class="aBlack">CSStransition未生效?检查初始状态写法</a></div> <div class="opt"> <span><i class="view"></i>190</span> <span class="collectBtn user_collection" data-id="475419" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  9分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/475417.html" title="CSS徽章颜色统一技巧" class="aBlack">CSS徽章颜色统一技巧</a></div> <div class="opt"> <span><i class="view"></i>103</span> <span class="collectBtn user_collection" data-id="475417" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  9分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/475415.html" title="响应式列布局,自动调整显示列数" class="aBlack">响应式列布局,自动调整显示列数</a></div> <div class="opt"> <span><i class="view"></i>433</span> <span class="collectBtn user_collection" data-id="475415" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  13分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/475412.html" title="HTML5XML解析优化技巧分享" class="aBlack">HTML5XML解析优化技巧分享</a></div> <div class="opt"> <span><i class="view"></i>373</span> <span class="collectBtn user_collection" data-id="475412" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  13分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/475411.html" title="Map与Set在JS中的使用详解" class="aBlack">Map与Set在JS中的使用详解</a></div> <div class="opt"> <span><i class="view"></i>368</span> <span class="collectBtn user_collection" data-id="475411" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  18分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/475403.html" title="Angular高效数据缓存服务技巧" class="aBlack">Angular高效数据缓存服务技巧</a></div> <div class="opt"> <span><i class="view"></i>296</span> <span class="collectBtn user_collection" data-id="475403" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  21分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/475398.html" title="高饱和度HSL色提升CSS强调效果" class="aBlack">高饱和度HSL色提升CSS强调效果</a></div> <div class="opt"> <span><i class="view"></i>432</span> <span class="collectBtn user_collection" data-id="475398" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  22分钟前  |   <a href="javascript:;" class="aLightGray" title="数组遍历">数组遍历</a> </div> <div class="tit lineOverflow"><a href="/article/475397.html" title="JS数组map方法使用教程" class="aBlack">JS数组map方法使用教程</a></div> <div class="opt"> <span><i class="view"></i>442</span> <span class="collectBtn user_collection" data-id="475397" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  27分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/475390.html" title="HTML5视频播放器安装与浏览器设置教程" class="aBlack">HTML5视频播放器安装与浏览器设置教程</a></div> <div class="opt"> <span><i class="view"></i>476</span> <span class="collectBtn user_collection" data-id="475390" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/88_new_0_1.html" class="aLightGray" title="前端">前端</a>   |  30分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/475388.html" title="CSS宽高不生效怎么办?详解定位元素问题" class="aBlack">CSS宽高不生效怎么办?详解定位元素问题</a></div> <div class="opt"> <span><i class="view"></i>411</span> <span class="collectBtn user_collection" data-id="475388" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> </ul> </div> <!-- 课程推荐 --> <div class="contBoxNor"> <div class="contTit"> <div class="tit">课程推荐</div> <a href="/courselist.html" class="more">更多></a> </div> <ul class="classRecomList"> <li> <a href="/course/9.html" title="前端进阶之JavaScript设计模式" class="img_box"> <img src="/uploads/20221222/52fd0f23a454c71029c2c72d206ed815.jpg" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="前端进阶之JavaScript设计模式"> </a> <dl> <dt class="lineOverflow"> 前端进阶之JavaScript设计模式 </dt> <dd class="cont1 lineOverflow">设计模式是开发人员在软件开发过程中面临一般问题时的解决方案,代表了最佳的实践。本课程的主打内容包括JS常见设计模式以及具体应用场景,打造一站式知识长龙服务,适合有JS基础的同学学习。</dd> <dd class="cont2"> <a href="/course/9.html" title="前端进阶之JavaScript设计模式" class="toStudy">立即学习</a> <span>543次学习</span> </dd> </dl> </li> <li> <a href="/course/2.html" title="GO语言核心编程课程" class="img_box"> <img src="/uploads/20221221/634ad7404159bfefc6a54a564d437b5f.png" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="GO语言核心编程课程"> </a> <dl> <dt class="lineOverflow"> GO语言核心编程课程 </dt> <dd class="cont1 lineOverflow">本课程采用真实案例,全面具体可落地,从理论到实践,一步一步将GO核心编程技术、编程思想、底层实现融会贯通,使学习者贴近时代脉搏,做IT互联网时代的弄潮儿。</dd> <dd class="cont2"> <a href="/course/2.html" title="GO语言核心编程课程" class="toStudy">立即学习</a> <span>516次学习</span> </dd> </dl> </li> <li> <a href="/course/74.html" title="简单聊聊mysql8与网络通信" class="img_box"> <img src="/uploads/20240103/bad35fe14edbd214bee16f88343ac57c.png" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="简单聊聊mysql8与网络通信"> </a> <dl> <dt class="lineOverflow"> 简单聊聊mysql8与网络通信 </dt> <dd class="cont1 lineOverflow">如有问题加微信:Le-studyg;在课程中,我们将首先介绍MySQL8的新特性,包括性能优化、安全增强、新数据类型等,帮助学生快速熟悉MySQL8的最新功能。接着,我们将深入解析MySQL的网络通信机制,包括协议、连接管理、数据传输等,让</dd> <dd class="cont2"> <a href="/course/74.html" title="简单聊聊mysql8与网络通信" class="toStudy">立即学习</a> <span>500次学习</span> </dd> </dl> </li> <li> <a href="/course/57.html" title="JavaScript正则表达式基础与实战" class="img_box"> <img src="/uploads/20221226/bbe4083bb3cb0dd135fb02c31c3785fb.jpg" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="JavaScript正则表达式基础与实战"> </a> <dl> <dt class="lineOverflow"> JavaScript正则表达式基础与实战 </dt> <dd class="cont1 lineOverflow">在任何一门编程语言中,正则表达式,都是一项重要的知识,它提供了高效的字符串匹配与捕获机制,可以极大的简化程序设计。</dd> <dd class="cont2"> <a href="/course/57.html" title="JavaScript正则表达式基础与实战" class="toStudy">立即学习</a> <span>487次学习</span> </dd> </dl> </li> <li> <a href="/course/28.html" title="从零制作响应式网站—Grid布局" class="img_box"> <img src="/uploads/20221223/ac110f88206daeab6c0cf38ebf5fe9ed.jpg" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="从零制作响应式网站—Grid布局"> </a> <dl> <dt class="lineOverflow"> 从零制作响应式网站—Grid布局 </dt> <dd class="cont1 lineOverflow">本系列教程将展示从零制作一个假想的网络科技公司官网,分为导航,轮播,关于我们,成功案例,服务流程,团队介绍,数据部分,公司动态,底部信息等内容区块。网站整体采用CSSGrid布局,支持响应式,有流畅过渡和展现动画。</dd> <dd class="cont2"> <a href="/course/28.html" title="从零制作响应式网站—Grid布局" class="toStudy">立即学习</a> <span>485次学习</span> </dd> </dl> </li> </ul> </div> </div> <!-- footer --> <link href="https://fonts.googleapis.com/icon?family=Material+Icons" rel="stylesheet"> <div class="footer"> <ul> <li ><a href="/" class="aLightGray"><em class="material-icons">home</em><span>首页</span></a></li> <li class="curr"><a href="/articlelist.html" class="aLightGray"><em class="material-icons">menu_book</em><span>阅读</span></a></li> <li ><a href="/courselist.html" class="aLightGray"><em class="material-icons">school</em><span>课程</span></a></li> <li ><a href="/ai.html" class="aLightGray"><em class="material-icons">smart_toy</em><span>AI助手</span></a></li> <li ><a href="/user.html" class="aLightGray"><em class="material-icons">person</em><span>我的</span></a></li> </ul> </div> <script src="/assets/js/require.js" data-main="/assets/js/require-frontend.js?v=1671101972"></script> <script> var _hmt = _hmt || []; (function() { var hm = document.createElement("script"); hm.src = "https://hm.baidu.com/hm.js?3dc5666f6478c7bf39cd5c91e597423d"; var s = document.getElementsByTagName("script")[0]; s.parentNode.insertBefore(hm, s); })(); </script> </body> </html>