登录
首页 >  文章 >  php教程

PHP自动翻页抓取数据教程详解

时间:2026-03-15 13:03:36 113浏览 收藏

本文深入剖析了PHP抓取分页数据的关键难点与实战要点,强调成功的核心不在于编写“自动翻页”的循环代码,而在于精准识别目标网站的分页机制(GET参数、AJAX接口或前端渲染的无限滚动),并针对性地构造请求、严谨处理curl异常响应、合理延时防封、以及用XPath等技术稳健提取结构化数据;文章直击开发者常见误区——盲目写循环却忽视网页真实逻辑和反爬变化,提醒读者:花10分钟读懂三页源码和网络请求,远胜于事后数小时调试。

怎么抓取分页内容_PHP自动翻页抓取列表数据方法【教程】

PHP 抓取分页内容,核心不是“自动翻页”,而是识别分页逻辑并构造正确请求——多数失败源于没看懂目标网站的分页机制,而非代码写得不够“智能”。

怎么判断分页是 GET 参数、AJAX 还是滚动加载

先打开浏览器开发者工具(F12),切到 Network 标签,手动点下一页或滚动到底部,观察触发了什么请求:

  • URL 变成 list.php?page=2index.html?p=3 → 是 GET 分页,直接拼 page 参数即可
  • 看到 XHRfetch 请求,响应是 JSON,地址像 /api/items?page=2 → 是 AJAX 分页,需模拟带 Accept: application/json 的请求,注意 Referer 和 Cookie 是否需要携带
  • 没新请求,但页面 DOM 增加了新条目 → 很可能是前端 JS 渲染的无限滚动,PHP 无法直接抓取,得换方案(如 Puppeteer)或找它背后的真实 API

curl 循环请求时必须处理的三个细节

curl 抓多页,光写个 for 循环远远不够:

  • 每次请求后检查 HTTP 状态码:curl_getinfo($ch, CURLINFO_HTTP_CODE) 不是 200 就该停,别硬刷
  • 解析 HTML 前先确认是否拿到完整内容:用 mb_strlen($html) < 1000 或匹配关键标签(如 </code>)防返回 403/503 页面被当正常数据</li> <li>加延时不是可选:<code>usleep(100000)</code>(100ms)比 <code>sleep(1)</code> 更合理,既防封又不拖慢整体速度;某些站对 User-Agent 敏感,记得设 <code>curl_setopt($ch, CURLOPT_USERAGENT, 'Mozilla/5.0')</code></li> </ul><h3>用 <code>DOMDocument</code> 提取列表时容易漏掉的边界情况</h3> <p>别一上来就 <code>$dom->getElementsByTagName('li')</code>,真实页面常有干扰:</p> <ul><li>分页导航栏本身也含 <code><li></code>,得用更精确的父容器定位,比如先找 <code>$dom->getElementById('content')</code> 或 <code>$dom->getElementsByClassName('item-list')[0]</code></li> <li>有些列表项是空的、被注释掉的、或用 <code><div class="item"></code> 而非 <code><li></code>,建议统一用 XPath:<code>$xpath->query('//div[contains(@class,"item") or contains(@class,"post")]')</code></li> <li>内容含 HTML 实体(如 <code> </code>)或 UTF-8 BOM,<code>textContent</code> 会带乱码,改用 <code>trim(html_entity_decode($node->nodeValue, ENT_QUOTES, 'UTF-8'))</code></li> </ul><p>真正难的从来不是写循环,而是每一页的 HTML 结构是否一致、反爬策略是否在第 5 页突然升级、目标 URL 是否用了动态 token。动手前花 10 分钟看三页源码和请求头,比写完再调试两小时更省时间。</p><p>以上就是《PHP自动翻页抓取数据教程详解》的详细内容,更多关于的资料请关注golang学习网公众号!</p> <div id="third-party-overlay" style="position:fixed;left:0;top:0;width:100%;height:100%;display:none;justify-content:center;align-items:center;background:rgba(0,0,0,0.4);z-index:9999;"> <div style="background:#FFF3CD;border:1px solid #FFEEBA;padding:16px;border-radius:6px;box-sizing:border-box;max-width:480px;width:90%;text-align:center;"> <div style="font-size:14px;color:#856404;margin-bottom:12px;">您即将跳转至第三方网站,请注意保护好个人信息和财产安全!</div> <a href="https://comicdow.pdlcomic.top/1273%2F%E5%9B%A7%E6%AC%A1%E5%85%83.apk" target="_blank" rel="nofollow noopener noreferrer" style="color:#2d8cf0;text-decoration:none;" onclick="showThirdParty('none');">继续访问</a> </div> </div> <script> function showThirdParty(mode){ var el = document.getElementById('third-party-overlay'); if (!el) return; el.style.display = (mode === 'none' ? 'none' : 'flex'); } </script> </div> <div class="labsList"> </div> </div> <div class="contBoxNor"> <div class="contTit"> <div class="tit">资料下载</div> </div> <ul class="classRecomList"> <li> <a href="https://pan.quark.cn/s/ba8ef670cabd" rel="nofollow" target="_blank" title="编程学习资料下载" class="img_box"> <img loading="lazy" src="/assets/images/xuexiziliao.jpeg" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="编程学习资料下载"> </a> <dl> <dt class="lineOverflow"> <a href="https://pan.quark.cn/s/ba8ef670cabd" rel="nofollow" target="_blank" class="aBlack" title="编程学习资料下载">编程学习资料下载</a> </dt> <dd class="cont1 lineTwoOverflow"> 精选 编程(Golang、Python、Java、C++、JavaScript等) 教程、电子书与示例源码,一键打包本地下载学习。 </dd> <dd class="cont2"> <a href="https://pan.quark.cn/s/ba8ef670cabd" rel="nofollow" target="_blank" class="toStudy">立即下载</a> </dd> </dl> </li> </ul> </div> <!-- 最新阅读 --> <div class="contBoxNor"> <div class="contTit"> <div class="tit">相关阅读</div> <a href="/articlelist.html" class="more">更多></a> </div> <ul class="latestReadList"> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  2年前  |   <a href="javascript:;" class="aLightGray" title="PHP技术">PHP技术</a> <a href="javascript:;" class="aLightGray" title="高薪回报">高薪回报</a> <a href="javascript:;" class="aLightGray" title="发展前景">发展前景</a> </div> <div class="tit lineOverflow"><a href="/article/61908.html" title="PHP技术的高薪回报与发展前景" class="aBlack">PHP技术的高薪回报与发展前景</a></div> <div class="opt"> <span><i class="view"></i>501</span> <span class="collectBtn user_collection" data-id="61908" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  2年前  |   <a href="javascript:;" class="aLightGray" title="php">php</a> <a href="javascript:;" class="aLightGray" title="优惠券">优惠券</a> <a href="javascript:;" class="aLightGray" title="商场">商场</a> </div> <div class="tit lineOverflow"><a href="/article/62538.html" title="基于 PHP 的商场优惠券系统开发中的常见问题解决方案" class="aBlack">基于 PHP 的商场优惠券系统开发中的常见问题解决方案</a></div> <div class="opt"> <span><i class="view"></i>501</span> <span class="collectBtn user_collection" data-id="62538" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  2年前  |   <a href="javascript:;" class="aLightGray" title="PHP支付功能">PHP支付功能</a> <a href="javascript:;" class="aLightGray" title="在线支付开发">在线支付开发</a> <a href="javascript:;" class="aLightGray" title="简单支付实现">简单支付实现</a> </div> <div class="tit lineOverflow"><a href="/article/62741.html" title="如何使用PHP开发简单的在线支付功能" class="aBlack">如何使用PHP开发简单的在线支付功能</a></div> <div class="opt"> <span><i class="view"></i>501</span> <span class="collectBtn user_collection" data-id="62741" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  2年前  |   <a href="javascript:;" class="aLightGray" title="分布式缓存">分布式缓存</a> <a href="javascript:;" class="aLightGray" title="PHP消息队列">PHP消息队列</a> <a href="javascript:;" class="aLightGray" title="缓存刷新器">缓存刷新器</a> </div> <div class="tit lineOverflow"><a href="/article/62881.html" title="PHP消息队列开发指南:实现分布式缓存刷新器" class="aBlack">PHP消息队列开发指南:实现分布式缓存刷新器</a></div> <div class="opt"> <span><i class="view"></i>501</span> <span class="collectBtn user_collection" data-id="62881" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  2年前  |   <a href="javascript:;" class="aLightGray" title="微服务">微服务</a> <a href="javascript:;" class="aLightGray" title="调度">调度</a> <a href="javascript:;" class="aLightGray" title="分布式任务">分布式任务</a> </div> <div class="tit lineOverflow"><a href="/article/63734.html" title="如何在PHP微服务中实现分布式任务分配和调度" class="aBlack">如何在PHP微服务中实现分布式任务分配和调度</a></div> <div class="opt"> <span><i class="view"></i>501</span> <span class="collectBtn user_collection" data-id="63734" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> </ul> </div> <!-- 最新阅读 --> <div class="contBoxNor"> <div class="contTit"> <div class="tit">最新阅读</div> <a href="/articlelist.html" class="more">更多></a> </div> <ul class="latestReadList"> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  3分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/531599.html" title="PHP计算两个日期间隔天数技巧" class="aBlack">PHP计算两个日期间隔天数技巧</a></div> <div class="opt"> <span><i class="view"></i>107</span> <span class="collectBtn user_collection" data-id="531599" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  18分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/531579.html" title="PHP双数组同时遍历技巧详解" class="aBlack">PHP双数组同时遍历技巧详解</a></div> <div class="opt"> <span><i class="view"></i>410</span> <span class="collectBtn user_collection" data-id="531579" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  25分钟前  |   <a href="javascript:;" class="aLightGray" title="php">php</a> </div> <div class="tit lineOverflow"><a href="/article/531570.html" title="如何打开PHP源码网页?手把手教学解析" class="aBlack">如何打开PHP源码网页?手把手教学解析</a></div> <div class="opt"> <span><i class="view"></i>291</span> <span class="collectBtn user_collection" data-id="531570" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  36分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/531557.html" title="PHPksort与asort有何不同?" class="aBlack">PHPksort与asort有何不同?</a></div> <div class="opt"> <span><i class="view"></i>455</span> <span class="collectBtn user_collection" data-id="531557" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  42分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/531547.html" title="PHP变量持久化恢复流程详解" class="aBlack">PHP变量持久化恢复流程详解</a></div> <div class="opt"> <span><i class="view"></i>137</span> <span class="collectBtn user_collection" data-id="531547" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  48分钟前  |   <a href="javascript:;" class="aLightGray" title="php">php</a> </div> <div class="tit lineOverflow"><a href="/article/531540.html" title="PHP性能优化技巧与提升方法" class="aBlack">PHP性能优化技巧与提升方法</a></div> <div class="opt"> <span><i class="view"></i>449</span> <span class="collectBtn user_collection" data-id="531540" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  1小时前  |   <a href="javascript:;" class="aLightGray" title="php">php</a> </div> <div class="tit lineOverflow"><a href="/article/531524.html" title="PHPcURL使用教程与实例详解" class="aBlack">PHPcURL使用教程与实例详解</a></div> <div class="opt"> <span><i class="view"></i>381</span> <span class="collectBtn user_collection" data-id="531524" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  1小时前  |   </div> <div class="tit lineOverflow"><a href="/article/531515.html" title="PHP数组遍历技巧大全" class="aBlack">PHP数组遍历技巧大全</a></div> <div class="opt"> <span><i class="view"></i>287</span> <span class="collectBtn user_collection" data-id="531515" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  1小时前  |   <a href="javascript:;" class="aLightGray" title="php">php</a> </div> <div class="tit lineOverflow"><a href="/article/531505.html" title="PHP异步下载文件方法详解" class="aBlack">PHP异步下载文件方法详解</a></div> <div class="opt"> <span><i class="view"></i>136</span> <span class="collectBtn user_collection" data-id="531505" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  1小时前  |   </div> <div class="tit lineOverflow"><a href="/article/531502.html" title="宝塔面板PHP超时设置教程" class="aBlack">宝塔面板PHP超时设置教程</a></div> <div class="opt"> <span><i class="view"></i>414</span> <span class="collectBtn user_collection" data-id="531502" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  1小时前  |   <a href="javascript:;" class="aLightGray" title="Symfony">Symfony</a> </div> <div class="tit lineOverflow"><a href="/article/531485.html" title="Symfony邮件发送配置详解" class="aBlack">Symfony邮件发送配置详解</a></div> <div class="opt"> <span><i class="view"></i>360</span> <span class="collectBtn user_collection" data-id="531485" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/19_new_0_1.html" class="aLightGray" title="文章">文章</a> · <a href="/articlelist/84_new_0_1.html" class="aLightGray" title="php教程">php教程</a>   |  1小时前  |   </div> <div class="tit lineOverflow"><a href="/article/531457.html" title="PHP获取URL参数的实用方法" class="aBlack">PHP获取URL参数的实用方法</a></div> <div class="opt"> <span><i class="view"></i>460</span> <span class="collectBtn user_collection" data-id="531457" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> </ul> </div> <!-- 课程推荐 --> <div class="contBoxNor"> <div class="contTit"> <div class="tit">课程推荐</div> <a href="/courselist.html" class="more">更多></a> </div> <ul class="classRecomList"> <li> <a href="/course/9.html" title="前端进阶之JavaScript设计模式" class="img_box"> <img loading="lazy" src="/uploads/20221222/52fd0f23a454c71029c2c72d206ed815.jpg" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="前端进阶之JavaScript设计模式"> </a> <dl> <dt class="lineOverflow"> 前端进阶之JavaScript设计模式 </dt> <dd class="cont1 lineOverflow">设计模式是开发人员在软件开发过程中面临一般问题时的解决方案,代表了最佳的实践。本课程的主打内容包括JS常见设计模式以及具体应用场景,打造一站式知识长龙服务,适合有JS基础的同学学习。</dd> <dd class="cont2"> <a href="/course/9.html" title="前端进阶之JavaScript设计模式" class="toStudy">立即学习</a> <span>543次学习</span> </dd> </dl> </li> <li> <a href="/course/2.html" title="GO语言核心编程课程" class="img_box"> <img loading="lazy" src="/uploads/20221221/634ad7404159bfefc6a54a564d437b5f.png" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="GO语言核心编程课程"> </a> <dl> <dt class="lineOverflow"> GO语言核心编程课程 </dt> <dd class="cont1 lineOverflow">本课程采用真实案例,全面具体可落地,从理论到实践,一步一步将GO核心编程技术、编程思想、底层实现融会贯通,使学习者贴近时代脉搏,做IT互联网时代的弄潮儿。</dd> <dd class="cont2"> <a href="/course/2.html" title="GO语言核心编程课程" class="toStudy">立即学习</a> <span>516次学习</span> </dd> </dl> </li> <li> <a href="/course/74.html" title="简单聊聊mysql8与网络通信" class="img_box"> <img loading="lazy" src="/uploads/20240103/bad35fe14edbd214bee16f88343ac57c.png" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="简单聊聊mysql8与网络通信"> </a> <dl> <dt class="lineOverflow"> 简单聊聊mysql8与网络通信 </dt> <dd class="cont1 lineOverflow">如有问题加微信:Le-studyg;在课程中,我们将首先介绍MySQL8的新特性,包括性能优化、安全增强、新数据类型等,帮助学生快速熟悉MySQL8的最新功能。接着,我们将深入解析MySQL的网络通信机制,包括协议、连接管理、数据传输等,让</dd> <dd class="cont2"> <a href="/course/74.html" title="简单聊聊mysql8与网络通信" class="toStudy">立即学习</a> <span>500次学习</span> </dd> </dl> </li> <li> <a href="/course/57.html" title="JavaScript正则表达式基础与实战" class="img_box"> <img loading="lazy" src="/uploads/20221226/bbe4083bb3cb0dd135fb02c31c3785fb.jpg" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="JavaScript正则表达式基础与实战"> </a> <dl> <dt class="lineOverflow"> JavaScript正则表达式基础与实战 </dt> <dd class="cont1 lineOverflow">在任何一门编程语言中,正则表达式,都是一项重要的知识,它提供了高效的字符串匹配与捕获机制,可以极大的简化程序设计。</dd> <dd class="cont2"> <a href="/course/57.html" title="JavaScript正则表达式基础与实战" class="toStudy">立即学习</a> <span>487次学习</span> </dd> </dl> </li> <li> <a href="/course/28.html" title="从零制作响应式网站—Grid布局" class="img_box"> <img loading="lazy" src="/uploads/20221223/ac110f88206daeab6c0cf38ebf5fe9ed.jpg" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="从零制作响应式网站—Grid布局"> </a> <dl> <dt class="lineOverflow"> 从零制作响应式网站—Grid布局 </dt> <dd class="cont1 lineOverflow">本系列教程将展示从零制作一个假想的网络科技公司官网,分为导航,轮播,关于我们,成功案例,服务流程,团队介绍,数据部分,公司动态,底部信息等内容区块。网站整体采用CSSGrid布局,支持响应式,有流畅过渡和展现动画。</dd> <dd class="cont2"> <a href="/course/28.html" title="从零制作响应式网站—Grid布局" class="toStudy">立即学习</a> <span>485次学习</span> </dd> </dl> </li> </ul> </div> </div> <!-- footer --> <link href="https://fonts.googleapis.com/icon?family=Material+Icons" rel="stylesheet"> <div class="footer"> <ul> <li ><a href="/" class="aLightGray"><em class="material-icons">home</em><span>首页</span></a></li> <li class="curr"><a href="/articlelist.html" class="aLightGray"><em class="material-icons">menu_book</em><span>阅读</span></a></li> <li ><a href="/courselist.html" class="aLightGray"><em class="material-icons">school</em><span>课程</span></a></li> <li ><a href="/ai.html" class="aLightGray"><em class="material-icons">smart_toy</em><span>AI助手</span></a></li> <li ><a href="/user.html" class="aLightGray"><em class="material-icons">person</em><span>我的</span></a></li> </ul> </div> <script src="/assets/js/require.js" data-main="/assets/js/require-frontend.js?v=1671101972"></script> <script> var _hmt = _hmt || []; (function() { var hm = document.createElement("script"); hm.src = "https://hm.baidu.com/hm.js?3dc5666f6478c7bf39cd5c91e597423d"; var s = document.getElementsByTagName("script")[0]; s.parentNode.insertBefore(hm, s); })(); </script> <script src="/assets/js/SyntaxHighlighter/shCore.js?3.1.1"></script> <script> document.addEventListener('DOMContentLoaded', function () { if (document.querySelector('.cont pre')) { SyntaxHighlighter.all(); } }); </script> </body> </html>