登录
首页 >  Golang >  Go教程

Golang爬虫入门:Colly框架实战教程

时间:2025-07-13 19:48:31 312浏览 收藏

从现在开始,我们要努力学习啦!今天我给大家带来《Golang爬虫入门:用Colly框架轻松实战》,感兴趣的朋友请继续看下去吧!下文中的内容我们主要会涉及到等等知识点,如果在阅读本文过程中有遇到不清楚的地方,欢迎留言呀!我们一起讨论,一起学习!

用Golang写爬虫不难,尤其使用colly框架时上手快。1. 安装colly并创建基础爬虫:执行go get github.com/gocolly/colly/v2,编写代码创建collector实例、设置回调函数、访问目标URL提取页面标题;2. 抓取列表页中的链接:通过c.OnHTML配合CSS选择器如.post-list a[href]提取详情链接,并可复用collector访问这些链接;3. 存储抓取数据:定义结构体如Article保存标题和URL,将结果存入变量后续导出为JSON或数据库;4. 防止重复抓取可通过WithDomain限制域名,设置User-Agent避免被屏蔽,控制并发和延迟提升大规模抓取效率,动态内容需结合chromedp等工具处理。掌握这些要点即可快速入门colly爬虫开发。

如何用Golang开发简易爬虫程序 使用colly框架入门实战

用Golang写爬虫其实不难,尤其当你使用colly这个框架时,上手非常快。Colly是Go语言中最受欢迎的爬虫库之一,结构清晰、性能好,适合做数据抓取类的小项目。如果你刚接触Go和网络爬虫,这篇文章可以帮你快速入门。

如何用Golang开发简易爬虫程序 使用colly框架入门实战

安装Colly并创建一个基础爬虫

首先你得确保Go环境已经配好,然后安装colly:

如何用Golang开发简易爬虫程序 使用colly框架入门实战
go get github.com/gocolly/colly/v2

接下来是最简单的例子:访问一个网页并打印页面标题。

package main

import (
    "fmt"
    "github.com/gocolly/colly/v2"
)

func main() {
    // 创建一个新的Collector实例
    c := colly.NewCollector()

    // 在每次抓取页面时触发
    c.OnHTML("html", func(e *colly.HTMLElement) {
        title := e.ChildText("title")
        fmt.Println("页面标题是:", title)
    })

    // 开始请求目标URL
    c.Visit("https://example.com")
}

这段代码会访问example.com,提取它的</code>标签内容并输出。看起来简单,但已经包含了colly的基本结构:创建collector → 设置回调函数 → 发起请求。</p><img src="/uploads/20250713/175240728768739cf760a59.jpg" alt="如何用Golang开发简易爬虫程序 使用colly框架入门实战"><hr><h3>抓取列表页中的链接</h3><p>实际开发中,我们经常需要从一个列表页里抓取多个条目的详情链接。比如新闻网站的首页,每条新闻都是一个链接。</p><p>假设你想抓取某个博客首页的所有文章链接,可以这样做:</p><pre class="brush:go;toolbar:false;">c.OnHTML(".post-list a[href]", func(e *colly.HTMLElement) { link := e.Attr("href") fmt.Println("发现文章链接:", link) })</pre><p>这里的关键点在于选择器要准确,<code>.post-list a[href]</code>表示在class为<code>post-list</code>的容器内找所有带<code>href</code>属性的<code>a</code>标签。你可以根据实际页面结构调整选择器。</p><p>如果想进一步访问这些链接,可以用另一个collector去处理详情页,或者复用当前collector,加上限制域名等设置。</p><hr><h3>存储抓取到的数据</h3><p>光打印出来不够实用,一般我们会把数据保存下来,比如JSON文件或数据库。</p><p>最简单的做法是定义一个结构体,把抓取结果存进去:</p><pre class="brush:go;toolbar:false;">type Article struct { Title string URL string } var articles []Article c.OnHTML(".post-list a[href]", func(e *colly.HTMLElement) { link := e.Attr("href") title := e.Text articles = append(articles, Article{ Title: title, URL: link, }) })</pre><p>之后你可以把这些数据导出成JSON,或者插入到SQLite、MySQL这样的数据库里。这部分就不展开讲了,重点还是放在爬虫本身逻辑上。</p><hr><h3>一些常见问题和建议</h3><ul><li><p><strong>防止重复抓取</strong>:可以用<code>colly.WithDomain("example.com")</code>限制域名,避免进入无关页面。</p></li><li><p><strong>设置User-Agent</strong>:有些网站会屏蔽默认的Go User-Agent,可以在初始化collector后加上:</p><pre class="brush:go;toolbar:false;">c.UserAgent = "Mozilla/5.0 (compatible; ExampleBot/1.0; +http://example.com/bot)"</pre></li><li><p><strong>控制并发和限速</strong>:对于大规模抓取,可以设置最大并发数和延迟:</p><pre class="brush:go;toolbar:false;">c.Limit(&colly.LimitRule{DomainGlob: "*", Parallelism: 2, Delay: 1 * time.Second})</pre></li><li><p><strong>处理JavaScript渲染页面</strong>:Colly本身只能抓静态HTML,无法执行JS。如果目标页面是动态加载的内容,就得考虑用其他工具配合,比如chromedp或selenium。</p></li></ul><hr><p>基本上就这些。用colly写个简易爬虫并不复杂,关键是熟悉HTML结构和CSS选择器的写法。多练几个小项目,就能掌握常见的抓取套路了。</p><p>终于介绍完啦!小伙伴们,这篇关于《Golang爬虫入门:Colly框架实战教程》的介绍应该让你收获多多了吧!欢迎大家收藏或分享给更多需要学习的朋友吧~golang学习网公众号也会发布Golang相关知识,快来关注吧!</p> </div> <div class="labsList"> </div> </div> <!-- 最新阅读 --> <div class="contBoxNor"> <div class="contTit"> <div class="tit">相关阅读</div> <a href="/articlelist.html" class="more">更多></a> </div> <ul class="latestReadList"> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  2年前  |   <a href="/articletag/56_new_0_1.html" class="aLightGray" title="map">map</a> · <a href="/articletag/993_new_0_1.html" class="aLightGray" title="实践">实践</a> · <a href="/articletag/994_new_0_1.html" class="aLightGray" title="实现原理">实现原理</a> · <a href="/special/3_new_0_1.html" target="_blank" class="aLightGray" title="golang">golang</a> </div> <div class="tit lineOverflow"><a href="/article/10762.html" title="Golangmap实践及实现原理解析" class="aBlack">Golangmap实践及实现原理解析</a></div> <div class="opt"> <span><i class="view"></i>505</span> <span class="collectBtn user_collection" data-id="10762" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  2年前  |   <a href="/articletag/1668_new_0_1.html" class="aLightGray" title="try">try</a> · <a href="/articletag/1669_new_0_1.html" class="aLightGray" title="catch">catch</a> · <a href="/special/3_new_0_1.html" target="_blank" class="aLightGray" title="golang">golang</a> </div> <div class="tit lineOverflow"><a href="/article/11451.html" title="试了下Golang实现try catch的方法" class="aBlack">试了下Golang实现try catch的方法</a></div> <div class="opt"> <span><i class="view"></i>502</span> <span class="collectBtn user_collection" data-id="11451" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  2年前  |   <a href="javascript:;" class="aLightGray" title="并发 (concurrency)">并发 (concurrency)</a> <a href="javascript:;" class="aLightGray" title="Go语言 (Go language)">Go语言 (Go language)</a> <a href="javascript:;" class="aLightGray" title="服务器架构 (Server Architecture)">服务器架构 (Server Architecture)</a> </div> <div class="tit lineOverflow"><a href="/article/53565.html" title="如何在go语言中实现高并发的服务器架构" class="aBlack">如何在go语言中实现高并发的服务器架构</a></div> <div class="opt"> <span><i class="view"></i>502</span> <span class="collectBtn user_collection" data-id="53565" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  1年前  |   <a href="/special/3_new_0_1.html" target="_blank" class="aLightGray" title="golang">golang</a> <a href="javascript:;" class="aLightGray" title="Go">Go</a> <a href="javascript:;" class="aLightGray" title="编程语言选择">编程语言选择</a> <a href="javascript:;" class="aLightGray" title="区别解析">区别解析</a> </div> <div class="tit lineOverflow"><a href="/article/80612.html" title="go和golang的区别解析:帮你选择合适的编程语言" class="aBlack">go和golang的区别解析:帮你选择合适的编程语言</a></div> <div class="opt"> <span><i class="view"></i>502</span> <span class="collectBtn user_collection" data-id="80612" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  1年前  |   <a href="javascript:;" class="aLightGray" title="工作效率">工作效率</a> <a href="javascript:;" class="aLightGray" title="Go语言">Go语言</a> <a href="javascript:;" class="aLightGray" title="项目开发">项目开发</a> </div> <div class="tit lineOverflow"><a href="/article/72902.html" title="提升工作效率的Go语言项目开发经验分享" class="aBlack">提升工作效率的Go语言项目开发经验分享</a></div> <div class="opt"> <span><i class="view"></i>502</span> <span class="collectBtn user_collection" data-id="72902" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> </ul> </div> <!-- 最新阅读 --> <div class="contBoxNor"> <div class="contTit"> <div class="tit">最新阅读</div> <a href="/articlelist.html" class="more">更多></a> </div> <ul class="latestReadList"> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  4分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/305781.html" title="Golang错误码设计与业务规范定义" class="aBlack">Golang错误码设计与业务规范定义</a></div> <div class="opt"> <span><i class="view"></i>411</span> <span class="collectBtn user_collection" data-id="305781" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  7分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/305777.html" title="Golang日志系统作用与FluentBit插件开发解析" class="aBlack">Golang日志系统作用与FluentBit插件开发解析</a></div> <div class="opt"> <span><i class="view"></i>183</span> <span class="collectBtn user_collection" data-id="305777" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  8分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/305775.html" title="Golang桥接模式解析:抽象与实现分离" class="aBlack">Golang桥接模式解析:抽象与实现分离</a></div> <div class="opt"> <span><i class="view"></i>362</span> <span class="collectBtn user_collection" data-id="305775" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  15分钟前  |   <a href="javascript:;" class="aLightGray" title="配置">配置</a> <a href="javascript:;" class="aLightGray" title="GoLand">GoLand</a> <a href="javascript:;" class="aLightGray" title="GoModules">GoModules</a> <a href="javascript:;" class="aLightGray" title="多版本管理">多版本管理</a> <a href="javascript:;" class="aLightGray" title="GoSDK">GoSDK</a> </div> <div class="tit lineOverflow"><a href="/article/305766.html" title="GoLand首次启动如何设置GolangSDK" class="aBlack">GoLand首次启动如何设置GolangSDK</a></div> <div class="opt"> <span><i class="view"></i>480</span> <span class="collectBtn user_collection" data-id="305766" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  16分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/305765.html" title="Golang适配器模式与接口转换技巧" class="aBlack">Golang适配器模式与接口转换技巧</a></div> <div class="opt"> <span><i class="view"></i>156</span> <span class="collectBtn user_collection" data-id="305765" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  22分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/305761.html" title="Golang值传递与返回拷贝详解" class="aBlack">Golang值传递与返回拷贝详解</a></div> <div class="opt"> <span><i class="view"></i>121</span> <span class="collectBtn user_collection" data-id="305761" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  22分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/305760.html" title="Golang文件权限与用户组操作全解析" class="aBlack">Golang文件权限与用户组操作全解析</a></div> <div class="opt"> <span><i class="view"></i>372</span> <span class="collectBtn user_collection" data-id="305760" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  23分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/305759.html" title="Golang大规模部署用Kustomize渲染模板方法" class="aBlack">Golang大规模部署用Kustomize渲染模板方法</a></div> <div class="opt"> <span><i class="view"></i>347</span> <span class="collectBtn user_collection" data-id="305759" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  27分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/305755.html" title="Golang空接口类型判断技巧" class="aBlack">Golang空接口类型判断技巧</a></div> <div class="opt"> <span><i class="view"></i>397</span> <span class="collectBtn user_collection" data-id="305755" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  29分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/305754.html" title="Golang项目如何打包成单文件二进制" class="aBlack">Golang项目如何打包成单文件二进制</a></div> <div class="opt"> <span><i class="view"></i>496</span> <span class="collectBtn user_collection" data-id="305754" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  38分钟前  |   <a href="javascript:;" class="aLightGray" title="内存安全">内存安全</a> <a href="javascript:;" class="aLightGray" title="Go指针">Go指针</a> <a href="javascript:;" class="aLightGray" title="unsafe.Pointer">unsafe.Pointer</a> <a href="javascript:;" class="aLightGray" title="C语言指针">C语言指针</a> <a href="javascript:;" class="aLightGray" title="指针算术运算">指针算术运算</a> </div> <div class="tit lineOverflow"><a href="/article/305745.html" title="Golang指针限制对比C语言指针差异" class="aBlack">Golang指针限制对比C语言指针差异</a></div> <div class="opt"> <span><i class="view"></i>357</span> <span class="collectBtn user_collection" data-id="305745" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> <li> <div class="info"> <a href="/articlelist/25_new_0_1.html" class="aLightGray" title="Golang">Golang</a> · <a href="/articlelist/44_new_0_1.html" class="aLightGray" title="Go教程">Go教程</a>   |  43分钟前  |   </div> <div class="tit lineOverflow"><a href="/article/305738.html" title="Go并发编程:Goroutine通信技巧与避坑指南" class="aBlack">Go并发编程:Goroutine通信技巧与避坑指南</a></div> <div class="opt"> <span><i class="view"></i>239</span> <span class="collectBtn user_collection" data-id="305738" data-type="article" title="收藏"><i class="collect"></i>收藏</span> </div> </li> </ul> </div> <!-- 课程推荐 --> <div class="contBoxNor"> <div class="contTit"> <div class="tit">课程推荐</div> <a href="/courselist.html" class="more">更多></a> </div> <ul class="classRecomList"> <li> <a href="/course/9.html" title="前端进阶之JavaScript设计模式" class="img_box"> <img src="/uploads/20221222/52fd0f23a454c71029c2c72d206ed815.jpg" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="前端进阶之JavaScript设计模式"> </a> <dl> <dt class="lineOverflow"> 前端进阶之JavaScript设计模式 </dt> <dd class="cont1 lineOverflow">设计模式是开发人员在软件开发过程中面临一般问题时的解决方案,代表了最佳的实践。本课程的主打内容包括JS常见设计模式以及具体应用场景,打造一站式知识长龙服务,适合有JS基础的同学学习。</dd> <dd class="cont2"> <a href="/course/9.html" title="前端进阶之JavaScript设计模式" class="toStudy">立即学习</a> <span>543次学习</span> </dd> </dl> </li> <li> <a href="/course/2.html" title="GO语言核心编程课程" class="img_box"> <img src="/uploads/20221221/634ad7404159bfefc6a54a564d437b5f.png" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="GO语言核心编程课程"> </a> <dl> <dt class="lineOverflow"> GO语言核心编程课程 </dt> <dd class="cont1 lineOverflow">本课程采用真实案例,全面具体可落地,从理论到实践,一步一步将GO核心编程技术、编程思想、底层实现融会贯通,使学习者贴近时代脉搏,做IT互联网时代的弄潮儿。</dd> <dd class="cont2"> <a href="/course/2.html" title="GO语言核心编程课程" class="toStudy">立即学习</a> <span>512次学习</span> </dd> </dl> </li> <li> <a href="/course/74.html" title="简单聊聊mysql8与网络通信" class="img_box"> <img src="/uploads/20240103/bad35fe14edbd214bee16f88343ac57c.png" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="简单聊聊mysql8与网络通信"> </a> <dl> <dt class="lineOverflow"> 简单聊聊mysql8与网络通信 </dt> <dd class="cont1 lineOverflow">如有问题加微信:Le-studyg;在课程中,我们将首先介绍MySQL8的新特性,包括性能优化、安全增强、新数据类型等,帮助学生快速熟悉MySQL8的最新功能。接着,我们将深入解析MySQL的网络通信机制,包括协议、连接管理、数据传输等,让</dd> <dd class="cont2"> <a href="/course/74.html" title="简单聊聊mysql8与网络通信" class="toStudy">立即学习</a> <span>499次学习</span> </dd> </dl> </li> <li> <a href="/course/57.html" title="JavaScript正则表达式基础与实战" class="img_box"> <img src="/uploads/20221226/bbe4083bb3cb0dd135fb02c31c3785fb.jpg" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="JavaScript正则表达式基础与实战"> </a> <dl> <dt class="lineOverflow"> JavaScript正则表达式基础与实战 </dt> <dd class="cont1 lineOverflow">在任何一门编程语言中,正则表达式,都是一项重要的知识,它提供了高效的字符串匹配与捕获机制,可以极大的简化程序设计。</dd> <dd class="cont2"> <a href="/course/57.html" title="JavaScript正则表达式基础与实战" class="toStudy">立即学习</a> <span>487次学习</span> </dd> </dl> </li> <li> <a href="/course/28.html" title="从零制作响应式网站—Grid布局" class="img_box"> <img src="/uploads/20221223/ac110f88206daeab6c0cf38ebf5fe9ed.jpg" onerror="this.onerror='';this.src='/assets/images/moren/morentu.png'" alt="从零制作响应式网站—Grid布局"> </a> <dl> <dt class="lineOverflow"> 从零制作响应式网站—Grid布局 </dt> <dd class="cont1 lineOverflow">本系列教程将展示从零制作一个假想的网络科技公司官网,分为导航,轮播,关于我们,成功案例,服务流程,团队介绍,数据部分,公司动态,底部信息等内容区块。网站整体采用CSSGrid布局,支持响应式,有流畅过渡和展现动画。</dd> <dd class="cont2"> <a href="/course/28.html" title="从零制作响应式网站—Grid布局" class="toStudy">立即学习</a> <span>484次学习</span> </dd> </dl> </li> </ul> </div> </div> <!-- footer --> <link href="https://fonts.googleapis.com/icon?family=Material+Icons" rel="stylesheet"> <div class="footer"> <ul> <li ><a href="/" class="aLightGray"><em class="material-icons">home</em><span>首页</span></a></li> <li class="curr"><a href="/articlelist.html" class="aLightGray"><em class="material-icons">menu_book</em><span>阅读</span></a></li> <li ><a href="/courselist.html" class="aLightGray"><em class="material-icons">school</em><span>课程</span></a></li> <li ><a href="/ai.html" class="aLightGray"><em class="material-icons">smart_toy</em><span>AI助手</span></a></li> <li ><a href="/user.html" class="aLightGray"><em class="material-icons">person</em><span>我的</span></a></li> </ul> </div> <script src="/assets/js/require.js" data-main="/assets/js/require-frontend.js?v=1671101972"></script> <script> var _hmt = _hmt || []; (function() { var hm = document.createElement("script"); hm.src = "https://hm.baidu.com/hm.js?3dc5666f6478c7bf39cd5c91e597423d"; var s = document.getElementsByTagName("script")[0]; s.parentNode.insertBefore(hm, s); })(); </script> </body> </html>