明确目标 目标网址 获取节点 模拟请求 解析内容 本地处理 过滤 获取当前页面并解析 绝大部分网站目前都是utf-8 编码格式 解析页面库beautiSoup 针对节点进行分析 基本就是遍历目标节点 基本结构12345678910111213141516171819202122import urllib.requestfrom bs4 import BeautifulSoupimport timeBASE_URL = 'https://xxx.com'if __name__ == '__main__': while 1: #获取当前页面 page = urllib.request.urlopen(url).read().decode('utf-8') soup = BeautifulSoup(page, 'lxml') #分页 lists = soup.find_all('div', class='pg') if len(lists) < 1: break for listone in lists: # 获取每页链接 aLists = listone.find_all('a', attrs={'class':''}) for a in aLists: url = BASE_URL + a['href'] time.sleep(1) Author: inpm.cy@gmail.com Link: https://inpm.top/2019/03/21/crawler/ Copyright: All articles in this blog are licensed under inpm unless stating additionally.< PreviousPost黑苹果安装NextPost >你永运不会准备好,so Just do it!
明确目标 目标网址 获取节点 模拟请求 解析内容 本地处理 过滤 获取当前页面并解析 绝大部分网站目前都是utf-8 编码格式 解析页面库beautiSoup 针对节点进行分析 基本就是遍历目标节点 基本结构12345678910111213141516171819202122import urllib.requestfrom bs4 import BeautifulSoupimport timeBASE_URL = 'https://xxx.com'if __name__ == '__main__': while 1: #获取当前页面 page = urllib.request.urlopen(url).read().decode('utf-8') soup = BeautifulSoup(page, 'lxml') #分页 lists = soup.find_all('div', class='pg') if len(lists) < 1: break for listone in lists: # 获取每页链接 aLists = listone.find_all('a', attrs={'class':''}) for a in aLists: url = BASE_URL + a['href'] time.sleep(1)