inpm
! dlrow olleH

爬虫

2019-03-21 python
Word count: 156 | Reading time: 1min

明确目标

  • 目标网址
  • 获取节点
  • 模拟请求
  • 解析内容
  • 本地处理
  • 过滤

获取当前页面并解析

  • 绝大部分网站目前都是utf-8 编码格式
  • 解析页面库beautiSoup
  • 针对节点进行分析
  • 基本就是遍历目标节点

基本结构

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
import urllib.request
from bs4 import BeautifulSoup
import time

BASE_URL = 'https://xxx.com'
if __name__ == '__main__':
while 1:
#获取当前页面
page = urllib.request.urlopen(url).read().decode('utf-8')
soup = BeautifulSoup(page, 'lxml')

#分页
lists = soup.find_all('div', class='pg')
if len(lists) < 1:
break
for listone in lists:
# 获取每页链接
aLists = listone.find_all('a', attrs={'class':''})
for a in aLists:
url = BASE_URL + a['href']

time.sleep(1)

Author: inpm.cy@gmail.com

Link: https://inpm.top/2019/03/21/crawler/

Copyright: All articles in this blog are licensed under inpm unless stating additionally.

< PreviousPost
黑苹果安装
NextPost >
你永运不会准备好,so Just do it!
CATALOG
  1. 1. 明确目标
  2. 2. 获取当前页面并解析
  3. 3. 基本结构