禹凕头像
关注
机器学习之数据解析(Machine Learning about Analysis of data)封面图

机器学习之数据解析(Machine Learning about Analysis of data)

二、爬虫自动化——数据解析(Analysis of data)

1、BeautifulSoup解析HTML数据

1.1 HTML解析

当使用 request 库获取网页内容时,返回 HTML 源码,HTML 是一种标记语言,相当于组织结构。

使用 HTML 能够快速地阅读和提取对应的文本信息。能够直接使用正则表达式对其进行字符串处理,但是若是当数据变得复杂,那么 HTML 结构也会变得复杂难以进行维护。

因此 BeautifulSoup 库函数的出现变成了 Python 最流行的 HTML/XML 解析库之一。其能够将 HTML 变成一棵 DOM 树(文档对象模型树),使我们能够向遍历文件一样遍历 HTML 结构。

💡BeautifulSoup库核心用途:解析 HTML/XML 文档,通过 DOM 树结构搜索和提取到数据,自动处理破损的 HTML

1.2 Beautiful下载库与基础用法

Beautiful下载库:

💡安装bs4库的控制台语句:pip install beautifulsoup4

基础使用流程:

import requests
import re
from bs4 import BeautifulSoup

# requests 获取网页,返回的就是 HTML 源码(字符串)
html = requests.get("https://example.com").text

# 数据简单时,正则也能处理:直接按字符串模式抓取
titles = re.findall(r'<h2[^>]*>(.*?)</h2>', html)
print(titles)

# 结构复杂后,BeautifulSoup 把 HTML 变成 DOM 树,逐层向下遍历
soup = BeautifulSoup(html, "html.parser")
for post in soup.select("div.post"):              # 先定位每个文章块
    print(post.select_one("h2.title").text)       # 再沿树向下取标题节点

1.3 文本解析器 lxml

💡文本解析器库下载:pip install lxml

日常推荐使用 lxml 解析器,速度快并且容错性高

html = '<div>test</div>'
soup = BeautifuiSoup(html, 'lxml')

print('lxml:', soup) # 查看解析结果

1.4 BeautifulSoup 库的核心方法详解

(1)soup.select() - 返回所有匹配元素的列表

from bs4 import BeautifulSoup

html = """
<div class="post">
    <h2 class="title">第一篇</h2>
    <h2 class="title">第二篇</h2>
</div>
"""
soup = BeautifulSoup(html, "lxml")

titles = soup.select("h2.title")
print(titles)          # [<h2 class="title">第一篇</h2>, <h2 class="title">第二篇</h2>]
print(len(titles))     # 2,可用于 for 循环遍历

(2)soup.select_one() - 返回匹配到的第一个元素(或 None)

from bs4 import BeautifulSoup

html = """
<div class="post">
    <h2 class="title">第一篇</h2>
    <h2 class="title">第二篇</h2>
</div>
"""
soup = BeautifulSoup(html, "html.parser")

first = soup.select_one("h2.title")
print(first.text)      # 第一篇

nothing = soup.select_one("h3")
print(nothing)         # 返回 None

1.5 提取元素属性和文本

方法作用示例
.text获取元素内的所有文本soup.select_one('span').text
.get('属性名')获取元素的某个属性值soup.select_one('span').get('href')
['属性名']与.get('属性名‘)效果一样soup.select_one('span')['href']
from bs4 import BeautifulSoup

html = """
<a class="link" href="https://example.com">
    <span>点击这里</span>
    去官网
</a>
"""
soup = BeautifulSoup(html, "html.parser")
a = soup.select_one("a.link")

# .text:取元素内的【所有】文本,含子元素的文本
print(a.text)              # 点击这里\n去官网(span 的文本也被拼进来)

# .get('href'):取属性值;属性不存在时返回 None,不报错
print(a.get("href"))       # https://example.com
print(a.get("target"))     # None

# ['href']:效果相同;但属性不存在时抛 KeyError
print(a["href"])           # https://example.com
# print(a["target"])       # KeyError: 'target'

1.6 嵌套处理

html = '''
<div class="item">
    <h3 class="name">商品A</h3>
    <span class="price">99元</span>
    <a href="/p/1">链接</a>
</div>
<div class="item">
    <h3 class="name">商品B</h3>
    <span class="price">199元</span>
    <a href="/p/2">链接</a>
</div>
'''
soup = BeautifulSoup(html, "lxml")

# 第一步:select() 一次性拿到所有外层商品块
items = soup.select("div.item")

# 第二步:循环内用 select_one() 相对当前块向下取字段
for item in items:
    title = item.select_one("h3.name").text
    price = item.select_one("span.price").text
    link = item.select_one("a")["href"]
    print(f"{title} - {price} - {link}")

# 输出:
# 商品A - 99元 - /p/1
# 商品B - 199元 - /p/2

1.7 数据处理

  • 处理文本空格与换行(默认空格和换行):strip()
  • 处理文本切割(不包含切割符号,默认空格):split()
  • 处理切割后文本链接:''.join(s.split())

💡常见的编码错误:

soup = BeautilfulSoup(html, 'lxml', encoding='utf-8-sig')


2、JSON数据解析

2.1 JSON基础知识

JSON(JavaScript Object Notation)是一种轻量级的数据交换格式。在大多数网站里API接口都是使用JSON数据作为返回结构

为什么使用JSON数据多?主要还是因为其轻量、易读、跨语言兼容性好。

JSON类型Python类型示例
{} 对象dict 字典{‘name’:'Bob'}
[] 数组list 对象[1,2,3]
'string' 对象str 字符串‘Hello, World’
12 整数int 整数12
12.5 浮点数float 浮点数12.5
true / falseTrue / Falsetrue
nullNonenull

2.2 JSON与Python之间的转换

import json
# ========== 1. JSON字符串 -> Python对象 ==========
json_str = '{"name": "Alice", "age": 25, "is_student": false}'
data = json.loads(json_str) # loads = load string
print(type(data)) # <class 'dict'>
print(data['name']) # Alice
print(data['age']) # 25
print(data['is_student']) # False
# ========== 2. Python对象 -> JSON字符串 ==========
data = {
    'name': 'Bob',
    'score': 92.5,
    'hobbies': ['读书', '跑步']
}
json_str = json.dumps(data, ensure_ascii=False, indent=2)
print(json_str)
💡ensure_ascii=False确保中文不被转义为\uXXXX;indent=2使输出格式化,便于阅读。

2.3 JSON文件的读写

# 写入JSON文件 json.dump
data = {'name': '测试', 'value': 100}
with open('data.json', 'w', encoding='utf-8') as f:
json.dump(data, f, ensure_ascii=False, indent=2)
# 读取JSON文件 json.load
with open('data.json', 'r', encoding='utf-8') as f:
loaded_data = json.load(f)
print(loaded_data)

2.4 处理 requests 响应的JSON

import requests
# requests库的Response对象有json()方法,自动解析JSON
resp = requests.get('https://jsonplaceholder.typicode.com/posts/1')
print(resp.status_code) # 200
# 如果响应是JSON格式,直接用resp.json()
data = resp.json()
print(data['title'])

2.5 JSON嵌套遍历

data = {
    "code": 200,
    "data": {
        "user": {
            "name": "李四",
            "email": "[email protected]"
        },
        "orders": [
            {"id": 1, "product": "iPhone", "price": 5999},
            {"id": 2, "product": "AirPods", "price": 1899}
        ],
        "statistics": {
            "total_orders": 2,
            "total_amount": 7898
        }
    }
}

# 第一步:逐层访问用户信息(字典套字典,路径 = 层级)
print('用户名:', data['data']['user']['name'])   # 李四
print('邮箱:', data['data']['user']['email'])    # [email protected]

# 第二步:遍历订单列表(字典套列表,循环取每个元素)
print('\n订单明细:')
for order in data['data']['orders']:
    order_id = order['id']
    product = order['product']
    price = order['price']
    print(f'  订单{order_id}: {product} - ¥{price}')
# 输出:
#   订单1: iPhone - ¥5999
#   订单2: AirPods - ¥1899

# 第三步:先取子字典再访问,避免重复写长路径
stats = data['data']['statistics']
print(f'总订单数: {stats["total_orders"]}')     # 2
print(f'消费总额: ¥{stats["total_amount"]}')    # ¥7898

2.6 JSON处理数据技巧

# 1. 使用get()方法安全访问,避免KeyError
user = data.get('data', {}).get('user', {})
name = user.get('name', '未知') # 如果不存在,返回默认值
# 2. 列表推导式提取特定字段
prices = [order['price'] for order in data['data']['orders']]
print('所有商品价格:', prices)
# 3. 条件筛选
expensive_orders = [o for o in data['data']['orders'] if o['price'] > 2000]
print('2000元以上的订单:', expensive_orders)
# 4. 求和计算
total = sum(order['price'] for order in data['data']['orders'])
print(f'消费总额: ¥{total}')

2.7 常见错误

  • 访问不存在的键
data = '{"name": "Bob"}'
parsed = json.loads(data)
# print(parsed['age']) # KeyError!
# 安全做法
print(parsed.get('age', 0)) # 输出: 0
  • resp.json不是json格式
resp = requests.get('https://example.com') # 返回HTML,不是JSON
# data = resp.json() # 抛出异常
# 先检查Content-Type
if 'application/json' in resp.headers.get('Content-Type', ''):
    data = resp.json()
else:
    print('不是JSON响应')

转载自 CSDN-专业IT技术社区

原文链接:https://blog.csdn.net/2301_81858235/article/details/166644637

文章来源转载

评论

赞0

评论列表

微信小程序
QQ小程序

关于作者

点赞数:0
关注数:0
粉丝:0
文章:0
关注标签:0
加入于:--