Python爬虫包BeautifulSoup异常处理（二） 2024/11/17 饿虎岗资源网

Python爬虫包BeautifulSoup异常处理（二）

编辑：jimmy

日期: 2024/11/17 浏览：1 次

面对网络不稳定，页面更新等问题，很可能出现程序异常的问题，所以我们要对程序进行一些异常处理。大家可能觉得处理异常是一个比较麻烦的活，但在面对复杂网页和任务的时候，无疑成为一个很好的代码习惯。

网页‘404'、‘500'等问题

try:
    html = urlopen('http://www.pmcaff.com/2221')
  except HTTPError as e:
    print(e)

返回的是空网页

if html is None:
    print('没有找到网页')

目标标签在网页中缺失

try:
    #不存在的标签
    content = bsObj.nonExistingTag.anotherTag 
  except AttributeError as e:
    print('没有找到你想要的标签')
  else:
    if content == None:
      print('没有找到你想要的标签')
    else:
      print(content)

实例

if sys.version_info[0] == 2:
  from urllib2 import urlopen # Python 2
  from urllib2 import HTTPError
else:
  from urllib.request import urlopen # Python3
  from urllib.error import HTTPError
from bs4 import BeautifulSoup
import sys


def getTitle(url):
  try:
    html = urlopen(url)
  except HTTPError as e:
    print(e)
    return None
  try:
    bsObj = BeautifulSoup(html.read())
    title = bsObj.body.h1
  except AttributeError as e:
    return None
  return title

title = getTitle("http://www.pythonscraping.com/exercises/exercise1.html")
if title == None:
  print("Title could not be found")
else:
  print(title)

以上全部为本篇文章的全部内容，希望对大家的学习有所帮助，也希望大家多多支持。

一句话新闻
Windows上运行安卓你用过了吗

在去年的5月23日，借助Intel Bridge Technology以及Intel Celadon两项技术的驱动，Intel为PC用户带来了Android On Windows（AOW）平台，并携手国内软件公司腾讯共同推出了腾讯应用宝电脑版，将Windows与安卓两大生态进行了融合，PC的使用体验随即被带入到了一个全新的阶段。

Python爬虫包BeautifulSoup异常处理（二）

最新资源