python提取页面内url列表的方法

yipeiwu_com6年前 (2020-03-06)Python基础

本文实例讲述了python提取页面内url列表的方法。分享给大家供大家参考。具体实现方法如下：

from bs4 import BeautifulSoup
import time,re,urllib2
t=time.time()
websiteurls={}
def scanpage(url):
  websiteurl=url
  t=time.time()
  n=0
  html=urllib2.urlopen(websiteurl).read()
  soup=BeautifulSoup(html)
  pageurls=[]
  Upageurls={}
  pageurls=soup.find_all("a",href=True)
  for links in pageurls:
    if websiteurl in links.get("href") and links.get("href") not in Upageurls and links.get("href") not in websiteurls:
      Upageurls[links.get("href")]=0
  for links in Upageurls.keys():
    try:
      urllib2.urlopen(links).getcode()
    except:
      print "connect failed"
    else:
      t2=time.time()
      Upageurls[links]=urllib2.urlopen(links).getcode()
      print n,
      print links,
      print Upageurls[links]
      t1=time.time()
      print t1-t2
    n+=1
  print ("total is "+repr(n)+" links")
  print time.time()-t
scanpage("http://news.163.com/")

希望本文所述对大家的Python程序设计有所帮助。

返回列表

上一篇：Python字符转换

下一篇：PHP生成静态页面详解

python3实现单目标粒子群算法

本文实例为大家分享了python3单目标粒子群算法的具体代码，供大家参考，具体内容如下关于PSO的基本知识......就说一下算法流程 1) 初始化粒子群； ...

python利用dlib获取人脸的68个landmark

（1）单人脸情况 import cv2 import dlib path = "1.jpg" img = cv2.imread(path) gray = cv2.cvtColo...

详解python中init方法和随机数方法

1、__init__方法的使用 2、random方法的使用在python中，有一些方法是特殊的，是以两个下划线开始，两个下划线结束，定义类，最常用的方法就是__init__()方法，这...

python实现换位加密算法的示例

如下所示： def translationCipher(msg,key): result = [""]*key for i in range(key):#把每一列元素按...

Python lambda函数基本用法实例分析

本文实例讲述了Python lambda函数基本用法。分享给大家供大家参考，具体如下：这里我们简单学习一下python lambda函数。首先，看一下python lambda函数的...

宜配屋

python提取页面内url列表的方法

相关文章

python3实现单目标粒子群算法

python利用dlib获取人脸的68个landmark

详解python中init方法和随机数方法

python实现换位加密算法的示例

Python lambda函数基本用法实例分析

© YiPeiWu.com 【宜配屋】粤ICP备17031333号

Powered By Z-BlogPHP. Theme by TOYEAN.

宜配屋

python提取页面内url列表的方法

相关文章

python3实现单目标粒子群算法

python利用dlib获取人脸的68个landmark

详解python中init方法和随机数方法

python实现换位加密算法的示例

Python lambda函数基本用法实例分析

© YiPeiWu.com 【宜配屋】 粤ICP备17031333号 var _hmt = _hmt || [];(function() { var hm = document.createElement("script"); hm.src = "https://hm.baidu.com/hm.js?8aa60ae04b767b2af31903508928acc0"; var s = document.getElementsByTagName("script")[0]; s.parentNode.insertBefore(hm, s);})();

Powered By Z-BlogPHP. Theme by TOYEAN.

© YiPeiWu.com 【宜配屋】粤ICP备17031333号