推荐学习书目
› Learn Python the Hard Way
Python Sites
› PyPI - Python Package Index
› http://diveintopython.org/toc/index.html
› Pocoo
值得关注的项目
› PyPy
› Celery
› Jinja2
› Read the Docs
› gevent
› pyenv
› virtualenv
› Stackless Python
› Beautiful Soup
› 结巴中文分词
› Green Unicorn
› Sentry
› Shovel
› Pyflakes
› pytest
Python 编程
› pep8 Checker
Styles
› PEP 8
› Google Python Style Guide
› Code Style from The Hitchhiker's Guide
sundays
V2EX  ›  Python

用循环爬网站子页面, requests,bs4,有什么要注意的么?

  •  
  •   sundays · Mar 9, 2017 · 3116 views
    This topic created in 3496 days ago, the information mentioned may be changed or developed.

    可能是我没配对环境,反正 scrapy 现在用不了,就只能写正常脚本爬,除了加 time.sleep 还有啥注意事项啊,没用过几次 requests 库...

    3 replies  •  2017-03-09 23:48:53 +08:00
    freeminder
        1
    freeminder  
       Mar 9, 2017
    换 UA ,有条件就换 IP ,最好把不是必要的 BS4 解析移出去,比如最终需要的那个页面先别做结构化解析,保留 html 就好。另外注意列表页的互斥,尽量在爬列表页的时候保证列表的内容也没什么重复最好。
    chendajun
        2
    chendajun  
       Mar 9, 2017
    配置 scrapy 环境, linux 比 windows 简单些。
    Akkuman
        3
    Akkuman  
       Mar 9, 2017 via Android
    soup 解析,一般可以一行写完,用列表推导式
    About   ·   Help   ·   Advertise   ·   Blog   ·   API   ·   FAQ   ·   Privacy   ·   Solana   ·   2321 Online   Highest 6679   ·     Select Language
    创意工作者们的社区
    World is powered by solitude
    VERSION: 3.9.8.5 · 29ms · UTC 08:27 · PVG 16:27 · LAX 01:27 · JFK 04:27
    ♥ Do have faith in what you're doing.