mirror of https://github.com/scrapy/scrapy.git
Merge remote-tracking branch 'upstream/master'
This commit is contained in:
commit
a152c066fa
|
|
@ -1,5 +1,5 @@
|
|||
[bumpversion]
|
||||
current_version = 1.2.0dev2
|
||||
current_version = 1.3.0
|
||||
commit = True
|
||||
tag = True
|
||||
tag_name = {new_version}
|
||||
|
|
|
|||
|
|
@ -8,7 +8,7 @@ branches:
|
|||
- /^\d\.\d+\.\d+(rc\d+|dev\d+)?$/
|
||||
env:
|
||||
- TOXENV=py27
|
||||
- TOXENV=precise
|
||||
- TOXENV=jessie
|
||||
- TOXENV=py33
|
||||
- TOXENV=py35
|
||||
- TOXENV=docs
|
||||
|
|
|
|||
|
|
@ -1,24 +1,41 @@
|
|||
# Contributor Code of Conduct
|
||||
# Contributor Covenant Code of Conduct
|
||||
|
||||
As contributors and maintainers of this project, and in the interest of
|
||||
fostering an open and welcoming community, we pledge to respect all people who
|
||||
contribute through reporting issues, posting feature requests, updating
|
||||
documentation, submitting pull requests or patches, and other activities.
|
||||
## Our Pledge
|
||||
|
||||
We are committed to making participation in this project a harassment-free
|
||||
experience for everyone, regardless of level of experience, gender, gender
|
||||
identity and expression, sexual orientation, disability, personal appearance,
|
||||
body size, race, ethnicity, age, religion, or nationality.
|
||||
In the interest of fostering an open and welcoming environment, we as
|
||||
contributors and maintainers pledge to making participation in our project and
|
||||
our community a harassment-free experience for everyone, regardless of age, body
|
||||
size, disability, ethnicity, gender identity and expression, level of experience,
|
||||
nationality, personal appearance, race, religion, or sexual identity and
|
||||
orientation.
|
||||
|
||||
## Our Standards
|
||||
|
||||
Examples of behavior that contributes to creating a positive environment
|
||||
include:
|
||||
|
||||
* Using welcoming and inclusive language
|
||||
* Being respectful of differing viewpoints and experiences
|
||||
* Gracefully accepting constructive criticism
|
||||
* Focusing on what is best for the community
|
||||
* Showing empathy towards other community members
|
||||
|
||||
Examples of unacceptable behavior by participants include:
|
||||
|
||||
* The use of sexualized language or imagery
|
||||
* Personal attacks
|
||||
* Trolling or insulting/derogatory comments
|
||||
* The use of sexualized language or imagery and unwelcome sexual attention or
|
||||
advances
|
||||
* Trolling, insulting/derogatory comments, and personal or political attacks
|
||||
* Public or private harassment
|
||||
* Publishing other's private information, such as physical or electronic
|
||||
addresses, without explicit permission
|
||||
* Other unethical or unprofessional conduct
|
||||
* Publishing others' private information, such as a physical or electronic
|
||||
address, without explicit permission
|
||||
* Other conduct which could reasonably be considered inappropriate in a
|
||||
professional setting
|
||||
|
||||
## Our Responsibilities
|
||||
|
||||
Project maintainers are responsible for clarifying the standards of acceptable
|
||||
behavior and are expected to take appropriate and fair corrective action in
|
||||
response to any instances of unacceptable behavior.
|
||||
|
||||
Project maintainers have the right and responsibility to remove, edit, or
|
||||
reject comments, commits, code, wiki edits, issues, and other contributions
|
||||
|
|
@ -26,25 +43,32 @@ that are not aligned to this Code of Conduct, or to ban temporarily or
|
|||
permanently any contributor for other behaviors that they deem inappropriate,
|
||||
threatening, offensive, or harmful.
|
||||
|
||||
By adopting this Code of Conduct, project maintainers commit themselves to
|
||||
fairly and consistently applying these principles to every aspect of managing
|
||||
this project. Project maintainers who do not follow or enforce the Code of
|
||||
Conduct may be permanently removed from the project team.
|
||||
## Scope
|
||||
|
||||
This Code of Conduct applies both within project spaces and in public spaces
|
||||
when an individual is representing the project or its community.
|
||||
when an individual is representing the project or its community. Examples of
|
||||
representing a project or community include using an official project e-mail
|
||||
address, posting via an official social media account, or acting as an appointed
|
||||
representative at an online or offline event. Representation of a project may be
|
||||
further defined and clarified by project maintainers.
|
||||
|
||||
## Enforcement
|
||||
|
||||
Instances of abusive, harassing, or otherwise unacceptable behavior may be
|
||||
reported by contacting a project maintainer at opensource@scrapinghub.com. All
|
||||
reported by contacting the project team at opensource@scrapinghub.com. All
|
||||
complaints will be reviewed and investigated and will result in a response that
|
||||
is deemed necessary and appropriate to the circumstances. Maintainers are
|
||||
obligated to maintain confidentiality with regard to the reporter of an
|
||||
incident.
|
||||
is deemed necessary and appropriate to the circumstances. The project team is
|
||||
obligated to maintain confidentiality with regard to the reporter of an incident.
|
||||
Further details of specific enforcement policies may be posted separately.
|
||||
|
||||
Project maintainers who do not follow or enforce the Code of Conduct in good
|
||||
faith may face temporary or permanent repercussions as determined by other
|
||||
members of the project's leadership.
|
||||
|
||||
This Code of Conduct is adapted from the [Contributor Covenant][homepage],
|
||||
version 1.3.0, available at
|
||||
[http://contributor-covenant.org/version/1/3/0/][version]
|
||||
## Attribution
|
||||
|
||||
This Code of Conduct is adapted from the [Contributor Covenant][homepage], version 1.4,
|
||||
available at [http://contributor-covenant.org/version/1/4][version]
|
||||
|
||||
[homepage]: http://contributor-covenant.org
|
||||
[version]: http://contributor-covenant.org/version/1/3/0/
|
||||
[version]: http://contributor-covenant.org/version/1/4/
|
||||
|
|
|
|||
|
|
@ -12,3 +12,4 @@ prune docs/build
|
|||
recursive-include extras *
|
||||
recursive-include bin *
|
||||
recursive-include tests *
|
||||
global-exclude __pycache__ *.py[cod]
|
||||
|
|
|
|||
11
README.rst
11
README.rst
|
|
@ -22,6 +22,10 @@ Scrapy
|
|||
:target: http://codecov.io/github/scrapy/scrapy?branch=master
|
||||
:alt: Coverage report
|
||||
|
||||
.. image:: https://anaconda.org/conda-forge/scrapy/badges/version.svg
|
||||
:target: https://anaconda.org/conda-forge/scrapy
|
||||
:alt: Conda Version
|
||||
|
||||
|
||||
Overview
|
||||
========
|
||||
|
|
@ -69,14 +73,17 @@ See http://scrapy.org/community/
|
|||
Contributing
|
||||
============
|
||||
|
||||
See http://doc.scrapy.org/en/master/contributing.html
|
||||
|
||||
Code of Conduct
|
||||
---------------
|
||||
|
||||
Please note that this project is released with a Contributor Code of Conduct
|
||||
(see https://github.com/scrapy/scrapy/blob/master/CODE_OF_CONDUCT.md).
|
||||
|
||||
By participating in this project you agree to abide by its terms.
|
||||
Please report unacceptable behavior to opensource@scrapinghub.com.
|
||||
|
||||
See http://doc.scrapy.org/en/master/contributing.html
|
||||
|
||||
Companies using Scrapy
|
||||
======================
|
||||
|
||||
|
|
|
|||
|
|
@ -317,6 +317,6 @@ I'm scraping a XML document and my XPath selector doesn't return any items
|
|||
You may need to remove namespaces. See :ref:`removing-namespaces`.
|
||||
|
||||
.. _user agents: https://en.wikipedia.org/wiki/User_agent
|
||||
.. _LIFO: https://en.wikipedia.org/wiki/LIFO
|
||||
.. _LIFO: https://en.wikipedia.org/wiki/Stack_(abstract_data_type)
|
||||
.. _DFO order: https://en.wikipedia.org/wiki/Depth-first_search
|
||||
.. _BFO order: https://en.wikipedia.org/wiki/Breadth-first_search
|
||||
|
|
|
|||
|
|
@ -13,13 +13,15 @@ Having trouble? We'd like to help!
|
|||
|
||||
* Try the :doc:`FAQ <faq>` -- it's got answers to some common questions.
|
||||
* Looking for specific information? Try the :ref:`genindex` or :ref:`modindex`.
|
||||
* Ask or search questions in `StackOverflow using the scrapy tag`_,
|
||||
* Search for information in the `archives of the scrapy-users mailing list`_, or
|
||||
`post a question`_.
|
||||
* Ask a question in the `#scrapy IRC channel`_.
|
||||
* Ask a question in the `#scrapy IRC channel`_,
|
||||
* Report bugs with Scrapy in our `issue tracker`_.
|
||||
|
||||
.. _archives of the scrapy-users mailing list: https://groups.google.com/forum/#!forum/scrapy-users
|
||||
.. _post a question: https://groups.google.com/forum/#!forum/scrapy-users
|
||||
.. _StackOverflow using the scrapy tag: https://stackoverflow.com/tags/scrapy
|
||||
.. _#scrapy IRC channel: irc://irc.freenode.net/scrapy
|
||||
.. _issue tracker: https://github.com/scrapy/scrapy/issues
|
||||
|
||||
|
|
@ -153,7 +155,6 @@ Solving specific problems
|
|||
topics/firebug
|
||||
topics/leaks
|
||||
topics/media-pipeline
|
||||
topics/ubuntu
|
||||
topics/deploy
|
||||
topics/autothrottle
|
||||
topics/benchmarking
|
||||
|
|
@ -186,9 +187,6 @@ Solving specific problems
|
|||
:doc:`topics/media-pipeline`
|
||||
Download files and/or images associated with your scraped items.
|
||||
|
||||
:doc:`topics/ubuntu`
|
||||
Install latest Scrapy packages easily on Ubuntu
|
||||
|
||||
:doc:`topics/deploy`
|
||||
Deploying your Scrapy spiders and run them in a remote server.
|
||||
|
||||
|
|
|
|||
|
|
@ -5,21 +5,16 @@ Examples
|
|||
========
|
||||
|
||||
The best way to learn is with examples, and Scrapy is no exception. For this
|
||||
reason, there is an example Scrapy project named dirbot_, that you can use to
|
||||
play and learn more about Scrapy. It contains the dmoz spider described in the
|
||||
tutorial.
|
||||
reason, there is an example Scrapy project named quotesbot_, that you can use to
|
||||
play and learn more about Scrapy. It contains two spiders for
|
||||
http://quotes.toscrape.com, one using CSS selectors and another one using XPath
|
||||
expressions.
|
||||
|
||||
This dirbot_ project is available at: https://github.com/scrapy/dirbot
|
||||
|
||||
It contains a README file with a detailed description of the project contents.
|
||||
The quotesbot_ project is available at: https://github.com/scrapy/quotesbot.
|
||||
You can find more information about it in the project's README.
|
||||
|
||||
If you're familiar with git, you can checkout the code. Otherwise you can
|
||||
download a tarball or zip file of the project by clicking on `Downloads`_.
|
||||
download the project as a zip file by clicking
|
||||
`here <https://github.com/scrapy/quotesbot/archive/master.zip>`_.
|
||||
|
||||
The `scrapy tag on Snipplr`_ is used for sharing code snippets such as spiders,
|
||||
middlewares, extensions, or scripts. Feel free (and encouraged!) to share any
|
||||
code there.
|
||||
|
||||
.. _dirbot: https://github.com/scrapy/dirbot
|
||||
.. _Downloads: https://github.com/scrapy/dirbot/downloads
|
||||
.. _scrapy tag on Snipplr: http://snipplr.com/all/tags/scrapy/
|
||||
.. _quotesbot: https://github.com/scrapy/quotesbot
|
||||
|
|
|
|||
|
|
@ -7,48 +7,104 @@ Installation guide
|
|||
Installing Scrapy
|
||||
=================
|
||||
|
||||
.. note:: Check :ref:`intro-install-platform-notes` first.
|
||||
Scrapy runs on Python 2.7 and Python 3.3 or above
|
||||
(except on Windows where Python 3 is not supported yet).
|
||||
|
||||
The installation steps assume that you have the following things installed:
|
||||
If you’re already familiar with installation of Python packages,
|
||||
you can install Scrapy and its dependencies from PyPI with::
|
||||
|
||||
* `Python`_ 2.7 or above 3.3
|
||||
pip install Scrapy
|
||||
|
||||
* `pip`_ and `setuptools`_ Python packages. Nowadays `pip`_ requires and
|
||||
installs `setuptools`_ if not installed. Python 2.7.9 and later include
|
||||
`pip`_ by default, so you may have it already.
|
||||
We strongly recommend that you install Scrapy in :ref:`a dedicated virtualenv <intro-using-virtualenv>`,
|
||||
to avoid conflicting with your system packages.
|
||||
|
||||
* `lxml`_. Most Linux distributions ships prepackaged versions of lxml.
|
||||
Otherwise refer to http://lxml.de/installation.html
|
||||
For more detailed and platform specifics instructions, read on.
|
||||
|
||||
* `OpenSSL`_. This comes preinstalled in all operating systems, except Windows
|
||||
where the Python installer ships it bundled.
|
||||
|
||||
You can install Scrapy using pip (which is the canonical way to install Python
|
||||
packages). To install using ``pip`` run::
|
||||
Things that are good to know
|
||||
----------------------------
|
||||
|
||||
Scrapy is written in pure Python and depends on a few key Python packages (among others):
|
||||
|
||||
* `lxml`_, an efficient XML and HTML parser
|
||||
* `parsel`_, an HTML/XML data extraction library written on top of lxml,
|
||||
* `w3lib`_, a multi-purpose helper for dealing with URLs and web page encodings
|
||||
* `twisted`_, an asynchronous networking framework
|
||||
* `cryptography`_ and `pyOpenSSL`_, to deal with various network-level security needs
|
||||
|
||||
The minimal versions which Scrapy is tested against are:
|
||||
|
||||
* Twisted 14.0
|
||||
* lxml 3.4
|
||||
* pyOpenSSL 0.14
|
||||
|
||||
Scrapy may work with older versions of these packages
|
||||
but it is not guaranteed it will continue working
|
||||
because it’s not being tested against them.
|
||||
|
||||
Some of these packages themselves depends on non-Python packages
|
||||
that might require additional installation steps depending on your platform.
|
||||
Please check :ref:`platform-specific guides below <intro-install-platform-notes>`.
|
||||
|
||||
In case of any trouble related to these dependencies,
|
||||
please refer to their respective installation instructions:
|
||||
|
||||
* `lxml installation`_
|
||||
* `cryptography installation`_
|
||||
|
||||
.. _lxml installation: http://lxml.de/installation.html
|
||||
.. _cryptography installation: https://cryptography.io/en/latest/installation/
|
||||
|
||||
|
||||
.. _intro-using-virtualenv:
|
||||
|
||||
Using a virtual environment (recommended)
|
||||
-----------------------------------------
|
||||
|
||||
TL;DR: We recommend installing Scrapy inside a virtual environment
|
||||
on all platforms.
|
||||
|
||||
Python packages can be installed either globally (a.k.a system wide),
|
||||
or in user-space. We do not recommend installing scrapy system wide.
|
||||
|
||||
Instead, we recommend that you install scrapy within a so-called
|
||||
"virtual environment" (`virtualenv`_).
|
||||
Virtualenvs allow you to not conflict with already-installed Python
|
||||
system packages (which could break some of your system tools and scripts),
|
||||
and still install packages normally with ``pip`` (without ``sudo`` and the likes).
|
||||
|
||||
To get started with virtual environments, see `virtualenv installation instructions`_.
|
||||
To install it globally (having it globally installed actually helps here),
|
||||
it should be a matter of running::
|
||||
|
||||
$ [sudo] pip install virtualenv
|
||||
|
||||
Check this `user guide`_ on how to create your virtualenv.
|
||||
|
||||
.. note::
|
||||
If you use Linux or OS X, `virtualenvwrapper`_ is a handy tool to create virtualenvs.
|
||||
|
||||
Once you have created a virtualenv, you can install scrapy inside it with ``pip``,
|
||||
just like any other Python package.
|
||||
(See :ref:`platform-specific guides <intro-install-platform-notes>`
|
||||
below for non-Python dependencies that you may need to install beforehand).
|
||||
|
||||
Python virtualenvs can be created to use Python 2 by default, or Python 3 by default.
|
||||
|
||||
* If you want to install scrapy with Python 3, install scrapy within a Python 3 virtualenv.
|
||||
* And if you want to install scrapy with Python 2, install scrapy within a Python 2 virtualenv.
|
||||
|
||||
.. _virtualenv: https://virtualenv.pypa.io
|
||||
.. _virtualenv installation instructions: https://virtualenv.pypa.io/en/stable/installation/
|
||||
.. _virtualenvwrapper: http://virtualenvwrapper.readthedocs.io/en/latest/install.html
|
||||
.. _user guide: https://virtualenv.pypa.io/en/stable/userguide/
|
||||
|
||||
pip install Scrapy
|
||||
|
||||
.. _intro-install-platform-notes:
|
||||
|
||||
Platform specific installation notes
|
||||
====================================
|
||||
|
||||
Anaconda
|
||||
--------
|
||||
|
||||
.. note::
|
||||
|
||||
For Windows users, or if you have issues installing through `pip`, this is
|
||||
the recommended way to install Scrapy.
|
||||
|
||||
If you already have installed `Anaconda`_ or `Miniconda`_, the company
|
||||
`Scrapinghub`_ maintains official conda packages for Linux, Windows and OS X.
|
||||
|
||||
To install Scrapy using ``conda``, run::
|
||||
|
||||
conda install -c scrapinghub scrapy
|
||||
|
||||
|
||||
Windows
|
||||
-------
|
||||
|
||||
|
|
@ -76,7 +132,7 @@ Windows
|
|||
* *(Only required for Python<2.7.9)* Install `pip`_ from
|
||||
https://pip.pypa.io/en/latest/installing/
|
||||
|
||||
Now open a Command prompt to check ``pip`` is installed correctly::
|
||||
Now open a Command prompt to check ``pip`` is installed correctly::
|
||||
|
||||
pip --version
|
||||
|
||||
|
|
@ -89,37 +145,40 @@ Windows
|
|||
Python 3 is not supported on Windows. This is because Scrapy core requirement Twisted does not support
|
||||
Python 3 on Windows.
|
||||
|
||||
Ubuntu 9.10 or above
|
||||
--------------------
|
||||
Ubuntu 12.04 or above
|
||||
---------------------
|
||||
|
||||
Scrapy is currently tested with recent-enough versions of lxml,
|
||||
twisted and pyOpenSSL, and is compatible with recent Ubuntu distributions.
|
||||
But it should support older versions of Ubuntu too, like Ubuntu 12.04,
|
||||
albeit with potential issues with TLS connections.
|
||||
|
||||
**Don't** use the ``python-scrapy`` package provided by Ubuntu, they are
|
||||
typically too old and slow to catch up with latest Scrapy.
|
||||
|
||||
Instead, use the official :ref:`Ubuntu Packages <topics-ubuntu>`, which already
|
||||
solve all dependencies for you and are continuously updated with the latest bug
|
||||
fixes.
|
||||
|
||||
If you prefer to build the python dependencies locally instead of relying on
|
||||
system packages you'll need to install their required non-python dependencies
|
||||
first::
|
||||
To install scrapy on Ubuntu (or Ubuntu-based) systems, you need to install
|
||||
these dependencies::
|
||||
|
||||
sudo apt-get install python-dev python-pip libxml2-dev libxslt1-dev zlib1g-dev libffi-dev libssl-dev
|
||||
|
||||
You can install Scrapy with ``pip`` after that::
|
||||
- ``python-dev``, ``zlib1g-dev``, ``libxml2-dev`` and ``libxslt1-dev``
|
||||
are required for ``lxml``
|
||||
- ``libssl-dev`` and ``libffi-dev`` are required for ``cryptography``
|
||||
|
||||
pip install Scrapy
|
||||
If you want to install scrapy on Python 3, you’ll also need Python 3 development headers::
|
||||
|
||||
sudo apt-get install python3 python3-dev
|
||||
|
||||
Inside a :ref:`virtualenv <intro-using-virtualenv>`,
|
||||
you can install Scrapy with ``pip`` after that::
|
||||
|
||||
pip install scrapy
|
||||
|
||||
.. note::
|
||||
|
||||
The same non-python dependencies can be used to install Scrapy in Debian
|
||||
Wheezy (7.0) and above.
|
||||
|
||||
Archlinux
|
||||
---------
|
||||
|
||||
You can follow the generic instructions or install Scrapy from `AUR Scrapy package`::
|
||||
|
||||
yaourt -S scrapy
|
||||
|
||||
Mac OS X
|
||||
--------
|
||||
|
|
@ -174,17 +233,39 @@ After any of these workarounds you should be able to install Scrapy::
|
|||
|
||||
pip install Scrapy
|
||||
|
||||
|
||||
Anaconda
|
||||
--------
|
||||
|
||||
|
||||
Using Anaconda is an alternative to using a virtualenv and installing with ``pip``.
|
||||
|
||||
.. note::
|
||||
|
||||
For Windows users, or if you have issues installing through ``pip``, this is
|
||||
the recommended way to install Scrapy.
|
||||
|
||||
If you already have `Anaconda`_ or `Miniconda`_ installed, the `conda-forge`_
|
||||
community have up-to-date packages for Linux, Windows and OS X.
|
||||
|
||||
To install Scrapy using ``conda``, run::
|
||||
|
||||
conda install -c conda-forge scrapy
|
||||
|
||||
.. _Python: https://www.python.org/
|
||||
.. _pip: https://pip.pypa.io/en/latest/installing/
|
||||
.. _easy_install: https://pypi.python.org/pypi/setuptools
|
||||
.. _Control Panel: https://www.microsoft.com/resources/documentation/windows/xp/all/proddocs/en-us/sysdm_advancd_environmnt_addchange_variable.mspx
|
||||
.. _lxml: http://lxml.de/
|
||||
.. _OpenSSL: https://pypi.python.org/pypi/pyOpenSSL
|
||||
.. _parsel: https://pypi.python.org/pypi/parsel
|
||||
.. _w3lib: https://pypi.python.org/pypi/w3lib
|
||||
.. _twisted: https://twistedmatrix.com/
|
||||
.. _cryptography: https://cryptography.io/
|
||||
.. _pyOpenSSL: https://pypi.python.org/pypi/pyOpenSSL
|
||||
.. _setuptools: https://pypi.python.org/pypi/setuptools
|
||||
.. _AUR Scrapy package: https://aur.archlinux.org/packages/scrapy/
|
||||
.. _homebrew: http://brew.sh/
|
||||
.. _zsh: http://www.zsh.org/
|
||||
.. _virtualenv: https://virtualenv.pypa.io/en/latest/
|
||||
.. _Scrapinghub: http://scrapinghub.com
|
||||
.. _Anaconda: http://docs.continuum.io/anaconda/index
|
||||
.. _Miniconda: http://conda.pydata.org/docs/install/quick.html
|
||||
.. _conda-forge: https://conda-forge.github.io/
|
||||
|
|
|
|||
|
|
@ -19,74 +19,69 @@ Walk-through of an example spider
|
|||
In order to show you what Scrapy brings to the table, we'll walk you through an
|
||||
example of a Scrapy Spider using the simplest way to run a spider.
|
||||
|
||||
So, here's the code for a spider that follows the links to the top
|
||||
voted questions on StackOverflow and scrapes some data from each page::
|
||||
Here's the code for a spider that scrapes famous quotes from website
|
||||
http://quotes.toscrape.com, following the pagination::
|
||||
|
||||
import scrapy
|
||||
|
||||
|
||||
class StackOverflowSpider(scrapy.Spider):
|
||||
name = 'stackoverflow'
|
||||
start_urls = ['http://stackoverflow.com/questions?sort=votes']
|
||||
class QuotesSpider(scrapy.Spider):
|
||||
name = "quotes"
|
||||
start_urls = [
|
||||
'http://quotes.toscrape.com/tag/humor/',
|
||||
]
|
||||
|
||||
def parse(self, response):
|
||||
for href in response.css('.question-summary h3 a::attr(href)'):
|
||||
full_url = response.urljoin(href.extract())
|
||||
yield scrapy.Request(full_url, callback=self.parse_question)
|
||||
for quote in response.css('div.quote'):
|
||||
yield {
|
||||
'text': quote.css('span.text::text').extract_first(),
|
||||
'author': quote.xpath('span/small/text()').extract_first(),
|
||||
}
|
||||
|
||||
def parse_question(self, response):
|
||||
yield {
|
||||
'title': response.css('h1 a::text').extract_first(),
|
||||
'votes': response.css('.question .vote-count-post::text').extract_first(),
|
||||
'body': response.css('.question .post-text').extract_first(),
|
||||
'tags': response.css('.question .post-tag::text').extract(),
|
||||
'link': response.url,
|
||||
}
|
||||
next_page = response.css('li.next a::attr("href")').extract_first()
|
||||
if next_page is not None:
|
||||
next_page = response.urljoin(next_page)
|
||||
yield scrapy.Request(next_page, callback=self.parse)
|
||||
|
||||
|
||||
Put this in a file, name it to something like ``stackoverflow_spider.py``
|
||||
Put this in a text file, name it to something like ``quotes_spider.py``
|
||||
and run the spider using the :command:`runspider` command::
|
||||
|
||||
scrapy runspider stackoverflow_spider.py -o top-stackoverflow-questions.json
|
||||
scrapy runspider quotes_spider.py -o quotes.json
|
||||
|
||||
|
||||
When this finishes you will have in the ``top-stackoverflow-questions.json`` file
|
||||
a list of the most upvoted questions in StackOverflow in JSON format, containing the
|
||||
title, link, number of upvotes, a list of the tags and the question content in HTML,
|
||||
looking like this (reformatted for easier reading)::
|
||||
When this finishes you will have in the ``quotes.json`` file a list of the
|
||||
quotes in JSON format, containing text and author, looking like this (reformatted
|
||||
here for better readability)::
|
||||
|
||||
[{
|
||||
"body": "... LONG HTML HERE ...",
|
||||
"link": "http://stackoverflow.com/questions/11227809/why-is-processing-a-sorted-array-faster-than-an-unsorted-array",
|
||||
"tags": ["java", "c++", "performance", "optimization"],
|
||||
"title": "Why is processing a sorted array faster than an unsorted array?",
|
||||
"votes": "9924"
|
||||
"author": "Jane Austen",
|
||||
"text": "\u201cThe person, be it gentleman or lady, who has not pleasure in a good novel, must be intolerably stupid.\u201d"
|
||||
},
|
||||
{
|
||||
"body": "... LONG HTML HERE ...",
|
||||
"link": "http://stackoverflow.com/questions/1260748/how-do-i-remove-a-git-submodule",
|
||||
"tags": ["git", "git-submodules"],
|
||||
"title": "How do I remove a Git submodule?",
|
||||
"votes": "1764"
|
||||
"author": "Groucho Marx",
|
||||
"text": "\u201cOutside of a dog, a book is man's best friend. Inside of a dog it's too dark to read.\u201d"
|
||||
},
|
||||
{
|
||||
"author": "Steve Martin",
|
||||
"text": "\u201cA day without sunshine is like, you know, night.\u201d"
|
||||
},
|
||||
...]
|
||||
|
||||
|
||||
|
||||
What just happened?
|
||||
-------------------
|
||||
|
||||
When you ran the command ``scrapy runspider somefile.py``, Scrapy looked for a
|
||||
When you ran the command ``scrapy runspider quotes_spider.py``, Scrapy looked for a
|
||||
Spider definition inside it and ran it through its crawler engine.
|
||||
|
||||
The crawl started by making requests to the URLs defined in the ``start_urls``
|
||||
attribute (in this case, only the URL for StackOverflow top questions page)
|
||||
attribute (in this case, only the URL for quotes in *humor* category)
|
||||
and called the default callback method ``parse``, passing the response object as
|
||||
an argument. In the ``parse`` callback we extract the links to the
|
||||
question pages using a CSS Selector with a custom extension that allows to get
|
||||
the value for an attribute. Then we yield a few more requests to be sent,
|
||||
registering the method ``parse_question`` as the callback to be called for each
|
||||
of them as they finish.
|
||||
an argument. In the ``parse`` callback, we loop through the quote elements
|
||||
using a CSS Selector, yield a Python dict with the extracted quote text and author,
|
||||
look for a link to the next page and schedule another request using the same
|
||||
``parse`` method as callback.
|
||||
|
||||
Here you notice one of the main advantages about Scrapy: requests are
|
||||
:ref:`scheduled and processed asynchronously <topics-architecture>`. This
|
||||
|
|
@ -103,10 +98,6 @@ each request, limiting amount of concurrent requests per domain or per IP, and
|
|||
even :ref:`using an auto-throttling extension <topics-autothrottle>` that tries
|
||||
to figure out these automatically.
|
||||
|
||||
Finally, the ``parse_question`` callback scrapes the question data for each
|
||||
page yielding a dict, which Scrapy then collects and writes to a JSON file as
|
||||
requested in the command line.
|
||||
|
||||
.. note::
|
||||
|
||||
This is using :ref:`feed exports <topics-feed-exports>` to generate the
|
||||
|
|
@ -145,12 +136,13 @@ scraping easy and efficient, such as:
|
|||
:ref:`pipelines <topics-item-pipeline>`).
|
||||
|
||||
* Wide range of built-in extensions and middlewares for handling:
|
||||
* cookies and session handling
|
||||
* HTTP features like compression, authentication, caching
|
||||
* user-agent spoofing
|
||||
* robots.txt
|
||||
* crawl depth restriction
|
||||
* and more
|
||||
|
||||
- cookies and session handling
|
||||
- HTTP features like compression, authentication, caching
|
||||
- user-agent spoofing
|
||||
- robots.txt
|
||||
- crawl depth restriction
|
||||
- and more
|
||||
|
||||
* A :ref:`Telnet console <topics-telnetconsole>` for hooking into a Python
|
||||
console running inside your Scrapy process, to introspect and debug your
|
||||
|
|
@ -165,8 +157,8 @@ What's next?
|
|||
============
|
||||
|
||||
The next steps for you are to :ref:`install Scrapy <intro-install>`,
|
||||
:ref:`follow through the tutorial <intro-tutorial>` to learn how to organize
|
||||
your code in Scrapy projects and `join the community`_. Thanks for your
|
||||
:ref:`follow through the tutorial <intro-tutorial>` to learn how to create
|
||||
a full-blown Scrapy project and `join the community`_. Thanks for your
|
||||
interest!
|
||||
|
||||
.. _join the community: http://scrapy.org/community/
|
||||
|
|
|
|||
File diff suppressed because it is too large
Load Diff
361
docs/news.rst
361
docs/news.rst
|
|
@ -3,8 +3,193 @@
|
|||
Release notes
|
||||
=============
|
||||
|
||||
1.1.2 (2016-08-18)
|
||||
------------------
|
||||
Scrapy 1.3.0 (2016-12-21)
|
||||
-------------------------
|
||||
|
||||
This release comes rather soon after 1.2.2 for one main reason:
|
||||
it was found out that releases since 0.18 up to 1.2.2 (included) use
|
||||
some backported code from Twisted (``scrapy.xlib.tx.*``),
|
||||
even if newer Twisted modules are available.
|
||||
Scrapy now uses ``twisted.web.client`` and ``twisted.internet.endpoints`` directly.
|
||||
(See also cleanups below.)
|
||||
|
||||
As it is a major change, we wanted to get the bug fix out quickly
|
||||
while not breaking any projects using the 1.2 series.
|
||||
|
||||
New Features
|
||||
~~~~~~~~~~~~
|
||||
|
||||
- ``MailSender`` now accepts single strings as values for ``to`` and ``cc``
|
||||
arguments (:issue:`2272`)
|
||||
- ``scrapy fetch url``, ``scrapy shell url`` and ``fetch(url)`` inside
|
||||
scrapy shell now follow HTTP redirections by default (:issue:`2290`);
|
||||
See :command:`fetch` and :command:`shell` for details.
|
||||
- ``HttpErrorMiddleware`` now logs errors with ``INFO`` level instead of ``DEBUG``;
|
||||
this is technically **backwards incompatible** so please check your log parsers.
|
||||
- By default, logger names now use a long-form path, e.g. ``[scrapy.extensions.logstats]``,
|
||||
instead of the shorter "top-level" variant of prior releases (e.g. ``[scrapy]``);
|
||||
this is **backwards incompatible** if you have log parsers expecting the short
|
||||
logger name part. You can switch back to short logger names using :setting:`LOG_SHORT_NAMES`
|
||||
set to ``True``.
|
||||
|
||||
Dependencies & Cleanups
|
||||
~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
- Scrapy now requires Twisted >= 13.1 which is the case for many Linux
|
||||
distributions already.
|
||||
- As a consequence, we got rid of ``scrapy.xlib.tx.*`` modules, which
|
||||
copied some of Twisted code for users stuck with an "old" Twisted version
|
||||
- ``ChunkedTransferMiddleware`` is deprecated and removed from the default
|
||||
downloader middlewares.
|
||||
|
||||
|
||||
Scrapy 1.2.2 (2016-12-06)
|
||||
-------------------------
|
||||
|
||||
Bug fixes
|
||||
~~~~~~~~~
|
||||
|
||||
- Fix a cryptic traceback when a pipeline fails on ``open_spider()`` (:issue:`2011`)
|
||||
- Fix embedded IPython shell variables (fixing :issue:`396` that re-appeared
|
||||
in 1.2.0, fixed in :issue:`2418`)
|
||||
- A couple of patches when dealing with robots.txt:
|
||||
|
||||
- handle (non-standard) relative sitemap URLs (:issue:`2390`)
|
||||
- handle non-ASCII URLs and User-Agents in Python 2 (:issue:`2373`)
|
||||
|
||||
Documentation
|
||||
~~~~~~~~~~~~~
|
||||
|
||||
- Document ``"download_latency"`` key in ``Request``'s ``meta`` dict (:issue:`2033`)
|
||||
- Remove page on (deprecated & unsupported) Ubuntu packages from ToC (:issue:`2335`)
|
||||
- A few fixed typos (:issue:`2346`, :issue:`2369`, :issue:`2369`, :issue:`2380`)
|
||||
and clarifications (:issue:`2354`, :issue:`2325`, :issue:`2414`)
|
||||
|
||||
Other changes
|
||||
~~~~~~~~~~~~~
|
||||
|
||||
- Advertize `conda-forge`_ as Scrapy's official conda channel (:issue:`2387`)
|
||||
- More helpful error messages when trying to use ``.css()`` or ``.xpath()``
|
||||
on non-Text Responses (:issue:`2264`)
|
||||
- ``startproject`` command now generates a sample ``middlewares.py`` file (:issue:`2335`)
|
||||
- Add more dependencies' version info in ``scrapy version`` verbose output (:issue:`2404`)
|
||||
- Remove all ``*.pyc`` files from source distribution (:issue:`2386`)
|
||||
|
||||
.. _conda-forge: https://anaconda.org/conda-forge/scrapy
|
||||
|
||||
|
||||
Scrapy 1.2.1 (2016-10-21)
|
||||
-------------------------
|
||||
|
||||
Bug fixes
|
||||
~~~~~~~~~
|
||||
|
||||
- Include OpenSSL's more permissive default ciphers when establishing
|
||||
TLS/SSL connections (:issue:`2314`).
|
||||
- Fix "Location" HTTP header decoding on non-ASCII URL redirects (:issue:`2321`).
|
||||
|
||||
Documentation
|
||||
~~~~~~~~~~~~~
|
||||
|
||||
- Fix JsonWriterPipeline example (:issue:`2302`).
|
||||
- Various notes: :issue:`2330` on spider names,
|
||||
:issue:`2329` on middleware methods processing order,
|
||||
:issue:`2327` on getting multi-valued HTTP headers as lists.
|
||||
|
||||
Other changes
|
||||
~~~~~~~~~~~~~
|
||||
|
||||
- Removed ``www.`` from ``start_urls`` in built-in spider templates (:issue:`2299`).
|
||||
|
||||
|
||||
Scrapy 1.2.0 (2016-10-03)
|
||||
-------------------------
|
||||
|
||||
New Features
|
||||
~~~~~~~~~~~~
|
||||
|
||||
- New :setting:`FEED_EXPORT_ENCODING` setting to customize the encoding
|
||||
used when writing items to a file.
|
||||
This can be used to turn off ``\uXXXX`` escapes in JSON output.
|
||||
This is also useful for those wanting something else than UTF-8
|
||||
for XML or CSV output (:issue:`2034`).
|
||||
- ``startproject`` command now supports an optional destination directory
|
||||
to override the default one based on the project name (:issue:`2005`).
|
||||
- New :setting:`SCHEDULER_DEBUG` setting to log requests serialization
|
||||
failures (:issue:`1610`).
|
||||
- JSON encoder now supports serialization of ``set`` instances (:issue:`2058`).
|
||||
- Interpret ``application/json-amazonui-streaming`` as ``TextResponse`` (:issue:`1503`).
|
||||
- ``scrapy`` is imported by default when using shell tools (:command:`shell`,
|
||||
:ref:`inspect_response <topics-shell-inspect-response>`) (:issue:`2248`).
|
||||
|
||||
Bug fixes
|
||||
~~~~~~~~~
|
||||
|
||||
- DefaultRequestHeaders middleware now runs before UserAgent middleware
|
||||
(:issue:`2088`). **Warning: this is technically backwards incompatible**,
|
||||
though we consider this a bug fix.
|
||||
- HTTP cache extension and plugins that use the ``.scrapy`` data directory now
|
||||
work outside projects (:issue:`1581`). **Warning: this is technically
|
||||
backwards incompatible**, though we consider this a bug fix.
|
||||
- ``Selector`` does not allow passing both ``response`` and ``text`` anymore
|
||||
(:issue:`2153`).
|
||||
- Fixed logging of wrong callback name with ``scrapy parse`` (:issue:`2169`).
|
||||
- Fix for an odd gzip decompression bug (:issue:`1606`).
|
||||
- Fix for selected callbacks when using ``CrawlSpider`` with :command:`scrapy parse <parse>`
|
||||
(:issue:`2225`).
|
||||
- Fix for invalid JSON and XML files when spider yields no items (:issue:`872`).
|
||||
- Implement ``flush()`` fpr ``StreamLogger`` avoiding a warning in logs (:issue:`2125`).
|
||||
|
||||
Refactoring
|
||||
~~~~~~~~~~~
|
||||
|
||||
- ``canonicalize_url`` has been moved to `w3lib.url`_ (:issue:`2168`).
|
||||
|
||||
.. _w3lib.url: http://w3lib.readthedocs.io/en/latest/w3lib.html#w3lib.url.canonicalize_url
|
||||
|
||||
Tests & Requirements
|
||||
~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Scrapy's new requirements baseline is Debian 8 "Jessie". It was previously
|
||||
Ubuntu 12.04 Precise.
|
||||
What this means in practice is that we run continuous integration tests
|
||||
with these (main) packages versions at a minimum:
|
||||
Twisted 14.0, pyOpenSSL 0.14, lxml 3.4.
|
||||
|
||||
Scrapy may very well work with older versions of these packages
|
||||
(the code base still has switches for older Twisted versions for example)
|
||||
but it is not guaranteed (because it's not tested anymore).
|
||||
|
||||
Documentation
|
||||
~~~~~~~~~~~~~
|
||||
|
||||
- Grammar fixes: :issue:`2128`, :issue:`1566`.
|
||||
- Download stats badge removed from README (:issue:`2160`).
|
||||
- New scrapy :ref:`architecture diagram <topics-architecture>` (:issue:`2165`).
|
||||
- Updated ``Response`` parameters documentation (:issue:`2197`).
|
||||
- Reworded misleading :setting:`RANDOMIZE_DOWNLOAD_DELAY` description (:issue:`2190`).
|
||||
- Add StackOverflow as a support channel (:issue:`2257`).
|
||||
|
||||
|
||||
Scrapy 1.1.3 (2016-09-22)
|
||||
-------------------------
|
||||
|
||||
Bug fixes
|
||||
~~~~~~~~~
|
||||
|
||||
- Class attributes for subclasses of ``ImagesPipeline`` and ``FilesPipeline``
|
||||
work as they did before 1.1.1 (:issue:`2243`, fixes :issue:`2198`)
|
||||
|
||||
Documentation
|
||||
~~~~~~~~~~~~~
|
||||
|
||||
- :ref:`Overview <intro-overview>` and :ref:`tutorial <intro-tutorial>`
|
||||
rewritten to use http://toscrape.com websites
|
||||
(:issue:`2236`, :issue:`2249`, :issue:`2252`).
|
||||
|
||||
|
||||
Scrapy 1.1.2 (2016-08-18)
|
||||
-------------------------
|
||||
|
||||
Bug fixes
|
||||
~~~~~~~~~
|
||||
|
|
@ -17,8 +202,8 @@ Bug fixes
|
|||
(the regression was introduced in 1.1.1)
|
||||
|
||||
|
||||
1.1.1 (2016-07-13)
|
||||
------------------
|
||||
Scrapy 1.1.1 (2016-07-13)
|
||||
-------------------------
|
||||
|
||||
Bug fixes
|
||||
~~~~~~~~~
|
||||
|
|
@ -69,8 +254,8 @@ Tests
|
|||
- Upgrade py.test requirement on Travis CI and Pin pytest-cov to 2.2.1 (:issue:`2095`)
|
||||
|
||||
|
||||
1.1.0 (2016-05-11)
|
||||
------------------
|
||||
Scrapy 1.1.0 (2016-05-11)
|
||||
-------------------------
|
||||
|
||||
This 1.1 release brings a lot of interesting features and bug fixes:
|
||||
|
||||
|
|
@ -258,8 +443,8 @@ Bugfixes
|
|||
to same remote host (:issue:`1912`).
|
||||
|
||||
|
||||
1.0.6 (2016-05-04)
|
||||
------------------
|
||||
Scrapy 1.0.6 (2016-05-04)
|
||||
-------------------------
|
||||
|
||||
- FIX: RetryMiddleware is now robust to non-standard HTTP status codes (:issue:`1857`)
|
||||
- FIX: Filestorage HTTP cache was checking wrong modified time (:issue:`1875`)
|
||||
|
|
@ -267,8 +452,8 @@ Bugfixes
|
|||
- DOC: Consistency in selectors examples (:issue:`1869`)
|
||||
|
||||
|
||||
1.0.5 (2016-02-04)
|
||||
------------------
|
||||
Scrapy 1.0.5 (2016-02-04)
|
||||
-------------------------
|
||||
|
||||
- FIX: [Backport] Ignore bogus links in LinkExtractors (fixes :issue:`907`, :commit:`108195e`)
|
||||
- TST: Changed buildbot makefile to use 'pytest' (:commit:`1f3d90a`)
|
||||
|
|
@ -276,8 +461,8 @@ Bugfixes
|
|||
- DOC: Add AjaxCrawlMiddleware to DOWNLOADER_MIDDLEWARES_BASE in settings docs (:commit:`aa94121`)
|
||||
|
||||
|
||||
1.0.4 (2015-12-30)
|
||||
------------------
|
||||
Scrapy 1.0.4 (2015-12-30)
|
||||
-------------------------
|
||||
|
||||
- Ignoring xlib/tx folder, depending on Twisted version. (:commit:`7dfa979`)
|
||||
- Run on new travis-ci infra (:commit:`6e42f0b`)
|
||||
|
|
@ -328,14 +513,14 @@ Bugfixes
|
|||
- Small grammatical change (:commit:`8752294`)
|
||||
- Add openssl version to version command (:commit:`13c45ac`)
|
||||
|
||||
1.0.3 (2015-08-11)
|
||||
------------------
|
||||
Scrapy 1.0.3 (2015-08-11)
|
||||
-------------------------
|
||||
|
||||
- add service_identity to scrapy install_requires (:commit:`cbc2501`)
|
||||
- Workaround for travis#296 (:commit:`66af9cd`)
|
||||
|
||||
1.0.2 (2015-08-06)
|
||||
------------------
|
||||
Scrapy 1.0.2 (2015-08-06)
|
||||
-------------------------
|
||||
|
||||
- Twisted 15.3.0 does not raises PicklingError serializing lambda functions (:commit:`b04dd7d`)
|
||||
- Minor method name fix (:commit:`6f85c7f`)
|
||||
|
|
@ -344,8 +529,8 @@ Bugfixes
|
|||
- Fixed typos (:commit:`a9ae7b0`)
|
||||
- Fix doc reference. (:commit:`7c8a4fe`)
|
||||
|
||||
1.0.1 (2015-07-01)
|
||||
------------------
|
||||
Scrapy 1.0.1 (2015-07-01)
|
||||
-------------------------
|
||||
|
||||
- Unquote request path before passing to FTPClient, it already escape paths (:commit:`cc00ad2`)
|
||||
- include tests/ to source distribution in MANIFEST.in (:commit:`eca227e`)
|
||||
|
|
@ -354,8 +539,8 @@ Bugfixes
|
|||
- DOC remove version suffix from ubuntu package (:commit:`5303c66`)
|
||||
- DOC Update release date for 1.0 (:commit:`c89fa29`)
|
||||
|
||||
1.0.0 (2015-06-19)
|
||||
------------------
|
||||
Scrapy 1.0.0 (2015-06-19)
|
||||
-------------------------
|
||||
|
||||
You will find a lot of new features and bugfixes in this major release. Make
|
||||
sure to check our updated :ref:`overview <intro-overview>` to get a glance of
|
||||
|
|
@ -720,8 +905,8 @@ Code refactoring
|
|||
(:issue:`805`)
|
||||
- rename "sflo" local variables to less cryptic "log_observer" (:issue:`775`)
|
||||
|
||||
0.24.6 (2015-04-20)
|
||||
-------------------
|
||||
Scrapy 0.24.6 (2015-04-20)
|
||||
--------------------------
|
||||
|
||||
- encode invalid xpath with unicode_escape under PY2 (:commit:`07cb3e5`)
|
||||
- fix IPython shell scope issue and load IPython user config (:commit:`2c8e573`)
|
||||
|
|
@ -730,8 +915,8 @@ Code refactoring
|
|||
- Converted sel.xpath() calls to response.xpath() in Extracting the data (:commit:`c2c6d15`)
|
||||
|
||||
|
||||
0.24.5 (2015-02-25)
|
||||
-------------------
|
||||
Scrapy 0.24.5 (2015-02-25)
|
||||
--------------------------
|
||||
|
||||
- Support new _getEndpoint Agent signatures on Twisted 15.0.0 (:commit:`540b9bc`)
|
||||
- DOC a couple more references are fixed (:commit:`b4c454b`)
|
||||
|
|
@ -756,14 +941,14 @@ Code refactoring
|
|||
- Update request-response.rst (:commit:`3f3263d`)
|
||||
- SgmlLinkExtractor - fix for parsing <area> tag with Unicode present (:commit:`49b40f0`)
|
||||
|
||||
0.24.4 (2014-08-09)
|
||||
-------------------
|
||||
Scrapy 0.24.4 (2014-08-09)
|
||||
--------------------------
|
||||
|
||||
- pem file is used by mockserver and required by scrapy bench (:commit:`5eddc68`)
|
||||
- scrapy bench needs scrapy.tests* (:commit:`d6cb999`)
|
||||
|
||||
0.24.3 (2014-08-09)
|
||||
-------------------
|
||||
Scrapy 0.24.3 (2014-08-09)
|
||||
--------------------------
|
||||
|
||||
- no need to waste travis-ci time on py3 for 0.24 (:commit:`8e080c1`)
|
||||
- Update installation docs (:commit:`1d0c096`)
|
||||
|
|
@ -795,23 +980,23 @@ Code refactoring
|
|||
- better testcase for settings.overrides.setdefault (:commit:`e22daaf`)
|
||||
- Using CRLF as line marker according to http 1.1 definition (:commit:`5ec430b`)
|
||||
|
||||
0.24.2 (2014-07-08)
|
||||
-------------------
|
||||
Scrapy 0.24.2 (2014-07-08)
|
||||
--------------------------
|
||||
|
||||
- Use a mutable mapping to proxy deprecated settings.overrides and settings.defaults attribute (:commit:`e5e8133`)
|
||||
- there is not support for python3 yet (:commit:`3cd6146`)
|
||||
- Update python compatible version set to debian packages (:commit:`fa5d76b`)
|
||||
- DOC fix formatting in release notes (:commit:`c6a9e20`)
|
||||
|
||||
0.24.1 (2014-06-27)
|
||||
-------------------
|
||||
Scrapy 0.24.1 (2014-06-27)
|
||||
--------------------------
|
||||
|
||||
- Fix deprecated CrawlerSettings and increase backwards compatibility with
|
||||
.defaults attribute (:commit:`8e3f20a`)
|
||||
|
||||
|
||||
0.24.0 (2014-06-26)
|
||||
-------------------
|
||||
Scrapy 0.24.0 (2014-06-26)
|
||||
--------------------------
|
||||
|
||||
Enhancements
|
||||
~~~~~~~~~~~~
|
||||
|
|
@ -890,15 +1075,15 @@ Bugfixes
|
|||
- Testsuite doesn't require PIL anymore (:issue:`585`)
|
||||
|
||||
|
||||
0.22.2 (released 2014-02-14)
|
||||
----------------------------
|
||||
Scrapy 0.22.2 (released 2014-02-14)
|
||||
-----------------------------------
|
||||
|
||||
- fix a reference to unexistent engine.slots. closes #593 (:commit:`13c099a`)
|
||||
- downloaderMW doc typo (spiderMW doc copy remnant) (:commit:`8ae11bf`)
|
||||
- Correct typos (:commit:`1346037`)
|
||||
|
||||
0.22.1 (released 2014-02-08)
|
||||
----------------------------
|
||||
Scrapy 0.22.1 (released 2014-02-08)
|
||||
-----------------------------------
|
||||
|
||||
- localhost666 can resolve under certain circumstances (:commit:`2ec2279`)
|
||||
- test inspect.stack failure (:commit:`cc3eda3`)
|
||||
|
|
@ -926,8 +1111,8 @@ Bugfixes
|
|||
- fix 0.22.0 release date (:commit:`af0219a`)
|
||||
- fix typos in news.rst and remove (not released yet) header (:commit:`b7f58f4`)
|
||||
|
||||
0.22.0 (released 2014-01-17)
|
||||
----------------------------
|
||||
Scrapy 0.22.0 (released 2014-01-17)
|
||||
-----------------------------------
|
||||
|
||||
Enhancements
|
||||
~~~~~~~~~~~~
|
||||
|
|
@ -970,20 +1155,20 @@ Fixes
|
|||
- Fix tests runner under pip 1.5 (:issue:`513`)
|
||||
- Fix logging error when spider name is unicode (:issue:`479`)
|
||||
|
||||
0.20.2 (released 2013-12-09)
|
||||
----------------------------
|
||||
Scrapy 0.20.2 (released 2013-12-09)
|
||||
-----------------------------------
|
||||
|
||||
- Update CrawlSpider Template with Selector changes (:commit:`6d1457d`)
|
||||
- fix method name in tutorial. closes GH-480 (:commit:`b4fc359`
|
||||
|
||||
0.20.1 (released 2013-11-28)
|
||||
----------------------------
|
||||
Scrapy 0.20.1 (released 2013-11-28)
|
||||
-----------------------------------
|
||||
|
||||
- include_package_data is required to build wheels from published sources (:commit:`5ba1ad5`)
|
||||
- process_parallel was leaking the failures on its internal deferreds. closes #458 (:commit:`419a780`)
|
||||
|
||||
0.20.0 (released 2013-11-08)
|
||||
----------------------------
|
||||
Scrapy 0.20.0 (released 2013-11-08)
|
||||
-----------------------------------
|
||||
|
||||
Enhancements
|
||||
~~~~~~~~~~~~
|
||||
|
|
@ -1070,15 +1255,15 @@ List of contributors sorted by number of commits::
|
|||
1 cacovsky <amarquesferraz@...>
|
||||
1 Berend Iwema <berend@...>
|
||||
|
||||
0.18.4 (released 2013-10-10)
|
||||
----------------------------
|
||||
Scrapy 0.18.4 (released 2013-10-10)
|
||||
-----------------------------------
|
||||
|
||||
- IPython refuses to update the namespace. fix #396 (:commit:`3d32c4f`)
|
||||
- Fix AlreadyCalledError replacing a request in shell command. closes #407 (:commit:`b1d8919`)
|
||||
- Fix start_requests laziness and early hangs (:commit:`89faf52`)
|
||||
|
||||
0.18.3 (released 2013-10-03)
|
||||
----------------------------
|
||||
Scrapy 0.18.3 (released 2013-10-03)
|
||||
-----------------------------------
|
||||
|
||||
- fix regression on lazy evaluation of start requests (:commit:`12693a5`)
|
||||
- forms: do not submit reset inputs (:commit:`e429f63`)
|
||||
|
|
@ -1086,14 +1271,14 @@ List of contributors sorted by number of commits::
|
|||
- backport master fixes to json exporter (:commit:`cfc2d46`)
|
||||
- Fix permission and set umask before generating sdist tarball (:commit:`06149e0`)
|
||||
|
||||
0.18.2 (released 2013-09-03)
|
||||
----------------------------
|
||||
Scrapy 0.18.2 (released 2013-09-03)
|
||||
-----------------------------------
|
||||
|
||||
- Backport `scrapy check` command fixes and backward compatible multi
|
||||
crawler process(:issue:`339`)
|
||||
|
||||
0.18.1 (released 2013-08-27)
|
||||
----------------------------
|
||||
Scrapy 0.18.1 (released 2013-08-27)
|
||||
-----------------------------------
|
||||
|
||||
- remove extra import added by cherry picked changes (:commit:`d20304e`)
|
||||
- fix crawling tests under twisted pre 11.0.0 (:commit:`1994f38`)
|
||||
|
|
@ -1111,8 +1296,8 @@ List of contributors sorted by number of commits::
|
|||
- minor updates to 0.18 release notes (:commit:`c45e5f1`)
|
||||
- fix contributters list format (:commit:`0b60031`)
|
||||
|
||||
0.18.0 (released 2013-08-09)
|
||||
----------------------------
|
||||
Scrapy 0.18.0 (released 2013-08-09)
|
||||
-----------------------------------
|
||||
|
||||
- Lot of improvements to testsuite run using Tox, including a way to test on pypi
|
||||
- Handle GET parameters for AJAX crawleable urls (:commit:`3fe2a32`)
|
||||
|
|
@ -1203,8 +1388,8 @@ contributors sorted by number of commits::
|
|||
1 Berend Iwema <berend@...>
|
||||
|
||||
|
||||
0.16.5 (released 2013-05-30)
|
||||
----------------------------
|
||||
Scrapy 0.16.5 (released 2013-05-30)
|
||||
-----------------------------------
|
||||
|
||||
- obey request method when scrapy deploy is redirected to a new endpoint (:commit:`8c4fcee`)
|
||||
- fix inaccurate downloader middleware documentation. refs #280 (:commit:`40667cb`)
|
||||
|
|
@ -1212,8 +1397,8 @@ contributors sorted by number of commits::
|
|||
- Find form nodes in invalid html5 documents (:commit:`e3d6945`)
|
||||
- Fix typo labeling attrs type bool instead of list (:commit:`a274276`)
|
||||
|
||||
0.16.4 (released 2013-01-23)
|
||||
----------------------------
|
||||
Scrapy 0.16.4 (released 2013-01-23)
|
||||
-----------------------------------
|
||||
|
||||
- fixes spelling errors in documentation (:commit:`6d2b3aa`)
|
||||
- add doc about disabling an extension. refs #132 (:commit:`c90de33`)
|
||||
|
|
@ -1224,8 +1409,8 @@ contributors sorted by number of commits::
|
|||
- fix bug in scrapy parse command when spider is not specified explicitly. closes #209 (:commit:`c72e682`)
|
||||
- Update docs/topics/commands.rst (:commit:`28eac7a`)
|
||||
|
||||
0.16.3 (released 2012-12-07)
|
||||
----------------------------
|
||||
Scrapy 0.16.3 (released 2012-12-07)
|
||||
-----------------------------------
|
||||
|
||||
- Remove concurrency limitation when using download delays and still ensure inter-request delays are enforced (:commit:`487b9b5`)
|
||||
- add error details when image pipeline fails (:commit:`8232569`)
|
||||
|
|
@ -1237,8 +1422,8 @@ contributors sorted by number of commits::
|
|||
- Fixed docs typo in SpiderOpenCloseLogging example (:commit:`7184094`)
|
||||
|
||||
|
||||
0.16.2 (released 2012-11-09)
|
||||
----------------------------
|
||||
Scrapy 0.16.2 (released 2012-11-09)
|
||||
-----------------------------------
|
||||
|
||||
- scrapy contracts: python2.6 compat (:commit:`a4a9199`)
|
||||
- scrapy contracts verbose option (:commit:`ec41673`)
|
||||
|
|
@ -1248,8 +1433,8 @@ contributors sorted by number of commits::
|
|||
- Fix SpiderState bug in Windows platforms (:commit:`58998f4`)
|
||||
|
||||
|
||||
0.16.1 (released 2012-10-26)
|
||||
----------------------------
|
||||
Scrapy 0.16.1 (released 2012-10-26)
|
||||
-----------------------------------
|
||||
|
||||
- fixed LogStats extension, which got broken after a wrong merge before the 0.16 release (:commit:`8c780fd`)
|
||||
- better backwards compatibility for scrapy.conf.settings (:commit:`3403089`)
|
||||
|
|
@ -1259,8 +1444,8 @@ contributors sorted by number of commits::
|
|||
- set release date for 0.16.0 in news (:commit:`e292246`)
|
||||
|
||||
|
||||
0.16.0 (released 2012-10-18)
|
||||
----------------------------
|
||||
Scrapy 0.16.0 (released 2012-10-18)
|
||||
-----------------------------------
|
||||
|
||||
Scrapy changes:
|
||||
|
||||
|
|
@ -1301,8 +1486,8 @@ Scrapy changes:
|
|||
- number received responses are now tracked through Scrapy stats (stat name: ``response_received_count``)
|
||||
- removed ``scrapy.log.started`` attribute
|
||||
|
||||
0.14.4
|
||||
------
|
||||
Scrapy 0.14.4
|
||||
-------------
|
||||
|
||||
- added precise to supported ubuntu distros (:commit:`b7e46df`)
|
||||
- fixed bug in json-rpc webservice reported in https://groups.google.com/forum/#!topic/scrapy-users/qgVBmFybNAQ/discussion. also removed no longer supported 'run' command from extras/scrapy-ws.py (:commit:`340fbdb`)
|
||||
|
|
@ -1310,8 +1495,8 @@ Scrapy changes:
|
|||
- replace "import Image" by more standard "from PIL import Image". closes #88 (:commit:`4d17048`)
|
||||
- return trial status as bin/runtests.sh exit value. #118 (:commit:`b7b2e7f`)
|
||||
|
||||
0.14.3
|
||||
------
|
||||
Scrapy 0.14.3
|
||||
-------------
|
||||
|
||||
- forgot to include pydispatch license. #118 (:commit:`fd85f9c`)
|
||||
- include egg files used by testsuite in source distribution. #118 (:commit:`c897793`)
|
||||
|
|
@ -1323,8 +1508,8 @@ Scrapy changes:
|
|||
- fixed minor defect in link extractors documentation (:commit:`ba14f38`)
|
||||
- removed some obsolete remaining code related to sqlite support in scrapy (:commit:`0665175`)
|
||||
|
||||
0.14.2
|
||||
------
|
||||
Scrapy 0.14.2
|
||||
-------------
|
||||
|
||||
- move buffer pointing to start of file before computing checksum. refs #92 (:commit:`6a5bef2`)
|
||||
- Compute image checksum before persisting images. closes #92 (:commit:`9817df1`)
|
||||
|
|
@ -1338,8 +1523,8 @@ Scrapy changes:
|
|||
- scrapyd: fixed documentation link (:commit:`2b4e4c3`)
|
||||
- extras/makedeb.py: no longer obtaining version from git (:commit:`caffe0e`)
|
||||
|
||||
0.14.1
|
||||
------
|
||||
Scrapy 0.14.1
|
||||
-------------
|
||||
|
||||
- extras/makedeb.py: no longer obtaining version from git (:commit:`caffe0e`)
|
||||
- bumped version to 0.14.1 (:commit:`6cb9e1c`)
|
||||
|
|
@ -1353,8 +1538,8 @@ Scrapy changes:
|
|||
- Avoid _disconnectedDeferred AttributeError exception in Twisted>=11.1.0 (:commit:`98f3f87`)
|
||||
- allow spider to set autothrottle max concurrency (:commit:`175a4b5`)
|
||||
|
||||
0.14
|
||||
----
|
||||
Scrapy 0.14
|
||||
-----------
|
||||
|
||||
New features and settings
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
|
@ -1421,8 +1606,8 @@ Code rearranged and removed
|
|||
- Renamed attributes of core components: downloader.sites -> downloader.slots, scraper.sites -> scraper.slots (:rev:`2717`, :rev:`2718`)
|
||||
- Renamed setting ``CLOSESPIDER_ITEMPASSED`` to :setting:`CLOSESPIDER_ITEMCOUNT` (:rev:`2655`). Backwards compatibility kept.
|
||||
|
||||
0.12
|
||||
----
|
||||
Scrapy 0.12
|
||||
-----------
|
||||
|
||||
The numbers like #NNN reference tickets in the old issue tracker (Trac) which is no longer available.
|
||||
|
||||
|
|
@ -1464,8 +1649,8 @@ Deprecated/obsoleted functionality
|
|||
- Deprecated ``queue`` command in favor of using Scrapyd ``schedule.json`` API. See also: Scrapyd changes
|
||||
- Removed the !LxmlItemLoader (experimental contrib which never graduated to main contrib)
|
||||
|
||||
0.10
|
||||
----
|
||||
Scrapy 0.10
|
||||
-----------
|
||||
|
||||
The numbers like #NNN reference tickets in the old issue tracker (Trac) which is no longer available.
|
||||
|
||||
|
|
@ -1537,8 +1722,8 @@ Changes to settings
|
|||
- Removed ``COMMANDS_SETTINGS_MODULE`` setting (#201)
|
||||
- Renamed ``REQUEST_HANDLERS`` to ``DOWNLOAD_HANDLERS`` and make download handlers classes (instead of functions)
|
||||
|
||||
0.9
|
||||
---
|
||||
Scrapy 0.9
|
||||
----------
|
||||
|
||||
The numbers like #NNN reference tickets in the old issue tracker (Trac) which is no longer available.
|
||||
|
||||
|
|
@ -1577,8 +1762,8 @@ Changes to default settings
|
|||
|
||||
- Changed default ``SCHEDULER_ORDER`` to ``DFO`` (:rev:`1939`)
|
||||
|
||||
0.8
|
||||
---
|
||||
Scrapy 0.8
|
||||
----------
|
||||
|
||||
The numbers like #NNN reference tickets in the old issue tracker (Trac) which is no longer available.
|
||||
|
||||
|
|
@ -1623,8 +1808,8 @@ Backwards-incompatible changes
|
|||
- Renamed extension: ``DelayedCloseDomain`` to ``SpiderCloseDelay`` (:rev:`1861` | #121)
|
||||
- Removed obsolete ``scrapy.utils.markup.remove_escape_chars`` function - use ``scrapy.utils.markup.replace_escape_chars`` instead (:rev:`1865`)
|
||||
|
||||
0.7
|
||||
---
|
||||
Scrapy 0.7
|
||||
----------
|
||||
|
||||
First release of Scrapy.
|
||||
|
||||
|
|
|
|||
Binary file not shown.
|
Before Width: | Height: | Size: 34 KiB After Width: | Height: | Size: 53 KiB |
|
|
@ -171,7 +171,8 @@ SpiderLoader API
|
|||
|
||||
This class method is used by Scrapy to create an instance of the class.
|
||||
It's called with the current project settings, and it loads the spiders
|
||||
found in the modules of the :setting:`SPIDER_MODULES` setting.
|
||||
found recursively in the modules of the :setting:`SPIDER_MODULES`
|
||||
setting.
|
||||
|
||||
:param settings: project settings
|
||||
:type settings: :class:`~scrapy.settings.Settings` instance
|
||||
|
|
|
|||
|
|
@ -12,24 +12,77 @@ Overview
|
|||
|
||||
The following diagram shows an overview of the Scrapy architecture with its
|
||||
components and an outline of the data flow that takes place inside the system
|
||||
(shown by the green arrows). A brief description of the components is included
|
||||
(shown by the red arrows). A brief description of the components is included
|
||||
below with links for more detailed information about them. The data flow is
|
||||
also described below.
|
||||
|
||||
.. _data-flow:
|
||||
|
||||
Data flow
|
||||
=========
|
||||
|
||||
.. image:: _images/scrapy_architecture_02.png
|
||||
:width: 700
|
||||
:height: 470
|
||||
:alt: Scrapy architecture
|
||||
|
||||
The data flow in Scrapy is controlled by the execution engine, and goes like
|
||||
this:
|
||||
|
||||
1. The :ref:`Engine <component-engine>` gets the initial Requests to crawl from the
|
||||
:ref:`Spider <component-spiders>`.
|
||||
|
||||
2. The :ref:`Engine <component-engine>` schedules the Requests in the
|
||||
:ref:`Scheduler <component-scheduler>` and asks for the
|
||||
next Requests to crawl.
|
||||
|
||||
3. The :ref:`Scheduler <component-scheduler>` returns the next Requests
|
||||
to the :ref:`Engine <component-engine>`.
|
||||
|
||||
4. The :ref:`Engine <component-engine>` sends the Requests to the
|
||||
:ref:`Downloader <component-downloader>`, passing through the
|
||||
:ref:`Downloader Middlewares <component-downloader-middleware>` (see
|
||||
:meth:`~scrapy.downloadermiddlewares.DownloaderMiddleware.process_request`).
|
||||
|
||||
5. Once the page finishes downloading the
|
||||
:ref:`Downloader <component-downloader>` generates a Response (with
|
||||
that page) and sends it to the Engine, passing through the
|
||||
:ref:`Downloader Middlewares <component-downloader-middleware>` (see
|
||||
:meth:`~scrapy.downloadermiddlewares.DownloaderMiddleware.process_response`).
|
||||
|
||||
6. The :ref:`Engine <component-engine>` receives the Response from the
|
||||
:ref:`Downloader <component-downloader>` and sends it to the
|
||||
:ref:`Spider <component-spiders>` for processing, passing
|
||||
through the :ref:`Spider Middleware <component-spider-middleware>` (see
|
||||
:meth:`~scrapy.spidermiddlewares.SpiderMiddleware.process_spider_input`).
|
||||
|
||||
7. The :ref:`Spider <component-spiders>` processes the Response and returns
|
||||
scraped items and new Requests (to follow) to the
|
||||
:ref:`Engine <component-engine>`, passing through the
|
||||
:ref:`Spider Middleware <component-spider-middleware>` (see
|
||||
:meth:`~scrapy.spidermiddlewares.SpiderMiddleware.process_spider_output`).
|
||||
|
||||
8. The :ref:`Engine <component-engine>` sends processed items to
|
||||
:ref:`Item Pipelines <component-pipelines>`, then send processed Requests to
|
||||
the :ref:`Scheduler <component-scheduler>` and asks for possible next Requests
|
||||
to crawl.
|
||||
|
||||
9. The process repeats (from step 1) until there are no more requests from the
|
||||
:ref:`Scheduler <component-scheduler>`.
|
||||
|
||||
Components
|
||||
==========
|
||||
|
||||
.. _component-engine:
|
||||
|
||||
Scrapy Engine
|
||||
-------------
|
||||
|
||||
The engine is responsible for controlling the data flow between all components
|
||||
of the system, and triggering events when certain actions occur. See the Data
|
||||
Flow section below for more details.
|
||||
of the system, and triggering events when certain actions occur. See the
|
||||
:ref:`Data Flow <data-flow>` section above for more details.
|
||||
|
||||
.. _component-scheduler:
|
||||
|
||||
Scheduler
|
||||
---------
|
||||
|
|
@ -37,19 +90,25 @@ Scheduler
|
|||
The Scheduler receives requests from the engine and enqueues them for feeding
|
||||
them later (also to the engine) when the engine requests them.
|
||||
|
||||
.. _component-downloader:
|
||||
|
||||
Downloader
|
||||
----------
|
||||
|
||||
The Downloader is responsible for fetching web pages and feeding them to the
|
||||
engine which, in turn, feeds them to the spiders.
|
||||
|
||||
.. _component-spiders:
|
||||
|
||||
Spiders
|
||||
-------
|
||||
|
||||
Spiders are custom classes written by Scrapy users to parse responses and
|
||||
extract items (aka scraped items) from them or additional URLs (requests) to
|
||||
extract items (aka scraped items) from them or additional requests to
|
||||
follow. For more information see :ref:`topics-spiders`.
|
||||
|
||||
.. _component-pipelines:
|
||||
|
||||
Item Pipeline
|
||||
-------------
|
||||
|
||||
|
|
@ -58,6 +117,8 @@ extracted (or scraped) by the spiders. Typical tasks include cleansing,
|
|||
validation and persistence (like storing the item in a database). For more
|
||||
information see :ref:`topics-item-pipeline`.
|
||||
|
||||
.. _component-downloader-middleware:
|
||||
|
||||
Downloader middlewares
|
||||
----------------------
|
||||
|
||||
|
|
@ -76,6 +137,8 @@ Use a Downloader middleware if you need to do one of the following:
|
|||
|
||||
For more information see :ref:`topics-downloader-middleware`.
|
||||
|
||||
.. _component-spider-middleware:
|
||||
|
||||
Spider middlewares
|
||||
------------------
|
||||
|
||||
|
|
@ -93,39 +156,6 @@ Use a Spider middleware if you need to
|
|||
|
||||
For more information see :ref:`topics-spider-middleware`.
|
||||
|
||||
Data flow
|
||||
=========
|
||||
|
||||
The data flow in Scrapy is controlled by the execution engine, and goes like
|
||||
this:
|
||||
|
||||
1. The Engine gets the first URLs to crawl from the Spider.
|
||||
|
||||
2. The Engine schedules the URLs in the Scheduler as Requests and asks for the
|
||||
next URLs to crawl.
|
||||
|
||||
3. The Scheduler returns the next URLs to crawl to the Engine.
|
||||
|
||||
4. The Engine sends the URLs to the Downloader, passing through the
|
||||
Downloader Middleware (request direction).
|
||||
|
||||
5. Once the page finishes downloading the Downloader generates a Response (with
|
||||
that page) and sends it to the Engine, passing through the Downloader
|
||||
Middleware (response direction).
|
||||
|
||||
6. The Engine receives the Response from the Downloader and sends it to the
|
||||
Spider for processing, passing through the Spider Middleware (input direction).
|
||||
|
||||
7. The Spider processes the Response and returns scraped items and new Requests
|
||||
(to follow) to the Engine, passing through the Spider Middleware
|
||||
(output direction).
|
||||
|
||||
8. The Engine sends processed items to Item Pipelines and processed Requests to
|
||||
the Scheduler.
|
||||
|
||||
9. The process repeats (from step 1) until there are no more requests from the
|
||||
Scheduler.
|
||||
|
||||
Event-driven networking
|
||||
=======================
|
||||
|
||||
|
|
|
|||
|
|
@ -18,40 +18,66 @@ To run it use::
|
|||
|
||||
You should see an output like this::
|
||||
|
||||
2013-05-16 13:08:46-0300 [scrapy] INFO: Scrapy 0.17.0 started (bot: scrapybot)
|
||||
2013-05-16 13:08:47-0300 [scrapy] INFO: Spider opened
|
||||
2013-05-16 13:08:47-0300 [scrapy] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:48-0300 [scrapy] INFO: Crawled 74 pages (at 4440 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:49-0300 [scrapy] INFO: Crawled 143 pages (at 4140 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:50-0300 [scrapy] INFO: Crawled 210 pages (at 4020 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:51-0300 [scrapy] INFO: Crawled 274 pages (at 3840 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:52-0300 [scrapy] INFO: Crawled 343 pages (at 4140 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:53-0300 [scrapy] INFO: Crawled 410 pages (at 4020 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:54-0300 [scrapy] INFO: Crawled 474 pages (at 3840 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:55-0300 [scrapy] INFO: Crawled 538 pages (at 3840 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:56-0300 [scrapy] INFO: Crawled 602 pages (at 3840 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:57-0300 [scrapy] INFO: Closing spider (closespider_timeout)
|
||||
2013-05-16 13:08:57-0300 [scrapy] INFO: Crawled 666 pages (at 3840 pages/min), scraped 0 items (at 0 items/min)
|
||||
2013-05-16 13:08:57-0300 [scrapy] INFO: Dumping Scrapy stats:
|
||||
{'downloader/request_bytes': 231508,
|
||||
'downloader/request_count': 682,
|
||||
'downloader/request_method_count/GET': 682,
|
||||
'downloader/response_bytes': 1172802,
|
||||
'downloader/response_count': 682,
|
||||
'downloader/response_status_count/200': 682,
|
||||
'finish_reason': 'closespider_timeout',
|
||||
'finish_time': datetime.datetime(2013, 5, 16, 16, 8, 57, 985539),
|
||||
'log_count/INFO': 14,
|
||||
'request_depth_max': 34,
|
||||
'response_received_count': 682,
|
||||
'scheduler/dequeued': 682,
|
||||
'scheduler/dequeued/memory': 682,
|
||||
'scheduler/enqueued': 12767,
|
||||
'scheduler/enqueued/memory': 12767,
|
||||
'start_time': datetime.datetime(2013, 5, 16, 16, 8, 47, 676539)}
|
||||
2013-05-16 13:08:57-0300 [scrapy] INFO: Spider closed (closespider_timeout)
|
||||
2016-12-16 21:18:48 [scrapy.utils.log] INFO: Scrapy 1.2.2 started (bot: quotesbot)
|
||||
2016-12-16 21:18:48 [scrapy.utils.log] INFO: Overridden settings: {'CLOSESPIDER_TIMEOUT': 10, 'ROBOTSTXT_OBEY': True, 'SPIDER_MODULES': ['quotesbot.spiders'], 'LOGSTATS_INTERVAL': 1, 'BOT_NAME': 'quotesbot', 'LOG_LEVEL': 'INFO', 'NEWSPIDER_MODULE': 'quotesbot.spiders'}
|
||||
2016-12-16 21:18:49 [scrapy.middleware] INFO: Enabled extensions:
|
||||
['scrapy.extensions.closespider.CloseSpider',
|
||||
'scrapy.extensions.logstats.LogStats',
|
||||
'scrapy.extensions.telnet.TelnetConsole',
|
||||
'scrapy.extensions.corestats.CoreStats']
|
||||
2016-12-16 21:18:49 [scrapy.middleware] INFO: Enabled downloader middlewares:
|
||||
['scrapy.downloadermiddlewares.robotstxt.RobotsTxtMiddleware',
|
||||
'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware',
|
||||
'scrapy.downloadermiddlewares.downloadtimeout.DownloadTimeoutMiddleware',
|
||||
'scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware',
|
||||
'scrapy.downloadermiddlewares.useragent.UserAgentMiddleware',
|
||||
'scrapy.downloadermiddlewares.retry.RetryMiddleware',
|
||||
'scrapy.downloadermiddlewares.redirect.MetaRefreshMiddleware',
|
||||
'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware',
|
||||
'scrapy.downloadermiddlewares.redirect.RedirectMiddleware',
|
||||
'scrapy.downloadermiddlewares.cookies.CookiesMiddleware',
|
||||
'scrapy.downloadermiddlewares.stats.DownloaderStats']
|
||||
2016-12-16 21:18:49 [scrapy.middleware] INFO: Enabled spider middlewares:
|
||||
['scrapy.spidermiddlewares.httperror.HttpErrorMiddleware',
|
||||
'scrapy.spidermiddlewares.offsite.OffsiteMiddleware',
|
||||
'scrapy.spidermiddlewares.referer.RefererMiddleware',
|
||||
'scrapy.spidermiddlewares.urllength.UrlLengthMiddleware',
|
||||
'scrapy.spidermiddlewares.depth.DepthMiddleware']
|
||||
2016-12-16 21:18:49 [scrapy.middleware] INFO: Enabled item pipelines:
|
||||
[]
|
||||
2016-12-16 21:18:49 [scrapy.core.engine] INFO: Spider opened
|
||||
2016-12-16 21:18:49 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:18:50 [scrapy.extensions.logstats] INFO: Crawled 70 pages (at 4200 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:18:51 [scrapy.extensions.logstats] INFO: Crawled 134 pages (at 3840 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:18:52 [scrapy.extensions.logstats] INFO: Crawled 198 pages (at 3840 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:18:53 [scrapy.extensions.logstats] INFO: Crawled 254 pages (at 3360 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:18:54 [scrapy.extensions.logstats] INFO: Crawled 302 pages (at 2880 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:18:55 [scrapy.extensions.logstats] INFO: Crawled 358 pages (at 3360 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:18:56 [scrapy.extensions.logstats] INFO: Crawled 406 pages (at 2880 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:18:57 [scrapy.extensions.logstats] INFO: Crawled 438 pages (at 1920 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:18:58 [scrapy.extensions.logstats] INFO: Crawled 470 pages (at 1920 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:18:59 [scrapy.core.engine] INFO: Closing spider (closespider_timeout)
|
||||
2016-12-16 21:18:59 [scrapy.extensions.logstats] INFO: Crawled 518 pages (at 2880 pages/min), scraped 0 items (at 0 items/min)
|
||||
2016-12-16 21:19:00 [scrapy.statscollectors] INFO: Dumping Scrapy stats:
|
||||
{'downloader/request_bytes': 229995,
|
||||
'downloader/request_count': 534,
|
||||
'downloader/request_method_count/GET': 534,
|
||||
'downloader/response_bytes': 1565504,
|
||||
'downloader/response_count': 534,
|
||||
'downloader/response_status_count/200': 534,
|
||||
'finish_reason': 'closespider_timeout',
|
||||
'finish_time': datetime.datetime(2016, 12, 16, 16, 19, 0, 647725),
|
||||
'log_count/INFO': 17,
|
||||
'request_depth_max': 19,
|
||||
'response_received_count': 534,
|
||||
'scheduler/dequeued': 533,
|
||||
'scheduler/dequeued/memory': 533,
|
||||
'scheduler/enqueued': 10661,
|
||||
'scheduler/enqueued/memory': 10661,
|
||||
'start_time': datetime.datetime(2016, 12, 16, 16, 18, 49, 799869)}
|
||||
2016-12-16 21:19:00 [scrapy.core.engine] INFO: Spider closed (closespider_timeout)
|
||||
|
||||
That tells you that Scrapy is able to crawl about 3900 pages per minute in the
|
||||
That tells you that Scrapy is able to crawl about 3000 pages per minute in the
|
||||
hardware where you run it. Note that this is a very simple spider intended to
|
||||
follow links, any custom spider you write will probably do more stuff which
|
||||
results in slower crawl rates. How slower depends on how much your spider does
|
||||
|
|
|
|||
|
|
@ -322,6 +322,14 @@ So this command can be used to "see" how your spider would fetch a certain page.
|
|||
If used outside a project, no particular per-spider behaviour would be applied
|
||||
and it will just use the default Scrapy downloader settings.
|
||||
|
||||
Supported options:
|
||||
|
||||
* ``--spider=SPIDER``: bypass spider autodetection and force use of specific spider
|
||||
|
||||
* ``--headers``: print the response's HTTP headers instead of the response's body
|
||||
|
||||
* ``--no-redirect``: do not follow HTTP 3xx redirects (default is to follow them)
|
||||
|
||||
Usage examples::
|
||||
|
||||
$ scrapy fetch --nolog http://www.example.com/some/page.html
|
||||
|
|
@ -368,11 +376,34 @@ given. Also supports UNIX-style local file paths, either relative with
|
|||
``./`` or ``../`` prefixes or absolute file paths.
|
||||
See :ref:`topics-shell` for more info.
|
||||
|
||||
Supported options:
|
||||
|
||||
* ``--spider=SPIDER``: bypass spider autodetection and force use of specific spider
|
||||
|
||||
* ``-c code``: evaluate the code in the shell, print the result and exit
|
||||
|
||||
* ``--no-redirect``: do not follow HTTP 3xx redirects (default is to follow them);
|
||||
this only affects the URL you may pass as argument on the command line;
|
||||
once you are inside the shell, ``fetch(url)`` will still follow HTTP redirects by default.
|
||||
|
||||
Usage example::
|
||||
|
||||
$ scrapy shell http://www.example.com/some/page.html
|
||||
[ ... scrapy shell starts ... ]
|
||||
|
||||
$ scrapy shell --nolog http://www.example.com/ -c '(response.status, response.url)'
|
||||
(200, 'http://www.example.com/')
|
||||
|
||||
# shell follows HTTP redirects by default
|
||||
$ scrapy shell --nolog http://httpbin.org/redirect-to?url=http%3A%2F%2Fexample.com%2F -c '(response.status, response.url)'
|
||||
(200, 'http://example.com/')
|
||||
|
||||
# you can disable this with --no-redirect
|
||||
# (only for the URL passed as command line argument)
|
||||
$ scrapy shell --no-redirect --nolog http://httpbin.org/redirect-to?url=http%3A%2F%2Fexample.com%2F -c '(response.status, response.url)'
|
||||
(302, 'http://httpbin.org/redirect-to?url=http%3A%2F%2Fexample.com%2F')
|
||||
|
||||
|
||||
.. command:: parse
|
||||
|
||||
parse
|
||||
|
|
|
|||
|
|
@ -27,7 +27,11 @@ The :setting:`DOWNLOADER_MIDDLEWARES` setting is merged with the
|
|||
:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting defined in Scrapy (and not meant
|
||||
to be overridden) and then sorted by order to get the final sorted list of
|
||||
enabled middlewares: the first middleware is the one closer to the engine and
|
||||
the last is the one closer to the downloader.
|
||||
the last is the one closer to the downloader. In other words,
|
||||
the :meth:`~scrapy.downloadermiddlewares.DownloaderMiddleware.process_request`
|
||||
method of each middleware will be invoked in increasing
|
||||
middleware order (100, 200, 300, ...) and the :meth:`~scrapy.downloadermiddlewares.DownloaderMiddleware.process_response` method
|
||||
of each middleware will be invoked in decreasing order.
|
||||
|
||||
To decide which order to assign to your middleware see the
|
||||
:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting and pick a value according to
|
||||
|
|
@ -234,14 +238,14 @@ header) and all cookies received in responses (ie. ``Set-Cookie`` header).
|
|||
|
||||
Here's an example of a log with :setting:`COOKIES_DEBUG` enabled::
|
||||
|
||||
2011-04-06 14:35:10-0300 [scrapy] INFO: Spider opened
|
||||
2011-04-06 14:35:10-0300 [scrapy] DEBUG: Sending cookies to: <GET http://www.diningcity.com/netherlands/index.html>
|
||||
2011-04-06 14:35:10-0300 [scrapy.core.engine] INFO: Spider opened
|
||||
2011-04-06 14:35:10-0300 [scrapy.downloadermiddlewares.cookies] DEBUG: Sending cookies to: <GET http://www.diningcity.com/netherlands/index.html>
|
||||
Cookie: clientlanguage_nl=en_EN
|
||||
2011-04-06 14:35:14-0300 [scrapy] DEBUG: Received cookies from: <200 http://www.diningcity.com/netherlands/index.html>
|
||||
2011-04-06 14:35:14-0300 [scrapy.downloadermiddlewares.cookies] DEBUG: Received cookies from: <200 http://www.diningcity.com/netherlands/index.html>
|
||||
Set-Cookie: JSESSIONID=B~FA4DC0C496C8762AE4F1A620EAB34F38; Path=/
|
||||
Set-Cookie: ip_isocode=US
|
||||
Set-Cookie: clientlanguage_nl=en_EN; Expires=Thu, 07-Apr-2011 21:21:34 GMT; Path=/
|
||||
2011-04-06 14:49:50-0300 [scrapy] DEBUG: Crawled (200) <GET http://www.diningcity.com/netherlands/index.html> (referer: None)
|
||||
2011-04-06 14:49:50-0300 [scrapy.core.engine] DEBUG: Crawled (200) <GET http://www.diningcity.com/netherlands/index.html> (referer: None)
|
||||
[...]
|
||||
|
||||
|
||||
|
|
@ -653,16 +657,6 @@ Default: ``True``
|
|||
Whether the Compression middleware will be enabled.
|
||||
|
||||
|
||||
ChunkedTransferMiddleware
|
||||
-------------------------
|
||||
|
||||
.. module:: scrapy.downloadermiddlewares.chunked
|
||||
:synopsis: Chunked Transfer Middleware
|
||||
|
||||
.. class:: ChunkedTransferMiddleware
|
||||
|
||||
This middleware adds support for `chunked transfer encoding`_
|
||||
|
||||
HttpProxyMiddleware
|
||||
-------------------
|
||||
|
||||
|
|
@ -966,4 +960,3 @@ The default encoding for proxy authentication on :class:`HttpProxyMiddleware`.
|
|||
|
||||
.. _DBM: https://en.wikipedia.org/wiki/Dbm
|
||||
.. _anydbm: https://docs.python.org/2/library/anydbm.html
|
||||
.. _chunked transfer encoding: https://en.wikipedia.org/wiki/Chunked_transfer_encoding
|
||||
|
|
|
|||
|
|
@ -81,13 +81,13 @@ uses `Twisted non-blocking IO`_, like the rest of the framework.
|
|||
Send email to the given recipients.
|
||||
|
||||
:param to: the e-mail recipients
|
||||
:type to: list
|
||||
:type to: str or list of str
|
||||
|
||||
:param subject: the subject of the e-mail
|
||||
:type subject: str
|
||||
|
||||
:param cc: the e-mails to CC
|
||||
:type cc: list
|
||||
:type cc: str or list of str
|
||||
|
||||
:param body: the e-mail body
|
||||
:type body: str
|
||||
|
|
|
|||
|
|
@ -280,7 +280,7 @@ Whether to export empty feeds (ie. feeds with no items).
|
|||
FEED_STORAGES
|
||||
-------------
|
||||
|
||||
Default:: ``{}``
|
||||
Default: ``{}``
|
||||
|
||||
A dict containing additional feed storage backends supported by your project.
|
||||
The keys are URI schemes and the values are paths to storage classes.
|
||||
|
|
@ -314,7 +314,7 @@ can disable any of these backends by assigning ``None`` to their URI scheme in
|
|||
FEED_EXPORTERS
|
||||
--------------
|
||||
|
||||
Default:: ``{}``
|
||||
Default: ``{}``
|
||||
|
||||
A dict containing additional exporters supported by your project. The keys are
|
||||
serialization formats and the values are paths to :ref:`Item exporter
|
||||
|
|
|
|||
|
|
@ -17,7 +17,7 @@ when inspecting the page source is not the original HTML, but a modified one
|
|||
after applying some browser clean up and executing Javascript code. Firefox,
|
||||
in particular, is known for adding ``<tbody>`` elements to tables. Scrapy, on
|
||||
the other hand, does not modify the original page HTML, so you won't be able to
|
||||
extract any data if you use ``<tbody`` in your XPath expressions.
|
||||
extract any data if you use ``<tbody>`` in your XPath expressions.
|
||||
|
||||
Therefore, you should keep in mind the following things when working with
|
||||
Firefox and XPath:
|
||||
|
|
|
|||
|
|
@ -27,9 +27,10 @@ Each item pipeline component is a Python class that must implement the following
|
|||
|
||||
.. method:: process_item(self, item, spider)
|
||||
|
||||
This method is called for every item pipeline component and must either return
|
||||
a dict with data, :class:`~scrapy.item.Item` (or any descendant class) object
|
||||
or raise a :exc:`~scrapy.exceptions.DropItem` exception. Dropped items are no longer
|
||||
This method is called for every item pipeline component. :meth:`process_item`
|
||||
must either: return a dict with data, return an :class:`~scrapy.item.Item`
|
||||
(or any descendant class) object, return a `Twisted Deferred`_ or raise
|
||||
:exc:`~scrapy.exceptions.DropItem` exception. Dropped items are no longer
|
||||
processed by further pipeline components.
|
||||
|
||||
:param item: the item scraped
|
||||
|
|
@ -66,6 +67,8 @@ Additionally, they may also implement the following methods:
|
|||
:type crawler: :class:`~scrapy.crawler.Crawler` object
|
||||
|
||||
|
||||
.. _Twisted Deferred: https://twistedmatrix.com/documents/current/core/howto/defer.html
|
||||
|
||||
Item pipeline example
|
||||
=====================
|
||||
|
||||
|
|
@ -103,9 +106,12 @@ format::
|
|||
|
||||
class JsonWriterPipeline(object):
|
||||
|
||||
def __init__(self):
|
||||
def open_spider(self, spider):
|
||||
self.file = open('items.jl', 'wb')
|
||||
|
||||
def close_spider(self, spider):
|
||||
self.file.close()
|
||||
|
||||
def process_item(self, item, spider):
|
||||
line = json.dumps(dict(item)) + "\n"
|
||||
self.file.write(line)
|
||||
|
|
@ -123,14 +129,7 @@ MongoDB address and database name are specified in Scrapy settings;
|
|||
MongoDB collection is named after item class.
|
||||
|
||||
The main point of this example is to show how to use :meth:`from_crawler`
|
||||
method and how to clean up the resources properly.
|
||||
|
||||
.. note::
|
||||
|
||||
Previous example (JsonWriterPipeline) doesn't clean up resources properly.
|
||||
Fixing it is left as an exercise for the reader.
|
||||
|
||||
::
|
||||
method and how to clean up the resources properly.::
|
||||
|
||||
import pymongo
|
||||
|
||||
|
|
@ -163,6 +162,55 @@ method and how to clean up the resources properly.
|
|||
.. _MongoDB: https://www.mongodb.org/
|
||||
.. _pymongo: https://api.mongodb.org/python/current/
|
||||
|
||||
|
||||
Take screenshot of item
|
||||
-----------------------
|
||||
|
||||
This example demonstrates how to return Deferred_ from :meth:`process_item` method.
|
||||
It uses Splash_ to render screenshot of item url. Pipeline
|
||||
makes request to locally running instance of Splash_. After request is downloaded
|
||||
and Deferred callback fires, it saves item to a file and adds filename to an item.
|
||||
|
||||
::
|
||||
|
||||
import scrapy
|
||||
import hashlib
|
||||
from urllib.parse import quote
|
||||
|
||||
|
||||
class ScreenshotPipeline(object):
|
||||
"""Pipeline that uses Splash to render screenshot of
|
||||
every Scrapy item."""
|
||||
|
||||
SPLASH_URL = "http://localhost:8050/render.png?url={}"
|
||||
|
||||
def process_item(self, item, spider):
|
||||
encoded_item_url = quote(item["url"])
|
||||
screenshot_url = self.SPLASH_URL.format(encoded_item_url)
|
||||
request = scrapy.Request(screenshot_url)
|
||||
dfd = spider.crawler.engine.download(request, spider)
|
||||
dfd.addBoth(self.return_item, item)
|
||||
return dfd
|
||||
|
||||
def return_item(self, response, item):
|
||||
if response.status != 200:
|
||||
# Error happened, return item.
|
||||
return item
|
||||
|
||||
# Save screenshot to file, filename will be hash of url.
|
||||
url = item["url"]
|
||||
url_hash = hashlib.md5(url.encode("utf8")).hexdigest()
|
||||
filename = "{}.png".format(url_hash)
|
||||
with open(filename, "wb") as f:
|
||||
f.write(response.body)
|
||||
|
||||
# Store filename in item.
|
||||
item["screenshot_filename"] = filename
|
||||
return item
|
||||
|
||||
.. _Splash: http://splash.readthedocs.io/en/stable/
|
||||
.. _Deferred: https://twistedmatrix.com/documents/current/core/howto/defer.html
|
||||
|
||||
Duplicates filter
|
||||
-----------------
|
||||
|
||||
|
|
|
|||
|
|
@ -10,7 +10,7 @@ Logging
|
|||
about the new logging system.
|
||||
|
||||
Scrapy uses `Python's builtin logging system
|
||||
<https://docs.python.org/2/library/logging.html>`_ for event logging. We'll
|
||||
<https://docs.python.org/3/library/logging.html>`_ for event logging. We'll
|
||||
provide some simple examples to get you started, but for more advanced
|
||||
use-cases it's strongly suggested to read thoroughly its documentation.
|
||||
|
||||
|
|
@ -150,6 +150,7 @@ These settings can be used to configure the logging:
|
|||
* :setting:`LOG_FORMAT`
|
||||
* :setting:`LOG_DATEFORMAT`
|
||||
* :setting:`LOG_STDOUT`
|
||||
* :setting:`LOG_SHORT_NAMES`
|
||||
|
||||
The first couple of settings define a destination for log messages. If
|
||||
:setting:`LOG_FILE` is set, messages sent through the root logger will be
|
||||
|
|
@ -170,6 +171,10 @@ listed in `logging's logrecord attributes docs
|
|||
<https://docs.python.org/2/library/datetime.html#strftime-and-strptime-behavior>`_
|
||||
respectively.
|
||||
|
||||
If :setting:`LOG_SHORT_NAMES` is set, then the logs will not display the scrapy
|
||||
component that prints the log. It is unset by default, hence logs contain the
|
||||
scrapy component responsible for that log output.
|
||||
|
||||
Command-line options
|
||||
--------------------
|
||||
|
||||
|
|
@ -188,6 +193,43 @@ to override some of the Scrapy settings regarding logging.
|
|||
Module `logging.handlers <https://docs.python.org/2/library/logging.handlers.html>`_
|
||||
Further documentation on available handlers
|
||||
|
||||
Advanced customization
|
||||
----------------------
|
||||
|
||||
Because Scrapy uses stdlib logging module, you can customize logging using
|
||||
all features of stdlib logging.
|
||||
|
||||
For example, let's say you're scraping a website which returns many
|
||||
HTTP 404 and 500 responses, and you want to hide all messages like this::
|
||||
|
||||
2016-12-16 22:00:06 [scrapy.spidermiddlewares.httperror] INFO: Ignoring
|
||||
response <500 http://quotes.toscrape.com/page/1-34/>: HTTP status code
|
||||
is not handled or not allowed
|
||||
|
||||
The first thing to note is a logger name - it is in brackets:
|
||||
``[scrapy.spidermiddlewares.httperror]``. If you get just ``[scrapy]`` then
|
||||
:setting:`LOG_SHORT_NAMES` is likely set to True; set it to False and re-run
|
||||
the crawl.
|
||||
|
||||
Next, we can see that the message has INFO level. To hide it
|
||||
we should set logging level for ``scrapy.spidermiddlewares.httperror``
|
||||
higher than INFO; next level after INFO is WARNING. It could be done
|
||||
e.g. in the spider's ``__init__`` method::
|
||||
|
||||
import logging
|
||||
import scrapy
|
||||
|
||||
|
||||
class MySpider(scrapy.Spider):
|
||||
# ...
|
||||
def __init__(self, *args, **kwargs):
|
||||
logger = logging.getLogger('scrapy.spidermiddlewares.httperror')
|
||||
logger.setLevel(logging.WARNING)
|
||||
super().__init__(*args, **kwargs)
|
||||
|
||||
If you run this spider again then INFO messages from
|
||||
``scrapy.spidermiddlewares.httperror`` logger will be gone.
|
||||
|
||||
scrapy.utils.log module
|
||||
=======================
|
||||
|
||||
|
|
|
|||
|
|
@ -379,7 +379,7 @@ See here the methods that you can override in your custom Files Pipeline:
|
|||
By default the :meth:`get_media_requests` method returns ``None`` which
|
||||
means there are no files to download for the item.
|
||||
|
||||
.. method:: FilesPipeline.item_completed(results, items, info)
|
||||
.. method:: FilesPipeline.item_completed(results, item, info)
|
||||
|
||||
The :meth:`FilesPipeline.item_completed` method called when all file
|
||||
requests for a single item have completed (either finished downloading, or
|
||||
|
|
@ -422,7 +422,7 @@ See here the methods that you can override in your custom Images Pipeline:
|
|||
|
||||
Must return a Request for each image URL.
|
||||
|
||||
.. method:: ImagesPipeline.item_completed(results, items, info)
|
||||
.. method:: ImagesPipeline.item_completed(results, item, info)
|
||||
|
||||
The :meth:`ImagesPipeline.item_completed` method is called when all image
|
||||
requests for a single item have completed (either finished downloading, or
|
||||
|
|
|
|||
|
|
@ -238,7 +238,8 @@ Here are some tips to keep in mind when dealing with these kinds of sites:
|
|||
* if possible, use `Google cache`_ to fetch pages, instead of hitting the sites
|
||||
directly
|
||||
* use a pool of rotating IPs. For example, the free `Tor project`_ or paid
|
||||
services like `ProxyMesh`_
|
||||
services like `ProxyMesh`_. An open source alterantive is `scrapoxy`_, a
|
||||
super proxy that you can attach your own proxies to.
|
||||
* use a highly distributed downloader that circumvents bans internally, so you
|
||||
can just focus on parsing clean pages. One example of such downloaders is
|
||||
`Crawlera`_
|
||||
|
|
@ -253,3 +254,4 @@ If you are still unable to prevent your bot getting banned, consider contacting
|
|||
.. _testspiders: https://github.com/scrapinghub/testspiders
|
||||
.. _Twisted Reactor Overview: https://twistedmatrix.com/documents/current/core/howto/reactor-basics.html
|
||||
.. _Crawlera: http://scrapinghub.com/crawlera
|
||||
.. _scrapoxy: http://scrapoxy.io/
|
||||
|
|
|
|||
|
|
@ -299,6 +299,7 @@ Those are:
|
|||
* :reqmeta:`dont_obey_robotstxt`
|
||||
* :reqmeta:`download_timeout`
|
||||
* :reqmeta:`download_maxsize`
|
||||
* :reqmeta:`download_latency`
|
||||
* :reqmeta:`proxy`
|
||||
|
||||
.. reqmeta:: bindaddress
|
||||
|
|
@ -316,6 +317,15 @@ download_timeout
|
|||
The amount of time (in secs) that the downloader will wait before timing out.
|
||||
See also: :setting:`DOWNLOAD_TIMEOUT`.
|
||||
|
||||
.. reqmeta:: download_latency
|
||||
|
||||
download_latency
|
||||
----------------
|
||||
|
||||
The amount of time spent to fetch the response, since the request has been
|
||||
started, i.e. HTTP message sent over the network. This meta key only becomes
|
||||
available when the response has been downloaded. While most other meta keys are
|
||||
used to control Scrapy behavior, this one is supposed to be read-only.
|
||||
|
||||
.. _topics-request-response-ref-request-subclasses:
|
||||
|
||||
|
|
@ -507,7 +517,13 @@ Response objects
|
|||
|
||||
.. attribute:: Response.headers
|
||||
|
||||
A dictionary-like object which contains the response headers.
|
||||
A dictionary-like object which contains the response headers. Values can
|
||||
be accessed using :meth:`get` to return the first header value with the
|
||||
specified name or :meth:`getlist` to return all header values with the
|
||||
specified name. For example, this call will give you all cookies in the
|
||||
headers::
|
||||
|
||||
response.headers.getlist('Set-Cookie')
|
||||
|
||||
.. attribute:: Response.body
|
||||
|
||||
|
|
|
|||
|
|
@ -468,7 +468,6 @@ Default::
|
|||
'scrapy.downloadermiddlewares.redirect.RedirectMiddleware': 600,
|
||||
'scrapy.downloadermiddlewares.cookies.CookiesMiddleware': 700,
|
||||
'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware': 750,
|
||||
'scrapy.downloadermiddlewares.chunked.ChunkedTransferMiddleware': 830,
|
||||
'scrapy.downloadermiddlewares.stats.DownloaderStats': 850,
|
||||
'scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware': 900,
|
||||
}
|
||||
|
|
@ -789,6 +788,16 @@ If ``True``, all standard output (and error) of your process will be redirected
|
|||
to the log. For example if you ``print 'hello'`` it will appear in the Scrapy
|
||||
log.
|
||||
|
||||
.. setting:: LOG_SHORT_NAMES
|
||||
|
||||
LOG_SHORT_NAMES
|
||||
---------------
|
||||
|
||||
Default: ``False``
|
||||
|
||||
If ``True``, the logs will just contain the root path. If it is set to ``False``
|
||||
then it displays the component responsible for the log output
|
||||
|
||||
.. setting:: MEMDEBUG_ENABLED
|
||||
|
||||
MEMDEBUG_ENABLED
|
||||
|
|
@ -1028,11 +1037,40 @@ Stats counter (``scheduler/unserializable``) tracks the number of times this hap
|
|||
|
||||
Example entry in logs::
|
||||
|
||||
1956-01-31 00:00:00+0800 [scrapy] ERROR: Unable to serialize request:
|
||||
1956-01-31 00:00:00+0800 [scrapy.core.scheduler] ERROR: Unable to serialize request:
|
||||
<GET http://example.com> - reason: cannot serialize <Request at 0x9a7c7ec>
|
||||
(type Request)> - no more unserializable requests will be logged
|
||||
(see 'scheduler/unserializable' stats counter)
|
||||
|
||||
|
||||
.. setting:: SCHEDULER_DISK_QUEUE
|
||||
|
||||
SCHEDULER_DISK_QUEUE
|
||||
--------------------
|
||||
|
||||
Default: ``'scrapy.squeues.PickleLifoDiskQueue'``
|
||||
|
||||
Type of disk queue that will be used by scheduler. Other available types are
|
||||
``scrapy.squeues.PickleFifoDiskQueue``, ``scrapy.squeues.MarshalFifoDiskQueue``,
|
||||
``scrapy.squeues.MarshalLifoDiskQueue``.
|
||||
|
||||
.. setting:: SCHEDULER_MEMORY_QUEUE
|
||||
|
||||
SCHEDULER_MEMORY_QUEUE
|
||||
----------------------
|
||||
Default: ``'scrapy.squeues.LifoMemoryQueue'``
|
||||
|
||||
Type of in-memory queue used by scheduler. Other available type is:
|
||||
``scrapy.squeues.FifoMemoryQueue``.
|
||||
|
||||
.. setting:: SCHEDULER_PRIORITY_QUEUE
|
||||
|
||||
SCHEDULER_PRIORITY_QUEUE
|
||||
------------------------
|
||||
Default: ``'queuelib.PriorityQueue'``
|
||||
|
||||
Type of priority queue used by scheduler.
|
||||
|
||||
.. setting:: SPIDER_CONTRACTS
|
||||
|
||||
SPIDER_CONTRACTS
|
||||
|
|
|
|||
|
|
@ -97,8 +97,12 @@ Available Shortcuts
|
|||
|
||||
* ``shelp()`` - print a help with the list of available objects and shortcuts
|
||||
|
||||
* ``fetch(request_or_url)`` - fetch a new response from the given request or
|
||||
URL and update all related objects accordingly.
|
||||
* ``fetch(url[, redirect=True])`` - fetch a new response from the given
|
||||
URL and update all related objects accordingly. You can optionaly ask for
|
||||
HTTP 3xx redirections to not be followed by passing ``redirect=False``
|
||||
|
||||
* ``fetch(request)`` - fetch a new response from the given request and
|
||||
update all related objects accordingly.
|
||||
|
||||
* ``view(response)`` - open the given response in your local web browser, for
|
||||
inspection. This will add a `\<base\> tag`_ to the response body in order
|
||||
|
|
@ -157,49 +161,65 @@ list of available objects and useful shortcuts (you'll notice that these lines
|
|||
all start with the ``[s]`` prefix)::
|
||||
|
||||
[s] Available Scrapy objects:
|
||||
[s] crawler <scrapy.crawler.Crawler object at 0x1e16b50>
|
||||
[s] scrapy scrapy module (contains scrapy.Request, scrapy.Selector, etc)
|
||||
[s] crawler <scrapy.crawler.Crawler object at 0x7f07395dd690>
|
||||
[s] item {}
|
||||
[s] request <GET http://scrapy.org>
|
||||
[s] response <200 http://scrapy.org>
|
||||
[s] settings <scrapy.settings.Settings object at 0x2bfd650>
|
||||
[s] spider <Spider 'default' at 0x20c6f50>
|
||||
[s] response <200 https://scrapy.org/>
|
||||
[s] settings <scrapy.settings.Settings object at 0x7f07395dd710>
|
||||
[s] spider <DefaultSpider 'default' at 0x7f0735891690>
|
||||
[s] Useful shortcuts:
|
||||
[s] fetch(url[, redirect=True]) Fetch URL and update local objects (by default, redirects are followed)
|
||||
[s] fetch(req) Fetch a scrapy.Request and update local objects
|
||||
[s] shelp() Shell help (print this help)
|
||||
[s] fetch(req_or_url) Fetch request (or URL) and update local objects
|
||||
[s] view(response) View response in a browser
|
||||
|
||||
>>>
|
||||
|
||||
|
||||
After that, we can start playing with the objects::
|
||||
|
||||
>>> response.xpath('//title/text()').extract_first()
|
||||
u'Scrapy | A Fast and Powerful Scraping and Web Crawling Framework'
|
||||
'Scrapy | A Fast and Powerful Scraping and Web Crawling Framework'
|
||||
|
||||
>>> fetch("http://reddit.com")
|
||||
[s] Available Scrapy objects:
|
||||
[s] crawler <scrapy.crawler.Crawler object at 0x7fb3ed9c9c90>
|
||||
[s] item {}
|
||||
[s] request <GET http://reddit.com>
|
||||
[s] response <200 https://www.reddit.com/>
|
||||
[s] settings <scrapy.settings.Settings object at 0x7fb3ed9c9c10>
|
||||
[s] spider <DefaultSpider 'default' at 0x7fb3ecdd3390>
|
||||
[s] Useful shortcuts:
|
||||
[s] shelp() Shell help (print this help)
|
||||
[s] fetch(req_or_url) Fetch request (or URL) and update local objects
|
||||
[s] view(response) View response in a browser
|
||||
|
||||
>>> response.xpath('//title/text()').extract()
|
||||
[u'reddit: the front page of the internet']
|
||||
['reddit: the front page of the internet']
|
||||
|
||||
>>> request = request.replace(method="POST")
|
||||
|
||||
>>> fetch(request)
|
||||
[s] Available Scrapy objects:
|
||||
[s] crawler <scrapy.crawler.Crawler object at 0x1e16b50>
|
||||
...
|
||||
|
||||
>>> response.status
|
||||
404
|
||||
|
||||
>>> from pprint import pprint
|
||||
|
||||
>>> pprint(response.headers)
|
||||
{'Accept-Ranges': ['bytes'],
|
||||
'Cache-Control': ['max-age=0, must-revalidate'],
|
||||
'Content-Type': ['text/html; charset=UTF-8'],
|
||||
'Date': ['Thu, 08 Dec 2016 16:21:19 GMT'],
|
||||
'Server': ['snooserv'],
|
||||
'Set-Cookie': ['loid=KqNLou0V9SKMX4qb4n; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure',
|
||||
'loidcreated=2016-12-08T16%3A21%3A19.445Z; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure',
|
||||
'loid=vi0ZVe4NkxNWdlH7r7; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure',
|
||||
'loidcreated=2016-12-08T16%3A21%3A19.459Z; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure'],
|
||||
'Vary': ['accept-encoding'],
|
||||
'Via': ['1.1 varnish'],
|
||||
'X-Cache': ['MISS'],
|
||||
'X-Cache-Hits': ['0'],
|
||||
'X-Content-Type-Options': ['nosniff'],
|
||||
'X-Frame-Options': ['SAMEORIGIN'],
|
||||
'X-Moose': ['majestic'],
|
||||
'X-Served-By': ['cache-cdg8730-CDG'],
|
||||
'X-Timer': ['S1481214079.394283,VS0,VE159'],
|
||||
'X-Ua-Compatible': ['IE=edge'],
|
||||
'X-Xss-Protection': ['1; mode=block']}
|
||||
>>>
|
||||
|
||||
|
||||
.. _topics-shell-inspect-response:
|
||||
|
||||
Invoking the shell from spiders to inspect responses
|
||||
|
|
@ -234,8 +254,8 @@ Here's an example of how you would call it from your spider::
|
|||
|
||||
When you run the spider, you will get something similar to this::
|
||||
|
||||
2014-01-23 17:48:31-0400 [scrapy] DEBUG: Crawled (200) <GET http://example.com> (referer: None)
|
||||
2014-01-23 17:48:31-0400 [scrapy] DEBUG: Crawled (200) <GET http://example.org> (referer: None)
|
||||
2014-01-23 17:48:31-0400 [scrapy.core.engine] DEBUG: Crawled (200) <GET http://example.com> (referer: None)
|
||||
2014-01-23 17:48:31-0400 [scrapy.core.engine] DEBUG: Crawled (200) <GET http://example.org> (referer: None)
|
||||
[s] Available Scrapy objects:
|
||||
[s] crawler <scrapy.crawler.Crawler object at 0x1e16b50>
|
||||
...
|
||||
|
|
@ -258,7 +278,7 @@ Finally you hit Ctrl-D (or Ctrl-Z in Windows) to exit the shell and resume the
|
|||
crawling::
|
||||
|
||||
>>> ^D
|
||||
2014-01-23 17:50:03-0400 [scrapy] DEBUG: Crawled (200) <GET http://example.net> (referer: None)
|
||||
2014-01-23 17:50:03-0400 [scrapy.core.engine] DEBUG: Crawled (200) <GET http://example.net> (referer: None)
|
||||
...
|
||||
|
||||
Note that you can't use the ``fetch`` shortcut here since the Scrapy engine is
|
||||
|
|
|
|||
|
|
@ -28,7 +28,12 @@ The :setting:`SPIDER_MIDDLEWARES` setting is merged with the
|
|||
:setting:`SPIDER_MIDDLEWARES_BASE` setting defined in Scrapy (and not meant to
|
||||
be overridden) and then sorted by order to get the final sorted list of enabled
|
||||
middlewares: the first middleware is the one closer to the engine and the last
|
||||
is the one closer to the spider.
|
||||
is the one closer to the spider. In other words,
|
||||
the :meth:`~scrapy.spidermiddlewares.SpiderMiddleware.process_spider_input`
|
||||
method of each middleware will be invoked in increasing
|
||||
middleware order (100, 200, 300, ...), and the
|
||||
:meth:`~scrapy.spidermiddlewares.SpiderMiddleware.process_spider_output` method
|
||||
of each middleware will be invoked in decreasing order.
|
||||
|
||||
To decide which order to assign to your middleware see the
|
||||
:setting:`SPIDER_MIDDLEWARES_BASE` setting and pick a value according to where
|
||||
|
|
|
|||
|
|
@ -24,8 +24,8 @@ For spiders, the scraping cycle goes through something like this:
|
|||
Requests.
|
||||
|
||||
2. In the callback function, you parse the response (web page) and return either
|
||||
dicts with extracted data, :class:`~scrapy.item.Item` objects,
|
||||
:class:`~scrapy.http.Request` objects, or an iterable of these objects.
|
||||
dicts with extracted data, :class:`~scrapy.item.Item` objects,
|
||||
:class:`~scrapy.http.Request` objects, or an iterable of these objects.
|
||||
Those Requests will also contain a callback (maybe
|
||||
the same) and will then be downloaded by Scrapy and then their
|
||||
response handled by the specified callback.
|
||||
|
|
@ -56,7 +56,7 @@ scrapy.Spider
|
|||
must inherit (including spiders that come bundled with Scrapy, as well as spiders
|
||||
that you write yourself). It doesn't provide any special functionality. It just
|
||||
provides a default :meth:`start_requests` implementation which sends requests from
|
||||
the :attr:`start_urls` spider attribute and calls the spider's method ``parse``
|
||||
the :attr:`start_urls` spider attribute and calls the spider's method ``parse``
|
||||
for each of the resulting responses.
|
||||
|
||||
.. attribute:: name
|
||||
|
|
@ -72,6 +72,8 @@ scrapy.Spider
|
|||
spider that crawls ``mywebsite.com`` would often be called
|
||||
``mywebsite``.
|
||||
|
||||
.. note:: In Python 2 this must be ASCII only.
|
||||
|
||||
.. attribute:: allowed_domains
|
||||
|
||||
An optional list of strings containing domains that this spider is
|
||||
|
|
@ -159,7 +161,7 @@ scrapy.Spider
|
|||
|
||||
class MySpider(scrapy.Spider):
|
||||
name = 'myspider'
|
||||
|
||||
|
||||
def start_requests(self):
|
||||
return [scrapy.FormRequest("http://www.example.com/login",
|
||||
formdata={'user': 'john', 'pass': 'secret'},
|
||||
|
|
@ -245,8 +247,8 @@ Return multiple Requests and items from a single callback::
|
|||
|
||||
for url in response.xpath('//a/@href').extract():
|
||||
yield scrapy.Request(url, callback=self.parse)
|
||||
|
||||
Instead of :attr:`~.start_urls` you can use :meth:`~.start_requests` directly;
|
||||
|
||||
Instead of :attr:`~.start_urls` you can use :meth:`~.start_requests` directly;
|
||||
to give data more structure you can use :ref:`topics-items`::
|
||||
|
||||
import scrapy
|
||||
|
|
@ -255,7 +257,7 @@ to give data more structure you can use :ref:`topics-items`::
|
|||
class MySpider(scrapy.Spider):
|
||||
name = 'example.com'
|
||||
allowed_domains = ['example.com']
|
||||
|
||||
|
||||
def start_requests(self):
|
||||
yield scrapy.Request('http://www.example.com/1.html', self.parse)
|
||||
yield scrapy.Request('http://www.example.com/2.html', self.parse)
|
||||
|
|
@ -267,7 +269,7 @@ to give data more structure you can use :ref:`topics-items`::
|
|||
|
||||
for url in response.xpath('//a/@href').extract():
|
||||
yield scrapy.Request(url, callback=self.parse)
|
||||
|
||||
|
||||
.. _spiderargs:
|
||||
|
||||
Spider arguments
|
||||
|
|
@ -283,7 +285,7 @@ Spider arguments are passed through the :command:`crawl` command using the
|
|||
|
||||
scrapy crawl myspider -a category=electronics
|
||||
|
||||
Spiders receive arguments in their constructors::
|
||||
Spiders can access arguments in their `__init__` methods::
|
||||
|
||||
import scrapy
|
||||
|
||||
|
|
@ -295,11 +297,42 @@ Spiders receive arguments in their constructors::
|
|||
self.start_urls = ['http://www.example.com/categories/%s' % category]
|
||||
# ...
|
||||
|
||||
The default `__init__` method will take any spider arguments
|
||||
and copy them to the spider as attributes.
|
||||
The above example can also be written as follows::
|
||||
|
||||
import scrapy
|
||||
|
||||
class MySpider(scrapy.Spider):
|
||||
name = 'myspider'
|
||||
|
||||
def start_requests(self):
|
||||
yield scrapy.Request('http://www.example.com/categories/%s' % self.category)
|
||||
|
||||
Keep in mind that spider arguments are only strings.
|
||||
The spider will not do any parsing on its own.
|
||||
If you were to set the `start_urls` attribute from the command line,
|
||||
you would have to parse it on your own into a list
|
||||
using something like
|
||||
`ast.literal_eval <https://docs.python.org/library/ast.html#ast.literal_eval>`_
|
||||
or `json.loads <https://docs.python.org/library/json.html#json.loads>`_
|
||||
and then set it as an attribute.
|
||||
Otherwise, you would cause iteration over a `start_urls` string
|
||||
(a very common python pitfall)
|
||||
resulting in each character being seen as a separate url.
|
||||
|
||||
A valid use case is to set the http auth credentials
|
||||
used by :class:`~scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware`
|
||||
or the user agent
|
||||
used by :class:`~scrapy.downloadermiddlewares.useragent.UserAgentMiddleware`::
|
||||
|
||||
scrapy crawl myspider -a http_user=myuser -a http_pass=mypassword -a user_agent=mybot
|
||||
|
||||
Spider arguments can also be passed through the Scrapyd ``schedule.json`` API.
|
||||
See `Scrapyd documentation`_.
|
||||
|
||||
.. _builtin-spiders:
|
||||
|
||||
|
||||
Generic Spiders
|
||||
===============
|
||||
|
||||
|
|
|
|||
|
|
@ -1,3 +1,5 @@
|
|||
:orphan: Ubuntu packages are obsolete
|
||||
|
||||
.. _topics-ubuntu:
|
||||
|
||||
===============
|
||||
|
|
@ -11,6 +13,9 @@ those in Ubuntu, and more stable too since they're continuously built from
|
|||
`GitHub repo`_ (master & stable branches) and so they contain the latest bug
|
||||
fixes.
|
||||
|
||||
.. caution:: These packages are currently not updated and may not work on
|
||||
Ubuntu 16.04 and above, see :issue:`2076` and :issue:`2137`.
|
||||
|
||||
To use the packages:
|
||||
|
||||
1. Import the GPG key used to sign Scrapy packages into APT keyring::
|
||||
|
|
|
|||
|
|
@ -7,22 +7,33 @@ Versioning and API Stability
|
|||
Versioning
|
||||
==========
|
||||
|
||||
Scrapy uses the `odd-numbered versions for development releases`_.
|
||||
|
||||
There are 3 numbers in a Scrapy version: *A.B.C*
|
||||
|
||||
* *A* is the major version. This will rarely change and will signify very
|
||||
large changes.
|
||||
* *B* is the release number. This will include many changes including features
|
||||
and things that possibly break backwards compatibility. Even Bs will be
|
||||
stable branches, and odd Bs will be development.
|
||||
and things that possibly break backwards compatibility, although we strive to
|
||||
keep theses cases at a minimum.
|
||||
* *C* is the bugfix release number.
|
||||
|
||||
Backward-incompatibilities are explicitly mentioned in the :ref:`release notes <news>`,
|
||||
and may require special attention before upgrading.
|
||||
|
||||
Development releases do not follow 3-numbers version and are generally
|
||||
released as ``dev`` suffixed versions, e.g. ``1.3dev``.
|
||||
|
||||
.. note::
|
||||
With Scrapy 0.* series, Scrapy used `odd-numbered versions for development releases`_.
|
||||
This is not the case anymore from Scrapy 1.0 onwards.
|
||||
|
||||
Starting with Scrapy 1.0, all releases should be considered production-ready.
|
||||
|
||||
For example:
|
||||
|
||||
* *0.14.1* is the first bugfix release of the *0.14* series (safe to use in
|
||||
* *1.1.1* is the first bugfix release of the *1.1* series (safe to use in
|
||||
production)
|
||||
|
||||
|
||||
API Stability
|
||||
=============
|
||||
|
||||
|
|
|
|||
|
|
@ -1,4 +1,4 @@
|
|||
Twisted>=10.0.0
|
||||
Twisted>=13.1.0
|
||||
lxml
|
||||
pyOpenSSL
|
||||
cssselect>=0.9
|
||||
|
|
|
|||
|
|
@ -1 +1 @@
|
|||
1.2.0dev2
|
||||
1.3.0
|
||||
|
|
|
|||
|
|
@ -5,6 +5,7 @@ from w3lib.url import is_url
|
|||
from scrapy.commands import ScrapyCommand
|
||||
from scrapy.http import Request
|
||||
from scrapy.exceptions import UsageError
|
||||
from scrapy.utils.datatypes import SequenceExclude
|
||||
from scrapy.utils.spider import spidercls_for_request, DefaultSpider
|
||||
|
||||
class Command(ScrapyCommand):
|
||||
|
|
@ -27,6 +28,8 @@ class Command(ScrapyCommand):
|
|||
help="use this spider")
|
||||
parser.add_option("--headers", dest="headers", action="store_true", \
|
||||
help="print response HTTP headers instead of body")
|
||||
parser.add_option("--no-redirect", dest="no_redirect", action="store_true", \
|
||||
default=False, help="do not handle HTTP 3xx status codes and print response as-is")
|
||||
|
||||
def _print_headers(self, headers, prefix):
|
||||
for key, values in headers.items():
|
||||
|
|
@ -50,7 +53,12 @@ class Command(ScrapyCommand):
|
|||
raise UsageError()
|
||||
cb = lambda x: self._print_response(x, opts)
|
||||
request = Request(args[0], callback=cb, dont_filter=True)
|
||||
request.meta['handle_httpstatus_all'] = True
|
||||
# by default, let the framework handle redirects,
|
||||
# i.e. command handles all codes expect 3xx
|
||||
if not opts.no_redirect:
|
||||
request.meta['handle_httpstatus_list'] = SequenceExclude(range(300, 400))
|
||||
else:
|
||||
request.meta['handle_httpstatus_all'] = True
|
||||
|
||||
spidercls = DefaultSpider
|
||||
spider_loader = self.crawler_process.spider_loader
|
||||
|
|
|
|||
|
|
@ -121,8 +121,8 @@ class Command(ScrapyCommand):
|
|||
def get_callback_from_rules(self, spider, response):
|
||||
if getattr(spider, 'rules', None):
|
||||
for rule in spider.rules:
|
||||
if rule.link_extractor.matches(response.url) and rule.callback:
|
||||
return rule.callback
|
||||
if rule.link_extractor.matches(response.url):
|
||||
return rule.callback or "parse"
|
||||
else:
|
||||
logger.error('No CrawlSpider rules found in spider %(spider)r, '
|
||||
'please specify a callback to use for parsing',
|
||||
|
|
@ -166,6 +166,11 @@ class Command(ScrapyCommand):
|
|||
if not cb:
|
||||
if opts.rules and self.first_response == response:
|
||||
cb = self.get_callback_from_rules(spider, response)
|
||||
|
||||
if not cb:
|
||||
logger.error('Cannot find a rule that matches %(url)r in spider: %(spider)s',
|
||||
{'url': response.url, 'spider': spider.name})
|
||||
return
|
||||
else:
|
||||
cb = 'parse'
|
||||
|
||||
|
|
|
|||
|
|
@ -36,6 +36,8 @@ class Command(ScrapyCommand):
|
|||
help="evaluate the code in the shell, print the result and exit")
|
||||
parser.add_option("--spider", dest="spider",
|
||||
help="use this spider")
|
||||
parser.add_option("--no-redirect", dest="no_redirect", action="store_true", \
|
||||
default=False, help="do not handle HTTP 3xx status codes and print response as-is")
|
||||
|
||||
def update_vars(self, vars):
|
||||
"""You can use this function to update the Scrapy objects that will be
|
||||
|
|
@ -68,7 +70,7 @@ class Command(ScrapyCommand):
|
|||
self._start_crawler_thread()
|
||||
|
||||
shell = Shell(crawler, update_vars=self.update_vars, code=opts.code)
|
||||
shell.start(url=url)
|
||||
shell.start(url=url, redirect=not opts.no_redirect)
|
||||
|
||||
def _start_crawler_thread(self):
|
||||
t = Thread(target=self.crawler_process.start,
|
||||
|
|
|
|||
|
|
@ -17,6 +17,7 @@ TEMPLATES_TO_RENDER = (
|
|||
('${project_name}', 'settings.py.tmpl'),
|
||||
('${project_name}', 'items.py.tmpl'),
|
||||
('${project_name}', 'pipelines.py.tmpl'),
|
||||
('${project_name}', 'middlewares.py.tmpl'),
|
||||
)
|
||||
|
||||
IGNORE = ignore_patterns('*.pyc', '.svn')
|
||||
|
|
|
|||
|
|
@ -26,12 +26,25 @@ class Command(ScrapyCommand):
|
|||
|
||||
def run(self, args, opts):
|
||||
if opts.verbose:
|
||||
import cssselect
|
||||
import parsel
|
||||
import lxml.etree
|
||||
import w3lib
|
||||
|
||||
lxml_version = ".".join(map(str, lxml.etree.LXML_VERSION))
|
||||
libxml2_version = ".".join(map(str, lxml.etree.LIBXML_VERSION))
|
||||
|
||||
try:
|
||||
w3lib_version = w3lib.__version__
|
||||
except AttributeError:
|
||||
w3lib_version = "<1.14.3"
|
||||
|
||||
print("Scrapy : %s" % scrapy.__version__)
|
||||
print("lxml : %s" % lxml_version)
|
||||
print("libxml2 : %s" % libxml2_version)
|
||||
print("cssselect : %s" % cssselect.__version__)
|
||||
print("parsel : %s" % parsel.__version__)
|
||||
print("w3lib : %s" % w3lib_version)
|
||||
print("Twisted : %s" % twisted.version.short())
|
||||
print("Python : %s" % sys.version.replace("\n", "- "))
|
||||
print("pyOpenSSL : %s" % self._get_openssl_version())
|
||||
|
|
|
|||
|
|
@ -13,7 +13,7 @@ try:
|
|||
from twisted.web.client import BrowserLikePolicyForHTTPS
|
||||
from twisted.web.iweb import IPolicyForHTTPS
|
||||
|
||||
from scrapy.core.downloader.tls import ScrapyClientTLSOptions
|
||||
from scrapy.core.downloader.tls import ScrapyClientTLSOptions, DEFAULT_CIPHERS
|
||||
|
||||
|
||||
@implementer(IPolicyForHTTPS)
|
||||
|
|
@ -46,7 +46,9 @@ try:
|
|||
# not calling super(..., self).__init__
|
||||
return CertificateOptions(verify=False,
|
||||
method=getattr(self, 'method',
|
||||
getattr(self, '_ssl_method', None)))
|
||||
getattr(self, '_ssl_method', None)),
|
||||
fixBrokenPeers=True,
|
||||
acceptableCiphers=DEFAULT_CIPHERS)
|
||||
|
||||
# kept for old-style HTTP/1.0 downloader context twisted calls,
|
||||
# e.g. connectSSL()
|
||||
|
|
|
|||
|
|
@ -13,8 +13,9 @@ from twisted.web.http_headers import Headers as TxHeaders
|
|||
from twisted.web.iweb import IBodyProducer, UNKNOWN_LENGTH
|
||||
from twisted.internet.error import TimeoutError
|
||||
from twisted.web.http import PotentialDataLoss
|
||||
from scrapy.xlib.tx import Agent, ProxyAgent, ResponseDone, \
|
||||
HTTPConnectionPool, TCP4ClientEndpoint
|
||||
from twisted.web.client import Agent, ProxyAgent, ResponseDone, \
|
||||
HTTPConnectionPool
|
||||
from twisted.internet.endpoints import TCP4ClientEndpoint
|
||||
|
||||
from scrapy.http import Headers
|
||||
from scrapy.responsetypes import responsetypes
|
||||
|
|
@ -318,14 +319,13 @@ class ScrapyAgent(object):
|
|||
expected_size = txresponse.length if txresponse.length != UNKNOWN_LENGTH else -1
|
||||
|
||||
if maxsize and expected_size > maxsize:
|
||||
error_message = ("Cancelling download of {url}: expected response "
|
||||
"size ({size}) larger than "
|
||||
"download max size ({maxsize})."
|
||||
).format(url=request.url, size=expected_size, maxsize=maxsize)
|
||||
error_msg = ("Cancelling download of %(url)s: expected response "
|
||||
"size (%(size)s) larger than download max size (%(maxsize)s).")
|
||||
error_args = {'url': request.url, 'size': expected_size, 'maxsize': maxsize}
|
||||
|
||||
logger.error(error_message)
|
||||
logger.error(error_msg, error_args)
|
||||
txresponse._transport._producer.loseConnection()
|
||||
raise defer.CancelledError(error_message)
|
||||
raise defer.CancelledError(error_msg % error_args)
|
||||
|
||||
if warnsize and expected_size > warnsize:
|
||||
logger.warning("Expected response size (%(size)s) larger than "
|
||||
|
|
|
|||
|
|
@ -28,6 +28,7 @@ try:
|
|||
SSL_CB_HANDSHAKE_START = 0x10
|
||||
SSL_CB_HANDSHAKE_DONE = 0x20
|
||||
|
||||
from twisted.internet.ssl import AcceptableCiphers
|
||||
from twisted.internet._sslverify import (ClientTLSOptions,
|
||||
_maybeSetHostNameIndication,
|
||||
verifyHostname,
|
||||
|
|
@ -60,6 +61,8 @@ try:
|
|||
'from host "{}" (exception: {})'.format(
|
||||
self._hostnameASCII, repr(e)))
|
||||
|
||||
DEFAULT_CIPHERS = AcceptableCiphers.fromOpenSSLCipherString('DEFAULT')
|
||||
|
||||
except ImportError:
|
||||
# ImportError should not matter for older Twisted versions
|
||||
# as the above is not used in the fallback ScrapyClientContextFactory
|
||||
|
|
|
|||
|
|
@ -48,7 +48,8 @@ class Slot(object):
|
|||
if self.closing and not self.inprogress:
|
||||
if self.nextcall:
|
||||
self.nextcall.cancel()
|
||||
self.heartbeat.stop()
|
||||
if self.heartbeat.running:
|
||||
self.heartbeat.stop()
|
||||
self.closing.callback(None)
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -89,8 +89,8 @@ class Scheduler(object):
|
|||
msg = ("Unable to serialize request: %(request)s - reason:"
|
||||
" %(reason)s - no more unserializable requests will be"
|
||||
" logged (stats being collected)")
|
||||
logger.error(msg, {'request': request, 'reason': e},
|
||||
exc_info=True, extra={'spider': self.spider})
|
||||
logger.warning(msg, {'request': request, 'reason': e},
|
||||
exc_info=True, extra={'spider': self.spider})
|
||||
self.logunser = False
|
||||
self.stats.inc_value('scheduler/unserializable',
|
||||
spider=self.spider)
|
||||
|
|
|
|||
|
|
@ -119,7 +119,7 @@ class Scraper(object):
|
|||
self._scrape(response, request, spider).chainDeferred(deferred)
|
||||
|
||||
def _scrape(self, response, request, spider):
|
||||
"""Handle the downloaded response or failure trough the spider
|
||||
"""Handle the downloaded response or failure through the spider
|
||||
callback/errback"""
|
||||
assert isinstance(response, (Response, Failure))
|
||||
|
||||
|
|
|
|||
|
|
@ -1,6 +1,14 @@
|
|||
import warnings
|
||||
|
||||
from scrapy.exceptions import ScrapyDeprecationWarning
|
||||
from scrapy.utils.http import decode_chunked_transfer
|
||||
|
||||
|
||||
warnings.warn("Module `scrapy.downloadermiddlewares.chunked` is deprecated, "
|
||||
"chunked transfers are supported by default.",
|
||||
ScrapyDeprecationWarning, stacklevel=2)
|
||||
|
||||
|
||||
class ChunkedTransferMiddleware(object):
|
||||
"""This middleware adds support for chunked transfer encoding, as
|
||||
documented in: http://en.wikipedia.org/wiki/Chunked_transfer_encoding
|
||||
|
|
|
|||
|
|
@ -3,10 +3,10 @@ from twisted.internet import defer
|
|||
from twisted.internet.error import TimeoutError, DNSLookupError, \
|
||||
ConnectionRefusedError, ConnectionDone, ConnectError, \
|
||||
ConnectionLost, TCPTimedOutError
|
||||
from twisted.web.client import ResponseFailed
|
||||
from scrapy import signals
|
||||
from scrapy.exceptions import NotConfigured, IgnoreRequest
|
||||
from scrapy.utils.misc import load_object
|
||||
from scrapy.xlib.tx import ResponseFailed
|
||||
|
||||
|
||||
class HttpCacheMiddleware(object):
|
||||
|
|
|
|||
|
|
@ -1,9 +1,10 @@
|
|||
import logging
|
||||
from six.moves.urllib.parse import urljoin
|
||||
|
||||
from w3lib.url import safe_url_string
|
||||
|
||||
from scrapy.http import HtmlResponse
|
||||
from scrapy.utils.response import get_meta_refresh
|
||||
from scrapy.utils.python import to_native_str
|
||||
from scrapy.exceptions import IgnoreRequest, NotConfigured
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
|
@ -65,8 +66,7 @@ class RedirectMiddleware(BaseRedirectMiddleware):
|
|||
if 'Location' not in response.headers or response.status not in allowed_status:
|
||||
return response
|
||||
|
||||
# HTTP header is ascii or latin1, redirected url will be percent-encoded utf-8
|
||||
location = to_native_str(response.headers['location'].decode('latin1'))
|
||||
location = safe_url_string(response.headers['location'])
|
||||
|
||||
redirected_url = urljoin(request.url, location)
|
||||
|
||||
|
|
|
|||
|
|
@ -17,10 +17,10 @@ from twisted.internet import defer
|
|||
from twisted.internet.error import TimeoutError, DNSLookupError, \
|
||||
ConnectionRefusedError, ConnectionDone, ConnectError, \
|
||||
ConnectionLost, TCPTimedOutError
|
||||
from twisted.web.client import ResponseFailed
|
||||
|
||||
from scrapy.exceptions import NotConfigured
|
||||
from scrapy.utils.response import response_status_message
|
||||
from scrapy.xlib.tx import ResponseFailed
|
||||
from scrapy.core.downloader.handlers.http11 import TunnelError
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
|
|
|||
|
|
@ -13,6 +13,7 @@ from scrapy.exceptions import NotConfigured, IgnoreRequest
|
|||
from scrapy.http import Request
|
||||
from scrapy.utils.httpobj import urlparse_cached
|
||||
from scrapy.utils.log import failure_to_exc_info
|
||||
from scrapy.utils.python import to_native_str
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
|
@ -40,7 +41,8 @@ class RobotsTxtMiddleware(object):
|
|||
return d
|
||||
|
||||
def process_request_2(self, rp, request, spider):
|
||||
if rp is not None and not rp.can_fetch(self._useragent, request.url):
|
||||
if rp is not None and not rp.can_fetch(
|
||||
to_native_str(self._useragent), request.url):
|
||||
logger.debug("Forbidden by robots.txt: %(request)s",
|
||||
{'request': request}, extra={'spider': spider})
|
||||
raise IgnoreRequest()
|
||||
|
|
@ -94,7 +96,9 @@ class RobotsTxtMiddleware(object):
|
|||
# Running rp.parse() will set rp state from
|
||||
# 'disallow all' to 'allow any'.
|
||||
pass
|
||||
rp.parse(body.splitlines())
|
||||
# stdlib's robotparser expects native 'str' ;
|
||||
# with unicode input, non-ASCII encoded bytes decoding fails in Python2
|
||||
rp.parse(to_native_str(body).splitlines())
|
||||
|
||||
rp_dfd = self._parsers[netloc]
|
||||
self._parsers[netloc] = rp
|
||||
|
|
|
|||
|
|
@ -170,6 +170,7 @@ class FeedExporter(object):
|
|||
if not self._exporter_supported(self.format):
|
||||
raise NotConfigured
|
||||
self.store_empty = settings.getbool('FEED_STORE_EMPTY')
|
||||
self._exporting = False
|
||||
self.export_fields = settings.getlist('FEED_EXPORT_FIELDS') or None
|
||||
uripar = settings['FEED_URI_PARAMS']
|
||||
self._uripar = load_object(uripar) if uripar else lambda x, y: None
|
||||
|
|
@ -188,14 +189,18 @@ class FeedExporter(object):
|
|||
file = storage.open(spider)
|
||||
exporter = self._get_exporter(file, fields_to_export=self.export_fields,
|
||||
encoding=self.export_encoding)
|
||||
exporter.start_exporting()
|
||||
if self.store_empty:
|
||||
exporter.start_exporting()
|
||||
self._exporting = True
|
||||
self.slot = SpiderSlot(file, exporter, storage, uri)
|
||||
|
||||
def close_spider(self, spider):
|
||||
slot = self.slot
|
||||
if not slot.itemcount and not self.store_empty:
|
||||
return
|
||||
slot.exporter.finish_exporting()
|
||||
if self._exporting:
|
||||
slot.exporter.finish_exporting()
|
||||
self._exporting = False
|
||||
logfmt = "%s %%(format)s feed (%%(itemcount)d items) in: %%(uri)s"
|
||||
log_args = {'format': self.format,
|
||||
'itemcount': slot.itemcount,
|
||||
|
|
@ -210,6 +215,9 @@ class FeedExporter(object):
|
|||
|
||||
def item_scraped(self, item, spider):
|
||||
slot = self.slot
|
||||
if not self._exporting:
|
||||
slot.exporter.start_exporting()
|
||||
self._exporting = True
|
||||
slot.exporter.export_item(item)
|
||||
slot.itemcount += 1
|
||||
return item
|
||||
|
|
|
|||
|
|
@ -15,6 +15,7 @@ class LogStats(object):
|
|||
self.stats = stats
|
||||
self.interval = interval
|
||||
self.multiplier = 60.0 / self.interval
|
||||
self.task = None
|
||||
|
||||
@classmethod
|
||||
def from_crawler(cls, crawler):
|
||||
|
|
@ -47,5 +48,5 @@ class LogStats(object):
|
|||
logger.info(msg, log_args, extra={'spider': spider})
|
||||
|
||||
def spider_closed(self, spider, reason):
|
||||
if self.task.running:
|
||||
if self.task and self.task.running:
|
||||
self.task.stop()
|
||||
|
|
|
|||
|
|
@ -9,6 +9,8 @@ from six.moves.urllib.parse import urljoin
|
|||
from scrapy.http.headers import Headers
|
||||
from scrapy.utils.trackref import object_ref
|
||||
from scrapy.http.common import obsolete_setter
|
||||
from scrapy.exceptions import NotSupported
|
||||
|
||||
|
||||
class Response(object_ref):
|
||||
|
||||
|
|
@ -80,3 +82,22 @@ class Response(object_ref):
|
|||
"""Join this Response's url with a possible relative url to form an
|
||||
absolute interpretation of the latter."""
|
||||
return urljoin(self.url, url)
|
||||
|
||||
@property
|
||||
def text(self):
|
||||
"""For subclasses of TextResponse, this will return the body
|
||||
as text (unicode object in Python 2 and str in Python 3)
|
||||
"""
|
||||
raise AttributeError("Response content isn't text")
|
||||
|
||||
def css(self, *a, **kw):
|
||||
"""Shortcut method implemented only by responses whose content
|
||||
is text (subclasses of TextResponse).
|
||||
"""
|
||||
raise NotSupported("Response content isn't text")
|
||||
|
||||
def xpath(self, *a, **kw):
|
||||
"""Shortcut method implemented only by responses whose content
|
||||
is text (subclasses of TextResponse).
|
||||
"""
|
||||
raise NotSupported("Response content isn't text")
|
||||
|
|
|
|||
|
|
@ -3,6 +3,7 @@ HTMLParser-based link extractor
|
|||
"""
|
||||
|
||||
import warnings
|
||||
import six
|
||||
from six.moves.html_parser import HTMLParser
|
||||
from six.moves.urllib.parse import urljoin
|
||||
|
||||
|
|
@ -39,7 +40,7 @@ class HtmlParserLinkExtractor(HTMLParser):
|
|||
ret = []
|
||||
base_url = urljoin(response_url, self.base_url) if self.base_url else response_url
|
||||
for link in links:
|
||||
if isinstance(link.url, unicode):
|
||||
if isinstance(link.url, six.text_type):
|
||||
link.url = link.url.encode(response_encoding)
|
||||
try:
|
||||
link.url = urljoin(base_url, link.url)
|
||||
|
|
|
|||
|
|
@ -1,6 +1,7 @@
|
|||
"""
|
||||
SGMLParser-based Link extractors
|
||||
"""
|
||||
import six
|
||||
from six.moves.urllib.parse import urljoin
|
||||
import warnings
|
||||
from sgmllib import SGMLParser
|
||||
|
|
@ -40,7 +41,7 @@ class BaseSgmlLinkExtractor(SGMLParser):
|
|||
if base_url is None:
|
||||
base_url = urljoin(response_url, self.base_url) if self.base_url else response_url
|
||||
for link in self.links:
|
||||
if isinstance(link.url, unicode):
|
||||
if isinstance(link.url, six.text_type):
|
||||
link.url = link.url.encode(response_encoding)
|
||||
try:
|
||||
link.url = urljoin(base_url, link.url)
|
||||
|
|
|
|||
|
|
@ -21,6 +21,8 @@ else:
|
|||
|
||||
from twisted.internet import defer, reactor, ssl
|
||||
|
||||
from .utils.misc import arg_to_iter
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
|
|
@ -48,6 +50,10 @@ class MailSender(object):
|
|||
msg = MIMEMultipart()
|
||||
else:
|
||||
msg = MIMENonMultipart(*mimetype.split('/', 1))
|
||||
|
||||
to = list(arg_to_iter(to))
|
||||
cc = list(arg_to_iter(cc))
|
||||
|
||||
msg['From'] = self.mailfrom
|
||||
msg['To'] = COMMASPACE.join(to)
|
||||
msg['Date'] = formatdate(localtime=True)
|
||||
|
|
|
|||
|
|
@ -233,7 +233,8 @@ class FilesPipeline(MediaPipeline):
|
|||
cls_name = "FilesPipeline"
|
||||
self.store = self._get_store(store_uri)
|
||||
resolve = functools.partial(self._key_for_pipe,
|
||||
base_class_name=cls_name)
|
||||
base_class_name=cls_name,
|
||||
settings=settings)
|
||||
self.expires = settings.getint(
|
||||
resolve('FILES_EXPIRES'), self.EXPIRES
|
||||
)
|
||||
|
|
|
|||
|
|
@ -55,7 +55,8 @@ class ImagesPipeline(FilesPipeline):
|
|||
settings = Settings(settings)
|
||||
|
||||
resolve = functools.partial(self._key_for_pipe,
|
||||
base_class_name="ImagesPipeline")
|
||||
base_class_name="ImagesPipeline",
|
||||
settings=settings)
|
||||
self.expires = settings.getint(
|
||||
resolve("IMAGES_EXPIRES"), self.EXPIRES
|
||||
)
|
||||
|
|
|
|||
|
|
@ -28,7 +28,8 @@ class MediaPipeline(object):
|
|||
self.download_func = download_func
|
||||
|
||||
|
||||
def _key_for_pipe(self, key, base_class_name=None):
|
||||
def _key_for_pipe(self, key, base_class_name=None,
|
||||
settings=None):
|
||||
"""
|
||||
>>> MediaPipeline()._key_for_pipe("IMAGES")
|
||||
'IMAGES'
|
||||
|
|
@ -38,9 +39,11 @@ class MediaPipeline(object):
|
|||
'MYPIPE_IMAGES'
|
||||
"""
|
||||
class_name = self.__class__.__name__
|
||||
if class_name == base_class_name or not base_class_name:
|
||||
formatted_key = "{}_{}".format(class_name.upper(), key)
|
||||
if class_name == base_class_name or not base_class_name \
|
||||
or (settings and not settings.get(formatted_key)):
|
||||
return key
|
||||
return "{}_{}".format(class_name.upper(), key)
|
||||
return formatted_key
|
||||
|
||||
@classmethod
|
||||
def from_crawler(cls, crawler):
|
||||
|
|
|
|||
|
|
@ -102,7 +102,6 @@ DOWNLOADER_MIDDLEWARES_BASE = {
|
|||
'scrapy.downloadermiddlewares.redirect.RedirectMiddleware': 600,
|
||||
'scrapy.downloadermiddlewares.cookies.CookiesMiddleware': 700,
|
||||
'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware': 750,
|
||||
'scrapy.downloadermiddlewares.chunked.ChunkedTransferMiddleware': 830,
|
||||
'scrapy.downloadermiddlewares.stats.DownloaderStats': 850,
|
||||
'scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware': 900,
|
||||
# Downloader side
|
||||
|
|
@ -192,6 +191,7 @@ LOG_DATEFORMAT = '%Y-%m-%d %H:%M:%S'
|
|||
LOG_STDOUT = False
|
||||
LOG_LEVEL = 'DEBUG'
|
||||
LOG_FILE = None
|
||||
LOG_SHORT_NAMES = False
|
||||
|
||||
SCHEDULER_DEBUG = False
|
||||
|
||||
|
|
|
|||
|
|
@ -20,6 +20,7 @@ from scrapy.item import BaseItem
|
|||
from scrapy.settings import Settings
|
||||
from scrapy.spiders import Spider
|
||||
from scrapy.utils.console import start_python_console
|
||||
from scrapy.utils.datatypes import SequenceExclude
|
||||
from scrapy.utils.misc import load_object
|
||||
from scrapy.utils.response import open_in_browser
|
||||
from scrapy.utils.conf import get_config
|
||||
|
|
@ -40,11 +41,11 @@ class Shell(object):
|
|||
self.code = code
|
||||
self.vars = {}
|
||||
|
||||
def start(self, url=None, request=None, response=None, spider=None):
|
||||
def start(self, url=None, request=None, response=None, spider=None, redirect=True):
|
||||
# disable accidental Ctrl-C key press from shutting down the engine
|
||||
signal.signal(signal.SIGINT, signal.SIG_IGN)
|
||||
if url:
|
||||
self.fetch(url, spider)
|
||||
self.fetch(url, spider, redirect=redirect)
|
||||
elif request:
|
||||
self.fetch(request, spider)
|
||||
elif response:
|
||||
|
|
@ -98,14 +99,16 @@ class Shell(object):
|
|||
self.spider = spider
|
||||
return spider
|
||||
|
||||
def fetch(self, request_or_url, spider=None):
|
||||
def fetch(self, request_or_url, spider=None, redirect=True, **kwargs):
|
||||
if isinstance(request_or_url, Request):
|
||||
request = request_or_url
|
||||
url = request.url
|
||||
else:
|
||||
url = any_to_uri(request_or_url)
|
||||
request = Request(url, dont_filter=True)
|
||||
request.meta['handle_httpstatus_all'] = True
|
||||
request = Request(url, dont_filter=True, **kwargs)
|
||||
if redirect:
|
||||
request.meta['handle_httpstatus_list'] = SequenceExclude(range(300, 400))
|
||||
else:
|
||||
request.meta['handle_httpstatus_all'] = True
|
||||
response = None
|
||||
try:
|
||||
response, spider = threads.blockingCallFromThread(
|
||||
|
|
@ -115,6 +118,9 @@ class Shell(object):
|
|||
self.populate_vars(response, request, spider)
|
||||
|
||||
def populate_vars(self, response=None, request=None, spider=None):
|
||||
import scrapy
|
||||
|
||||
self.vars['scrapy'] = scrapy
|
||||
self.vars['crawler'] = self.crawler
|
||||
self.vars['item'] = self.item_class()
|
||||
self.vars['settings'] = self.crawler.settings
|
||||
|
|
@ -136,14 +142,18 @@ class Shell(object):
|
|||
def get_help(self):
|
||||
b = []
|
||||
b.append("Available Scrapy objects:")
|
||||
b.append(" scrapy scrapy module (contains scrapy.Request, scrapy.Selector, etc)")
|
||||
for k, v in sorted(self.vars.items()):
|
||||
if self._is_relevant(v):
|
||||
b.append(" %-10s %s" % (k, v))
|
||||
b.append("Useful shortcuts:")
|
||||
b.append(" shelp() Shell help (print this help)")
|
||||
if self.inthread:
|
||||
b.append(" fetch(req_or_url) Fetch request (or URL) and "
|
||||
"update local objects")
|
||||
b.append(" fetch(url[, redirect=True]) "
|
||||
"Fetch URL and update local objects "
|
||||
"(by default, redirects are followed)")
|
||||
b.append(" fetch(req) "
|
||||
"Fetch a scrapy.Request and update local objects ")
|
||||
b.append(" shelp() Shell help (print this help)")
|
||||
b.append(" view(response) View response in a browser")
|
||||
|
||||
return "\n".join("[s] %s" % l for l in b)
|
||||
|
|
|
|||
|
|
@ -1,5 +1,7 @@
|
|||
# -*- coding: utf-8 -*-
|
||||
from __future__ import absolute_import
|
||||
import traceback
|
||||
import warnings
|
||||
|
||||
from zope.interface import implementer
|
||||
|
||||
|
|
@ -18,15 +20,21 @@ class SpiderLoader(object):
|
|||
self.spider_modules = settings.getlist('SPIDER_MODULES')
|
||||
self._spiders = {}
|
||||
self._load_all_spiders()
|
||||
|
||||
|
||||
def _load_spiders(self, module):
|
||||
for spcls in iter_spider_classes(module):
|
||||
self._spiders[spcls.name] = spcls
|
||||
|
||||
def _load_all_spiders(self):
|
||||
for name in self.spider_modules:
|
||||
for module in walk_modules(name):
|
||||
self._load_spiders(module)
|
||||
try:
|
||||
for module in walk_modules(name):
|
||||
self._load_spiders(module)
|
||||
except ImportError as e:
|
||||
msg = ("\n{tb}Could not load spiders from module '{modname}'. "
|
||||
"Check SPIDER_MODULES setting".format(
|
||||
modname=name, tb=traceback.format_exc()))
|
||||
warnings.warn(msg, RuntimeWarning)
|
||||
|
||||
@classmethod
|
||||
def from_settings(cls, settings):
|
||||
|
|
|
|||
|
|
@ -46,7 +46,7 @@ class HttpErrorMiddleware(object):
|
|||
|
||||
def process_spider_exception(self, response, exception, spider):
|
||||
if isinstance(exception, HttpError):
|
||||
logger.debug(
|
||||
logger.info(
|
||||
"Ignoring response %(response)r: HTTP status code is not handled or not allowed",
|
||||
{'response': response}, extra={'spider': spider},
|
||||
)
|
||||
|
|
|
|||
|
|
@ -32,7 +32,7 @@ class SitemapSpider(Spider):
|
|||
|
||||
def _parse_sitemap(self, response):
|
||||
if response.url.endswith('/robots.txt'):
|
||||
for url in sitemap_urls_from_robots(response.text):
|
||||
for url in sitemap_urls_from_robots(response.text, base_url=response.url):
|
||||
yield Request(url, callback=self._parse_sitemap)
|
||||
else:
|
||||
body = self._get_sitemap_body(response)
|
||||
|
|
|
|||
|
|
@ -0,0 +1,56 @@
|
|||
# -*- coding: utf-8 -*-
|
||||
|
||||
# Define here the models for your spider middleware
|
||||
#
|
||||
# See documentation in:
|
||||
# http://doc.scrapy.org/en/latest/topics/spider-middleware.html
|
||||
|
||||
from scrapy import signals
|
||||
|
||||
|
||||
class ${ProjectName}SpiderMiddleware(object):
|
||||
# Not all methods need to be defined. If a method is not defined,
|
||||
# scrapy acts as if the spider middleware does not modify the
|
||||
# passed objects.
|
||||
|
||||
@classmethod
|
||||
def from_crawler(cls, crawler):
|
||||
# This method is used by Scrapy to create your spiders.
|
||||
s = cls()
|
||||
crawler.signals.connect(s.spider_opened, signal=signals.spider_opened)
|
||||
return s
|
||||
|
||||
def process_spider_input(response, spider):
|
||||
# Called for each response that goes through the spider
|
||||
# middleware and into the spider.
|
||||
|
||||
# Should return None or raise an exception.
|
||||
return None
|
||||
|
||||
def process_spider_output(response, result, spider):
|
||||
# Called with the results returned from the Spider, after
|
||||
# it has processed the response.
|
||||
|
||||
# Must return an iterable of Request, dict or Item objects.
|
||||
for i in result:
|
||||
yield i
|
||||
|
||||
def process_spider_exception(response, exception, spider):
|
||||
# Called when a spider or process_spider_input() method
|
||||
# (from other spider middleware) raises an exception.
|
||||
|
||||
# Should return either None or an iterable of Response, dict
|
||||
# or Item objects.
|
||||
pass
|
||||
|
||||
def process_start_requests(start_requests, spider):
|
||||
# Called with the start requests of the spider, and works
|
||||
# similarly to the process_spider_output() method, except
|
||||
# that it doesn’t have a response associated.
|
||||
|
||||
# Must return only requests (not items).
|
||||
for r in start_requests:
|
||||
yield r
|
||||
|
||||
def spider_opened(self, spider):
|
||||
spider.logger.info('Spider opened: %s' % spider.name)
|
||||
|
|
@ -47,7 +47,7 @@ ROBOTSTXT_OBEY = True
|
|||
# Enable or disable spider middlewares
|
||||
# See http://scrapy.readthedocs.org/en/latest/topics/spider-middleware.html
|
||||
#SPIDER_MIDDLEWARES = {
|
||||
# '$project_name.middlewares.MyCustomSpiderMiddleware': 543,
|
||||
# '$project_name.middlewares.${ProjectName}SpiderMiddleware': 543,
|
||||
#}
|
||||
|
||||
# Enable or disable downloader middlewares
|
||||
|
|
|
|||
|
|
@ -5,9 +5,7 @@ import scrapy
|
|||
class $classname(scrapy.Spider):
|
||||
name = "$name"
|
||||
allowed_domains = ["$domain"]
|
||||
start_urls = (
|
||||
'http://www.$domain/',
|
||||
)
|
||||
start_urls = ['http://$domain/']
|
||||
|
||||
def parse(self, response):
|
||||
pass
|
||||
|
|
|
|||
|
|
@ -7,7 +7,7 @@ from scrapy.spiders import CrawlSpider, Rule
|
|||
class $classname(CrawlSpider):
|
||||
name = '$name'
|
||||
allowed_domains = ['$domain']
|
||||
start_urls = ['http://www.$domain/']
|
||||
start_urls = ['http://$domain/']
|
||||
|
||||
rules = (
|
||||
Rule(LinkExtractor(allow=r'Items/'), callback='parse_item', follow=True),
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@ from scrapy.spiders import CSVFeedSpider
|
|||
class $classname(CSVFeedSpider):
|
||||
name = '$name'
|
||||
allowed_domains = ['$domain']
|
||||
start_urls = ['http://www.$domain/feed.csv']
|
||||
start_urls = ['http://$domain/feed.csv']
|
||||
# headers = ['id', 'name', 'description', 'image_link']
|
||||
# delimiter = '\t'
|
||||
|
||||
|
|
|
|||
|
|
@ -5,7 +5,7 @@ from scrapy.spiders import XMLFeedSpider
|
|||
class $classname(XMLFeedSpider):
|
||||
name = '$name'
|
||||
allowed_domains = ['$domain']
|
||||
start_urls = ['http://www.$domain/feed.xml']
|
||||
start_urls = ['http://$domain/feed.xml']
|
||||
iterator = 'iternodes' # you can change this; see the docs
|
||||
itertag = 'item' # change it accordingly
|
||||
|
||||
|
|
|
|||
|
|
@ -13,7 +13,12 @@ def _embed_ipython_shell(namespace={}, banner=''):
|
|||
@wraps(_embed_ipython_shell)
|
||||
def wrapper(namespace=namespace, banner=''):
|
||||
config = load_default_config()
|
||||
shell = InteractiveShellEmbed(
|
||||
# Always use .instace() to ensure _instance propagation to all parents
|
||||
# this is needed for <TAB> completion works well for new imports
|
||||
# and clear the instance to always have the fresh env
|
||||
# on repeated breaks like with inspect_response()
|
||||
InteractiveShellEmbed.clear_instance()
|
||||
shell = InteractiveShellEmbed.instance(
|
||||
banner1=banner, user_ns=namespace, config=config)
|
||||
shell()
|
||||
return wrapper
|
||||
|
|
|
|||
|
|
@ -304,3 +304,13 @@ class LocalCache(OrderedDict):
|
|||
while len(self) >= self.limit:
|
||||
self.popitem(last=False)
|
||||
super(LocalCache, self).__setitem__(key, value)
|
||||
|
||||
|
||||
class SequenceExclude(object):
|
||||
"""Object to test if an item is NOT within some sequence."""
|
||||
|
||||
def __init__(self, seq):
|
||||
self.seq = seq
|
||||
|
||||
def __contains__(self, item):
|
||||
return item not in self.seq
|
||||
|
|
|
|||
|
|
@ -131,7 +131,8 @@ def _get_handler(settings):
|
|||
)
|
||||
handler.setFormatter(formatter)
|
||||
handler.setLevel(settings.get('LOG_LEVEL'))
|
||||
handler.addFilter(TopLevelFormatter(['scrapy']))
|
||||
if settings.getbool('LOG_SHORT_NAMES'):
|
||||
handler.addFilter(TopLevelFormatter(['scrapy']))
|
||||
return handler
|
||||
|
||||
|
||||
|
|
@ -158,6 +159,10 @@ class StreamLogger(object):
|
|||
for line in buf.rstrip().splitlines():
|
||||
self.logger.log(self.log_level, line.rstrip())
|
||||
|
||||
def flush(self):
|
||||
for h in self.logger.handlers:
|
||||
h.flush()
|
||||
|
||||
|
||||
class LogCounterHandler(logging.Handler):
|
||||
"""Record log levels count into a crawler stats"""
|
||||
|
|
|
|||
|
|
@ -12,6 +12,7 @@ from scrapy.exceptions import NotConfigured
|
|||
ENVVAR = 'SCRAPY_SETTINGS_MODULE'
|
||||
DATADIR_CFG_SECTION = 'datadir'
|
||||
|
||||
|
||||
def inside_project():
|
||||
scrapy_module = os.environ.get('SCRAPY_SETTINGS_MODULE')
|
||||
if scrapy_module is not None:
|
||||
|
|
@ -23,6 +24,7 @@ def inside_project():
|
|||
return True
|
||||
return bool(closest_scrapy_cfg())
|
||||
|
||||
|
||||
def project_data_dir(project='default'):
|
||||
"""Return the current project data dir, creating it if it doesn't exist"""
|
||||
if not inside_project():
|
||||
|
|
@ -39,16 +41,22 @@ def project_data_dir(project='default'):
|
|||
os.makedirs(d)
|
||||
return d
|
||||
|
||||
|
||||
def data_path(path, createdir=False):
|
||||
"""If path is relative, return the given path inside the project data dir,
|
||||
otherwise return the path unmodified
|
||||
"""
|
||||
Return the given path joined with the .scrapy data directory.
|
||||
If given an absolute path, return it unmodified.
|
||||
"""
|
||||
if not isabs(path):
|
||||
path = join(project_data_dir(), path)
|
||||
if inside_project():
|
||||
path = join(project_data_dir(), path)
|
||||
else:
|
||||
path = join('.scrapy', path)
|
||||
if createdir and not exists(path):
|
||||
os.makedirs(path)
|
||||
return path
|
||||
|
||||
|
||||
def get_project_settings():
|
||||
if ENVVAR not in os.environ:
|
||||
project = os.environ.get('SCRAPY_PROJECT', 'default')
|
||||
|
|
|
|||
|
|
@ -4,7 +4,9 @@ Module for processing Sitemaps.
|
|||
Note: The main purpose of this module is to provide support for the
|
||||
SitemapSpider, its API is subject to change without notice.
|
||||
"""
|
||||
|
||||
import lxml.etree
|
||||
from six.moves.urllib.parse import urljoin
|
||||
|
||||
|
||||
class Sitemap(object):
|
||||
|
|
@ -34,10 +36,11 @@ class Sitemap(object):
|
|||
yield d
|
||||
|
||||
|
||||
def sitemap_urls_from_robots(robots_text):
|
||||
def sitemap_urls_from_robots(robots_text, base_url=None):
|
||||
"""Return an iterator over all sitemap urls contained in the given
|
||||
robots.txt file
|
||||
"""
|
||||
for line in robots_text.splitlines():
|
||||
if line.lstrip().lower().startswith('sitemap:'):
|
||||
yield line.split(':', 1)[1].strip()
|
||||
url = line.split(':', 1)[1].strip()
|
||||
yield urljoin(base_url, url)
|
||||
|
|
|
|||
|
|
@ -20,12 +20,20 @@ class SiteTest(object):
|
|||
return urljoin(self.baseurl, path)
|
||||
|
||||
|
||||
class NoMetaRefreshRedirect(util.Redirect):
|
||||
def render(self, request):
|
||||
content = util.Redirect.render(self, request)
|
||||
return content.replace(b'http-equiv=\"refresh\"',
|
||||
b'http-no-equiv=\"do-not-refresh-me\"')
|
||||
|
||||
|
||||
def test_site():
|
||||
r = resource.Resource()
|
||||
r.putChild(b"text", static.Data(b"Works", "text/plain"))
|
||||
r.putChild(b"html", static.Data(b"<body><p class='one'>Works</p><p class='two'>World</p></body>", "text/html"))
|
||||
r.putChild(b"enc-gb18030", static.Data(b"<p>gb18030 encoding</p>", "text/html; charset=gb18030"))
|
||||
r.putChild(b"redirect", util.Redirect(b"/redirected"))
|
||||
r.putChild(b"redirect-no-meta-refresh", NoMetaRefreshRedirect(b"/redirected"))
|
||||
r.putChild(b"redirected", static.Data(b"Redirected here", "text/plain"))
|
||||
return server.Site(r)
|
||||
|
||||
|
|
|
|||
|
|
@ -0,0 +1,19 @@
|
|||
from __future__ import absolute_import
|
||||
|
||||
import warnings
|
||||
from scrapy.exceptions import ScrapyDeprecationWarning
|
||||
|
||||
from twisted.web import client
|
||||
from twisted.internet import endpoints
|
||||
|
||||
Agent = client.Agent # since < 11.1
|
||||
ProxyAgent = client.ProxyAgent # since 11.1
|
||||
ResponseDone = client.ResponseDone # since 11.1
|
||||
ResponseFailed = client.ResponseFailed # since 11.1
|
||||
HTTPConnectionPool = client.HTTPConnectionPool # since 12.1
|
||||
TCP4ClientEndpoint = endpoints.TCP4ClientEndpoint # since 10.1
|
||||
|
||||
warnings.warn("Importing from scrapy.xlib.tx is deprecated and will"
|
||||
" no longer be supported in future Scrapy versions."
|
||||
" Update your code to import from twisted proper.",
|
||||
ScrapyDeprecationWarning, stacklevel=2)
|
||||
|
|
@ -1,57 +0,0 @@
|
|||
Copyright (c) 2001-2013
|
||||
Allen Short
|
||||
Andy Gayton
|
||||
Andrew Bennetts
|
||||
Antoine Pitrou
|
||||
Apple Computer, Inc.
|
||||
Benjamin Bruheim
|
||||
Bob Ippolito
|
||||
Canonical Limited
|
||||
Christopher Armstrong
|
||||
David Reid
|
||||
Donovan Preston
|
||||
Eric Mangold
|
||||
Eyal Lotem
|
||||
Itamar Turner-Trauring
|
||||
James Knight
|
||||
Jason A. Mobarak
|
||||
Jean-Paul Calderone
|
||||
Jessica McKellar
|
||||
Jonathan Jacobs
|
||||
Jonathan Lange
|
||||
Jonathan D. Simms
|
||||
Jürgen Hermann
|
||||
Kevin Horn
|
||||
Kevin Turner
|
||||
Mary Gardiner
|
||||
Matthew Lefkowitz
|
||||
Massachusetts Institute of Technology
|
||||
Moshe Zadka
|
||||
Paul Swartz
|
||||
Pavel Pergamenshchik
|
||||
Ralph Meijer
|
||||
Sean Riley
|
||||
Software Freedom Conservancy
|
||||
Travis B. Hartwell
|
||||
Thijs Triemstra
|
||||
Thomas Herve
|
||||
Timothy Allen
|
||||
|
||||
Permission is hereby granted, free of charge, to any person obtaining
|
||||
a copy of this software and associated documentation files (the
|
||||
"Software"), to deal in the Software without restriction, including
|
||||
without limitation the rights to use, copy, modify, merge, publish,
|
||||
distribute, sublicense, and/or sell copies of the Software, and to
|
||||
permit persons to whom the Software is furnished to do so, subject to
|
||||
the following conditions:
|
||||
|
||||
The above copyright notice and this permission notice shall be
|
||||
included in all copies or substantial portions of the Software.
|
||||
|
||||
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND,
|
||||
EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
|
||||
MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND
|
||||
NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE
|
||||
LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION
|
||||
OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION
|
||||
WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
|
||||
|
|
@ -1,2 +0,0 @@
|
|||
This source files are adapted copies from Twisted trunk to support HTTP1.1
|
||||
handler under Twisted >= 11.1 and Twisted <= 13.0.0
|
||||
|
|
@ -1,23 +0,0 @@
|
|||
from scrapy import twisted_version
|
||||
if twisted_version > (13, 0, 0):
|
||||
from twisted.web import client
|
||||
from twisted.internet import endpoints
|
||||
if twisted_version >= (11, 1, 0):
|
||||
from . import client, endpoints
|
||||
else:
|
||||
from scrapy.exceptions import NotSupported
|
||||
class _Mocked(object):
|
||||
def __init__(self, *args, **kw):
|
||||
raise NotSupported('HTTP1.1 not supported')
|
||||
class _Mock(object):
|
||||
def __getattr__(self, name):
|
||||
return _Mocked
|
||||
client = endpoints = _Mock()
|
||||
|
||||
|
||||
Agent = client.Agent
|
||||
ProxyAgent = client.ProxyAgent
|
||||
ResponseDone = client.ResponseDone
|
||||
ResponseFailed = client.ResponseFailed
|
||||
HTTPConnectionPool = client.HTTPConnectionPool
|
||||
TCP4ClientEndpoint = endpoints.TCP4ClientEndpoint
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
|
|
@ -1,587 +0,0 @@
|
|||
# -*- test-case-name: twisted.web.test -*-
|
||||
# Copyright (c) Twisted Matrix Laboratories.
|
||||
# See LICENSE for details.
|
||||
|
||||
"""
|
||||
Interface definitions for L{twisted.web}.
|
||||
|
||||
@var UNKNOWN_LENGTH: An opaque object which may be used as the value of
|
||||
L{IBodyProducer.length} to indicate that the length of the entity
|
||||
body is not known in advance.
|
||||
"""
|
||||
|
||||
from zope.interface import Interface, Attribute
|
||||
|
||||
from twisted.internet.interfaces import IPushProducer
|
||||
|
||||
|
||||
class IRequest(Interface):
|
||||
"""
|
||||
An HTTP request.
|
||||
|
||||
@since: 9.0
|
||||
"""
|
||||
|
||||
method = Attribute("A C{str} giving the HTTP method that was used.")
|
||||
uri = Attribute(
|
||||
"A C{str} giving the full encoded URI which was requested (including "
|
||||
"query arguments).")
|
||||
path = Attribute(
|
||||
"A C{str} giving the encoded query path of the request URI.")
|
||||
args = Attribute(
|
||||
"A mapping of decoded query argument names as C{str} to "
|
||||
"corresponding query argument values as C{list}s of C{str}. "
|
||||
"For example, for a URI with C{'foo=bar&foo=baz&quux=spam'} "
|
||||
"for its query part, C{args} will be C{{'foo': ['bar', 'baz'], "
|
||||
"'quux': ['spam']}}.")
|
||||
|
||||
received_headers = Attribute(
|
||||
"Backwards-compatibility access to C{requestHeaders}. Use "
|
||||
"C{requestHeaders} instead. C{received_headers} behaves mostly "
|
||||
"like a C{dict} and does not provide access to all header values.")
|
||||
|
||||
requestHeaders = Attribute(
|
||||
"A L{http_headers.Headers} instance giving all received HTTP request "
|
||||
"headers.")
|
||||
|
||||
content = Attribute(
|
||||
"A file-like object giving the request body. This may be a file on "
|
||||
"disk, a C{StringIO}, or some other type. The implementation is free "
|
||||
"to decide on a per-request basis.")
|
||||
|
||||
headers = Attribute(
|
||||
"Backwards-compatibility access to C{responseHeaders}. Use"
|
||||
"C{responseHeaders} instead. C{headers} behaves mostly like a "
|
||||
"C{dict} and does not provide access to all header values nor "
|
||||
"does it allow multiple values for one header to be set.")
|
||||
|
||||
responseHeaders = Attribute(
|
||||
"A L{http_headers.Headers} instance holding all HTTP response "
|
||||
"headers to be sent.")
|
||||
|
||||
def getHeader(key):
|
||||
"""
|
||||
Get an HTTP request header.
|
||||
|
||||
@type key: C{str}
|
||||
@param key: The name of the header to get the value of.
|
||||
|
||||
@rtype: C{str} or C{NoneType}
|
||||
@return: The value of the specified header, or C{None} if that header
|
||||
was not present in the request.
|
||||
"""
|
||||
|
||||
|
||||
def getCookie(key):
|
||||
"""
|
||||
Get a cookie that was sent from the network.
|
||||
"""
|
||||
|
||||
|
||||
def getAllHeaders():
|
||||
"""
|
||||
Return dictionary mapping the names of all received headers to the last
|
||||
value received for each.
|
||||
|
||||
Since this method does not return all header information,
|
||||
C{requestHeaders.getAllRawHeaders()} may be preferred.
|
||||
"""
|
||||
|
||||
|
||||
def getRequestHostname():
|
||||
"""
|
||||
Get the hostname that the user passed in to the request.
|
||||
|
||||
This will either use the Host: header (if it is available) or the
|
||||
host we are listening on if the header is unavailable.
|
||||
|
||||
@returns: the requested hostname
|
||||
@rtype: C{str}
|
||||
"""
|
||||
|
||||
|
||||
def getHost():
|
||||
"""
|
||||
Get my originally requesting transport's host.
|
||||
|
||||
@return: An L{IAddress<twisted.internet.interfaces.IAddress>}.
|
||||
"""
|
||||
|
||||
|
||||
def getClientIP():
|
||||
"""
|
||||
Return the IP address of the client who submitted this request.
|
||||
|
||||
@returns: the client IP address or C{None} if the request was submitted
|
||||
over a transport where IP addresses do not make sense.
|
||||
@rtype: L{str} or C{NoneType}
|
||||
"""
|
||||
|
||||
|
||||
def getClient():
|
||||
"""
|
||||
Return the hostname of the IP address of the client who submitted this
|
||||
request, if possible.
|
||||
|
||||
This method is B{deprecated}. See L{getClientIP} instead.
|
||||
|
||||
@rtype: C{NoneType} or L{str}
|
||||
@return: The canonical hostname of the client, as determined by
|
||||
performing a name lookup on the IP address of the client.
|
||||
"""
|
||||
|
||||
|
||||
def getUser():
|
||||
"""
|
||||
Return the HTTP user sent with this request, if any.
|
||||
|
||||
If no user was supplied, return the empty string.
|
||||
|
||||
@returns: the HTTP user, if any
|
||||
@rtype: C{str}
|
||||
"""
|
||||
|
||||
|
||||
def getPassword():
|
||||
"""
|
||||
Return the HTTP password sent with this request, if any.
|
||||
|
||||
If no password was supplied, return the empty string.
|
||||
|
||||
@returns: the HTTP password, if any
|
||||
@rtype: C{str}
|
||||
"""
|
||||
|
||||
|
||||
def isSecure():
|
||||
"""
|
||||
Return True if this request is using a secure transport.
|
||||
|
||||
Normally this method returns True if this request's HTTPChannel
|
||||
instance is using a transport that implements ISSLTransport.
|
||||
|
||||
This will also return True if setHost() has been called
|
||||
with ssl=True.
|
||||
|
||||
@returns: True if this request is secure
|
||||
@rtype: C{bool}
|
||||
"""
|
||||
|
||||
|
||||
def getSession(sessionInterface=None):
|
||||
"""
|
||||
Look up the session associated with this request or create a new one if
|
||||
there is not one.
|
||||
|
||||
@return: The L{Session} instance identified by the session cookie in
|
||||
the request, or the C{sessionInterface} component of that session
|
||||
if C{sessionInterface} is specified.
|
||||
"""
|
||||
|
||||
|
||||
def URLPath():
|
||||
"""
|
||||
@return: A L{URLPath} instance which identifies the URL for which this
|
||||
request is.
|
||||
"""
|
||||
|
||||
|
||||
def prePathURL():
|
||||
"""
|
||||
@return: At any time during resource traversal, a L{str} giving an
|
||||
absolute URL to the most nested resource which has yet been
|
||||
reached.
|
||||
"""
|
||||
|
||||
|
||||
def rememberRootURL():
|
||||
"""
|
||||
Remember the currently-processed part of the URL for later
|
||||
recalling.
|
||||
"""
|
||||
|
||||
|
||||
def getRootURL():
|
||||
"""
|
||||
Get a previously-remembered URL.
|
||||
"""
|
||||
|
||||
|
||||
# Methods for outgoing response
|
||||
def finish():
|
||||
"""
|
||||
Indicate that the response to this request is complete.
|
||||
"""
|
||||
|
||||
|
||||
def write(data):
|
||||
"""
|
||||
Write some data to the body of the response to this request. Response
|
||||
headers are written the first time this method is called, after which
|
||||
new response headers may not be added.
|
||||
"""
|
||||
|
||||
|
||||
def addCookie(k, v, expires=None, domain=None, path=None, max_age=None, comment=None, secure=None):
|
||||
"""
|
||||
Set an outgoing HTTP cookie.
|
||||
|
||||
In general, you should consider using sessions instead of cookies, see
|
||||
L{twisted.web.server.Request.getSession} and the
|
||||
L{twisted.web.server.Session} class for details.
|
||||
"""
|
||||
|
||||
|
||||
def setResponseCode(code, message=None):
|
||||
"""
|
||||
Set the HTTP response code.
|
||||
"""
|
||||
|
||||
|
||||
def setHeader(k, v):
|
||||
"""
|
||||
Set an HTTP response header. Overrides any previously set values for
|
||||
this header.
|
||||
|
||||
@type name: C{str}
|
||||
@param name: The name of the header for which to set the value.
|
||||
|
||||
@type value: C{str}
|
||||
@param value: The value to set for the named header.
|
||||
"""
|
||||
|
||||
|
||||
def redirect(url):
|
||||
"""
|
||||
Utility function that does a redirect.
|
||||
|
||||
The request should have finish() called after this.
|
||||
"""
|
||||
|
||||
|
||||
def setLastModified(when):
|
||||
"""
|
||||
Set the C{Last-Modified} time for the response to this request.
|
||||
|
||||
If I am called more than once, I ignore attempts to set Last-Modified
|
||||
earlier, only replacing the Last-Modified time if it is to a later
|
||||
value.
|
||||
|
||||
If I am a conditional request, I may modify my response code to
|
||||
L{NOT_MODIFIED<http.NOT_MODIFIED>} if appropriate for the time given.
|
||||
|
||||
@param when: The last time the resource being returned was modified, in
|
||||
seconds since the epoch.
|
||||
@type when: L{int}, L{long} or L{float}
|
||||
|
||||
@return: If I am a C{If-Modified-Since} conditional request and the time
|
||||
given is not newer than the condition, I return
|
||||
L{CACHED<http.CACHED>} to indicate that you should write no body.
|
||||
Otherwise, I return a false value.
|
||||
"""
|
||||
|
||||
|
||||
def setETag(etag):
|
||||
"""
|
||||
Set an C{entity tag} for the outgoing response.
|
||||
|
||||
That's "entity tag" as in the HTTP/1.1 I{ETag} header, "used for
|
||||
comparing two or more entities from the same requested resource."
|
||||
|
||||
If I am a conditional request, I may modify my response code to
|
||||
L{NOT_MODIFIED<http.NOT_MODIFIED>} or
|
||||
L{PRECONDITION_FAILED<http.PRECONDITION_FAILED>}, if appropriate for the
|
||||
tag given.
|
||||
|
||||
@param etag: The entity tag for the resource being returned.
|
||||
@type etag: C{str}
|
||||
|
||||
@return: If I am a C{If-None-Match} conditional request and the tag
|
||||
matches one in the request, I return L{CACHED<http.CACHED>} to
|
||||
indicate that you should write no body. Otherwise, I return a
|
||||
false value.
|
||||
"""
|
||||
|
||||
|
||||
def setHost(host, port, ssl=0):
|
||||
"""
|
||||
Change the host and port the request thinks it's using.
|
||||
|
||||
This method is useful for working with reverse HTTP proxies (e.g. both
|
||||
Squid and Apache's mod_proxy can do this), when the address the HTTP
|
||||
client is using is different than the one we're listening on.
|
||||
|
||||
For example, Apache may be listening on https://www.example.com, and
|
||||
then forwarding requests to http://localhost:8080, but we don't want
|
||||
HTML produced by Twisted to say 'http://localhost:8080', they should
|
||||
say 'https://www.example.com', so we do::
|
||||
|
||||
request.setHost('www.example.com', 443, ssl=1)
|
||||
"""
|
||||
|
||||
|
||||
|
||||
class ICredentialFactory(Interface):
|
||||
"""
|
||||
A credential factory defines a way to generate a particular kind of
|
||||
authentication challenge and a way to interpret the responses to these
|
||||
challenges. It creates
|
||||
L{ICredentials<twisted.cred.credentials.ICredentials>} providers from
|
||||
responses. These objects will be used with L{twisted.cred} to authenticate
|
||||
an authorize requests.
|
||||
"""
|
||||
scheme = Attribute(
|
||||
"A C{str} giving the name of the authentication scheme with which "
|
||||
"this factory is associated. For example, C{'basic'} or C{'digest'}.")
|
||||
|
||||
|
||||
def getChallenge(request):
|
||||
"""
|
||||
Generate a new challenge to be sent to a client.
|
||||
|
||||
@type peer: L{twisted.web.http.Request}
|
||||
@param peer: The request the response to which this challenge will be
|
||||
included.
|
||||
|
||||
@rtype: C{dict}
|
||||
@return: A mapping from C{str} challenge fields to associated C{str}
|
||||
values.
|
||||
"""
|
||||
|
||||
|
||||
def decode(response, request):
|
||||
"""
|
||||
Create a credentials object from the given response.
|
||||
|
||||
@type response: C{str}
|
||||
@param response: scheme specific response string
|
||||
|
||||
@type request: L{twisted.web.http.Request}
|
||||
@param request: The request being processed (from which the response
|
||||
was taken).
|
||||
|
||||
@raise twisted.cred.error.LoginFailed: If the response is invalid.
|
||||
|
||||
@rtype: L{twisted.cred.credentials.ICredentials} provider
|
||||
@return: The credentials represented by the given response.
|
||||
"""
|
||||
|
||||
|
||||
|
||||
class IBodyProducer(IPushProducer):
|
||||
"""
|
||||
Objects which provide L{IBodyProducer} write bytes to an object which
|
||||
provides L{IConsumer<twisted.internet.interfaces.IConsumer>} by calling its
|
||||
C{write} method repeatedly.
|
||||
|
||||
L{IBodyProducer} providers may start producing as soon as they have an
|
||||
L{IConsumer<twisted.internet.interfaces.IConsumer>} provider. That is, they
|
||||
should not wait for a C{resumeProducing} call to begin writing data.
|
||||
|
||||
L{IConsumer.unregisterProducer<twisted.internet.interfaces.IConsumer.unregisterProducer>}
|
||||
must not be called. Instead, the
|
||||
L{Deferred<twisted.internet.defer.Deferred>} returned from C{startProducing}
|
||||
must be fired when all bytes have been written.
|
||||
|
||||
L{IConsumer.write<twisted.internet.interfaces.IConsumer.write>} may
|
||||
synchronously invoke any of C{pauseProducing}, C{resumeProducing}, or
|
||||
C{stopProducing}. These methods must be implemented with this in mind.
|
||||
|
||||
@since: 9.0
|
||||
"""
|
||||
|
||||
# Despite the restrictions above and the additional requirements of
|
||||
# stopProducing documented below, this interface still needs to be an
|
||||
# IPushProducer subclass. Providers of it will be passed to IConsumer
|
||||
# providers which only know about IPushProducer and IPullProducer, not
|
||||
# about this interface. This interface needs to remain close enough to one
|
||||
# of those interfaces for consumers to work with it.
|
||||
|
||||
length = Attribute(
|
||||
"""
|
||||
C{length} is a C{int} indicating how many bytes in total this
|
||||
L{IBodyProducer} will write to the consumer or L{UNKNOWN_LENGTH}
|
||||
if this is not known in advance.
|
||||
""")
|
||||
|
||||
def startProducing(consumer):
|
||||
"""
|
||||
Start producing to the given
|
||||
L{IConsumer<twisted.internet.interfaces.IConsumer>} provider.
|
||||
|
||||
@return: A L{Deferred<twisted.internet.defer.Deferred>} which fires with
|
||||
C{None} when all bytes have been produced or with a
|
||||
L{Failure<twisted.python.failure.Failure>} if there is any problem
|
||||
before all bytes have been produced.
|
||||
"""
|
||||
|
||||
|
||||
def stopProducing():
|
||||
"""
|
||||
In addition to the standard behavior of
|
||||
L{IProducer.stopProducing<twisted.internet.interfaces.IProducer.stopProducing>}
|
||||
(stop producing data), make sure the
|
||||
L{Deferred<twisted.internet.defer.Deferred>} returned by
|
||||
C{startProducing} is never fired.
|
||||
"""
|
||||
|
||||
|
||||
|
||||
class IRenderable(Interface):
|
||||
"""
|
||||
An L{IRenderable} is an object that may be rendered by the
|
||||
L{twisted.web.template} templating system.
|
||||
"""
|
||||
|
||||
def lookupRenderMethod(name):
|
||||
"""
|
||||
Look up and return the render method associated with the given name.
|
||||
|
||||
@type name: C{str}
|
||||
@param name: The value of a render directive encountered in the
|
||||
document returned by a call to L{IRenderable.render}.
|
||||
|
||||
@return: A two-argument callable which will be invoked with the request
|
||||
being responded to and the tag object on which the render directive
|
||||
was encountered.
|
||||
"""
|
||||
|
||||
|
||||
def render(request):
|
||||
"""
|
||||
Get the document for this L{IRenderable}.
|
||||
|
||||
@type request: L{IRequest} provider or C{NoneType}
|
||||
@param request: The request in response to which this method is being
|
||||
invoked.
|
||||
|
||||
@return: An object which can be flattened.
|
||||
"""
|
||||
|
||||
|
||||
|
||||
class ITemplateLoader(Interface):
|
||||
"""
|
||||
A loader for templates; something usable as a value for
|
||||
L{twisted.web.template.Element}'s C{loader} attribute.
|
||||
"""
|
||||
|
||||
def load():
|
||||
"""
|
||||
Load a template suitable for rendering.
|
||||
|
||||
@return: a C{list} of C{list}s, C{unicode} objects, C{Element}s and
|
||||
other L{IRenderable} providers.
|
||||
"""
|
||||
|
||||
|
||||
|
||||
class IResponse(Interface):
|
||||
"""
|
||||
An object representing an HTTP response received from an HTTP server.
|
||||
|
||||
@since: 11.1
|
||||
"""
|
||||
|
||||
version = Attribute(
|
||||
"A three-tuple describing the protocol and protocol version "
|
||||
"of the response. The first element is of type C{str}, the second "
|
||||
"and third are of type C{int}. For example, C{('HTTP', 1, 1)}.")
|
||||
|
||||
|
||||
code = Attribute("The HTTP status code of this response, as a C{int}.")
|
||||
|
||||
|
||||
phrase = Attribute(
|
||||
"The HTTP reason phrase of this response, as a C{str}.")
|
||||
|
||||
|
||||
headers = Attribute("The HTTP response L{Headers} of this response.")
|
||||
|
||||
|
||||
length = Attribute(
|
||||
"The C{int} number of bytes expected to be in the body of this "
|
||||
"response or L{UNKNOWN_LENGTH} if the server did not indicate how "
|
||||
"many bytes to expect. For I{HEAD} responses, this will be 0; if "
|
||||
"the response includes a I{Content-Length} header, it will be "
|
||||
"available in C{headers}.")
|
||||
|
||||
|
||||
def deliverBody(protocol):
|
||||
"""
|
||||
Register an L{IProtocol<twisted.internet.interfaces.IProtocol>} provider
|
||||
to receive the response body.
|
||||
|
||||
The protocol will be connected to a transport which provides
|
||||
L{IPushProducer}. The protocol's C{connectionLost} method will be
|
||||
called with:
|
||||
|
||||
- ResponseDone, which indicates that all bytes from the response
|
||||
have been successfully delivered.
|
||||
|
||||
- PotentialDataLoss, which indicates that it cannot be determined
|
||||
if the entire response body has been delivered. This only occurs
|
||||
when making requests to HTTP servers which do not set
|
||||
I{Content-Length} or a I{Transfer-Encoding} in the response.
|
||||
|
||||
- ResponseFailed, which indicates that some bytes from the response
|
||||
were lost. The C{reasons} attribute of the exception may provide
|
||||
more specific indications as to why.
|
||||
"""
|
||||
|
||||
|
||||
|
||||
class _IRequestEncoder(Interface):
|
||||
"""
|
||||
An object encoding data passed to L{IRequest.write}, for example for
|
||||
compression purpose.
|
||||
|
||||
@since: 12.3
|
||||
"""
|
||||
|
||||
def encode(data):
|
||||
"""
|
||||
Encode the data given and return the result.
|
||||
|
||||
@param data: The content to encode.
|
||||
@type data: C{str}
|
||||
|
||||
@return: The encoded data.
|
||||
@rtype: C{str}
|
||||
"""
|
||||
|
||||
|
||||
def finish():
|
||||
"""
|
||||
Callback called when the request is closing.
|
||||
|
||||
@return: If necessary, the pending data accumulated from previous
|
||||
C{encode} calls.
|
||||
@rtype: C{str}
|
||||
"""
|
||||
|
||||
|
||||
|
||||
class _IRequestEncoderFactory(Interface):
|
||||
"""
|
||||
A factory for returing L{_IRequestEncoder} instances.
|
||||
|
||||
@since: 12.3
|
||||
"""
|
||||
|
||||
def encoderForRequest(request):
|
||||
"""
|
||||
If applicable, returns a L{_IRequestEncoder} instance which will encode
|
||||
the request.
|
||||
"""
|
||||
|
||||
|
||||
|
||||
UNKNOWN_LENGTH = u"twisted.web.iweb.UNKNOWN_LENGTH"
|
||||
|
||||
__all__ = [
|
||||
"ICredentialFactory", "IRequest",
|
||||
"IBodyProducer", "IRenderable", "IResponse", "_IRequestEncoder",
|
||||
"_IRequestEncoderFactory",
|
||||
|
||||
"UNKNOWN_LENGTH"]
|
||||
4
setup.py
4
setup.py
|
|
@ -41,14 +41,14 @@ setup(
|
|||
'Topic :: Software Development :: Libraries :: Python Modules',
|
||||
],
|
||||
install_requires=[
|
||||
'Twisted>=10.0.0',
|
||||
'Twisted>=13.1.0',
|
||||
'w3lib>=1.15.0',
|
||||
'queuelib',
|
||||
'lxml',
|
||||
'pyOpenSSL',
|
||||
'cssselect>=0.9',
|
||||
'six>=1.5.2',
|
||||
'parsel>=0.9.3',
|
||||
'parsel>=0.9.5',
|
||||
'PyDispatcher>=2.0.5',
|
||||
'service_identity',
|
||||
],
|
||||
|
|
|
|||
|
|
@ -0,0 +1,11 @@
|
|||
"""
|
||||
Some pipelines used for testing
|
||||
"""
|
||||
|
||||
class ZeroDivisionErrorPipeline(object):
|
||||
|
||||
def open_spider(self, spider):
|
||||
a = 1/0
|
||||
|
||||
def process_item(self, item, spider):
|
||||
return item
|
||||
|
|
@ -14,6 +14,18 @@ class FetchTest(ProcessTest, SiteTest, unittest.TestCase):
|
|||
_, out, _ = yield self.execute([self.url('/text')])
|
||||
self.assertEqual(out.strip(), b'Works')
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_redirect_default(self):
|
||||
_, out, _ = yield self.execute([self.url('/redirect')])
|
||||
self.assertEqual(out.strip(), b'Redirected here')
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_redirect_disabled(self):
|
||||
_, out, err = yield self.execute(['--no-redirect', self.url('/redirect-no-meta-refresh')])
|
||||
err = err.strip()
|
||||
self.assertIn(b'downloader/response_status_count/302', err, err)
|
||||
self.assertNotIn(b'downloader/response_status_count/200', err, err)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_headers(self):
|
||||
_, out, _ = yield self.execute([self.url('/text'), '--headers'])
|
||||
|
|
|
|||
|
|
@ -0,0 +1,156 @@
|
|||
from os.path import join, abspath
|
||||
from twisted.trial import unittest
|
||||
from twisted.internet import defer
|
||||
from scrapy.utils.testsite import SiteTest
|
||||
from scrapy.utils.testproc import ProcessTest
|
||||
from scrapy.utils.python import to_native_str
|
||||
from tests.test_commands import CommandTest
|
||||
|
||||
|
||||
class ParseCommandTest(ProcessTest, SiteTest, CommandTest):
|
||||
command = 'parse'
|
||||
|
||||
def setUp(self):
|
||||
super(ParseCommandTest, self).setUp()
|
||||
self.spider_name = 'parse_spider'
|
||||
fname = abspath(join(self.proj_mod_path, 'spiders', 'myspider.py'))
|
||||
with open(fname, 'w') as f:
|
||||
f.write("""
|
||||
import scrapy
|
||||
from scrapy.linkextractors import LinkExtractor
|
||||
from scrapy.spiders import CrawlSpider, Rule
|
||||
|
||||
|
||||
class MySpider(scrapy.Spider):
|
||||
name = '{0}'
|
||||
|
||||
def parse(self, response):
|
||||
if getattr(self, 'test_arg', None):
|
||||
self.logger.debug('It Works!')
|
||||
return [scrapy.Item(), dict(foo='bar')]
|
||||
|
||||
|
||||
class MyGoodCrawlSpider(CrawlSpider):
|
||||
name = 'goodcrawl{0}'
|
||||
|
||||
rules = (
|
||||
Rule(LinkExtractor(allow=r'/html'), callback='parse_item', follow=True),
|
||||
Rule(LinkExtractor(allow=r'/text'), follow=True),
|
||||
)
|
||||
|
||||
def parse_item(self, response):
|
||||
return [scrapy.Item(), dict(foo='bar')]
|
||||
|
||||
def parse(self, response):
|
||||
return [scrapy.Item(), dict(nomatch='default')]
|
||||
|
||||
|
||||
class MyBadCrawlSpider(CrawlSpider):
|
||||
'''Spider which doesn't define a parse_item callback while using it in a rule.'''
|
||||
name = 'badcrawl{0}'
|
||||
|
||||
rules = (
|
||||
Rule(LinkExtractor(allow=r'/html'), callback='parse_item', follow=True),
|
||||
)
|
||||
|
||||
def parse(self, response):
|
||||
return [scrapy.Item(), dict(foo='bar')]
|
||||
""".format(self.spider_name))
|
||||
|
||||
fname = abspath(join(self.proj_mod_path, 'pipelines.py'))
|
||||
with open(fname, 'w') as f:
|
||||
f.write("""
|
||||
import logging
|
||||
|
||||
class MyPipeline(object):
|
||||
component_name = 'my_pipeline'
|
||||
|
||||
def process_item(self, item, spider):
|
||||
logging.info('It Works!')
|
||||
return item
|
||||
""")
|
||||
|
||||
fname = abspath(join(self.proj_mod_path, 'settings.py'))
|
||||
with open(fname, 'a') as f:
|
||||
f.write("""
|
||||
ITEM_PIPELINES = {'%s.pipelines.MyPipeline': 1}
|
||||
""" % self.project_name)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_spider_arguments(self):
|
||||
_, _, stderr = yield self.execute(['--spider', self.spider_name,
|
||||
'-a', 'test_arg=1',
|
||||
'-c', 'parse',
|
||||
self.url('/html')])
|
||||
self.assertIn("DEBUG: It Works!", to_native_str(stderr))
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_pipelines(self):
|
||||
_, _, stderr = yield self.execute(['--spider', self.spider_name,
|
||||
'--pipelines',
|
||||
'-c', 'parse',
|
||||
self.url('/html')])
|
||||
self.assertIn("INFO: It Works!", to_native_str(stderr))
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_parse_items(self):
|
||||
status, out, stderr = yield self.execute(
|
||||
['--spider', self.spider_name, '-c', 'parse', self.url('/html')]
|
||||
)
|
||||
self.assertIn("""[{}, {'foo': 'bar'}]""", to_native_str(out))
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_parse_items_no_callback_passed(self):
|
||||
status, out, stderr = yield self.execute(
|
||||
['--spider', self.spider_name, self.url('/html')]
|
||||
)
|
||||
self.assertIn("""[{}, {'foo': 'bar'}]""", to_native_str(out))
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_wrong_callback_passed(self):
|
||||
status, out, stderr = yield self.execute(
|
||||
['--spider', self.spider_name, '-c', 'dummy', self.url('/html')]
|
||||
)
|
||||
self.assertRegexpMatches(to_native_str(out), """# Scraped Items -+\n\[\]""")
|
||||
self.assertIn("""Cannot find callback""", to_native_str(stderr))
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_crawlspider_matching_rule_callback_set(self):
|
||||
"""If a rule matches the URL, use it's defined callback."""
|
||||
status, out, stderr = yield self.execute(
|
||||
['--spider', 'goodcrawl'+self.spider_name, '-r', self.url('/html')]
|
||||
)
|
||||
self.assertIn("""[{}, {'foo': 'bar'}]""", to_native_str(out))
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_crawlspider_matching_rule_default_callback(self):
|
||||
"""If a rule match but it has no callback set, use the 'parse' callback."""
|
||||
status, out, stderr = yield self.execute(
|
||||
['--spider', 'goodcrawl'+self.spider_name, '-r', self.url('/text')]
|
||||
)
|
||||
self.assertIn("""[{}, {'nomatch': 'default'}]""", to_native_str(out))
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_spider_with_no_rules_attribute(self):
|
||||
"""Using -r with a spider with no rule should not produce items."""
|
||||
status, out, stderr = yield self.execute(
|
||||
['--spider', self.spider_name, '-r', self.url('/html')]
|
||||
)
|
||||
self.assertRegexpMatches(to_native_str(out), """# Scraped Items -+\n\[\]""")
|
||||
self.assertIn("""No CrawlSpider rules found""", to_native_str(stderr))
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_crawlspider_missing_callback(self):
|
||||
status, out, stderr = yield self.execute(
|
||||
['--spider', 'badcrawl'+self.spider_name, '-r', self.url('/html')]
|
||||
)
|
||||
self.assertRegexpMatches(to_native_str(out), """# Scraped Items -+\n\[\]""")
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_crawlspider_no_matching_rule(self):
|
||||
"""The requested URL has no matching rule, so no items should be scraped"""
|
||||
status, out, stderr = yield self.execute(
|
||||
['--spider', 'badcrawl'+self.spider_name, '-r', self.url('/enc-gb18030')]
|
||||
)
|
||||
self.assertRegexpMatches(to_native_str(out), """# Scraped Items -+\n\[\]""")
|
||||
self.assertIn("""Cannot find a rule that matches""", to_native_str(stderr))
|
||||
|
|
@ -49,6 +49,35 @@ class ShellTest(ProcessTest, SiteTest, unittest.TestCase):
|
|||
_, out, _ = yield self.execute([self.url('/redirect'), '-c', 'response.url'])
|
||||
assert out.strip().endswith(b'/redirected')
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_redirect_follow_302(self):
|
||||
_, out, _ = yield self.execute([self.url('/redirect-no-meta-refresh'), '-c', 'response.status'])
|
||||
assert out.strip().endswith(b'200')
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_redirect_not_follow_302(self):
|
||||
_, out, _ = yield self.execute(['--no-redirect', self.url('/redirect-no-meta-refresh'), '-c', 'response.status'])
|
||||
assert out.strip().endswith(b'302')
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_fetch_redirect_follow_302(self):
|
||||
"""Test that calling `fetch(url)` follows HTTP redirects by default."""
|
||||
url = self.url('/redirect-no-meta-refresh')
|
||||
code = "fetch('{0}')"
|
||||
errcode, out, errout = yield self.execute(['-c', code.format(url)])
|
||||
self.assertEqual(errcode, 0, out)
|
||||
assert b'Redirecting (302)' in errout
|
||||
assert b'Crawled (200)' in errout
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_fetch_redirect_not_follow_302(self):
|
||||
"""Test that calling `fetch(url, redirect=False)` disables automatic redirects."""
|
||||
url = self.url('/redirect-no-meta-refresh')
|
||||
code = "fetch('{0}', redirect=False)"
|
||||
errcode, out, errout = yield self.execute(['-c', code.format(url)])
|
||||
self.assertEqual(errcode, 0, out)
|
||||
assert b'Crawled (302)' in errout
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_request_replace(self):
|
||||
url = self.url('/text')
|
||||
|
|
@ -56,6 +85,13 @@ class ShellTest(ProcessTest, SiteTest, unittest.TestCase):
|
|||
errcode, out, _ = yield self.execute(['-c', code.format(url)])
|
||||
self.assertEqual(errcode, 0, out)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_scrapy_import(self):
|
||||
url = self.url('/text')
|
||||
code = "fetch(scrapy.Request('{0}'))"
|
||||
errcode, out, _ = yield self.execute(['-c', code.format(url)])
|
||||
self.assertEqual(errcode, 0, out)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_local_file(self):
|
||||
filepath = join(tests_datadir, 'test_site/index.html')
|
||||
|
|
|
|||
|
|
@ -25,5 +25,7 @@ class VersionTest(ProcessTest, unittest.TestCase):
|
|||
_, out, _ = yield self.execute(['-v'])
|
||||
headers = [l.partition(":")[0].strip()
|
||||
for l in out.strip().decode(encoding).splitlines()]
|
||||
self.assertEqual(headers, ['Scrapy', 'lxml', 'libxml2', 'Twisted',
|
||||
'Python', 'pyOpenSSL', 'Platform'])
|
||||
self.assertEqual(headers, ['Scrapy', 'lxml', 'libxml2',
|
||||
'cssselect', 'parsel', 'w3lib',
|
||||
'Twisted', 'Python', 'pyOpenSSL',
|
||||
'Platform'])
|
||||
|
|
|
|||
|
|
@ -38,11 +38,11 @@ class ProjectTest(unittest.TestCase):
|
|||
return subprocess.call(args, stdout=out, stderr=out, cwd=self.cwd,
|
||||
env=self.env, **kwargs)
|
||||
|
||||
def proc(self, *new_args, **kwargs):
|
||||
def proc(self, *new_args, **popen_kwargs):
|
||||
args = (sys.executable, '-m', 'scrapy.cmdline') + new_args
|
||||
p = subprocess.Popen(args, cwd=self.cwd, env=self.env,
|
||||
stdout=subprocess.PIPE, stderr=subprocess.PIPE,
|
||||
**kwargs)
|
||||
**popen_kwargs)
|
||||
|
||||
waited = 0
|
||||
interval = 0.2
|
||||
|
|
@ -182,6 +182,17 @@ class MiscCommandsTest(CommandTest):
|
|||
|
||||
class RunSpiderCommandTest(CommandTest):
|
||||
|
||||
debug_log_spider = """
|
||||
import scrapy
|
||||
|
||||
class MySpider(scrapy.Spider):
|
||||
name = 'myspider'
|
||||
|
||||
def start_requests(self):
|
||||
self.logger.debug("It Works!")
|
||||
return []
|
||||
"""
|
||||
|
||||
@contextmanager
|
||||
def _create_file(self, content, name):
|
||||
tmpdir = self.mktemp()
|
||||
|
|
@ -194,32 +205,44 @@ class RunSpiderCommandTest(CommandTest):
|
|||
finally:
|
||||
rmtree(tmpdir)
|
||||
|
||||
def runspider(self, code, name='myspider.py'):
|
||||
def runspider(self, code, name='myspider.py', args=()):
|
||||
with self._create_file(code, name) as fname:
|
||||
return self.proc('runspider', fname)
|
||||
return self.proc('runspider', fname, *args)
|
||||
|
||||
def get_log(self, code, name='myspider.py', args=()):
|
||||
p = self.runspider(code, name=name, args=args)
|
||||
return to_native_str(p.stderr.read())
|
||||
|
||||
def test_runspider(self):
|
||||
spider = """
|
||||
import scrapy
|
||||
|
||||
class MySpider(scrapy.Spider):
|
||||
name = 'myspider'
|
||||
|
||||
def start_requests(self):
|
||||
self.logger.debug("It Works!")
|
||||
return []
|
||||
"""
|
||||
p = self.runspider(spider)
|
||||
log = to_native_str(p.stderr.read())
|
||||
|
||||
log = self.get_log(self.debug_log_spider)
|
||||
self.assertIn("DEBUG: It Works!", log)
|
||||
self.assertIn("INFO: Spider opened", log)
|
||||
self.assertIn("INFO: Closing spider (finished)", log)
|
||||
self.assertIn("INFO: Spider closed (finished)", log)
|
||||
|
||||
def test_runspider_log_level(self):
|
||||
log = self.get_log(self.debug_log_spider,
|
||||
args=('-s', 'LOG_LEVEL=INFO'))
|
||||
self.assertNotIn("DEBUG: It Works!", log)
|
||||
self.assertIn("INFO: Spider opened", log)
|
||||
|
||||
def test_runspider_log_short_names(self):
|
||||
log1 = self.get_log(self.debug_log_spider,
|
||||
args=('-s', 'LOG_SHORT_NAMES=1'))
|
||||
print(log1)
|
||||
self.assertIn("[myspider] DEBUG: It Works!", log1)
|
||||
self.assertIn("[scrapy]", log1)
|
||||
self.assertNotIn("[scrapy.core.engine]", log1)
|
||||
|
||||
log2 = self.get_log(self.debug_log_spider,
|
||||
args=('-s', 'LOG_SHORT_NAMES=0'))
|
||||
print(log2)
|
||||
self.assertIn("[myspider] DEBUG: It Works!", log2)
|
||||
self.assertNotIn("[scrapy]", log2)
|
||||
self.assertIn("[scrapy.core.engine]", log2)
|
||||
|
||||
def test_runspider_no_spider_found(self):
|
||||
p = self.runspider("from scrapy.spiders import Spider\n")
|
||||
log = to_native_str(p.stderr.read())
|
||||
log = self.get_log("from scrapy.spiders import Spider\n")
|
||||
self.assertIn("No spider found in file", log)
|
||||
|
||||
def test_runspider_file_not_found(self):
|
||||
|
|
@ -228,12 +251,11 @@ class MySpider(scrapy.Spider):
|
|||
self.assertIn("File not found: some_non_existent_file", log)
|
||||
|
||||
def test_runspider_unable_to_load(self):
|
||||
p = self.runspider('', 'myspider.txt')
|
||||
log = to_native_str(p.stderr.read())
|
||||
log = self.get_log('', name='myspider.txt')
|
||||
self.assertIn('Unable to load', log)
|
||||
|
||||
def test_start_requests_errors(self):
|
||||
p = self.runspider("""
|
||||
log = self.get_log("""
|
||||
import scrapy
|
||||
|
||||
class BadSpider(scrapy.Spider):
|
||||
|
|
@ -241,76 +263,11 @@ class BadSpider(scrapy.Spider):
|
|||
def start_requests(self):
|
||||
raise Exception("oops!")
|
||||
""", name="badspider.py")
|
||||
log = to_native_str(p.stderr.read())
|
||||
print(log)
|
||||
self.assertIn("start_requests", log)
|
||||
self.assertIn("badspider.py", log)
|
||||
|
||||
|
||||
class ParseCommandTest(ProcessTest, SiteTest, CommandTest):
|
||||
command = 'parse'
|
||||
|
||||
def setUp(self):
|
||||
super(ParseCommandTest, self).setUp()
|
||||
self.spider_name = 'parse_spider'
|
||||
fname = abspath(join(self.proj_mod_path, 'spiders', 'myspider.py'))
|
||||
with open(fname, 'w') as f:
|
||||
f.write("""
|
||||
import scrapy
|
||||
|
||||
class MySpider(scrapy.Spider):
|
||||
name = '{0}'
|
||||
|
||||
def parse(self, response):
|
||||
if getattr(self, 'test_arg', None):
|
||||
self.logger.debug('It Works!')
|
||||
return [scrapy.Item(), dict(foo='bar')]
|
||||
""".format(self.spider_name))
|
||||
|
||||
fname = abspath(join(self.proj_mod_path, 'pipelines.py'))
|
||||
with open(fname, 'w') as f:
|
||||
f.write("""
|
||||
import logging
|
||||
|
||||
class MyPipeline(object):
|
||||
component_name = 'my_pipeline'
|
||||
|
||||
def process_item(self, item, spider):
|
||||
logging.info('It Works!')
|
||||
return item
|
||||
""")
|
||||
|
||||
fname = abspath(join(self.proj_mod_path, 'settings.py'))
|
||||
with open(fname, 'a') as f:
|
||||
f.write("""
|
||||
ITEM_PIPELINES = {'%s.pipelines.MyPipeline': 1}
|
||||
""" % self.project_name)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_spider_arguments(self):
|
||||
_, _, stderr = yield self.execute(['--spider', self.spider_name,
|
||||
'-a', 'test_arg=1',
|
||||
'-c', 'parse',
|
||||
self.url('/html')])
|
||||
self.assertIn("DEBUG: It Works!", to_native_str(stderr))
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_pipelines(self):
|
||||
_, _, stderr = yield self.execute(['--spider', self.spider_name,
|
||||
'--pipelines',
|
||||
'-c', 'parse',
|
||||
self.url('/html')])
|
||||
self.assertIn("INFO: It Works!", to_native_str(stderr))
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_parse_items(self):
|
||||
status, out, stderr = yield self.execute(
|
||||
['--spider', self.spider_name, '-c', 'parse', self.url('/html')]
|
||||
)
|
||||
self.assertIn("""[{}, {'foo': 'bar'}]""", to_native_str(out))
|
||||
|
||||
|
||||
|
||||
class BenchCommandTest(CommandTest):
|
||||
|
||||
def test_run(self):
|
||||
|
|
|
|||
|
|
@ -250,6 +250,19 @@ with multiples lines
|
|||
yield self.assertFailure(crawler.crawl(), TestError)
|
||||
self.assertFalse(crawler.crawling)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_open_spider_error_on_faulty_pipeline(self):
|
||||
settings = {
|
||||
"ITEM_PIPELINES": {
|
||||
"tests.pipelines.ZeroDivisionErrorPipeline": 300,
|
||||
}
|
||||
}
|
||||
crawler = CrawlerRunner(settings).create_crawler(SimpleSpider)
|
||||
yield self.assertFailure(
|
||||
self.runner.crawl(crawler, "http://localhost:8998/status?n=200"),
|
||||
ZeroDivisionError)
|
||||
self.assertFalse(crawler.crawling)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_crawlerrunner_accepts_crawler(self):
|
||||
crawler = self.runner.create_crawler(SimpleSpider)
|
||||
|
|
|
|||
|
|
@ -157,15 +157,15 @@ class RedirectMiddlewareTest(unittest.TestCase):
|
|||
latin1_location = u'/ação'.encode('latin1') # HTTP historically supports latin1
|
||||
resp = Response('http://scrapytest.org/first', headers={'Location': latin1_location}, status=302)
|
||||
req_result = self.mw.process_response(req, resp, self.spider)
|
||||
perc_encoded_utf8_url = 'http://scrapytest.org/a%C3%A7%C3%A3o'
|
||||
perc_encoded_utf8_url = 'http://scrapytest.org/a%E7%E3o'
|
||||
self.assertEquals(perc_encoded_utf8_url, req_result.url)
|
||||
|
||||
def test_location_with_wrong_encoding(self):
|
||||
def test_utf8_location(self):
|
||||
req = Request('http://scrapytest.org/first')
|
||||
utf8_location = u'/ação' # header with wrong encoding (utf-8)
|
||||
utf8_location = u'/ação'.encode('utf-8') # header using UTF-8 encoding
|
||||
resp = Response('http://scrapytest.org/first', headers={'Location': utf8_location}, status=302)
|
||||
req_result = self.mw.process_response(req, resp, self.spider)
|
||||
perc_encoded_utf8_url = 'http://scrapytest.org/a%C3%83%C2%A7%C3%83%C2%A3o'
|
||||
perc_encoded_utf8_url = 'http://scrapytest.org/a%C3%A7%C3%A3o'
|
||||
self.assertEquals(perc_encoded_utf8_url, req_result.url)
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -3,10 +3,10 @@ from twisted.internet import defer
|
|||
from twisted.internet.error import TimeoutError, DNSLookupError, \
|
||||
ConnectionRefusedError, ConnectionDone, ConnectError, \
|
||||
ConnectionLost, TCPTimedOutError
|
||||
from twisted.web.client import ResponseFailed
|
||||
|
||||
from scrapy import twisted_version
|
||||
from scrapy.downloadermiddlewares.retry import RetryMiddleware
|
||||
from scrapy.xlib.tx import ResponseFailed
|
||||
from scrapy.spiders import Spider
|
||||
from scrapy.http import Request, Response
|
||||
from scrapy.utils.test import get_crawler
|
||||
|
|
|
|||
|
|
@ -1,3 +1,4 @@
|
|||
# -*- coding: utf-8 -*-
|
||||
from __future__ import absolute_import
|
||||
import re
|
||||
from twisted.internet import reactor, error
|
||||
|
|
@ -30,11 +31,18 @@ class RobotsTxtMiddlewareTest(unittest.TestCase):
|
|||
def _get_successful_crawler(self):
|
||||
crawler = self.crawler
|
||||
crawler.settings.set('ROBOTSTXT_OBEY', True)
|
||||
ROBOTS = re.sub(b'^\s+(?m)', b'', b'''
|
||||
ROBOTS = re.sub(b'^\s+(?m)', b'', u'''
|
||||
User-Agent: *
|
||||
Disallow: /admin/
|
||||
Disallow: /static/
|
||||
''')
|
||||
|
||||
# taken from https://en.wikipedia.org/robots.txt
|
||||
Disallow: /wiki/K%C3%A4ytt%C3%A4j%C3%A4:
|
||||
Disallow: /wiki/Käyttäjä:
|
||||
|
||||
User-Agent: UnicödeBöt
|
||||
Disallow: /some/randome/page.html
|
||||
'''.encode('utf-8'))
|
||||
response = TextResponse('http://site.local/robots.txt', body=ROBOTS)
|
||||
def return_response(request, spider):
|
||||
deferred = Deferred()
|
||||
|
|
@ -48,7 +56,9 @@ class RobotsTxtMiddlewareTest(unittest.TestCase):
|
|||
return DeferredList([
|
||||
self.assertNotIgnored(Request('http://site.local/allowed'), middleware),
|
||||
self.assertIgnored(Request('http://site.local/admin/main'), middleware),
|
||||
self.assertIgnored(Request('http://site.local/static/'), middleware)
|
||||
self.assertIgnored(Request('http://site.local/static/'), middleware),
|
||||
self.assertIgnored(Request('http://site.local/wiki/K%C3%A4ytt%C3%A4j%C3%A4:'), middleware),
|
||||
self.assertIgnored(Request(u'http://site.local/wiki/Käyttäjä:'), middleware)
|
||||
], fireOnOneErrback=True)
|
||||
|
||||
def test_robotstxt_ready_parser(self):
|
||||
|
|
|
|||
|
|
@ -197,6 +197,21 @@ class FeedExportTest(unittest.TestCase):
|
|||
data = yield self.run_and_export(TestSpider, settings)
|
||||
defer.returnValue(data)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def exported_no_data(self, settings):
|
||||
"""
|
||||
Return exported data which a spider yielding no ``items`` would return.
|
||||
"""
|
||||
class TestSpider(scrapy.Spider):
|
||||
name = 'testspider'
|
||||
start_urls = ['http://localhost:8998/']
|
||||
|
||||
def parse(self, response):
|
||||
pass
|
||||
|
||||
data = yield self.run_and_export(TestSpider, settings)
|
||||
defer.returnValue(data)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def assertExportedCsv(self, items, header, rows, settings=None, ordered=True):
|
||||
settings = settings or {}
|
||||
|
|
@ -283,6 +298,32 @@ class FeedExportTest(unittest.TestCase):
|
|||
header = self.MyItem.fields.keys()
|
||||
yield self.assertExported(items, header, rows, ordered=False)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_export_no_items_not_store_empty(self):
|
||||
formats = ('json',
|
||||
'jsonlines',
|
||||
'xml',
|
||||
'csv',)
|
||||
|
||||
for fmt in formats:
|
||||
settings = {'FEED_FORMAT': fmt}
|
||||
data = yield self.exported_no_data(settings)
|
||||
self.assertEqual(data, b'')
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_export_no_items_store_empty(self):
|
||||
formats = (
|
||||
('json', b'[\n\n]'),
|
||||
('jsonlines', b''),
|
||||
('xml', b'<?xml version="1.0" encoding="utf-8"?>\n<items></items>'),
|
||||
('csv', b''),
|
||||
)
|
||||
|
||||
for fmt, expctd in formats:
|
||||
settings = {'FEED_FORMAT': fmt, 'FEED_STORE_EMPTY': True}
|
||||
data = yield self.exported_no_data(settings)
|
||||
self.assertEqual(data, expctd)
|
||||
|
||||
@defer.inlineCallbacks
|
||||
def test_export_multiple_item_classes(self):
|
||||
|
||||
|
|
@ -376,26 +417,26 @@ class FeedExportTest(unittest.TestCase):
|
|||
def test_export_encoding(self):
|
||||
items = [dict({'foo': u'Test\xd6'})]
|
||||
header = ['foo']
|
||||
|
||||
|
||||
formats = {
|
||||
'json': u'[\n{"foo": "Test\\u00d6"}\n]'.encode('utf-8'),
|
||||
'jsonlines': u'{"foo": "Test\\u00d6"}\n'.encode('utf-8'),
|
||||
'xml': u'<?xml version="1.0" encoding="utf-8"?>\n<items><item><foo>Test\xd6</foo></item></items>'.encode('utf-8'),
|
||||
'csv': u'foo\r\nTest\xd6\r\n'.encode('utf-8'),
|
||||
}
|
||||
|
||||
|
||||
for format in formats:
|
||||
settings = {'FEED_FORMAT': format}
|
||||
data = yield self.exported_data(items, settings)
|
||||
self.assertEqual(formats[format], data)
|
||||
|
||||
|
||||
formats = {
|
||||
'json': u'[\n{"foo": "Test\xd6"}\n]'.encode('latin-1'),
|
||||
'jsonlines': u'{"foo": "Test\xd6"}\n'.encode('latin-1'),
|
||||
'xml': u'<?xml version="1.0" encoding="latin-1"?>\n<items><item><foo>Test\xd6</foo></item></items>'.encode('latin-1'),
|
||||
'csv': u'foo\r\nTest\xd6\r\n'.encode('latin-1'),
|
||||
}
|
||||
|
||||
|
||||
for format in formats:
|
||||
settings = {'FEED_FORMAT': format, 'FEED_EXPORT_ENCODING': 'latin-1'}
|
||||
data = yield self.exported_data(items, settings)
|
||||
|
|
|
|||
|
|
@ -7,6 +7,7 @@ from scrapy.http import (Request, Response, TextResponse, HtmlResponse,
|
|||
XmlResponse, Headers)
|
||||
from scrapy.selector import Selector
|
||||
from scrapy.utils.python import to_native_str
|
||||
from scrapy.exceptions import NotSupported
|
||||
|
||||
|
||||
class BaseResponseTest(unittest.TestCase):
|
||||
|
|
@ -127,6 +128,18 @@ class BaseResponseTest(unittest.TestCase):
|
|||
absolute = 'http://www.example.com/test'
|
||||
self.assertEqual(joined, absolute)
|
||||
|
||||
def test_shortcut_attributes(self):
|
||||
r = self.response_class("http://example.com", body=b'hello')
|
||||
if self.response_class == Response:
|
||||
msg = "Response content isn't text"
|
||||
self.assertRaisesRegexp(AttributeError, msg, getattr, r, 'text')
|
||||
self.assertRaisesRegexp(NotSupported, msg, r.css, 'body')
|
||||
self.assertRaisesRegexp(NotSupported, msg, r.xpath, '//body')
|
||||
else:
|
||||
r.text
|
||||
r.css('body')
|
||||
r.xpath('//body')
|
||||
|
||||
|
||||
class TextResponseTest(BaseResponseTest):
|
||||
|
||||
|
|
|
|||
|
|
@ -10,7 +10,8 @@ class MailSenderTest(unittest.TestCase):
|
|||
|
||||
def test_send(self):
|
||||
mailsender = MailSender(debug=True)
|
||||
mailsender.send(to=['test@scrapy.org'], subject='subject', body='body', _callback=self._catch_mail_sent)
|
||||
mailsender.send(to=['test@scrapy.org'], subject='subject', body='body',
|
||||
_callback=self._catch_mail_sent)
|
||||
|
||||
assert self.catched_msg
|
||||
|
||||
|
|
@ -24,9 +25,16 @@ class MailSenderTest(unittest.TestCase):
|
|||
self.assertEqual(msg.get_payload(), 'body')
|
||||
self.assertEqual(msg.get('Content-Type'), 'text/plain')
|
||||
|
||||
def test_send_single_values_to_and_cc(self):
|
||||
mailsender = MailSender(debug=True)
|
||||
mailsender.send(to='test@scrapy.org', subject='subject', body='body',
|
||||
cc='test@scrapy.org', _callback=self._catch_mail_sent)
|
||||
|
||||
def test_send_html(self):
|
||||
mailsender = MailSender(debug=True)
|
||||
mailsender.send(to=['test@scrapy.org'], subject='subject', body='<p>body</p>', mimetype='text/html', _callback=self._catch_mail_sent)
|
||||
mailsender.send(to=['test@scrapy.org'], subject='subject',
|
||||
body='<p>body</p>', mimetype='text/html',
|
||||
_callback=self._catch_mail_sent)
|
||||
|
||||
msg = self.catched_msg['msg']
|
||||
self.assertEqual(msg.get_payload(), '<p>body</p>')
|
||||
|
|
@ -90,7 +98,8 @@ class MailSenderTest(unittest.TestCase):
|
|||
|
||||
mailsender = MailSender(debug=True)
|
||||
mailsender.send(to=['test@scrapy.org'], subject=subject, body=body,
|
||||
attachs=attachs, charset='utf-8', _callback=self._catch_mail_sent)
|
||||
attachs=attachs, charset='utf-8',
|
||||
_callback=self._catch_mail_sent)
|
||||
|
||||
assert self.catched_msg
|
||||
self.assertEqual(self.catched_msg['subject'], subject)
|
||||
|
|
@ -99,7 +108,8 @@ class MailSenderTest(unittest.TestCase):
|
|||
msg = self.catched_msg['msg']
|
||||
self.assertEqual(msg['subject'], subject)
|
||||
self.assertEqual(msg.get_charset(), Charset('utf-8'))
|
||||
self.assertEqual(msg.get('Content-Type'), 'multipart/mixed; charset="utf-8"')
|
||||
self.assertEqual(msg.get('Content-Type'),
|
||||
'multipart/mixed; charset="utf-8"')
|
||||
|
||||
payload = msg.get_payload()
|
||||
assert isinstance(payload, list)
|
||||
|
|
|
|||
|
|
@ -208,7 +208,7 @@ class FilesPipelineTestCaseCustomSettings(unittest.TestCase):
|
|||
return "".join([chr(random.randint(97, 123)) for _ in range(10)])
|
||||
|
||||
settings = {
|
||||
"FILES_EXPIRES": random.randint(1, 1000),
|
||||
"FILES_EXPIRES": random.randint(100, 1000),
|
||||
"FILES_URLS_FIELD": random_string(),
|
||||
"FILES_RESULT_FIELD": random_string(),
|
||||
"FILES_STORE": self.tempdir
|
||||
|
|
@ -255,16 +255,17 @@ class FilesPipelineTestCaseCustomSettings(unittest.TestCase):
|
|||
|
||||
def test_subclass_attrs_preserved_custom_settings(self):
|
||||
"""
|
||||
If file settings are defined but they are not defined for subclass class attributes
|
||||
should be preserved.
|
||||
If file settings are defined but they are not defined for subclass
|
||||
settings should be preserved.
|
||||
"""
|
||||
pipeline_cls = self._generate_fake_pipeline()
|
||||
settings = self._generate_fake_settings()
|
||||
pipeline = pipeline_cls.from_settings(Settings(settings))
|
||||
for pipe_attr, settings_attr, pipe_ins_attr in self.file_cls_attr_settings_map:
|
||||
value = getattr(pipeline, pipe_ins_attr)
|
||||
setting_value = settings.get(settings_attr)
|
||||
self.assertNotEqual(value, self.default_cls_settings[pipe_attr])
|
||||
self.assertEqual(value, getattr(pipeline, pipe_attr))
|
||||
self.assertEqual(value, setting_value)
|
||||
|
||||
def test_no_custom_settings_for_subclasses(self):
|
||||
"""
|
||||
|
|
@ -321,6 +322,24 @@ class FilesPipelineTestCaseCustomSettings(unittest.TestCase):
|
|||
self.assertEqual(pipeline.files_urls_field, "that")
|
||||
|
||||
|
||||
def test_user_defined_subclass_default_key_names(self):
|
||||
"""Test situation when user defines subclass of FilesPipeline,
|
||||
but uses attribute names for default pipeline (without prefixing
|
||||
them with pipeline class name).
|
||||
"""
|
||||
settings = self._generate_fake_settings()
|
||||
|
||||
class UserPipe(FilesPipeline):
|
||||
pass
|
||||
|
||||
pipeline_cls = UserPipe.from_settings(Settings(settings))
|
||||
|
||||
for pipe_attr, settings_attr, pipe_inst_attr in self.file_cls_attr_settings_map:
|
||||
expected_value = settings.get(settings_attr)
|
||||
self.assertEqual(getattr(pipeline_cls, pipe_inst_attr),
|
||||
expected_value)
|
||||
|
||||
|
||||
class TestS3FilesStore(unittest.TestCase):
|
||||
@defer.inlineCallbacks
|
||||
def test_persist(self):
|
||||
|
|
|
|||
Some files were not shown because too many files have changed in this diff Show More
Loading…
Reference in New Issue