mirror of https://github.com/scrapy/scrapy.git
Merge pull request #1106 from eliasdorneles/overview-page-improvements
[MRG+1] some improvements to overview page
This commit is contained in:
commit
f4e241a018
|
|
@ -13,172 +13,108 @@ precisely, `web scraping`_), it can also be used to extract data using APIs
|
|||
(such as `Amazon Associates Web Services`_) or as a general purpose web
|
||||
crawler.
|
||||
|
||||
The purpose of this document is to introduce you to the concepts behind Scrapy
|
||||
so you can get an idea of how it works and decide if Scrapy is what you need.
|
||||
|
||||
When you're ready to start a project, you can :ref:`start with the tutorial
|
||||
<intro-tutorial>`.
|
||||
Walk-through of an example spider
|
||||
=================================
|
||||
|
||||
Pick a website
|
||||
==============
|
||||
In order to show you what Scrapy brings to the table, we'll walk you through an
|
||||
example of a Scrapy Spider using the simplest way to run a spider.
|
||||
|
||||
So you need to extract some information from a website, but the website doesn't
|
||||
provide any API or mechanism to access that info programmatically. Scrapy can
|
||||
help you extract that information.
|
||||
|
||||
Let's say we want to extract the URL, name, description and size of all torrent
|
||||
files added today in the `Mininova`_ site.
|
||||
|
||||
The list of all torrents added today can be found on this page:
|
||||
|
||||
http://www.mininova.org/today
|
||||
|
||||
.. _intro-overview-item:
|
||||
|
||||
Define the data you want to scrape
|
||||
==================================
|
||||
|
||||
The first thing is to define the data we want to scrape. In Scrapy, this is
|
||||
done through :ref:`Scrapy Items <topics-items>` (Torrent files, in this case).
|
||||
|
||||
This would be our Item::
|
||||
So, here's the code for a spider that follows the links to the top
|
||||
voted questions on StackOverflow and scrapes some data from each page::
|
||||
|
||||
import scrapy
|
||||
|
||||
class TorrentItem(scrapy.Item):
|
||||
url = scrapy.Field()
|
||||
name = scrapy.Field()
|
||||
description = scrapy.Field()
|
||||
size = scrapy.Field()
|
||||
|
||||
Write a Spider to extract the data
|
||||
==================================
|
||||
class StackOverflowSpider(scrapy.Spider):
|
||||
name = 'stackoverflow'
|
||||
start_urls = ['http://stackoverflow.com/questions?sort=votes']
|
||||
|
||||
The next thing is to write a Spider which defines the start URL
|
||||
(http://www.mininova.org/today), the rules for following links and the rules
|
||||
for extracting the data from pages.
|
||||
def parse(self, response):
|
||||
for href in response.css('.question-summary h3 a::attr(href)'):
|
||||
full_url = response.urljoin(href.extract())
|
||||
yield scrapy.Request(full_url, callback=self.parse_question)
|
||||
|
||||
If we take a look at that page content we'll see that all torrent URLs are like
|
||||
``http://www.mininova.org/tor/NUMBER`` where ``NUMBER`` is an integer. We'll use
|
||||
that to construct the regular expression for the links to follow: ``/tor/\d+``.
|
||||
|
||||
We'll use `XPath`_ for selecting the data to extract from the web page HTML
|
||||
source. Let's take one of those torrent pages:
|
||||
|
||||
http://www.mininova.org/tor/2676093
|
||||
|
||||
And look at the page HTML source to construct the XPath to select the data we
|
||||
want which is: torrent name, description and size.
|
||||
|
||||
.. highlight:: html
|
||||
|
||||
By looking at the page HTML source we can see that the file name is contained
|
||||
inside a ``<h1>`` tag::
|
||||
|
||||
<h1>Darwin - The Evolution Of An Exhibition</h1>
|
||||
|
||||
.. highlight:: none
|
||||
|
||||
An XPath expression to extract the name could be::
|
||||
|
||||
//h1/text()
|
||||
|
||||
.. highlight:: html
|
||||
|
||||
And the description is contained inside a ``<div>`` tag with ``id="description"``::
|
||||
|
||||
<h2>Description:</h2>
|
||||
|
||||
<div id="description">
|
||||
Short documentary made for Plymouth City Museum and Art Gallery regarding the setup of an exhibit about Charles Darwin in conjunction with the 200th anniversary of his birth.
|
||||
|
||||
...
|
||||
|
||||
.. highlight:: none
|
||||
|
||||
An XPath expression to select the description could be::
|
||||
|
||||
//div[@id='description']
|
||||
|
||||
.. highlight:: html
|
||||
|
||||
Finally, the file size is contained in the second ``<p>`` tag inside the ``<div>``
|
||||
tag with ``id=specifications``::
|
||||
|
||||
<div id="specifications">
|
||||
|
||||
<p>
|
||||
<strong>Category:</strong>
|
||||
<a href="/cat/4">Movies</a> > <a href="/sub/35">Documentary</a>
|
||||
</p>
|
||||
|
||||
<p>
|
||||
<strong>Total size:</strong>
|
||||
150.62 megabyte</p>
|
||||
def parse_question(self, response):
|
||||
yield {
|
||||
'title': response.css('h1 a::text').extract()[0],
|
||||
'votes': response.css('.question .vote-count-post::text').extract()[0],
|
||||
'body': response.css('.question .post-text').extract()[0],
|
||||
'tags': response.css('.question .post-tag::text').extract(),
|
||||
'link': response.url,
|
||||
}
|
||||
|
||||
|
||||
.. highlight:: none
|
||||
Put this in a file, name it to something like ``stackoverflow_spider.py``
|
||||
and run the spider using the :command:`runspider` command::
|
||||
|
||||
An XPath expression to select the file size could be::
|
||||
scrapy runspider stackoverflow_spider.py -o top-stackoverflow-questions.json
|
||||
|
||||
//div[@id='specifications']/p[2]/text()[2]
|
||||
|
||||
.. highlight:: python
|
||||
When this finishes you will have in the ``top-stackoverflow-questions.json`` file
|
||||
a list of the most upvoted questions in StackOverflow in JSON format, containing the
|
||||
title, link, number of upvotes, a list of the tags and the question content in HTML,
|
||||
looking like this (reformatted for easier reading)::
|
||||
|
||||
For more information about XPath see the `XPath reference`_.
|
||||
[{
|
||||
"body": "... LONG HTML HERE ...",
|
||||
"link": "http://stackoverflow.com/questions/11227809/why-is-processing-a-sorted-array-faster-than-an-unsorted-array",
|
||||
"tags": ["java", "c++", "performance", "optimization"],
|
||||
"title": "Why is processing a sorted array faster than an unsorted array?",
|
||||
"votes": "9924"
|
||||
},
|
||||
{
|
||||
"body": "... LONG HTML HERE ...",
|
||||
"link": "http://stackoverflow.com/questions/1260748/how-do-i-remove-a-git-submodule",
|
||||
"tags": ["git", "git-submodules"],
|
||||
"title": "How do I remove a Git submodule?",
|
||||
"votes": "1764"
|
||||
},
|
||||
...]
|
||||
|
||||
Finally, here's the spider code::
|
||||
|
||||
from scrapy.contrib.spiders import CrawlSpider, Rule
|
||||
from scrapy.contrib.linkextractors import LinkExtractor
|
||||
|
||||
class MininovaSpider(CrawlSpider):
|
||||
What just happened?
|
||||
-------------------
|
||||
|
||||
name = 'mininova'
|
||||
allowed_domains = ['mininova.org']
|
||||
start_urls = ['http://www.mininova.org/today']
|
||||
rules = [Rule(LinkExtractor(allow=['/tor/\d+']), 'parse_torrent')]
|
||||
When you ran the command ``scrapy runspider somefile.py``, Scrapy looked for a
|
||||
Spider definition inside it and ran it through its crawler engine.
|
||||
|
||||
def parse_torrent(self, response):
|
||||
torrent = TorrentItem()
|
||||
torrent['url'] = response.url
|
||||
torrent['name'] = response.xpath("//h1/text()").extract()
|
||||
torrent['description'] = response.xpath("//div[@id='description']").extract()
|
||||
torrent['size'] = response.xpath("//div[@id='info-left']/p[2]/text()[2]").extract()
|
||||
return torrent
|
||||
The crawl started by making requests to the URLs defined in the ``start_urls``
|
||||
attribute (in this case, only the URL for StackOverflow top questions page),
|
||||
and called the default callback method ``parse`` passing the response object as
|
||||
an argument. In the ``parse`` callback, we extract the links to the
|
||||
question pages using a CSS Selector with a custom extension that allows to get
|
||||
the value for an attribute. Then, we yield a few more requests to be sent,
|
||||
registering the method ``parse_question`` as the callback to be called for each
|
||||
of them as they finish.
|
||||
|
||||
The ``TorrentItem`` class is :ref:`defined above <intro-overview-item>`.
|
||||
Here you notice one of the main advantages about Scrapy: requests are
|
||||
:ref:`scheduled and processed asynchronously <topics-architecture>`. This
|
||||
means that Scrapy doesn't need to wait for a request to be finished and
|
||||
processed, it can send another request or do other things in the meantime. This
|
||||
also means that other requests can keep going even if some request fails or an
|
||||
error happens while handling it.
|
||||
|
||||
Run the spider to extract the data
|
||||
==================================
|
||||
While this enables you to do very fast crawlings sending multiple concurrent
|
||||
requests at the same time in a fault-tolerant way, Scrapy also gives you
|
||||
control over the politeness of the crawl through :ref:`a few settings
|
||||
<topics-settings-ref>`. You can do things like setting a download delay between
|
||||
each request, limit amount of concurrent requests per domain or per IP, and
|
||||
even :ref:`use an auto-throttling extension <topics-autothrottle>` that tries
|
||||
to figure out these automatically.
|
||||
|
||||
Finally, we'll run the spider to crawl the site and output the file
|
||||
``scraped_data.json`` with the scraped data in JSON format::
|
||||
Finally, the ``parse_question`` callback scrapes the question data for each
|
||||
page yielding a dict, which Scrapy then collects and writes to a JSON file as
|
||||
requested in the command line.
|
||||
|
||||
scrapy crawl mininova -o scraped_data.json
|
||||
.. note::
|
||||
|
||||
This uses :ref:`feed exports <topics-feed-exports>` to generate the JSON file.
|
||||
You can easily change the export format (XML or CSV, for example) or the
|
||||
storage backend (FTP or `Amazon S3`_, for example).
|
||||
This is using :ref:`feed exports <topics-feed-exports>` to generate the
|
||||
JSON file, you can easily change the export format (XML or CSV, for example) or the
|
||||
storage backend (FTP or `Amazon S3`_, for example). You can also write an
|
||||
:ref:`item pipeline <topics-item-pipeline>` to store the items in a database.
|
||||
|
||||
You can also write an :ref:`item pipeline <topics-item-pipeline>` to store the
|
||||
items in a database very easily.
|
||||
|
||||
Review scraped data
|
||||
===================
|
||||
|
||||
If you check the ``scraped_data.json`` file after the process finishes, you'll
|
||||
see the scraped items there::
|
||||
|
||||
[{"url": "http://www.mininova.org/tor/2676093", "name": ["Darwin - The Evolution Of An Exhibition"], "description": ["Short documentary made for Plymouth ..."], "size": ["150.62 megabyte"]},
|
||||
# ... other items ...
|
||||
]
|
||||
|
||||
You'll notice that all field values (except for the ``url`` which was assigned
|
||||
directly) are actually lists. This is because the :ref:`selectors
|
||||
<topics-selectors>` return lists. You may want to store single values, or
|
||||
perform some additional parsing/cleansing to the values. That's what
|
||||
:ref:`Item Loaders <topics-loaders>` are for.
|
||||
|
||||
.. _topics-whatelse:
|
||||
|
||||
|
|
@ -190,77 +126,53 @@ this is just the surface. Scrapy provides a lot of powerful features for making
|
|||
scraping easy and efficient, such as:
|
||||
|
||||
* Built-in support for :ref:`selecting and extracting <topics-selectors>` data
|
||||
from HTML and XML sources
|
||||
from HTML/XML sources using CSS selectors extended and XPath expressions,
|
||||
with helper methods to extract using regular expressions.
|
||||
|
||||
* Built-in support for cleaning and sanitizing the scraped data using a
|
||||
collection of reusable filters (called :ref:`Item Loaders <topics-loaders>`)
|
||||
shared between all the spiders.
|
||||
* An :ref:`interactive shell console <topics-shell>` (IPython aware) for trying
|
||||
out the CSS and XPath expressions to scrape data, very useful when writing or
|
||||
debugging your spiders.
|
||||
|
||||
* Built-in support for :ref:`generating feed exports <topics-feed-exports>` in
|
||||
multiple formats (JSON, CSV, XML) and storing them in multiple backends (FTP,
|
||||
S3, local filesystem)
|
||||
|
||||
* A media pipeline for :ref:`automatically downloading images <topics-images>`
|
||||
(or any other media) associated with the scraped items
|
||||
|
||||
* Support for :ref:`extending Scrapy <extending-scrapy>` by plugging
|
||||
your own functionality using :ref:`signals <topics-signals>` and a
|
||||
well-defined API (middlewares, :ref:`extensions <topics-extensions>`, and
|
||||
:ref:`pipelines <topics-item-pipeline>`).
|
||||
|
||||
* Wide range of built-in middlewares and extensions for:
|
||||
|
||||
* cookies and session handling
|
||||
* HTTP compression
|
||||
* HTTP authentication
|
||||
* HTTP cache
|
||||
* user-agent spoofing
|
||||
* robots.txt
|
||||
* crawl depth restriction
|
||||
* and more
|
||||
|
||||
* Robust encoding support and auto-detection, for dealing with foreign,
|
||||
non-standard and broken encoding declarations.
|
||||
|
||||
* Support for creating spiders based on pre-defined templates, to speed up
|
||||
spider creation and make their code more consistent on large projects. See
|
||||
:command:`genspider` command for more details.
|
||||
* :ref:`Strong extensibility support <extending-scrapy>`, allowing you to plug
|
||||
in your own functionality using :ref:`signals <topics-signals>` and a
|
||||
well-defined API (middlewares, :ref:`extensions <topics-extensions>`, and
|
||||
:ref:`pipelines <topics-item-pipeline>`).
|
||||
|
||||
* Extensible :ref:`stats collection <topics-stats>` for multiple spider
|
||||
metrics, useful for monitoring the performance of your spiders and detecting
|
||||
when they get broken
|
||||
|
||||
* An :ref:`Interactive shell console <topics-shell>` for trying XPaths, very
|
||||
useful for writing and debugging your spiders
|
||||
|
||||
* A :ref:`System service <topics-scrapyd>` designed to ease the deployment and
|
||||
run of your spiders in production.
|
||||
* Wide range of built-in extensions and middlewares for handling:
|
||||
* cookies and session handling
|
||||
* HTTP features like compression, authentication, caching
|
||||
* user-agent spoofing
|
||||
* robots.txt
|
||||
* crawl depth restriction
|
||||
* and more
|
||||
|
||||
* A :ref:`Telnet console <topics-telnetconsole>` for hooking into a Python
|
||||
console running inside your Scrapy process, to introspect and debug your
|
||||
crawler
|
||||
|
||||
* :ref:`Logging <topics-logging>` facility that you can hook on to for catching
|
||||
errors during the scraping process.
|
||||
|
||||
* Support for crawling based on URLs discovered through `Sitemaps`_
|
||||
|
||||
* A caching DNS resolver
|
||||
* Plus other goodies like reusable spiders to crawl sites from `Sitemaps`_ and
|
||||
XML/CSV feeds, a media pipeline for :ref:`automatically downloading images <topics-images>`
|
||||
(or any other media) associated with the scraped items, a caching DNS resolver,
|
||||
and much more!
|
||||
|
||||
What's next?
|
||||
============
|
||||
|
||||
The next obvious steps are for you to `download Scrapy`_, read :ref:`the
|
||||
tutorial <intro-tutorial>` and join `the community`_. Thanks for your
|
||||
The next steps for you are to :ref:`install Scrapy <intro-install>`,
|
||||
:ref:`follow through the tutorial <intro-tutorial>` to learn how to organize
|
||||
your code in Scrapy projects and `join the community`_. Thanks for your
|
||||
interest!
|
||||
|
||||
.. _download Scrapy: http://scrapy.org/download/
|
||||
.. _the community: http://scrapy.org/community/
|
||||
.. _join the community: http://scrapy.org/community/
|
||||
.. _screen scraping: http://en.wikipedia.org/wiki/Screen_scraping
|
||||
.. _web scraping: http://en.wikipedia.org/wiki/Web_scraping
|
||||
.. _Amazon Associates Web Services: https://affiliate-program.amazon.com/gp/advertising/api/detail/main.html
|
||||
.. _Mininova: http://www.mininova.org
|
||||
.. _XPath: http://www.w3.org/TR/xpath
|
||||
.. _XPath reference: http://www.w3.org/TR/xpath
|
||||
.. _Amazon Associates Web Services: http://aws.amazon.com/associates/
|
||||
.. _Amazon S3: http://aws.amazon.com/s3/
|
||||
.. _Sitemaps: http://www.sitemaps.org
|
||||
|
|
|
|||
|
|
@ -1,3 +1,5 @@
|
|||
.. _topics-autothrottle:
|
||||
|
||||
======================
|
||||
AutoThrottle extension
|
||||
======================
|
||||
|
|
|
|||
Loading…
Reference in New Issue