improved overview doc. closes #44

--HG--
extra : convert_revision : svn%3Ab85faa78-f9eb-468e-a121-7cced6da292c%40668
This commit is contained in:
Pablo Hoffman 2009-01-07 03:58:15 +00:00
parent a155ad26a9
commit 0319373c61
1 changed files with 175 additions and 21 deletions

View File

@ -1,30 +1,184 @@
.. _overview:
.. _intro-overview:
Overview
========
==================
Scrapy at a glance
==================
Scrapy is a framework designed for retrieving information from websites.
The basic idea of scrapy is to be a robot that goes through websites, crawling pages, and extracting information from them.
Scrapy a is an application framework for crawling web sites and extracting
structured data which can be used for a wide range of useful applications, like
data mining, information processing or historical archival.
The framework is formed by components that take care of different activities.
These components are basically:
Even though Scrapy was originally designed for `screen scraping`_, it can also
be used to extract data using APIs (such as `Amazon Associates Web Services`_)
or as a general purpose web crawler.
* :ref:`topics-spiders`
* :ref:`topics-selectors`
* Items
* Adaptors
.. _screen scraping: http://en.wikipedia.org/wiki/Screen_scraping
.. _Amazon Associates Web Services: http://aws.amazon.com/associates/
The purpose of this document is to introduce you to the concepts behind Scrapy
so you can get an idea of how it works and decide if Scrapy is what you need.
When you're ready to start a project, you can :ref:`start with the tutorial
<intro-tutorial1>`. For more detailed information you can take a look at the
:ref:`documentation contents <index>`.
Pick a website
==============
So you need to extract some information from a website, but the website doesn't
provide any API or mechanism to access that info from a computer program.
Scrapy can help you extract that information. Let's say we want to extract
information about all torrent files added today in the `mininova`_ torrent
site.
.. _mininova: http://www.mininova.org
The list of all torrents added today can be found in this page:
http://www.mininova.org/today
Define the Item
===============
First of all we need to define a class for the items we're going to extract, so
let's define a Torrent class, which must inherit from ScrapedItem::
from scrapy.item import ScrapedItem
class Torrent(ScrapedItem):
pass
Write a Spider to extract the Items
===================================
Now we'll write a Spider which defines the start URL
(http://www.mininova.org/today), the rules for following links and extracting
data from pages.
If we take a look at that page content we'll see that all torrent URLs are like
http://www.mininova.org/tor/NUMBER where NUMBER is a integer. We'll use that to
construct the regular expression for links to follow: `/tor/\d+`.
For extracting data we'll use `XPath`_ to select the part of the document where
the data is to be extracted. Let's take one of those torrent pages:
http://www.mininova.org/tor/2004522
.. _XPath: http://www.w3.org/TR/xpath
And look at the page HTML source to construct the XPath to select the data we
want to extract which is: torrent name, description and size.
.. highlight:: html
By looking at the page HTML source we can see that the file name is contained
inside a ``<h1>`` tag::
<h1>The Dark Knight[2008]DvDrip-aXXo</h1>
.. highlight:: none
An XPath expression to extract the name could be::
//h1/text()
.. highlight:: html
And the description is contained inside a ``<div>`` tag with ``id="description"``::
<h2>Description:</h2>
<div id="description">
&gt; F i L E i N F O <br />
&gt;<br />
&gt; TiTLE......[ The Dark Knight<br />
&gt; AKA........[ Batman Begins 2<br />
...
.. highlight:: none
An XPath expression to select the description could be::
//div[@id='description']
.. highlight:: html
Finally, the file size is contained in the second ``<p>`` tag inside the ``<div>``
tag with ``id=info-left``::
<div id="info-left">
<p>
<strong>Category:</strong>
<a href="/cat/4">Movies</a> &gt; <a href="/sub/1">Action</a>
</p>
<p>
<strong>Total size:</strong>
801.44&nbsp;megabyte</p>
.. highlight:: none
An XPath expression to select the description could be::
//div[@id='info-left']/p[2]/text()[2]
.. highlight:: python
For more information about XPath see the `XPath reference`_.
.. _XPath reference: http://www.w3.org/TR/xpath
Finally, here's the spider code::
class MininovaSpider(CrawlSpider):
domain_name = 'mininova.org'
start_urls = ['http://www.mininova.org/today']
rules = [Rule(RegexLinkExtractor(allow=['/tor/\d+']), 'parse_torrent')]
def parse_torrent(self, response):
x = HtmlXPathSelector(response)
torrent = Torrent()
torrent.url = response.url
torrent.name = x.x("//h1/text()").extract()
torrent.description = x.x("//div[@id='description']").extract()
torrent.size = x.x("//div[@id='info-left']/p[2]/text()[2]").extract()
return [torrent]
Features
--------
For brevity sake, we intentionally left out the import statements and the
Torrent class definition (which is included some paragraphs above).
Scrapy includes many interesting features that make the scraping process much more easier and faster. These include:
Write a pipeline to store the items extracted
=============================================
* Asynchronous crawling/parsing on top of the Twisted framework.
* Easily configurable crawling through sets of rules.
* Ability for parsing HTML, XML, and CSV files.
* Media pipeline useful for scraping items with images or any other media files.
* *Very* extensible thanks to pipelines, middlewares, downloader-middlewares, and extensions.
* Automatic handling of compression, cache, cookies, authentication and more through already-included middlewares.
* Interactive scraping shell console, very useful for developing.
Now let's write an :ref:`topics-item-pipeline` that serializes and stores the
extracted item into a file using `pickle`_::
import pickle
class StoreItemPipeline(object):
def process_item(self, domain, response, item):
torrent_id = item.url.split('/')[-1]
f = open("/tmp/torrent-%s" % torrent_id, "w")
pickle.dump(item, f)
f.close()
.. _pickle: http://docs.python.org/library/pickle.html
What next?
==========
You've seen how to extract and store items from a website using Scrapy, but
this is just the surface. Scrapy provides a lot of powerful features for making
scraping easy and efficient.
The next obvious steps are for you to `download Scrapy`_, read :ref:`the
tutorial <intro-tutorial1>` and join `the community`_. Thanks for your
interest!
.. _download Scrapy: http://scrapy.org/download/
.. _the community: http://scrapy.org/community/