From 1d36c3f0dd60b6a8281b610d9f27ce78c62d40e9 Mon Sep 17 00:00:00 2001 From: Pablo Hoffman Date: Tue, 6 Jan 2009 17:34:14 +0000 Subject: [PATCH] updated robotstxt, spidermw and downloadermw docs --HG-- extra : convert_revision : svn%3Ab85faa78-f9eb-468e-a121-7cced6da292c%40662 --- .../docs/topics/downloader-middleware.rst | 50 +++++++++++-------- scrapy/trunk/docs/topics/robotstxt.rst | 6 +-- .../trunk/docs/topics/spider-middleware.rst | 26 +++++----- 3 files changed, 44 insertions(+), 38 deletions(-) diff --git a/scrapy/trunk/docs/topics/downloader-middleware.rst b/scrapy/trunk/docs/topics/downloader-middleware.rst index be849744b..625a08cf9 100644 --- a/scrapy/trunk/docs/topics/downloader-middleware.rst +++ b/scrapy/trunk/docs/topics/downloader-middleware.rst @@ -8,10 +8,10 @@ The downloader middleware is a framework of hooks into Scrapy's request/response processing. It's a light, low-level system for globally altering Scrapy's input and/or output. -Activating a middleware -======================= +Activating a downloader middleware +================================== -To activate a middleware component, add it to the +To activate a downloader middleware component, add it to the :setting:`DOWNLOADER_MIDDLEWARES` list in your Scrapy settings. In :setting:`DOWNLOADER_MIDDLEWARES`, each middleware component is represented by a string: the full Python path to the middleware's class name. For example:: @@ -42,12 +42,12 @@ middleware. :class:`~scrapy.http.Response` object, or a :class:`~scrapy.http.Request` object. -If returns None, Scrapy will continue processing this request, executing all -other middlewares until, finally, the appropiate downloader handler is called +If returns ``None``, Scrapy will continue processing this request, executing all +other middlewares until, finally, the appropriate downloader handler is called the request performed (and its response downloaded). If returns a Response object, Scrapy won't bother calling ANY other request or -exception middleware, or the appropiate download function; it'll return that +exception middleware, or the appropriate download function; it'll return that Response. Response middleware is always called on every response. If returns a :class:`~scrapy.http.Request` object, returned request is used to @@ -61,32 +61,38 @@ original request don't finish until redirected request is completed. ``response`` is a :class:`~scrapy.http.Response` object ``spider`` is a BaseSpider object -process_response MUST return a Response object. It could alter the given -response, or it could create a brand-new Response. -To drop the response entirely an IgnoreRequest exception must be raised. +``process_response()`` should return a Response object or raise a +:exception:`IgnoreRequest` exception. -.. method:: process_exception(request, exception, spider) +If returns a Response (it could be the same given response, or a brand-new one) +that response will continue to be processed with the ``process_response()`` of +the next middleware in the pipeline. + +If returns an :exception:`IgnoreRequest` exception, the response will be +dropped completely and its callback never called. + +.. method:: process_download_exception(request, exception, spider) ``request`` is a :class:`~scrapy.http.Request` object. ``exception`` is an Exception object ``spider`` is a BaseSpider object -Scrapy calls process_exception() when a download handler or -process_request middleware raises an exception. +Scrapy calls ``process_download_exception()`` when a download handler or a +``process_request()`` (from a downloader middleware) raises an exception. -process_exception() should return either None, :class:`~scrapy.http.Response` -or :class:`~scrapy.http.Request` object. +``process_download_exception()`` should return either ``None``, +:class:`~scrapy.http.Response` or :class:`~scrapy.http.Request` object. -if it returns None, Scrapy will continue processing this exception, -executing any other exception middleware, until no middleware left and -default exception handling kicks in. +If it returns ``None``, Scrapy will continue processing this exception, +executing any other exception middleware, until no middleware is left and +the default exception handling kicks in. If it returns a :class:`~scrapy.http.Response` object, the response middleware kicks in, and won't bother calling any other exception middleware. -If it returns a :class:`~scrapy.http.Request` object, returned request is used to instruct a -immediate redirection. Redirection is handled inside middleware scope, -and original request don't finish until redirected request is -completed. This stop process_exception middleware as returning -Response does. +If it returns a :class:`~scrapy.http.Request` object, returned request is used +to instruct a immediate redirection. Redirection is handled inside middleware +scope, and the original request won't finish until redirected request is +completed. This stop ``process_download_exception()`` middleware as returning Response +would do. diff --git a/scrapy/trunk/docs/topics/robotstxt.rst b/scrapy/trunk/docs/topics/robotstxt.rst index fe12e316f..57b7718ac 100644 --- a/scrapy/trunk/docs/topics/robotstxt.rst +++ b/scrapy/trunk/docs/topics/robotstxt.rst @@ -1,8 +1,8 @@ .. _topics-robotstxt: -================== -Obeying robots.txt -================== +========== +robots.txt +========== Scrapy deals with robots.txt files using a :ref:`topics-downloader-middleware`. called `RobotsTxtMiddleware`. diff --git a/scrapy/trunk/docs/topics/spider-middleware.rst b/scrapy/trunk/docs/topics/spider-middleware.rst index 015707448..63278a09c 100644 --- a/scrapy/trunk/docs/topics/spider-middleware.rst +++ b/scrapy/trunk/docs/topics/spider-middleware.rst @@ -30,20 +30,20 @@ The first (top) middleware is the one closer to the engine, the last (bottom) middleware is the one closer to the spider. Writing your own spider middleware -====================================== +================================== Writing your own spider middleware is easy. Each middleware component is a single Python class that defines one or more of the following methods: -.. method:: process_scrape(response, spider) +.. method:: process_spider_input(response, spider) ``response`` is a :class:`~scrapy.http.Response` object ``spider`` is a :class:`~scrapy.spider.BaseSpider` object This method is called for each request that goes through the spider middleware. -``process_scrape()`` should return either ``None`` or an iterable of +``process_spider_input()`` should return either ``None`` or an iterable of :class:`~scrapy.http.Response` or :class:`~scrapy.http.ScrapedItem` objects. If returns ``None``, Scrapy will continue processing this response, executing all @@ -51,10 +51,10 @@ other middlewares until, finally, the response is handled to the spider for processing. If returns an iterable, Scrapy won't bother calling ANY other spider middleware -``process_scrape()`` and will return the iterable back in the other direction -for the ``process_exception()`` and ``process_result`` methods to hook it. +``process_spider_input()`` and will return the iterable back in the other direction +for the ``process_spider_exception()`` and ``process_spider_output()`` methods to hook it. -.. method:: process_result(response, result, spider) +.. method:: process_spider_output(response, result, spider) ``response`` is a :class:`~scrapy.http.Response` object ``result`` is an iterable of :class:`~scrapy.http.Request` or :class:`~scrapy.item.ScrapedItem` objects @@ -63,25 +63,25 @@ for the ``process_exception()`` and ``process_result`` methods to hook it. This method is called with the results that are returned from the Spider, after it has processed the response. -``process_result()`` must return an iterable of :class:`~scrapy.http.Request` +``process_spider_output()`` must return an iterable of :class:`~scrapy.http.Request` or :class:`~scrapy.item.ScrapedItem` objects. -.. method:: process_exception(request, exception, spider) +.. method:: process_spider_exception(request, exception, spider) ``request`` is a :class:`~scrapy.http.Request` object. ``exception`` is an Exception object ``spider`` is a BaseSpider object -Scrapy calls ``process_exception()`` when a spider or ``process_scrape()`` +Scrapy calls ``process_spider_exception()`` when a spider or ``process_spider_input()`` (from a spider middleware) raises an exception. -process_exception() should return either ``None`` or an iterable of +``process_spider_exception()`` should return either ``None`` or an iterable of :class:`~scrapy.http.Response` or :class:`~scrapy.item.ScrapedItem` objects. If it returns ``None``, Scrapy will continue processing this exception, -executing any other ``process_exception()`` in the middleware pipeline, until +executing any other ``process_spider_exception()`` in the middleware pipeline, until no middleware is left and the default exception handling kicks in. -If it returns an iterable the ``process_result()`` pipeline kicks in, and no -other ``process_exception()`` will be called. +If it returns an iterable the ``process_spider_output()`` pipeline kicks in, and no +other ``process_spider_exception()`` will be called.