diff --git a/docs/_templates/index.html b/docs/_templates/index.html index 323870d78..a8bf042d5 100644 --- a/docs/_templates/index.html +++ b/docs/_templates/index.html @@ -8,7 +8,7 @@ for an overview and tutorial

-
  • Home |
  • Getting Started |
  • Using Scrapy |
  • -
  • API reference |
  • +
  • API reference |
  • FAQ |
  • Search
  • diff --git a/docs/contents.rst b/docs/contents.rst index 605176260..6a748f364 100644 --- a/docs/contents.rst +++ b/docs/contents.rst @@ -8,6 +8,6 @@ Scrapy documentation contents intro/index topics/index - ref/index + reference faq experimental/index diff --git a/docs/faq.rst b/docs/faq.rst index 8a1fd8667..fb2a7bab0 100644 --- a/docs/faq.rst +++ b/docs/faq.rst @@ -69,7 +69,7 @@ You need to install `pywin32`_ because of `this Twisted bug`_. How can I simulate a user login in my spider? --------------------------------------------- -See :ref:`ref-request-userlogin`. +See :ref:`topics-request-response-ref-request-userlogin`. Can I crawl in depth-first order instead of breadth-first order? ---------------------------------------------------------------- diff --git a/docs/ref/downloader-middleware.rst b/docs/ref/downloader-middleware.rst deleted file mode 100644 index 976e32c04..000000000 --- a/docs/ref/downloader-middleware.rst +++ /dev/null @@ -1,71 +0,0 @@ -.. _ref-downloader-middleware: - -======================================== -Built-in downloader middleware reference -======================================== - -This page describes all downloader middleware components that come with -Scrapy. For information on how to use them and how to write your own downloader -middleware, see the :ref:`downloader middleware usage guide -`. - -For a list of the components enabled by default (and their orders) see the -:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting. - -Available downloader middlewares -================================ - -DefaultHeadersMiddleware ------------------------- - -.. module:: scrapy.contrib.downloadermiddleware.defaultheaders - :synopsis: Default Headers Downloader Middleware - -.. class:: DefaultHeadersMiddleware - - This middleware sets all default requests headers specified in the - :setting:`DEFAULT_REQUEST_HEADERS` setting. - -DebugMiddleware ---------------- - -.. module:: scrapy.contrib.downloadermiddleware.debug - :synopsis: Downloader middlewares for debugging - -.. class:: DebugMiddleware - - This is a convenient middleware to inspect what's passing through the - downloader middleware. It logs all requests and responses catched by the - middleware component methods. This middleware does not use any settings and - does not come enabled by default. Instead, it's meant to be inserted at the - point of the middleware that you want to inspect. - -HttpCacheMiddleware -------------------- - -.. module:: scrapy.contrib.downloadermiddleware.httpcache - :synopsis: HTTP Cache downloader middleware - -.. class:: HttpCacheMiddleware - - This middleware provides low-level cache to all HTTP requests and responses. - Every request and its corresponding response are cached and then, when that - same request is seen again, the response is returned without transferring - anything from the Internet. - - The HTTP cache is useful for testing spiders faster (without having to wait for - downloads every time) and for trying your spider off-line when you don't have - an Internet connection. - - The :class:`HttpCacheMiddleware` can be configured through the following - settings (see the settings documentation for more info): - - * :setting:`HTTPCACHE_DIR` - this one actually enables the cache besides - settings the cache dir - * :setting:`HTTPCACHE_IGNORE_MISSING` - ignoring missing requests instead - of downloading them - * :setting:`HTTPCACHE_SECTORIZE` - split HTTP cache in several directories - (for performance reasons) - * :setting:`HTTPCACHE_EXPIRATION_SECS` - how many secs until the cache is - considered out of date - diff --git a/docs/ref/extension-manager.rst b/docs/ref/extension-manager.rst deleted file mode 100644 index 2e0d18dee..000000000 --- a/docs/ref/extension-manager.rst +++ /dev/null @@ -1,62 +0,0 @@ -.. _ref-extension-manager: - -================= -Extension Manager -================= - -.. module:: scrapy.extension - :synopsis: The extension manager - -The Extension Manager is responsible for loading and keeping track of installed -extensions and it's configured through the :setting:`EXTENSIONS` setting which -contains a dictionary of all available extensions and their order similar to -how you :ref:`configure the downloader middlewares -`. - -.. class:: ExtensionManager - - The extension manager is a singleton object, which is instantiated at module - loading time and can be accessed like this:: - - from scrapy.extension import extensions - - .. attribute:: loaded - - A boolean which is True if extensions are already loaded or False if - they're not. - - .. attribute:: enabled - - A dict with the enabled extensions. The keys are the extension class names, - and the values are the extension objects. Example:: - - >>> from scrapy.extension import extensions - >>> extensions.load() - >>> print extensions.enabled - {'CoreStats': , - 'WebConsoke': , - ... - - .. attribute:: disabled - - A dict with the disabled extensions. The keys are the extension class names, - and the values are the extension class paths (because objects are never - instantiated for disabled extensions). Example:: - - >>> from scrapy.extension import extensions - >>> extensions.load() - >>> print extensions.disabled - {'MemoryDebugger': 'scrapy.contrib.webconsole.stats.MemoryDebugger', - 'SpiderProfiler': 'scrapy.contrib.spider.profiler.SpiderProfiler', - ... - - .. method:: load() - - Load the available extensions configured in the :setting:`EXTENSIONS` - setting. On a standard run, this method is usually called by the Execution - Manager, but you may need to call it explicitly if you're dealing with - code outside Scrapy. - - .. method:: reload() - - Reload the available extensions. See :meth:`load`. diff --git a/docs/ref/extensions.rst b/docs/ref/extensions.rst deleted file mode 100644 index d8c467959..000000000 --- a/docs/ref/extensions.rst +++ /dev/null @@ -1,247 +0,0 @@ -.. _ref-extensions: - -============================= -Built-in extensions reference -============================= - -This document explains all extensions that come with Scrapy. For information on -how to use them and how to write your own extensions, see the :ref:`extensions -usage guide `. - - -General purpose extensions -========================== - -Core Stats extension --------------------- - -.. module:: scrapy.stats.corestats - :synopsis: Core stats collection - -.. class:: scrapy.stats.corestats.CoreStats - -Enable the collection of core statistics, provided the stats collection are -enabled (see :ref:`topics-stats`). - -.. _ref-extensions-webconsole: - -Web console extension ---------------------- - -.. module:: scrapy.management.web - :synopsis: Web management console - -.. class:: scrapy.management.web.WebConsole - -Provides an extensible web server for managing a Scrapy process. It's enabled -by the :setting:`WEBCONSOLE_ENABLED` setting. The server will listen in the -port specified in :setting:`WEBCONSOLE_PORT`, and will log to the file -specified in :setting:`WEBCONSOLE_LOGFILE`. - -The web server is designed to be extended by other extensions which can add -their own management web interfaces. - -See also :ref:`topics-webconsole` for information on how to write your own web -console extension, and "Web console extensions" below for a list of available -built-in (web console) extensions. - -.. _ref-extensions-telnetconsole: - -Telnet console extension ------------------------- - -.. module:: scrapy.management.telnet - :synopsis: Telnet management console - -.. class:: scrapy.management.telnet.TelnetConsole - -Provides a telnet console for getting into a Python interpreter inside the -currently running Scrapy process, which can be very useful for debugging. - -The telnet console must be enabled by the :setting:`TELNETCONSOLE_ENABLED` -setting, and the server will listen in the port specified in -:setting:`WEBCONSOLE_PORT`. - -Spider reloader extension -------------------------- - -.. module:: scrapy.contrib.spider.reloader - :synopsis: Spider reloader extension - -.. class:: scrapy.contrib.spider.reloader.SpiderReloader - -Reload spider objects once they've finished scraping, to release the resources -and references to other objects they may hold. - -.. _ref-extensions-memusage: - -Memory usage extension ----------------------- - -.. module:: scrapy.contrib.memusage - :synopsis: Memory usage extension - -.. class:: scrapy.contrib.memusage.MemoryUsage - -Allows monitoring the memory used by a Scrapy process and: - -1, send a notification email when it exceeds a certain value -2. terminate the Scrapy process when it exceeds a certain value - -The notification emails can be triggered when a certain warning value is -reached (:setting:`MEMUSAGE_WARNING_MB`) and when the maximum value is reached -(:setting:`MEMUSAGE_LIMIT_MB`) which will also cause the Scrapy process to be -terminated. - -This extension is enabled by the :setting:`MEMUSAGE_ENABLED` setting and -can be configured with the following settings: - -* :setting:`MEMUSAGE_LIMIT_MB` -* :setting:`MEMUSAGE_WARNING_MB` -* :setting:`MEMUSAGE_NOTIFY_MAIL` -* :setting:`MEMUSAGE_REPORT` - -Memory debugger extension -------------------------- - -.. module:: scrapy.contrib.memdebug - :synopsis: Memory debugger extension - -.. class:: scrapy.contrib.memdebug.MemoryDebugger - -A memory debugger which collects some info about objects uncollected by the -garbage collector and libxml2 memory leaks. To enable this extension turn on -the :setting:`MEMDEBUG_ENABLED` setting. The report will be printed to standard -output. If the :setting:`MEMDEBUG_NOTIFY` setting contains a list of emails the -report will also be sent to those addresses. - -Close domain extension ----------------------- - -.. module:: scrapy.contrib.closedomain - :synopsis: Close domain extension - -.. class:: scrapy.contrib.closedomain.CloseDomain - -Closes a domain/spider automatically when some conditions are met, using a -specific closing reason for each condition. - -The conditions for closing a domain can be configured through the following -settings. Other conditions will be supported in the future. - -.. setting:: CLOSEDOMAIN_TIMEOUT - -CLOSEDOMAIN_TIMEOUT -~~~~~~~~~~~~~~~~~~~ - -Default: ``0`` - -An integer which specifies a number of seconds. If the domain remains open for -more than that number of second, it will be automatically closed with the -reason ``closedomain_timeout``. If zero (or non set) domains won't be closed by -timeout. - -.. setting:: CLOSEDOMAIN_ITEMPASSED - -CLOSEDOMAIN_ITEMPASSED -~~~~~~~~~~~~~~~~~~~~~~ - -Default: ``0`` - -An integer which specifies a number of items. If the spider scrapes more than -that amount if items and those items are passed by the item pipeline, the -domain will be closed with the reason ``closedomain_itempassed``. If zero (or -non set) domains won't be closed by number of passed items. - -Stack trace dump extension ---------------------------- - -.. module:: scrapy.contrib.debug - :synopsis: Extensions for debugging Scrapy - -.. class:: scrapy.contrib.debug.StackTraceDump - -Adds a `SIGUSR1`_ signal handler which dumps the stack trace of a runnning -Scrapy process when a ``SIGUSR1`` signal is catched. After the stack trace is -dumped, the Scrapy process continues to run normally. - -The stack trace is sent to standard output, or to the Scrapy log file if -:setting:`LOG_STDOUT` is enabled. - -This extension only works on POSIX-compliant platforms (ie. not Windows). - -.. _SIGUSR1: http://en.wikipedia.org/wiki/SIGUSR1_and_SIGUSR2 - -StatsMailer extension ---------------------- - -.. module:: scrapy.contrib.statsmailer - :synopsis: StatsMailer extension - -.. class:: scrapy.contrib.statsmailer.StatsMailer - -This simple extension can be used to send a notification email every time a -domain has finished scraping, including the Scrapy stats collected. The email -will be sent to all recipients specified in the :setting:`STATSMAILER_RCPTS` -setting. - -Web console extensions -====================== - -.. module:: scrapy.contrib.webconsole - :synopsis: Contains most built-in web console extensions - -Here is a list of built-in web console extensions. For clarity "web console -extension" is abbreviated as "WC extension". - -For more information see the see the :ref:`web console documentation -`. - -Scheduler queue WC extension ----------------------------- - -.. module:: scrapy.contrib.webconsole.scheduler - :synopsis: Scheduler queue web console extension - -.. class:: scrapy.contrib.webconsole.scheduler.SchedulerQueue - -Display a list of all pending Requests in the Scheduler queue, grouped by -domain/spider. - -Spider live stats WC extension ------------------------------- - -.. module:: scrapy.contrib.webconsole.livestats - :synopsis: Spider live stats web console extension - -.. class:: scrapy.contrib.webconsole.livestats.LiveStats - -Display a table with stats of all spider crawled by the current Scrapy run, -including: - -* Number of items scraped -* Number of pages crawled -* Number of pending requests in the scheduler -* Number of pending requests in the downloader queue -* Number of requests currently being downloaded - -Engine status WC extension ---------------------------- - -.. module:: scrapy.contrib.webconsole.enginestatus - :synopsis: Engine status web console extension - -.. class:: scrapy.contrib.webconsole.enginestatus.EngineStatus - -Display the current status of the Scrapy Engine, which is just the output of -the Scrapy engine ``getstatus()`` method. - -Stats collector dump WC extension ----------------------------------- - -.. module:: scrapy.contrib.webconsole.stats - :synopsis: Stats dump web console extension - -.. class:: scrapy.contrib.webconsole.stats.StatsDump - -Display the stats collected so far by the stats collector. diff --git a/docs/ref/index.rst b/docs/ref/index.rst deleted file mode 100644 index eb0adafde..000000000 --- a/docs/ref/index.rst +++ /dev/null @@ -1,26 +0,0 @@ -.. _ref-index: - -API Reference -============= - -This section documents the Scrapy |version| API. For more information see :ref:`misc-api-stability`. - -.. toctree:: - :maxdepth: 1 - - request-response - spiders - selectors - settings - signals - exceptions - logging - email - extension-manager - extensions - downloader-middleware - spider-middleware - scheduler-middleware - link-extractors - -* :ref:`topics-stats-api` diff --git a/docs/ref/link-extractors.rst b/docs/ref/link-extractors.rst deleted file mode 100644 index 103b1253a..000000000 --- a/docs/ref/link-extractors.rst +++ /dev/null @@ -1,121 +0,0 @@ -.. _ref-link-extractors: - -========================= -Available Link Extractors -========================= - -.. module:: scrapy.contrib.linkextractors - :synopsis: Link extractors classes - -All available link extractors classes bundled with Scrapy are provided in the -:mod:`scrapy.contrib.linkextractors` module. - -.. module:: scrapy.contrib.linkextractors.sgml - :synopsis: SGMLParser-based link extractors - -SgmlLinkExtractor -================= - -.. class:: SgmlLinkExtractor(allow=(), deny=(), allow_domains=(), deny_domains=(), restrict_xpaths(), tags=('a', 'area'), attrs=('href'), canonicalize=True, unique=True, process_value=None) - - The SgmlLinkExtractor extends the base :class:`BaseSgmlLinkExtractor` by - providing additional filters that you can specify to extract links, - including regular expressions patterns that the links must match to be - extracted. All those filters are configured through these constructor - parameters: - - :param allow: a single regular expression (or list of regular expressions) - that the (absolute) urls must match in order to be extracted. If not - given (or empty), it will match all links. - :type allow: a regular expression (or list of) - - :param deny: a single regular expression (or list of regular expressions) - that the (absolute) urls must match in order to be excluded (ie. not - extracted). It has precedence over the ``allow`` parameter. If not - given (or empty) it won't exclude any links. - :type allow: a regular expression (or list of) - - :param allow_domains: is single value or a list of string containing - domains which will be considered for extracting the links - :type allow: str or list - - :param deny_domains: is single value or a list of strings containing - domains which which won't be considered for extracting the links - :type allow: str or list - - :param restrict_xpaths: is a XPath (or list of XPath's) which defines - regions inside the response where links should be extracted from. - If given, only the text selected by those XPath will be scanned for - links. See examples below. - :type restrict_xpaths: str or list - - :param tags: a tag or a list of tags to consider when extracting links. - Defaults to ``('a', 'area')``. - :type tags: str or list - - :param attrs: list of attrbitues which should be considered when looking - for links to extract (only for those tags specified in the ``tags`` - parameter). Defaults to ``('href',)`` - :type attrs: boolean - - :param canonicalize: canonicalize each extracted url (using - scrapy.utils.url.canonicalize_url). Defaults to ``True``. - :type canonicalize: boolean - - :param unique: whether duplicate filtering should be applied to extracted - links. - :type unique: boolean - - :param process_value: see ``process_value`` argument of - :class:`LinkExtractor` class constructor - :type process_value: boolean - -BaseSgmlLinkExtractor -===================== - -.. class:: BaseSgmlLinkExtractor(tag="a", href="href", unique=False, process_value=None) - - The purpose of this Link Extractor is only to serve as a base class for the - :class:`SgmlLinkExtractor`. You should use that one instead. - - The constructor arguments are: - - :param tag: either a string (with the name of a tag) or a function that - receives a tag name and returns ``True`` if links should be extracted - from those tag, or ``False`` if they shouldn't. Defaults to ``'a'``. - request (once its downloaded) as its first parameter. For more - information see :ref:`ref-request-callback-arguments` below. - :type tag: str or callable - - :param attr: either string (with the name of a tag attribute), or a - function that receives a an attribute name and returns ``True`` if - links should be extracted from it, or ``False`` if the shouldn't. - Defaults to ``href``. - :type attr: str or callable - - :param unique: is a boolean that specifies if a duplicate filtering should - be applied to links extracted. - :type unique: boolean - - :param process_value: a function which receives each value extracted from - the tag and attributes scanned and can modify the value and return a - new one, or return ``None`` to ignore the link altogether. If not - given, ``process_value`` defaults to ``lambda x: x``. - - .. highlight:: html - - For example, to extract links from this code:: - - Link text - - .. highlight:: python - - You can use the following function in ``process_value``:: - - def process_value(value): - m = re.search("javascript:goToPage\('(.*?)'", value) - if m: - return m.group(1) - - :type process_value: callable - diff --git a/docs/ref/selectors.rst b/docs/ref/selectors.rst deleted file mode 100644 index 5be263cdc..000000000 --- a/docs/ref/selectors.rst +++ /dev/null @@ -1,189 +0,0 @@ -.. _ref-selectors: - -=================== -XPath Selectors API -=================== - -.. module:: scrapy.xpath - :synopsis: XPath selectors classes - -There are two types of selectors bundled with Scrapy: -:class:`HtmlXPathSelector` and :class:`XmlXPathSelector`. Both of them -implement the same :class:`XPathSelector` interface. The only different is that -one is used to process HTML data and the other XML data. - -XPathSelector objects -===================== - -.. class:: XPathSelector(response) - - A :class:`XPathSelector` object is a wrapper over response to select - certain parts of its content. - - A :class:`Request` object represents an HTTP request, which is usually - generated in the Spider and executed by the Downloader, and thus generating - a :class:`Response`. - - ``url`` is a :class:`~scrapy.http.Response` object that will be used for - selecting and extracting data - - -XPathSelector Methods ---------------------- - -.. method:: XPathSelector.select(xpath) - - Apply the given XPath relative to this XPathSelector and return a list - of :class:`XPathSelector` objects (ie. a :class:`XPathSelectorList`) with - the result. - - ``xpath`` is a string containing the XPath to apply - -.. method:: XPathSelector.re(regex) - - Apply the given regex and return a list of unicode strings with the - matches. - - ``regex`` can be either a compiled regular expression or a string which - will be compiled to a regular expression using ``re.compile(regex)`` - -.. method:: XPathSelector.extract() - - Return a unicode string with the content of this :class:`XPathSelector` - object. - -.. method:: XPathSelector.extract_unquoted() - - Return a unicode string with the content of this :class:`XPathSelector` - without entities or CDATA. This method is intended to be use for text-only - selectors, like ``//h1/text()`` (but not ``//h1``). If it's used for - :class:`XPathSelector` objects which don't select a textual content (ie. if - they contain tags), the output of this method is undefined. - -.. method:: XPathSelector.register_namespace(prefix, uri) - - Register the given namespace to be used in this :class:`XPathSelector`. - Without registering namespaces you can't select or extract data from - non-standard namespaces. See examples below. - -.. method:: XPathSelector.__nonzero__() - - Returns ``True`` if there is any real content selected by this - :class:`XPathSelector` or ``False`` otherwise. In other words, the boolean - value of an XPathSelector is given by the contents it selects. - -XPathSelectorList objects -========================= - -.. class:: XPathSelectorList - - The :class:`XPathSelectorList` class is subclass of the builtin ``list`` - class, which provides a few additional methods. - - -XPathSelectorList Methods -------------------------- - -.. method:: XPathSelectorList.select(xpath) - - Call the :meth:`XPathSelector.re` method for all :class:`XPathSelector` - objects in this list and return their results flattened, as new - :class:`XPathSelectorList`. - - ``xpath`` is the same argument as the one in :meth:`XPathSelector.x` - -.. method:: XPathSelector.re(regex) - - Call the :meth:`XPathSelector.re` method for all :class:`XPathSelector` - objects in this list and return their results flattened, as a list of - unicode strings. - - ``regex`` is the same argument as the one in :meth:`XPathSelector.re` - -.. method:: XPathSelector.extract() - - Call the :meth:`XPathSelector.re` method for all :class:`XPathSelector` - objects in this list and return their results flattened, as a list of - unicode strings. - -.. method:: XPathSelector.extract_unquoted() - - Call the :meth:`XPathSelector.extract_unoquoted` method for all - :class:`XPathSelector` objects in this list and return their results - flattened, as a list of unicode strings. This method should not be applied - to all kinds of XPathSelectors. For more info see - :meth:`XPathSelector.extract_unoquoted`. - -HtmlXPathSelector objects -========================= - -.. class:: HtmlXPathSelector(response) - - A subclass of :class:`XPathSelector` for working with HTML content. It uses - the `libxml2`_ HTML parser. See the :class:`XPathSelector` API for more info. - -.. _libxml2: http://xmlsoft.org/ - -HtmlXPathSelector examples --------------------------- - -Here's a couple of :class:`HtmlXPathSelector` examples to illustrate several -concepts. In all cases we assume there is already a :class:`HtmlPathSelector` -instanced with a :class:`~scrapy.http.Response` object like this:: - - x = HtmlXPathSelector(html_response) - -1. Select all ``

    `` elements from a HTML response body, returning a list of - :class:`XPathSelector` objects (ie. a :class:`XPathSelectorList` object):: - - x.select("//h1") - -2. Extract the text of all ``

    `` elements from a HTML response body, - returning a list of unicode strings:: - - x.select("//h1").extract() # this includes the h1 tag - x.select("//h1/text()").extract() # this excludes the h1 tag - -3. Iterate over all ``

    `` tags and print their class attribute:: - - for node in x.select("//p"): - ... print node.select("@href") - -4. Extract textual data from all ``

    `` tags without entities, as a list of - unicode strings:: - - x.select("//p/text()").extract_unquoted() - - # the following line is wrong. extract_unquoted() should only be used - # with textual XPathSelectors - x.select("//p").extract_unquoted() # it may work but output is unpredictable - -XmlXPathSelector objects -======================== - -.. class:: XmlXPathSelector(response) - - A subclass of :class:`XPathSelector` for working with XML content. It uses - the `libxml2`_ XML parser. See the :class:`XPathSelector` API for more info. - -XmlXPathSelector examples -------------------------- - -Here's a couple of :class:`XmlXPathSelector` examples to illustrate several -concepts. In all cases we assume there is already a :class:`XmlPathSelector` -instanced with a :class:`~scrapy.http.Response` object like this:: - - x = HtmlXPathSelector(xml_response) - -1. Select all ```` elements from a XML response body, returning a list of - :class:`XPathSelector` objects (ie. a :class:`XPathSelectorList` object):: - - x.select("//h1") - -2. Extract all prices from a `Google Base XML feed`_ which requires registering - a namespace:: - - x.register_namespace("g", "http://base.google.com/ns/1.0") - x.select("//g:price").extract() - -.. _Google Base XML feed: http://base.google.com/support/bin/answer.py?hl=en&answer=59461 diff --git a/docs/ref/settings.rst b/docs/ref/settings.rst deleted file mode 100644 index b3bc77840..000000000 --- a/docs/ref/settings.rst +++ /dev/null @@ -1,886 +0,0 @@ -.. _settings: - -Available Settings -================== - -Here's a list of all available Scrapy settings, in alphabetical order, along -with their default values and the scope where they apply. - -The scope, where available, shows where the setting is being used, if it's tied -to any particular component. In that case the module of that component will be -shown, typically an extension, middleware or pipeline. It also means that the -component must be enabled in order for the setting to have any effect. - -.. setting:: BOT_NAME - -BOT_NAME --------- - -Default: ``scrapybot`` - -The name of the bot implemented by this Scrapy project. This will be used to -construct the User-Agent by default, and also for logging. - -.. setting:: BOT_VERSION - -BOT_VERSION ------------ - -Default: ``1.0`` - -The version of the bot implemented by this Scrapy project. This will be used to -construct the User-Agent by default. - -.. setting:: HTTPCACHE_DIR - -HTTPCACHE_DIR -------------- - -Default: ``''`` (empty string) - -The directory to use for storing the (low-level) HTTP cache. If empty the HTTP -cache will be disabled. - -.. setting:: HTTPCACHE_EXPIRATION_SECS - -HTTPCACHE_EXPIRATION_SECS -------------------------- - -Default: ``0`` - -Number of seconds to use for HTTP cache expiration. Requests that were cached -before this time will be re-downloaded. If zero, cached requests will always -expire. Negative numbers means requests will never expire. - -.. setting:: HTTPCACHE_IGNORE_MISSING - -HTTPCACHE_IGNORE_MISSING ------------------------- - -Default: ``False`` - -If enabled, requests not found in the cache will be ignored instead of downloaded. - -.. setting:: HTTPCACHE_SECTORIZE - -HTTPCACHE_SECTORIZE -------------------- - -Default: ``True`` - -Whether to split HTTP cache storage in several dirs for performance. - -.. setting:: COMMANDS_MODULE - -COMMANDS_MODULE ---------------- - -Default: ``''`` (empty string) - -A module to use for looking for custom Scrapy commands. This is used to add -custom command for your Scrapy project. - -Example:: - - COMMANDS_MODULE = 'mybot.commands' - -.. setting:: COMMANDS_SETTINGS_MODULE - -COMMANDS_SETTINGS_MODULE ------------------------- - -Default: ``''`` (empty string) - -A module to use for looking for custom Scrapy command settings. - -Example:: - - COMMANDS_SETTINGS_MODULE = 'mybot.conf.commands' - -.. setting:: CONCURRENT_DOMAINS - -CONCURRENT_DOMAINS ------------------- - -Default: ``8`` - -Maximum number of domains to scrape in parallel. - -.. setting:: CONCURRENT_ITEMS - -CONCURRENT_ITEMS ----------------- - -Default: ``100`` - -Maximum number of concurrent items (per response) to process in parallel in the -Item Processor (also known as the Item Pipeline). - -.. setting:: COOKIES_DEBUG - -COOKIES_DEBUG -------------- - -Default: ``False`` - -Enable debugging message of Cookies Downloader Middleware. - -.. setting:: DEFAULT_ITEM_CLASS - -DEFAULT_ITEM_CLASS ------------------- - -Default: ``'scrapy.item.ScrapedItem'`` - -The default class that will be used for instantiating items in the :ref:`the -Scrapy shell `. - -.. setting:: DEFAULT_REQUEST_HEADERS - -DEFAULT_REQUEST_HEADERS ------------------------ - -Default:: - - { - 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', - 'Accept-Language': 'en', - } - -The default headers used for Scrapy HTTP Requests. They're populated in the -:class:`~scrapy.contrib.downloadermiddleware.defaultheaders.DefaultHeadersMiddleware`. - -.. setting:: DEFAULT_SPIDER - -DEFAULT_SPIDER --------------- - -Default: ``None`` - -The default spider class that will be instantiated for URLs for which no -specific spider is found. This class must have a constructor which receives as -only parameter the domain name of the given URL. - -.. setting:: DEPTH_LIMIT - -DEPTH_LIMIT ------------ - -Default: ``0`` - -The maximum depth that will be allowed to crawl for any site. If zero, no limit -will be imposed. - -.. setting:: DEPTH_STATS - -DEPTH_STATS ------------ - -Default: ``True`` - -Whether to collect depth stats. - -.. setting:: DOMAIN_SCHEDULER - -DOMAIN_SCHEDULER ----------------- - -Default: ``'scrapy.contrib.domainsch.FifoDomainScheduler'`` - -The Domain Scheduler to use. The domain scheduler returns the next domain -(spider) to scrape. - -.. setting:: DOWNLOADER_DEBUG - -DOWNLOADER_DEBUG ----------------- - -Default: ``False`` - -Whether to enable the Downloader debugging mode. - -.. setting:: DOWNLOADER_MIDDLEWARES - -DOWNLOADER_MIDDLEWARES ----------------------- - -Default:: ``{}`` - -A dict containing the downloader middlewares enabled in your project, and their -orders. For more info see :ref:`topics-downloader-middleware-setting`. - -.. setting:: DOWNLOADER_MIDDLEWARES_BASE - -DOWNLOADER_MIDDLEWARES_BASE ---------------------------- - -Default:: - - { - 'scrapy.contrib.downloadermiddleware.robotstxt.RobotsTxtMiddleware': 100, - 'scrapy.contrib.downloadermiddleware.httpauth.HttpAuthMiddleware': 300, - 'scrapy.contrib.downloadermiddleware.useragent.UserAgentMiddleware': 400, - 'scrapy.contrib.downloadermiddleware.retry.RetryMiddleware': 500, - 'scrapy.contrib.downloadermiddleware.defaultheaders.DefaultHeadersMiddleware': 550, - 'scrapy.contrib.downloadermiddleware.redirect.RedirectMiddleware': 600, - 'scrapy.contrib.downloadermiddleware.cookies.CookiesMiddleware': 700, - 'scrapy.contrib.downloadermiddleware.httpcompression.HttpCompressionMiddleware': 800, - 'scrapy.contrib.downloadermiddleware.stats.DownloaderStats': 850, - 'scrapy.contrib.downloadermiddleware.httpcache.HttpCacheMiddleware': 900, - } - -A dict containing the downloader middlewares enabled by default in Scrapy. You -should never modify this setting in your project, modify -:setting:`DOWNLOADER_MIDDLEWARES` instead. For more info see -:ref:`topics-downloader-middleware-setting`. - -.. setting:: DOWNLOADER_STATS - -DOWNLOADER_STATS ----------------- - -Default: ``True`` - -Whether to enable downloader stats collection. - -.. setting:: DOWNLOAD_DELAY - -DOWNLOAD_DELAY --------------- - -Default: ``0`` - -The amount of time (in secs) that the downloader should wait before downloading -consecutive pages from the same spider. This can be used to throttle the -crawling speed to avoid hitting servers too hard. Decimal numbers are -supported. Example:: - - DOWNLOAD_DELAY = 0.25 # 250 ms of delay - -.. setting:: DOWNLOAD_TIMEOUT - -DOWNLOAD_TIMEOUT ----------------- - -Default: ``180`` - -The amount of time (in secs) that the downloader will wait before timing out. - -.. setting:: DUPEFILTER_CLASS - -DUPEFILTER_CLASS ----------------- - -Default: ``'scrapy.contrib.dupefilter.RequestFingerprintDupeFilter'`` - -The class used to detect and filter duplicate requests. - -The default (``RequestFingerprintDupeFilter``) filters based on request fingerprint -(using ``scrapy.utils.request.request_fingerprint``) and grouping per domain. - -.. setting:: EXTENSIONS - -EXTENSIONS ----------- - -Default:: ``{}`` - -A dict containing the extensions enabled in your project, and their orders. - -.. setting:: EXTENSIONS_BASE - -EXTENSIONS_BASE ---------------- - -Default:: - - { - 'scrapy.stats.corestats.CoreStats': 0, - 'scrapy.management.web.WebConsole': 0, - 'scrapy.management.telnet.TelnetConsole': 0, - 'scrapy.contrib.webconsole.scheduler.SchedulerQueue': 0, - 'scrapy.contrib.webconsole.livestats.LiveStats': 0, - 'scrapy.contrib.webconsole.spiderctl.Spiderctl': 0, - 'scrapy.contrib.webconsole.enginestatus.EngineStatus': 0, - 'scrapy.contrib.webconsole.stats.StatsDump': 0, - 'scrapy.contrib.spider.reloader.SpiderReloader': 0, - 'scrapy.contrib.memusage.MemoryUsage': 0, - 'scrapy.contrib.memdebug.MemoryDebugger': 0, - 'scrapy.contrib.closedomain.CloseDomain': 0, - 'scrapy.contrib.debug.StackTraceDump': 0, - } - -The list of available extensions. Keep in mind that some of them need need to -be enabled through a setting. By default, this setting contains all stable -built-in extensions. - -For more information See the :ref:`extensions user guide ` -and the :ref:`list of available extensions `. - -.. setting:: GROUPSETTINGS_ENABLED - -GROUPSETTINGS_ENABLED ---------------------- - -Default: ``False`` - -Whether to enable group settings where spiders pull their settings from. - -.. setting:: GROUPSETTINGS_MODULE - -GROUPSETTINGS_MODULE --------------------- - -Default: ``''`` (empty string) - -The module to use for pulling settings from, if the group settings is enabled. - -.. setting:: IMAGES_DIR - -IMAGES_DIR ----------- - -Default: ``None`` - -Directory where :class:`ImagesPipeline` will store its images. - -For more information see :ref:`topics-images`. - -.. setting:: IMAGES_EXPIRES - -IMAGES_EXPIRES --------------- - -Default = 90 - -Number of days for an image to be considered `expired` (downloaded again) in -:class:`ImagesPipeline`. - -For more information see :ref:`topics-images`. - -.. setting:: IMAGES_MIN_HEIGHT - -IMAGES_MIN_HEIGHT ------------------ - -Default = 0 - -Minimum height that an image is allowed to have in :class:`ImagesPipeline`. - -For more information see :ref:`topics-images-size`. - -.. setting:: IMAGES_MIN_WIDTH - -IMAGES_MIN_WIDTH ------------------ - -Default = 0 - -Minimum width that an image is allowed to have in :class:`ImagesPipeline`. - -For more information see :ref:`topics-images-size`. - -.. setting:: ITEM_PIPELINES - -ITEM_PIPELINES --------------- - -Default: ``[]`` - -The item pipelines to use (a list of classes). - -Example:: - - ITEM_PIPELINES = [ - 'mybot.pipeline.validate.ValidateMyItem', - 'mybot.pipeline.validate.StoreMyItem' - ] - -.. setting:: LOG_ENABLED - -LOG_ENABLED ------------ - -Default: ``True`` - -Enable logging. - -.. setting:: LOG_STDOUT - -LOG_STDOUT ----------- - -Default: ``False`` - -If enabled logging will be sent to standard output, otherwise standard error -will be used. - -.. setting:: LOGFILE - -LOGFILE -------- - -Default: ``None`` - -File name to use for logging output. If None, standard input (or error) will be -used depending on the value of the LOG_STDOUT setting. - -.. setting:: LOGLEVEL - -LOGLEVEL --------- - -Default: ``'DEBUG'`` - -Minimum level to log. Available levels are: SILENT, CRITICAL, ERROR, WARNING, -INFO, DEBUG, TRACE - -.. setting:: MAIL_FROM - -MAIL_FROM ---------- - -Default: ``'scrapy@localhost'`` - -Email to use as sender address for sending emails using the :ref:`Scrapy e-mail -sending facility `. - -.. setting:: MAIL_HOST - -MAIL_HOST ---------- - -Default: ``'localhost'`` - -Host to use for sending emails using the :ref:`Scrapy e-mail sending facility -`. - -.. setting:: MEMDEBUG_ENABLED - -MEMDEBUG_ENABLED ----------------- - -Default: ``False`` - -Whether to enable memory debugging. - -.. setting:: MEMDEBUG_NOTIFY - -MEMDEBUG_NOTIFY ---------------- - -Default: ``[]`` - -When memory debugging is enabled a memory report will be sent to the specified -addresses if this setting is not empty, otherwise the report will be written to -the log. - -Example:: - - MEMDEBUG_NOTIFY = ['user@example.com'] - -.. setting:: MEMUSAGE_ENABLED - -MEMUSAGE_ENABLED ----------------- - -Default: ``False`` - -Scope: ``scrapy.contrib.memusage`` - -Whether to enable the memory usage extension that will shutdown the Scrapy -process when it exceeds a memory limit, and also notify by email when that -happened. - -See :ref:`ref-extensions-memusage`. - -.. setting:: MEMUSAGE_LIMIT_MB - -MEMUSAGE_LIMIT_MB ------------------ - -Default: ``0`` - -Scope: ``scrapy.contrib.memusage`` - -The maximum amount of memory to allow (in megabytes) before shutting down -Scrapy (if MEMUSAGE_ENABLED is True). If zero, no check will be performed. - -See :ref:`ref-extensions-memusage`. - -.. setting:: MEMUSAGE_NOTIFY_MAIL - -MEMUSAGE_NOTIFY_MAIL --------------------- - -Default: ``False`` - -Scope: ``scrapy.contrib.memusage`` - -A list of emails to notify if the memory limit has been reached. - -Example:: - - MEMUSAGE_NOTIFY_MAIL = ['user@example.com'] - -See :ref:`ref-extensions-memusage`. - -.. setting:: MEMUSAGE_REPORT - -MEMUSAGE_REPORT ---------------- - -Default: ``False`` - -Scope: ``scrapy.contrib.memusage`` - -Whether to send a memory usage report after each domain has been closed. - -See :ref:`ref-extensions-memusage`. - -.. setting:: MEMUSAGE_WARNING_MB - -MEMUSAGE_WARNING_MB -------------------- - -Default: ``0`` - -Scope: ``scrapy.contrib.memusage`` - -The maximum amount of memory to allow (in megabytes) before sending a warning -email notifying about it. If zero, no warning will be produced. - -.. setting:: MYSQL_CONNECTION_SETTINGS - -MYSQL_CONNECTION_SETTINGS -------------------------- - -Default: ``{}`` - -Scope: ``scrapy.utils.db.mysql_connect`` - -Settings to use for MySQL connections performed through -``scrapy.utils.db.mysql_connect`` - -.. setting:: NEWSPIDER_MODULE - -NEWSPIDER_MODULE ----------------- - -Default: ``''`` - -Module where to create new spiders using the ``genspider`` command. - -Example:: - - NEWSPIDER_MODULE = 'mybot.spiders_dev' - -.. setting:: PROJECT_NAME - -PROJECT_NAME ------------- - -Default: ``Not Defined`` - -The name of the current project. It matches the project module name as created -by ``startproject`` command, and is only defined by project settings file. - -.. setting:: REDIRECT_MAX_TIMES - -REDIRECT_MAX_TIMES ------------------- - -Default: ``20`` - -Defines the maximun times a request can be redirected. After this maximun the -request's response is returned as is. We used Firefox default value for the -same task. - -.. setting:: REDIRECT_MAX_METAREFRESH_DELAY - -REDIRECT_MAX_METAREFRESH_DELAY ------------------------------- - -Default: ``100`` - -Some sites use meta-refresh for redirecting to a session expired page, so we -restrict automatic redirection to a maximum delay (in seconds) - -.. setting:: REDIRECT_PRIORITY_ADJUST - -REDIRECT_PRIORITY_ADJUST ------------------------------- - -Default: ``+2`` - -Adjust redirect request priority relative to original request. -A negative priority adjust means more priority. - -.. setting:: REQUESTS_PER_DOMAIN - -REQUESTS_PER_DOMAIN -------------------- - -Default: ``8`` - -Specifies how many concurrent (ie. simultaneous) requests will be performed per -open spider. - -.. setting:: REQUESTS_QUEUE_SIZE - -REQUESTS_QUEUE_SIZE -------------------- - -Default: ``0`` - -Scope: ``scrapy.contrib.spidermiddleware.limit`` - -If non zero, it will be used as an upper limit for the amount of requests that -can be scheduled per domain. - -.. setting:: ROBOTSTXT_OBEY - -ROBOTSTXT_OBEY --------------- - -Default: ``False`` - -Scope: ``scrapy.contrib.downloadermiddleware.robotstxt`` - -If enabled, Scrapy will respect robots.txt policies. For more information see -:topic:`robotstxt` - -.. setting:: SCHEDULER - -SCHEDULER ---------- - -Default: ``'scrapy.core.scheduler.Scheduler'`` - -The scheduler to use for crawling. - -.. setting:: SCHEDULER_ORDER - -SCHEDULER_ORDER ---------------- - -Default: ``'BFO'`` - -Scope: ``scrapy.core.scheduler`` - -The order to use for the crawling scheduler. Available orders are: - -* ``'BFO'``: `Breadth-first order`_ - typically consumes more memory but - reaches most relevant pages earlier. - -* ``'DFO'``: `Depth-first order`_ - typically consumes less memory than - but takes longer to reach most relevant pages. - -.. _Breadth-first order: http://en.wikipedia.org/wiki/Breadth-first_search -.. _Depth-first order: http://en.wikipedia.org/wiki/Depth-first_search - -.. setting:: SCHEDULER_MIDDLEWARES - -SCHEDULER_MIDDLEWARES ---------------------- - -Default:: ``{}`` - -A dict containing the scheduler middlewares enabled in your project, and their -orders. - -.. setting:: SCHEDULER_MIDDLEWARES_BASE - -SCHEDULER_MIDDLEWARES_BASE --------------------------- - -Default:: - - SCHEDULER_MIDDLEWARES_BASE = { - 'scrapy.contrib.schedulermiddleware.duplicatesfilter.DuplicatesFilterMiddleware': 500, - } - -A dict containing the scheduler middlewares enabled by default in Scrapy. You -should never modify this setting in your project, modify -:setting:`SCHEDULER_MIDDLEWARES` instead. - -.. setting:: SPIDERPROFILER_ENABLED - -SPIDERPROFILER_ENABLED ----------------------- - -Default: ``False`` - -Enable the spider profiler. Warning: this could have a big impact in -performance. - -.. setting:: SPIDER_MIDDLEWARES - -SPIDER_MIDDLEWARES ------------------- - -Default:: ``{}`` - -A dict containing the spider middlewares enabled in your project, and their -orders. For more info see :ref:`topics-spider-middleware-setting`. - -.. setting:: SPIDER_MIDDLEWARES_BASE - -SPIDER_MIDDLEWARES_BASE ------------------------ - -Default:: - - { - 'scrapy.contrib.spidermiddleware.httperror.HttpErrorMiddleware': 50, - 'scrapy.contrib.itemsampler.ItemSamplerMiddleware': 100, - 'scrapy.contrib.spidermiddleware.requestlimit.RequestLimitMiddleware': 200, - 'scrapy.contrib.spidermiddleware.restrict.RestrictMiddleware': 300, - 'scrapy.contrib.spidermiddleware.offsite.OffsiteMiddleware': 500, - 'scrapy.contrib.spidermiddleware.referer.RefererMiddleware': 700, - 'scrapy.contrib.spidermiddleware.urllength.UrlLengthMiddleware': 800, - 'scrapy.contrib.spidermiddleware.depth.DepthMiddleware': 900, - } - -A dict containing the spider middlewares enabled by default in Scrapy. You -should never modify this setting in your project, modify -:setting:`SPIDER_MIDDLEWARES` instead. For more info see -:ref:`topics-spider-middleware-setting`. - -.. setting:: SPIDER_MODULES - -SPIDER_MODULES --------------- - -Default: ``[]`` - -A list of modules where Scrapy will look for spiders. - -Example:: - - SPIDER_MODULES = ['mybot.spiders_prod', 'mybot.spiders_dev'] - -.. setting:: STATS_CLASS - -STATS_CLASS ------------ - -Default: ``'scrapy.stats.collector.MemoryStatsCollector'`` - -The class to use for collecting stats (must implement the Stats Collector API, -or subclass the StatsCollector class). - -.. setting:: STATS_DUMP - -STATS_DUMP ----------- - -Default: ``False`` - -Dump (to log) domain-specific stats collected when a domain is closed, and all -global stats when the Scrapy process finishes (ie. when the engine is -shutdown). - -.. setting:: STATS_ENABLED - -STATS_ENABLED -------------- - -Default: ``True`` - -Enable stats collection. - -.. setting:: STATSMAILER_RCPTS - -STATSMAILER_RCPTS ------------------ - -Default: ``[]`` (empty list) - -Send Scrapy stats after domains finish scrapy. See -:class:`~scrapy.contrib.statsmailer.StatsMailer` for more info. - -.. setting:: TELNETCONSOLE_ENABLED - -TELNETCONSOLE_ENABLED ---------------------- - -Default: ``True`` - -Scope: ``scrapy.management.telnet`` - -A boolean which specifies if the telnet management console will be enabled -(provided its extension is also enabled). - -.. setting:: TELNETCONSOLE_PORT - -TELNETCONSOLE_PORT ------------------- - -Default: ``6023`` - -The port to use for the telnet console. If set to ``None`` or ``0``, a -dynamically assigned port is used. For more info see -:ref:`topics-telnetconsole`. - -.. setting:: TEMPLATES_DIR - -TEMPLATES_DIR -------------- - -Default: ``templates`` dir inside scrapy module - -The directory where to look for template when creating new projects with -scrapy-admin.py newproject. - -.. setting:: URLLENGTH_LIMIT - -URLLENGTH_LIMIT ---------------- - -Default: ``2083`` - -Scope: ``contrib.spidermiddleware.urllength`` - -The maximum URL length to allow for crawled URLs. For more information about -the default value for this setting see: http://www.boutell.com/newfaq/misc/urllength.html - -.. setting:: USER_AGENT - -USER_AGENT ----------- - -Default: ``"%s/%s" % (BOT_NAME, BOT_VERSION)`` - -The default User-Agent to use when crawling, unless overrided. - -.. setting:: WEBCONSOLE_ENABLED - -WEBCONSOLE_ENABLED ------------------- - -Default: True - -A boolean which specifies if the web management console will be enabled -(provided its extension is also enabled). - -.. setting:: WEBCONSOLE_LOGFILE - -WEBCONSOLE_LOGFILE ------------------- - -Default: ``None`` - -A file to use for logging HTTP requests made to the web console. If unset web -the log is sent to standard scrapy log. - -.. setting:: WEBCONSOLE_PORT - -WEBCONSOLE_PORT ---------------- - -Default: ``6080`` - -The port to use for the web console. If set to ``None`` or ``0``, a dynamically -assigned port is used. For more info see :ref:`topics-webconsole`. - diff --git a/docs/ref/spider-middleware.rst b/docs/ref/spider-middleware.rst deleted file mode 100644 index ca90a28fb..000000000 --- a/docs/ref/spider-middleware.rst +++ /dev/null @@ -1,117 +0,0 @@ -.. _ref-spider-middleware: - -==================================== -Built-in spider middleware reference -==================================== - -This page describes all spider middleware components that come with Scrapy. For -information on how to use them and how to write your own spider middleware, see -the :ref:`spider middleware usage guide `. - -For a list of the components enabled by default (and their orders) see the -:setting:`SPIDER_MIDDLEWARES_BASE` setting. - -Available spider middlewares -============================ - -DepthMiddleware ---------------- - -.. module:: scrapy.contrib.spidermiddleware.depth - -.. class:: DepthMiddleware - - DepthMiddleware is a scrape middleware used for tracking the depth of each - Request inside the site being scraped. It can be used to limit the maximum - depth to scrape or things like that. - - The :class:`DepthMiddleware` can be configured through the following - settings (see the settings documentation for more info): - - * :setting:`DEPTH_LIMIT` - The maximum depth that will be allowed to - crawl for any site. If zero, no limit will be imposed. - * :setting:`DEPTH_STATS` - Whether to collect depth stats. - -HttpErrorMiddleware -------------------- - -.. module:: scrapy.contrib.spidermiddleware.httperror - -.. class:: HttpErrorMiddleware - - Filter out response outside of a range of valid status codes. - - This middleware filters out every response with status outside of the range - 200<=status<300. Spiders can add more exceptions using - ``handle_httpstatus_list`` spider attribute. - -OffsiteMiddleware ------------------ - -.. module:: scrapy.contrib.spidermiddleware.offsite - -.. class:: OffsiteMiddleware - - Filters out Requests for URLs outside the domains covered by the spider. - - This middleware filters out every request whose host names doesn't match - :attr:`~scrapy.spider.BaseSpider.domain_name`, or the spider - :attr:`~scrapy.spider.BaseSpider.domain_name` prefixed by "www.". - Spider can add more domains to exclude using - :attr:`~scrapy.spider.BaseSpider.extra_domain_names` attribute. - -RequestLimitMiddleware ----------------------- - -.. module:: scrapy.contrib.spidermiddleware.requestlimit - -.. class:: RequestLimitMiddleware - - Limits the maximum number of requests in the scheduler for each spider. When - a spider tries to schedule more than the allowed amount of requests, the new - requests (returned by the spider) will be dropped. - - The :class:`RequestLimitMiddleware` can be configured through the following - settings (see the settings documentation for more info): - - * :setting:`REQUESTS_QUEUE_SIZE` - If non zero, it will be used as an - upper limit for the amount of requests that can be scheduled per - domain. Can be set per spider using ``requests_queue_size`` attribute. - -RestrictMiddleware ------------------- - -.. module:: scrapy.contrib.spidermiddleware.restrict - -.. class:: RestrictMiddleware - - Restricts crawling to fixed set of particular URLs. - - The :class:`RestrictMiddleware` can be configured through the following - settings (see the settings documentation for more info): - - * :setting:`RESTRICT_TO_URLS` - Set of URLs allowed to crawl. - -UrlFilterMiddleware -------------------- - -.. module:: scrapy.contrib.spidermiddleware.urlfilter - -.. class:: UrlFilterMiddleware - - Canonicalizes URLs to filter out duplicated ones - -UrlLengthMiddleware -------------------- - -.. module:: scrapy.contrib.spidermiddleware.urllength - -.. class:: UrlLengthMiddleware - - Filters out requests with URLs longer than URLLENGTH_LIMIT - - The :class:`UrlLengthMiddleware` can be configured through the following - settings (see the settings documentation for more info): - - * :setting:`URLLENGTH_LIMIT` - The maximum URL length to allow for crawled URLs. - diff --git a/docs/ref/spiders.rst b/docs/ref/spiders.rst deleted file mode 100644 index 2011562ed..000000000 --- a/docs/ref/spiders.rst +++ /dev/null @@ -1,377 +0,0 @@ -.. _ref-spiders: - -========================= -Available Generic Spiders -========================= - -.. module:: scrapy.spider - :synopsis: Spiders base class, spider manager and spider middleware - -BaseSpider -========== - -.. class:: BaseSpider() - -This is the simplest spider, and the one from which every other spider -must inherit from (either the ones that come bundled with Scrapy, or the ones -that you write yourself). It doesn't provide any special functionality. It just -requests the given ``start_urls``/``start_requests``, and calls the spider's -method ``parse`` for each of the resulting responses. - -.. attribute:: BaseSpider.domain_name - - A string which defines the domain name for this spider, which will also be - the unique identifier for this spider (which means you can't have two - spider with the same ``domain_name``). This is the most important spider - attribute and it's required, and it's the name by which Scrapy will known - the spider. - -.. attribute:: BaseSpider.extra_domain_names - - An optional list of strings containing additional domains that this spider - is allowed to crawl. Requests for URLs not belonging to the domain name - specified in :attr:`Spider.domain_name` or this list won't be followed. - -.. attribute:: BaseSpider.start_urls - - Is a list of URLs where the spider will begin to crawl from, when no - particular URLs are specified. So, the first pages downloaded will be those - listed here. The subsequent URLs will be generated successively from data - contained in the start URLs. - -.. method:: BaseSpider.start_requests() - - This method must return an iterable with the first Requests to crawl for - this spider. - - This is the method called by Scrapy when the spider is opened for scraping - when no particular URLs are specified. If particular URLs are specified, - the :meth:`BaseSpider.make_requests_from_url` is used instead to create the - Requests. This method is also called only once from Scrapy, so it's safe to - implement it as a generator. - - The default implementation uses :meth:`BaseSpider.make_requests_from_url` - to generate Requests for each url in :attr:`start_urls`. - - If you want to change the Requests used to start scraping a domain, this is - the method to override. For example, if you need to start by login in using - a POST request, you could do:: - - def start_requests(self): - return [FormRequest("http://www.example.com/login", - formdata={'user': 'john', 'pass': 'secret'}, - callback=self.logged_in)] - - def logged_in(self, response): - # here you would extract links to follow and return Requests for - # each of them, with another callback - pass - -.. method:: BaseSpider.make_requests_from_url(url) - - A method that receives a URL and returns a :class:`~scrapy.http.Request` - object (or a list of :class:`~scrapy.http.Request` objects) to scrape. This - method is used to construct the initial requests in the - :meth:`start_requests` method, and is typically used to convert urls to - requests. - - Unless overridden, this method returns Requests with the :meth:`parse` - method as their callback function, and with dont_filter parameter enabled - (see :class:`~scrapy.http.Request` class for more info). - -.. method:: BaseSpider.parse(response) - - This is the default callback used by the :meth:`start_requests` method, and - will be used to parse the first pages crawled by the spider. - - The ``parse`` method is in charge of processing the response and returning - scraped data and/or more URLs to follow, because of this, the method must - always return a list or at least an empty one. Other Requests callbacks - have the same requirements as the BaseSpider class. - -BaseSpider example ------------------- - -Let's see an example:: - - from scrapy import log # This module is useful for printing out debug information - from scrapy.spider import BaseSpider - - class MySpider(BaseSpider): - domain_name = 'http://www.example.com' - start_urls = [ - 'http://www.example.com/1.html', - 'http://www.example.com/2.html', - 'http://www.example.com/3.html', - ] - - def parse(self, response): - self.log('A response from %s just arrived!' % response.url) - return [] - - SPIDER = MySpider() - -.. module:: scrapy.contrib.spiders - :synopsis: Collection of generic spiders - -CrawlSpider -=========== - -.. class:: CrawlSpider - -This is the most commonly used spider for crawling regular websites, as it -provides a convenient mechanism for following links by defining a set of rules. -It may not be the best suited for your particular web sites or project, but -it's generic enough for several cases, so you can start from it and override it -as need more custom functionality, or just implement your own spider. - -Apart from the attributes inherited from BaseSpider (that you must -specify), this class supports a new attribute: - -.. attribute:: CrawlSpider.rules - - Which is a list of one (or more) :class:`Rule` objects. Each :class:`Rule` - defines a certain behaviour for crawling the site. Rules objects are - described below . - -Crawling rules --------------- - -.. class:: Rule(link_extractor, callback=None, cb_kwargs=None, follow=None, process_links=None) - -``link_extractor`` is a :ref:`Link Extractor ` object which -defines how links will be extracted from each crawled page. - -``callback`` is a callable or a string (in which case a method from the spider -object with that name will be used) to be called for each link extracted with -the specified link_extractor. This callback receives a response as its first -argument and must return a list containing either ScrapedItems and Requests (or -any subclass of them). - -``cb_kwargs`` is a dict containing the keyword arguments to be passed to the -callback function - -``follow`` is a boolean which specified if links should be followed from each -response extracted with this rule. If ``callback`` is None ``follow`` defaults -to ``True``, otherwise it default to ``False``. - -``process_links`` is a callable, or a string (in which case a method from the -spider object with that name will be used) which will be called for each list -of links extracted from each response using the specified ``link_extractor``. -This is mainly used for filtering purposes. - - -CrawlSpider example -------------------- - -Let's now take a look at an example CrawlSpider with rules:: - - from scrapy.contrib.spiders import CrawlSpider, Rule - from scrapy.contrib.linkextractors.sgml import SgmlLinkExtractor - from scrapy.xpath.selector import HtmlXPathSelector - from scrapy.item import ScrapedItem - - class MySpider(CrawlSpider): - domain_name = 'example.com' - start_urls = ['http://www.example.com'] - - rules = ( - # Extract links matching 'category.php' (but not matching 'subsection.php') - # and follow links from them (since no callback means follow=True by default). - Rule(SgmlLinkExtractor(allow=('category\.php', ), deny=('subsection\.php', ))), - - # Extract links matching 'item.php' and parse them with the spider's method parse_item - Rule(SgmlLinkExtractor(allow=('item\.php', )), callback='parse_item'), - ) - - def parse_item(self, response): - self.log('Hi, this is an item page! %s' % response.url) - - hxs = HtmlXPathSelector(response) - item = ScrapedItem() - item.id = hxs.select('//td[@id="item_id"]/text()').re(r'ID: (\d+)') - item.name = hxs.select('//td[@id="item_name"]/text()').extract() - item.description = hxs.select('//td[@id="item_description"]/text()').extract() - return [item] - - SPIDER = MySpider() - - -This spider would start crawling example.com's home page, collecting category -links, and item links, parsing the latter with the -:meth:`XMLFeedSpider.parse_item` method. For each item response, some data will -be extracted from the HTML using XPath, and a ScrapedItem will be filled with -it. - -XMLFeedSpider -============= - -.. class:: XMLFeedSpider - - XMLFeedSpider is designed for parsing XML feeds by iterating through them by a - certain node name. The iterator can be chosen from: ``iternodes``, ``xml``, - and ``html``. It's recommended to use the ``iternodes`` iterator for - performance reasons, since the ``xml`` and ``html`` iterators generate the - whole DOM at once in order to parse it. However, using ``html`` as the - iterator may be useful when parsing XML with bad markup. - - For setting the iterator and the tag name, you must define the following class - attributes: - - .. attribute:: iterator - - A string which defines the iterator to use. It can be either: - - - ``'iternodes'`` - a fast iterator based on regular expressions - - - ``'html'`` - an iterator which uses HtmlXPathSelector. Keep in mind - this uses DOM parsing and must load all DOM in memory which could be a - problem for big feeds - - - ``'xml'`` - an iterator which uses XmlXPathSelector. Keep in mind - this uses DOM parsing and must load all DOM in memory which could be a - problem for big feeds - - It defaults to: ``'iternodes'``. - - .. attribute:: itertag - - A string with the name of the node (or element) to iterate in. Example:: - - itertag = 'product' - - .. attribute:: namespaces - - A list of ``(prefix, uri)`` tuples which define the namespaces - available in that document that will be processed with this spider. The - ``prefix`` and ``uri`` will be used to automatically register - namespaces using the - :meth:`~scrapy.xpath.XPathSelector.register_namespace` method. - - You can then specify nodes with namespaces in the :attr:`itertag` - attribute. - - Example:: - - class YourSpider(XMLFeedSpider): - - namespaces = [('n', 'http://www.sitemaps.org/schemas/sitemap/0.9')] - itertag = 'n:url' - # ... - - Apart from these new attributes, this spider has the following overrideable - methods too: - - .. method:: adapt_response(response) - - A method that receives the response as soon as it arrives from the spider - middleware and before start parsing it. It can be used used for modifying - the response body before parsing it. This method receives a response and - returns response (it could be the same or another one). - - .. method:: parse_item(response, selector) - - This method is called for the nodes matching the provided tag name - (``itertag``). Receives the response and an XPathSelector for each node. - Overriding this method is mandatory. Otherwise, you spider won't work. - This method must return either a ScrapedItem, a Request, or a list - containing any of them. - - .. warning:: This method will soon change its name to ``parse_node`` - - .. method:: process_results(response, results) - - This method is called for each result (item or request) returned by the - spider, and it's intended to perform any last time processing required - before returning the results to the framework core, for example setting the - item IDs. It receives a list of results and the response which originated - that results. It must return a list of results (Items or Requests).""" - - -XMLFeedSpider example ---------------------- - -These spiders are pretty easy to use, let's have at one example:: - - from scrapy import log - from scrapy.contrib.spiders import XMLFeedSpider - from scrapy.item import ScrapedItem - - class MySpider(XMLFeedSpider): - domain_name = 'example.com' - start_urls = ['http://www.example.com/feed.xml'] - iterator = 'iternodes' # This is actually unnecesary, since it's the default value - itertag = 'item' - - def parse_item(self, response, node): - log.msg('Hi, this is a <%s> node!: %s' % (self.itertag, ''.join(node.extract()))) - - item = ScrapedItem() - item.id = node.select('@id').extract() - item.name = node.select('name').extract() - item.description = node.select('description').extract() - return item - - SPIDER = MySpider() - -Basically what we did up there was creating a spider that downloads a feed from -the given ``start_urls``, and then iterates through each of its ``item`` tags, -prints them out, and stores some random data in ScrapedItems. - -CSVFeedSpider -============= - -.. class:: CSVFeedSpider - -.. warning:: The API of the CSVFeedSpider is not yet stable. Use with caution. - -This spider is very similar to the XMLFeedSpider, although it iterates through -rows, instead of nodes. It also has other two different attributes: - -.. attribute:: CSVFeedSpider.delimiter - - A string with the separator character for each field in the CSV file - Defaults to ``','`` (comma). - -.. attribute:: CSVFeedSpider.headers - - A list of the rows contained in the file CSV feed which will be used for - extracting fields from it. - -In this spider, the method that gets called in each row iteration ``parse_row`` -instead of ``parse_item`` (like in :class:`XMLFeedSpider`). - -.. method:: CSVFeedSpider.parse_row(response, row) - - Receives a response and a dict (representing each row) with a key for each - provided (or detected) header of the CSV file. This spider also gives the - opportunity to override ``adapt_response`` and ``process_results`` methods - for pre and post-processing purposes. - -CSVFeedSpider example ---------------------- - -Let's see an example similar to the previous one, but using CSVFeedSpider:: - - from scrapy import log - from scrapy.contrib.spiders import CSVFeedSpider - from scrapy.item import ScrapedItem - - class MySpider(CSVFeedSpider): - domain_name = 'example.com' - start_urls = ['http://www.example.com/feed.csv'] - delimiter = ';' - headers = ['id', 'name', 'description'] - - def parse_row(self, response, row): - log.msg('Hi, this is a row!: %r' % row) - - item = ScrapedItem() - item.id = row['id'] - item.name = row['name'] - item.description = row['description'] - return item - - SPIDER = MySpider() - - diff --git a/docs/reference.rst b/docs/reference.rst new file mode 100644 index 000000000..16897148e --- /dev/null +++ b/docs/reference.rst @@ -0,0 +1,22 @@ +.. _ref: + +API Reference +============= + +This section documents the Scrapy |version| API. For more information see :ref:`misc-api-stability`. + +* :ref:`topics-request-response` +* :ref:`topics-spiders-ref` +* :ref:`topics-selectors-ref` +* :ref:`topics-settings-ref` +* :ref:`topics-signals-ref` +* :ref:`topics-exceptions-ref` +* :ref:`topics-logging` +* :ref:`topics-email` +* :ref:`topics-extensions-ref` +* :ref:`topics-downloader-middleware-ref` +* :ref:`topics-spider-middleware-ref` +* :ref:`topics-scheduler-middleware-ref` +* :ref:`topics-link-extractors-ref` +* :ref:`topics-stats-ref` + diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 22a4f9e89..9ea0a059f 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -128,3 +128,71 @@ scope, and the original request won't finish until redirected request is completed. This stop ``process_download_exception()`` middleware as returning Response would do. + +.. _topics-downloader-middleware-ref: + +Built-in downloader middleware reference +======================================== + +This page describes all downloader middleware components that come with +Scrapy. For information on how to use them and how to write your own downloader +middleware, see the :ref:`downloader middleware usage guide +`. + +For a list of the components enabled by default (and their orders) see the +:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting. + +DefaultHeadersMiddleware +------------------------ + +.. module:: scrapy.contrib.downloadermiddleware.defaultheaders + :synopsis: Default Headers Downloader Middleware + +.. class:: DefaultHeadersMiddleware + + This middleware sets all default requests headers specified in the + :setting:`DEFAULT_REQUEST_HEADERS` setting. + +DebugMiddleware +--------------- + +.. module:: scrapy.contrib.downloadermiddleware.debug + :synopsis: Downloader middlewares for debugging + +.. class:: DebugMiddleware + + This is a convenient middleware to inspect what's passing through the + downloader middleware. It logs all requests and responses catched by the + middleware component methods. This middleware does not use any settings and + does not come enabled by default. Instead, it's meant to be inserted at the + point of the middleware that you want to inspect. + +HttpCacheMiddleware +------------------- + +.. module:: scrapy.contrib.downloadermiddleware.httpcache + :synopsis: HTTP Cache downloader middleware + +.. class:: HttpCacheMiddleware + + This middleware provides low-level cache to all HTTP requests and responses. + Every request and its corresponding response are cached and then, when that + same request is seen again, the response is returned without transferring + anything from the Internet. + + The HTTP cache is useful for testing spiders faster (without having to wait for + downloads every time) and for trying your spider off-line when you don't have + an Internet connection. + + The :class:`HttpCacheMiddleware` can be configured through the following + settings (see the settings documentation for more info): + + * :setting:`HTTPCACHE_DIR` - this one actually enables the cache besides + settings the cache dir + * :setting:`HTTPCACHE_IGNORE_MISSING` - ignoring missing requests instead + of downloading them + * :setting:`HTTPCACHE_SECTORIZE` - split HTTP cache in several directories + (for performance reasons) + * :setting:`HTTPCACHE_EXPIRATION_SECS` - how many secs until the cache is + considered out of date + diff --git a/docs/ref/email.rst b/docs/topics/email.rst similarity index 99% rename from docs/ref/email.rst rename to docs/topics/email.rst index 76af6e1a1..7fc2f9d60 100644 --- a/docs/ref/email.rst +++ b/docs/topics/email.rst @@ -1,4 +1,4 @@ -.. _ref-email: +.. _topics-email: ============= Sending email diff --git a/docs/ref/exceptions.rst b/docs/topics/exceptions.rst similarity index 90% rename from docs/ref/exceptions.rst rename to docs/topics/exceptions.rst index 9a9288037..0fa6b524a 100644 --- a/docs/ref/exceptions.rst +++ b/docs/topics/exceptions.rst @@ -1,10 +1,16 @@ -.. _exceptions: +.. _topics-exceptions: + +========== +Exceptions +========== .. module:: scrapy.core.exceptions :synopsis: Core exceptions -Available Exceptions -==================== +.. _topics-exceptions-ref: + +Built-in Exceptions reference +============================= Here's a list of all exceptions included in Scrapy and their usage. diff --git a/docs/topics/extensions.rst b/docs/topics/extensions.rst index 3172887de..8a8ac01b6 100644 --- a/docs/topics/extensions.rst +++ b/docs/topics/extensions.rst @@ -65,21 +65,22 @@ Not all available extensions will be enabled. Some of them usually depend on a particular setting. For example, the HTTP Cache extension is available by default but disabled unless the :setting:`HTTPCACHE_DIR` setting is set. Both enabled and disabled extension can be accessed through the -:ref:`ref-extension-manager`. +:ref:`topics-extensions-ref-manager`. Accessing enabled extensions ============================ Even though it's not usually needed, you can access extension objects through -the :ref:`ref-extension-manager` which is populated when extensions are loaded. -For example, to access the ``WebConsole`` extension:: +the :ref:`topics-extensions-ref-manager` which is populated when extensions are +loaded. For example, to access the ``WebConsole`` extension:: from scrapy.extension import extensions webconsole_extension = extensions.enabled['WebConsole'] .. seealso:: - :ref:`ref-extension-manager`, for the complete Extension manager reference. + :ref:`topics-extensions-ref-manager`, for the complete Extension manager + reference. Writing your own extension ========================== @@ -110,8 +111,308 @@ everytime a domain/spider is opened and closed:: def domain_closed(self, domain, spider): log.msg("closed domain %s" % domain) -Built-in extensions -=================== -See :ref:`ref-extensions`. +.. _topics-extensions-ref: + +Built-in extensions reference +============================= + +.. _topics-extensions-ref-manager: + +Extension manager +----------------- + +.. module:: scrapy.extension + :synopsis: The extension manager + +The Extension Manager is responsible for loading and keeping track of installed +extensions and it's configured through the :setting:`EXTENSIONS` setting which +contains a dictionary of all available extensions and their order similar to +how you :ref:`configure the downloader middlewares +`. + +.. class:: ExtensionManager + + The extension manager is a singleton object, which is instantiated at module + loading time and can be accessed like this:: + + from scrapy.extension import extensions + + .. attribute:: loaded + + A boolean which is True if extensions are already loaded or False if + they're not. + + .. attribute:: enabled + + A dict with the enabled extensions. The keys are the extension class names, + and the values are the extension objects. Example:: + + >>> from scrapy.extension import extensions + >>> extensions.load() + >>> print extensions.enabled + {'CoreStats': , + 'WebConsoke': , + ... + + .. attribute:: disabled + + A dict with the disabled extensions. The keys are the extension class names, + and the values are the extension class paths (because objects are never + instantiated for disabled extensions). Example:: + + >>> from scrapy.extension import extensions + >>> extensions.load() + >>> print extensions.disabled + {'MemoryDebugger': 'scrapy.contrib.webconsole.stats.MemoryDebugger', + 'SpiderProfiler': 'scrapy.contrib.spider.profiler.SpiderProfiler', + ... + + .. method:: load() + + Load the available extensions configured in the :setting:`EXTENSIONS` + setting. On a standard run, this method is usually called by the Execution + Manager, but you may need to call it explicitly if you're dealing with + code outside Scrapy. + + .. method:: reload() + + Reload the available extensions. See :meth:`load`. + +General purpose extensions +-------------------------- + +Core Stats extension +~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.stats.corestats + :synopsis: Core stats collection + +.. class:: scrapy.stats.corestats.CoreStats + +Enable the collection of core statistics, provided the stats collection are +enabled (see :ref:`topics-stats`). + +.. _topics-extensions-ref-webconsole: + +Web console extension +~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.management.web + :synopsis: Web management console + +.. class:: scrapy.management.web.WebConsole + +Provides an extensible web server for managing a Scrapy process. It's enabled +by the :setting:`WEBCONSOLE_ENABLED` setting. The server will listen in the +port specified in :setting:`WEBCONSOLE_PORT`, and will log to the file +specified in :setting:`WEBCONSOLE_LOGFILE`. + +The web server is designed to be extended by other extensions which can add +their own management web interfaces. + +See also :ref:`topics-webconsole` for information on how to write your own web +console extension, and "Web console extensions" below for a list of available +built-in (web console) extensions. + +.. _topics-extensions-ref-telnetconsole: + +Telnet console extension +~~~~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.management.telnet + :synopsis: Telnet management console + +.. class:: scrapy.management.telnet.TelnetConsole + +Provides a telnet console for getting into a Python interpreter inside the +currently running Scrapy process, which can be very useful for debugging. + +The telnet console must be enabled by the :setting:`TELNETCONSOLE_ENABLED` +setting, and the server will listen in the port specified in +:setting:`WEBCONSOLE_PORT`. + +Spider reloader extension +~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.contrib.spider.reloader + :synopsis: Spider reloader extension + +.. class:: scrapy.contrib.spider.reloader.SpiderReloader + +Reload spider objects once they've finished scraping, to release the resources +and references to other objects they may hold. + +.. _topics-extensions-ref-memusage: + +Memory usage extension +~~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.contrib.memusage + :synopsis: Memory usage extension + +.. class:: scrapy.contrib.memusage.MemoryUsage + +Allows monitoring the memory used by a Scrapy process and: + +1, send a notification email when it exceeds a certain value +2. terminate the Scrapy process when it exceeds a certain value + +The notification emails can be triggered when a certain warning value is +reached (:setting:`MEMUSAGE_WARNING_MB`) and when the maximum value is reached +(:setting:`MEMUSAGE_LIMIT_MB`) which will also cause the Scrapy process to be +terminated. + +This extension is enabled by the :setting:`MEMUSAGE_ENABLED` setting and +can be configured with the following settings: + +* :setting:`MEMUSAGE_LIMIT_MB` +* :setting:`MEMUSAGE_WARNING_MB` +* :setting:`MEMUSAGE_NOTIFY_MAIL` +* :setting:`MEMUSAGE_REPORT` + +Memory debugger extension +~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.contrib.memdebug + :synopsis: Memory debugger extension + +.. class:: scrapy.contrib.memdebug.MemoryDebugger + +A memory debugger which collects some info about objects uncollected by the +garbage collector and libxml2 memory leaks. To enable this extension turn on +the :setting:`MEMDEBUG_ENABLED` setting. The report will be printed to standard +output. If the :setting:`MEMDEBUG_NOTIFY` setting contains a list of emails the +report will also be sent to those addresses. + +Close domain extension +~~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.contrib.closedomain + :synopsis: Close domain extension + +.. class:: scrapy.contrib.closedomain.CloseDomain + +Closes a domain/spider automatically when some conditions are met, using a +specific closing reason for each condition. + +The conditions for closing a domain can be configured through the following +settings. Other conditions will be supported in the future. + +.. setting:: CLOSEDOMAIN_TIMEOUT + +CLOSEDOMAIN_TIMEOUT +""""""""""""""""""" + +Default: ``0`` + +An integer which specifies a number of seconds. If the domain remains open for +more than that number of second, it will be automatically closed with the +reason ``closedomain_timeout``. If zero (or non set) domains won't be closed by +timeout. + +.. setting:: CLOSEDOMAIN_ITEMPASSED + +CLOSEDOMAIN_ITEMPASSED +"""""""""""""""""""""" + +Default: ``0`` + +An integer which specifies a number of items. If the spider scrapes more than +that amount if items and those items are passed by the item pipeline, the +domain will be closed with the reason ``closedomain_itempassed``. If zero (or +non set) domains won't be closed by number of passed items. + +Stack trace dump extension +~~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.contrib.debug + :synopsis: Extensions for debugging Scrapy + +.. class:: scrapy.contrib.debug.StackTraceDump + +Adds a `SIGUSR1`_ signal handler which dumps the stack trace of a runnning +Scrapy process when a ``SIGUSR1`` signal is catched. After the stack trace is +dumped, the Scrapy process continues to run normally. + +The stack trace is sent to standard output, or to the Scrapy log file if +:setting:`LOG_STDOUT` is enabled. + +This extension only works on POSIX-compliant platforms (ie. not Windows). + +.. _SIGUSR1: http://en.wikipedia.org/wiki/SIGUSR1_and_SIGUSR2 + +StatsMailer extension +~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.contrib.statsmailer + :synopsis: StatsMailer extension + +.. class:: scrapy.contrib.statsmailer.StatsMailer + +This simple extension can be used to send a notification email every time a +domain has finished scraping, including the Scrapy stats collected. The email +will be sent to all recipients specified in the :setting:`STATSMAILER_RCPTS` +setting. + +Web console extensions +---------------------- + +.. module:: scrapy.contrib.webconsole + :synopsis: Contains most built-in web console extensions + +Here is a list of built-in web console extensions. For clarity "web console +extension" is abbreviated as "WC extension". + +For more information see the see the :ref:`web console documentation +`. + +Scheduler queue WC extension +~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.contrib.webconsole.scheduler + :synopsis: Scheduler queue web console extension + +.. class:: scrapy.contrib.webconsole.scheduler.SchedulerQueue + +Display a list of all pending Requests in the Scheduler queue, grouped by +domain/spider. + +Spider live stats WC extension +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.contrib.webconsole.livestats + :synopsis: Spider live stats web console extension + +.. class:: scrapy.contrib.webconsole.livestats.LiveStats + +Display a table with stats of all spider crawled by the current Scrapy run, +including: + +* Number of items scraped +* Number of pages crawled +* Number of pending requests in the scheduler +* Number of pending requests in the downloader queue +* Number of requests currently being downloaded + +Engine status WC extension +~~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.contrib.webconsole.enginestatus + :synopsis: Engine status web console extension + +.. class:: scrapy.contrib.webconsole.enginestatus.EngineStatus + +Display the current status of the Scrapy Engine, which is just the output of +the Scrapy engine ``getstatus()`` method. + +Stats collector dump WC extension +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. module:: scrapy.contrib.webconsole.stats + :synopsis: Stats dump web console extension + +.. class:: scrapy.contrib.webconsole.stats.StatsDump + +Display the stats collected so far by the stats collector. diff --git a/docs/topics/index.rst b/docs/topics/index.rst index 0615aa748..49a115b96 100644 --- a/docs/topics/index.rst +++ b/docs/topics/index.rst @@ -25,3 +25,9 @@ This section describes all key concepts of Scrapy. robotstxt firefox firebug + signals + logging + scheduler-middleware + request-response + exceptions + email diff --git a/docs/topics/link-extractors.rst b/docs/topics/link-extractors.rst index 48d87f107..383568952 100644 --- a/docs/topics/link-extractors.rst +++ b/docs/topics/link-extractors.rst @@ -24,6 +24,124 @@ your spiders even if you don't subclass from :class:`~scrapy.contrib.spiders.CrawlSpider`, as its purpose is very simple: to extract links. -See :ref:`ref-link-extractors` for the list of available built-in Link -Extractors, including some examples. + +.. _topics-link-extractors-ref: + +Built-in link extractors reference +================================== + +.. module:: scrapy.contrib.linkextractors + :synopsis: Link extractors classes + +All available link extractors classes bundled with Scrapy are provided in the +:mod:`scrapy.contrib.linkextractors` module. + +.. module:: scrapy.contrib.linkextractors.sgml + :synopsis: SGMLParser-based link extractors + +SgmlLinkExtractor +----------------- + +.. class:: SgmlLinkExtractor(allow=(), deny=(), allow_domains=(), deny_domains=(), restrict_xpaths(), tags=('a', 'area'), attrs=('href'), canonicalize=True, unique=True, process_value=None) + + The SgmlLinkExtractor extends the base :class:`BaseSgmlLinkExtractor` by + providing additional filters that you can specify to extract links, + including regular expressions patterns that the links must match to be + extracted. All those filters are configured through these constructor + parameters: + + :param allow: a single regular expression (or list of regular expressions) + that the (absolute) urls must match in order to be extracted. If not + given (or empty), it will match all links. + :type allow: a regular expression (or list of) + + :param deny: a single regular expression (or list of regular expressions) + that the (absolute) urls must match in order to be excluded (ie. not + extracted). It has precedence over the ``allow`` parameter. If not + given (or empty) it won't exclude any links. + :type allow: a regular expression (or list of) + + :param allow_domains: is single value or a list of string containing + domains which will be considered for extracting the links + :type allow: str or list + + :param deny_domains: is single value or a list of strings containing + domains which which won't be considered for extracting the links + :type allow: str or list + + :param restrict_xpaths: is a XPath (or list of XPath's) which defines + regions inside the response where links should be extracted from. + If given, only the text selected by those XPath will be scanned for + links. See examples below. + :type restrict_xpaths: str or list + + :param tags: a tag or a list of tags to consider when extracting links. + Defaults to ``('a', 'area')``. + :type tags: str or list + + :param attrs: list of attrbitues which should be considered when looking + for links to extract (only for those tags specified in the ``tags`` + parameter). Defaults to ``('href',)`` + :type attrs: boolean + + :param canonicalize: canonicalize each extracted url (using + scrapy.utils.url.canonicalize_url). Defaults to ``True``. + :type canonicalize: boolean + + :param unique: whether duplicate filtering should be applied to extracted + links. + :type unique: boolean + + :param process_value: see ``process_value`` argument of + :class:`LinkExtractor` class constructor + :type process_value: boolean + +BaseSgmlLinkExtractor +--------------------- + +.. class:: BaseSgmlLinkExtractor(tag="a", href="href", unique=False, process_value=None) + + The purpose of this Link Extractor is only to serve as a base class for the + :class:`SgmlLinkExtractor`. You should use that one instead. + + The constructor arguments are: + + :param tag: either a string (with the name of a tag) or a function that + receives a tag name and returns ``True`` if links should be extracted from + those tag, or ``False`` if they shouldn't. Defaults to ``'a'``. request + (once its downloaded) as its first parameter. For more information see + :ref:`topics-request-response-ref-request-callback-arguments`. + :type tag: str or callable + + :param attr: either string (with the name of a tag attribute), or a + function that receives a an attribute name and returns ``True`` if + links should be extracted from it, or ``False`` if the shouldn't. + Defaults to ``href``. + :type attr: str or callable + + :param unique: is a boolean that specifies if a duplicate filtering should + be applied to links extracted. + :type unique: boolean + + :param process_value: a function which receives each value extracted from + the tag and attributes scanned and can modify the value and return a + new one, or return ``None`` to ignore the link altogether. If not + given, ``process_value`` defaults to ``lambda x: x``. + + .. highlight:: html + + For example, to extract links from this code:: + + Link text + + .. highlight:: python + + You can use the following function in ``process_value``:: + + def process_value(value): + m = re.search("javascript:goToPage\('(.*?)'", value) + if m: + return m.group(1) + + :type process_value: callable diff --git a/docs/ref/logging.rst b/docs/topics/logging.rst similarity index 99% rename from docs/ref/logging.rst rename to docs/topics/logging.rst index 734401d77..fe677f3c9 100644 --- a/docs/ref/logging.rst +++ b/docs/topics/logging.rst @@ -1,4 +1,4 @@ -.. _ref-logging: +.. _topics-logging: ======= Logging diff --git a/docs/ref/request-response.rst b/docs/topics/request-response.rst similarity index 93% rename from docs/ref/request-response.rst rename to docs/topics/request-response.rst index a94674c3b..49ab5a20f 100644 --- a/docs/ref/request-response.rst +++ b/docs/topics/request-response.rst @@ -1,15 +1,12 @@ -.. _ref-request-response: +.. _topics-request-response: -============================ -Request and Response objects -============================ +====================== +Requests and Responses +====================== .. module:: scrapy.http :synopsis: Request and Response classes -Quick overview -============== - Scrapy uses :class:`Request` and :class:`Response` objects for crawling web sites. @@ -20,7 +17,9 @@ issued the request. Both :class:`Request` and :class:`Response` classes have subclasses which adds additional functionality not required in the base classes. These are described -below in :ref:`ref-request-subclasses` and :ref:`ref-response-subclasses`. +below in :ref:`topics-request-response-ref-request-subclasses` and +:ref:`topics-request-response-ref-response-subclasses`. + Request objects =============== @@ -36,7 +35,7 @@ Request objects :param callback: the function that will be called with the response of this request (once its downloaded) as its first parameter. For more information - see :ref:`ref-request-callback-arguments` below. + see :ref:`topics-request-response-ref-request-callback-arguments` below. :type callback: callable :param method: the HTTP method of this request. Defaults to ``'GET'``. @@ -145,18 +144,19 @@ Request objects .. method:: Request.copy() Return a new Request which is a copy of this Request. The attribute - :attr:`Request.meta` is copied, while :attr:`Request.cache` is not. See also - :ref:`ref-request-callback-arguments`. + :attr:`Request.meta` is copied, while :attr:`Request.cache` is not. See + also :ref:`topics-request-response-ref-request-callback-arguments`. .. method:: Request.replace([url, callback, method, headers, body, cookies, meta, encoding, dont_filter]) Return a Request object with the same members, except for those members - given new values by whichever keyword arguments are specified. The attribute - :attr:`Request.meta` is copied by default (unless a new value is given - in the ``meta`` argument). The :attr:`Request.cache` attribute is always - cleared. See also :ref:`ref-request-callback-arguments`. + given new values by whichever keyword arguments are specified. The + attribute :attr:`Request.meta` is copied by default (unless a new value + is given in the ``meta`` argument). The :attr:`Request.cache` attribute + is always cleared. See also + :ref:`topics-request-response-ref-request-callback-arguments`. -.. _ref-request-callback-copy: +.. _topics-request-response-ref-callback-copy: Caveats with copying Requests and callbacks ------------------------------------------- @@ -178,7 +178,7 @@ In the above example, ``request2`` is a copy of ``request`` but it has no callback, while ``request3`` is a copy of ``request`` and also contains the callback. -.. _ref-request-callback-arguments: +.. _topics-request-response-ref-request-callback-arguments: Passing arguments to callback functions --------------------------------------- @@ -230,7 +230,7 @@ Using Request.meta:: referer_url = response.request.meta['referer_url'] self.log("Visited page %s from %s" % (response.url, referer_url)) -.. _ref-request-subclasses: +.. _topics-request-response-ref-request-subclasses: Request subclasses ================== @@ -266,7 +266,8 @@ objects. Returns a new :class:`FormRequest` object with its form field values pre-populated with those found in the HTML ``

    `` element contained - in the given response. For an example see :ref:`ref-request-userlogin`. + in the given response. For an example see + :ref:`topics-request-response-ref-request-userlogin`. :param response: the response containing a HTML form which will be used @@ -287,10 +288,10 @@ objects. Request usage examples -====================== +---------------------- Using FormRequest to send data via HTTP POST --------------------------------------------- +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ If you want to simulate a HTML Form POST in your spider, and send a couple of key-value fields you could return a :class:`FormRequest` object (from your @@ -300,10 +301,10 @@ spider) like this:: formdata={'name': 'John Doe', age: '27'}, callback=self.after_post)] -.. _ref-request-userlogin: +.. _topics-request-response-ref-request-userlogin: Using FormRequest.from_response() to simulate a user login ----------------------------------------------------------- +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ It is usual for web sites to provide pre-populated form fields through ```` elements, such as session related data or authentication @@ -349,8 +350,9 @@ Response objects :type status: integer :param body: the response body. It must be str, not unicode, unless you're - using a encoding-aware :ref:`Response subclass `, - such as :class:`TextResponse`. + using a encoding-aware :ref:`Response subclass + `, such as + :class:`TextResponse`. :type body: str :param meta: the initial values for the :attr:`Response.meta` attribute. If @@ -432,7 +434,7 @@ Response objects is given in the ``meta`` argument). The :attr:`Response.cache` attribute is always cleared. -.. _ref-response-subclasses: +.. _topics-request-response-ref-response-subclasses: Response subclasses =================== diff --git a/docs/ref/scheduler-middleware.rst b/docs/topics/scheduler-middleware.rst similarity index 84% rename from docs/ref/scheduler-middleware.rst rename to docs/topics/scheduler-middleware.rst index 2d9ed9e4d..f4b08faee 100644 --- a/docs/ref/scheduler-middleware.rst +++ b/docs/topics/scheduler-middleware.rst @@ -1,6 +1,12 @@ -.. _ref-scheduler-middleware: +.. _topics-scheduler-middleware: + +==================== +Scheduler middleware +==================== + + +.. _topics-scheduler-middleware-ref: -======================================== Built-in scheduler middleware reference ======================================== @@ -10,9 +16,6 @@ Scrapy. For a list of the components enabled by default (and their orders) see the :setting:`SCHEDULER_MIDDLEWARES_BASE` setting. -Available scheduler middlewares -=============================== - DuplicatesFilterMiddleware -------------------------- diff --git a/docs/topics/selectors.rst b/docs/topics/selectors.rst index 93211e912..5db60494d 100644 --- a/docs/topics/selectors.rst +++ b/docs/topics/selectors.rst @@ -4,9 +4,6 @@ XPath Selectors =============== -Introduction ------------- - When you're scraping web pages, the most common task you need to perform is to extract data from the HTML source. There are several libraries available to achieve this: @@ -36,7 +33,7 @@ small and simple, unlike the `lxml`_ API which is much bigger because the documents. For a complete reference of the selectors API see the :ref:`XPath selector -reference `. +reference `. .. _BeautifulSoup: http://www.crummy.com/software/BeautifulSoup/ .. _lxml: http://codespeak.net/lxml/ @@ -44,6 +41,9 @@ reference `. .. _libxml2: http://xmlsoft.org/ .. _XPath: http://www.w3.org/TR/xpath +Using selectors +=============== + Constructing selectors ---------------------- @@ -209,3 +209,186 @@ For more details about relative XPaths see the `Location Paths`_ section in the XPath specification. .. _Location Paths: http://www.w3.org/TR/xpath#location-paths + + +.. _topics-selectors-ref: + +Built-in XPath Selectors reference +================================== + +.. module:: scrapy.xpath + :synopsis: XPath selectors classes + +There are two types of selectors bundled with Scrapy: +:class:`HtmlXPathSelector` and :class:`XmlXPathSelector`. Both of them +implement the same :class:`XPathSelector` interface. The only different is that +one is used to process HTML data and the other XML data. + +XPathSelector objects +--------------------- + +.. class:: XPathSelector(response) + + A :class:`XPathSelector` object is a wrapper over response to select + certain parts of its content. + + A :class:`Request` object represents an HTTP request, which is usually + generated in the Spider and executed by the Downloader, and thus generating + a :class:`Response`. + + ``url`` is a :class:`~scrapy.http.Response` object that will be used for + selecting and extracting data + + .. method:: XPathSelector.select(xpath) + + Apply the given XPath relative to this XPathSelector and return a list + of :class:`XPathSelector` objects (ie. a :class:`XPathSelectorList`) with + the result. + + ``xpath`` is a string containing the XPath to apply + + .. method:: XPathSelector.re(regex) + + Apply the given regex and return a list of unicode strings with the + matches. + + ``regex`` can be either a compiled regular expression or a string which + will be compiled to a regular expression using ``re.compile(regex)`` + + .. method:: XPathSelector.extract() + + Return a unicode string with the content of this :class:`XPathSelector` + object. + + .. method:: XPathSelector.extract_unquoted() + + Return a unicode string with the content of this :class:`XPathSelector` + without entities or CDATA. This method is intended to be use for text-only + selectors, like ``//h1/text()`` (but not ``//h1``). If it's used for + :class:`XPathSelector` objects which don't select a textual content (ie. if + they contain tags), the output of this method is undefined. + + .. method:: XPathSelector.register_namespace(prefix, uri) + + Register the given namespace to be used in this :class:`XPathSelector`. + Without registering namespaces you can't select or extract data from + non-standard namespaces. See examples below. + + .. method:: XPathSelector.__nonzero__() + + Returns ``True`` if there is any real content selected by this + :class:`XPathSelector` or ``False`` otherwise. In other words, the boolean + value of an XPathSelector is given by the contents it selects. + +XPathSelectorList objects +------------------------- + +.. class:: XPathSelectorList + + The :class:`XPathSelectorList` class is subclass of the builtin ``list`` + class, which provides a few additional methods. + + + .. method:: XPathSelectorList.select(xpath) + + Call the :meth:`XPathSelector.re` method for all :class:`XPathSelector` + objects in this list and return their results flattened, as new + :class:`XPathSelectorList`. + + ``xpath`` is the same argument as the one in :meth:`XPathSelector.x` + + .. method:: XPathSelector.re(regex) + + Call the :meth:`XPathSelector.re` method for all :class:`XPathSelector` + objects in this list and return their results flattened, as a list of + unicode strings. + + ``regex`` is the same argument as the one in :meth:`XPathSelector.re` + + .. method:: XPathSelector.extract() + + Call the :meth:`XPathSelector.re` method for all :class:`XPathSelector` + objects in this list and return their results flattened, as a list of + unicode strings. + + .. method:: XPathSelector.extract_unquoted() + + Call the :meth:`XPathSelector.extract_unoquoted` method for all + :class:`XPathSelector` objects in this list and return their results + flattened, as a list of unicode strings. This method should not be applied + to all kinds of XPathSelectors. For more info see + :meth:`XPathSelector.extract_unoquoted`. + +HtmlXPathSelector objects +------------------------- + +.. class:: HtmlXPathSelector(response) + + A subclass of :class:`XPathSelector` for working with HTML content. It uses + the `libxml2`_ HTML parser. See the :class:`XPathSelector` API for more info. + +.. _libxml2: http://xmlsoft.org/ + +HtmlXPathSelector examples +~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Here's a couple of :class:`HtmlXPathSelector` examples to illustrate several +concepts. In all cases we assume there is already a :class:`HtmlPathSelector` +instanced with a :class:`~scrapy.http.Response` object like this:: + + x = HtmlXPathSelector(html_response) + +1. Select all ``

    `` elements from a HTML response body, returning a list of + :class:`XPathSelector` objects (ie. a :class:`XPathSelectorList` object):: + + x.select("//h1") + +2. Extract the text of all ``

    `` elements from a HTML response body, + returning a list of unicode strings:: + + x.select("//h1").extract() # this includes the h1 tag + x.select("//h1/text()").extract() # this excludes the h1 tag + +3. Iterate over all ``

    `` tags and print their class attribute:: + + for node in x.select("//p"): + ... print node.select("@href") + +4. Extract textual data from all ``

    `` tags without entities, as a list of + unicode strings:: + + x.select("//p/text()").extract_unquoted() + + # the following line is wrong. extract_unquoted() should only be used + # with textual XPathSelectors + x.select("//p").extract_unquoted() # it may work but output is unpredictable + +XmlXPathSelector objects +------------------------ + +.. class:: XmlXPathSelector(response) + + A subclass of :class:`XPathSelector` for working with XML content. It uses + the `libxml2`_ XML parser. See the :class:`XPathSelector` API for more info. + +XmlXPathSelector examples +~~~~~~~~~~~~~~~~~~~~~~~~~ + +Here's a couple of :class:`XmlXPathSelector` examples to illustrate several +concepts. In all cases we assume there is already a :class:`XmlPathSelector` +instanced with a :class:`~scrapy.http.Response` object like this:: + + x = HtmlXPathSelector(xml_response) + +1. Select all ```` elements from a XML response body, returning a list of + :class:`XPathSelector` objects (ie. a :class:`XPathSelectorList` object):: + + x.select("//h1") + +2. Extract all prices from a `Google Base XML feed`_ which requires registering + a namespace:: + + x.register_namespace("g", "http://base.google.com/ns/1.0") + x.select("//g:price").extract() + +.. _Google Base XML feed: http://base.google.com/support/bin/answer.py?hl=en&answer=59461 diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 4973c7ccb..02d2a868c 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -14,7 +14,7 @@ The settings infrastructure provides a global namespace of key-value mappings where the code can pull configuration values from. The settings can be populated through different mechanisms, which are described below. -Read :ref:`settings` for all supported entries. +Read :ref:`topics-settings-ref` for all supported entries. How to populate settings ======================== @@ -83,8 +83,7 @@ per-comand inside your project, by writing them in the module referenced by the -------------------------- The global defaults are located in scrapy.conf.default_settings and documented -in the :ref:`settings` page. - +in the :ref:`topics-settings-ref` section. How to access settings ====================== @@ -157,14 +156,852 @@ to do that you'll have to use one of the following methods: ``default`` is the value to return if no setting is found -Available built-in settings -=========================== - -See :ref:`settings`. - Rationale for setting names =========================== Setting names are usually prefixed with the component that they configure. For example, proper setting names for a fictional robots.txt extension would be ``ROBOTSTXT_ENABLED``, ``ROBOTSTXT_OBEY``, ``ROBOTSTXT_CACHEDIR``, etc. + + +.. _topics-settings-ref: + +Built-in settings reference +=========================== + +Here's a list of all available Scrapy settings, in alphabetical order, along +with their default values and the scope where they apply. + +The scope, where available, shows where the setting is being used, if it's tied +to any particular component. In that case the module of that component will be +shown, typically an extension, middleware or pipeline. It also means that the +component must be enabled in order for the setting to have any effect. + +.. setting:: BOT_NAME + +BOT_NAME +-------- + +Default: ``scrapybot`` + +The name of the bot implemented by this Scrapy project. This will be used to +construct the User-Agent by default, and also for logging. + +.. setting:: BOT_VERSION + +BOT_VERSION +----------- + +Default: ``1.0`` + +The version of the bot implemented by this Scrapy project. This will be used to +construct the User-Agent by default. + +.. setting:: HTTPCACHE_DIR + +HTTPCACHE_DIR +------------- + +Default: ``''`` (empty string) + +The directory to use for storing the (low-level) HTTP cache. If empty the HTTP +cache will be disabled. + +.. setting:: HTTPCACHE_EXPIRATION_SECS + +HTTPCACHE_EXPIRATION_SECS +------------------------- + +Default: ``0`` + +Number of seconds to use for HTTP cache expiration. Requests that were cached +before this time will be re-downloaded. If zero, cached requests will always +expire. Negative numbers means requests will never expire. + +.. setting:: HTTPCACHE_IGNORE_MISSING + +HTTPCACHE_IGNORE_MISSING +------------------------ + +Default: ``False`` + +If enabled, requests not found in the cache will be ignored instead of downloaded. + +.. setting:: HTTPCACHE_SECTORIZE + +HTTPCACHE_SECTORIZE +------------------- + +Default: ``True`` + +Whether to split HTTP cache storage in several dirs for performance. + +.. setting:: COMMANDS_MODULE + +COMMANDS_MODULE +--------------- + +Default: ``''`` (empty string) + +A module to use for looking for custom Scrapy commands. This is used to add +custom command for your Scrapy project. + +Example:: + + COMMANDS_MODULE = 'mybot.commands' + +.. setting:: COMMANDS_SETTINGS_MODULE + +COMMANDS_SETTINGS_MODULE +------------------------ + +Default: ``''`` (empty string) + +A module to use for looking for custom Scrapy command settings. + +Example:: + + COMMANDS_SETTINGS_MODULE = 'mybot.conf.commands' + +.. setting:: CONCURRENT_DOMAINS + +CONCURRENT_DOMAINS +------------------ + +Default: ``8`` + +Maximum number of domains to scrape in parallel. + +.. setting:: CONCURRENT_ITEMS + +CONCURRENT_ITEMS +---------------- + +Default: ``100`` + +Maximum number of concurrent items (per response) to process in parallel in the +Item Processor (also known as the Item Pipeline). + +.. setting:: COOKIES_DEBUG + +COOKIES_DEBUG +------------- + +Default: ``False`` + +Enable debugging message of Cookies Downloader Middleware. + +.. setting:: DEFAULT_ITEM_CLASS + +DEFAULT_ITEM_CLASS +------------------ + +Default: ``'scrapy.item.ScrapedItem'`` + +The default class that will be used for instantiating items in the :ref:`the +Scrapy shell `. + +.. setting:: DEFAULT_REQUEST_HEADERS + +DEFAULT_REQUEST_HEADERS +----------------------- + +Default:: + + { + 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', + 'Accept-Language': 'en', + } + +The default headers used for Scrapy HTTP Requests. They're populated in the +:class:`~scrapy.contrib.downloadermiddleware.defaultheaders.DefaultHeadersMiddleware`. + +.. setting:: DEFAULT_SPIDER + +DEFAULT_SPIDER +-------------- + +Default: ``None`` + +The default spider class that will be instantiated for URLs for which no +specific spider is found. This class must have a constructor which receives as +only parameter the domain name of the given URL. + +.. setting:: DEPTH_LIMIT + +DEPTH_LIMIT +----------- + +Default: ``0`` + +The maximum depth that will be allowed to crawl for any site. If zero, no limit +will be imposed. + +.. setting:: DEPTH_STATS + +DEPTH_STATS +----------- + +Default: ``True`` + +Whether to collect depth stats. + +.. setting:: DOMAIN_SCHEDULER + +DOMAIN_SCHEDULER +---------------- + +Default: ``'scrapy.contrib.domainsch.FifoDomainScheduler'`` + +The Domain Scheduler to use. The domain scheduler returns the next domain +(spider) to scrape. + +.. setting:: DOWNLOADER_DEBUG + +DOWNLOADER_DEBUG +---------------- + +Default: ``False`` + +Whether to enable the Downloader debugging mode. + +.. setting:: DOWNLOADER_MIDDLEWARES + +DOWNLOADER_MIDDLEWARES +---------------------- + +Default:: ``{}`` + +A dict containing the downloader middlewares enabled in your project, and their +orders. For more info see :ref:`topics-downloader-middleware-setting`. + +.. setting:: DOWNLOADER_MIDDLEWARES_BASE + +DOWNLOADER_MIDDLEWARES_BASE +--------------------------- + +Default:: + + { + 'scrapy.contrib.downloadermiddleware.robotstxt.RobotsTxtMiddleware': 100, + 'scrapy.contrib.downloadermiddleware.httpauth.HttpAuthMiddleware': 300, + 'scrapy.contrib.downloadermiddleware.useragent.UserAgentMiddleware': 400, + 'scrapy.contrib.downloadermiddleware.retry.RetryMiddleware': 500, + 'scrapy.contrib.downloadermiddleware.defaultheaders.DefaultHeadersMiddleware': 550, + 'scrapy.contrib.downloadermiddleware.redirect.RedirectMiddleware': 600, + 'scrapy.contrib.downloadermiddleware.cookies.CookiesMiddleware': 700, + 'scrapy.contrib.downloadermiddleware.httpcompression.HttpCompressionMiddleware': 800, + 'scrapy.contrib.downloadermiddleware.stats.DownloaderStats': 850, + 'scrapy.contrib.downloadermiddleware.httpcache.HttpCacheMiddleware': 900, + } + +A dict containing the downloader middlewares enabled by default in Scrapy. You +should never modify this setting in your project, modify +:setting:`DOWNLOADER_MIDDLEWARES` instead. For more info see +:ref:`topics-downloader-middleware-setting`. + +.. setting:: DOWNLOADER_STATS + +DOWNLOADER_STATS +---------------- + +Default: ``True`` + +Whether to enable downloader stats collection. + +.. setting:: DOWNLOAD_DELAY + +DOWNLOAD_DELAY +-------------- + +Default: ``0`` + +The amount of time (in secs) that the downloader should wait before downloading +consecutive pages from the same spider. This can be used to throttle the +crawling speed to avoid hitting servers too hard. Decimal numbers are +supported. Example:: + + DOWNLOAD_DELAY = 0.25 # 250 ms of delay + +.. setting:: DOWNLOAD_TIMEOUT + +DOWNLOAD_TIMEOUT +---------------- + +Default: ``180`` + +The amount of time (in secs) that the downloader will wait before timing out. + +.. setting:: DUPEFILTER_CLASS + +DUPEFILTER_CLASS +---------------- + +Default: ``'scrapy.contrib.dupefilter.RequestFingerprintDupeFilter'`` + +The class used to detect and filter duplicate requests. + +The default (``RequestFingerprintDupeFilter``) filters based on request fingerprint +(using ``scrapy.utils.request.request_fingerprint``) and grouping per domain. + +.. setting:: EXTENSIONS + +EXTENSIONS +---------- + +Default:: ``{}`` + +A dict containing the extensions enabled in your project, and their orders. + +.. setting:: EXTENSIONS_BASE + +EXTENSIONS_BASE +--------------- + +Default:: + + { + 'scrapy.stats.corestats.CoreStats': 0, + 'scrapy.management.web.WebConsole': 0, + 'scrapy.management.telnet.TelnetConsole': 0, + 'scrapy.contrib.webconsole.scheduler.SchedulerQueue': 0, + 'scrapy.contrib.webconsole.livestats.LiveStats': 0, + 'scrapy.contrib.webconsole.spiderctl.Spiderctl': 0, + 'scrapy.contrib.webconsole.enginestatus.EngineStatus': 0, + 'scrapy.contrib.webconsole.stats.StatsDump': 0, + 'scrapy.contrib.spider.reloader.SpiderReloader': 0, + 'scrapy.contrib.memusage.MemoryUsage': 0, + 'scrapy.contrib.memdebug.MemoryDebugger': 0, + 'scrapy.contrib.closedomain.CloseDomain': 0, + 'scrapy.contrib.debug.StackTraceDump': 0, + } + +The list of available extensions. Keep in mind that some of them need need to +be enabled through a setting. By default, this setting contains all stable +built-in extensions. + +For more information See the :ref:`extensions user guide ` +and the :ref:`list of available extensions `. + +.. setting:: GROUPSETTINGS_ENABLED + +GROUPSETTINGS_ENABLED +--------------------- + +Default: ``False`` + +Whether to enable group settings where spiders pull their settings from. + +.. setting:: GROUPSETTINGS_MODULE + +GROUPSETTINGS_MODULE +-------------------- + +Default: ``''`` (empty string) + +The module to use for pulling settings from, if the group settings is enabled. + +.. setting:: ITEM_PIPELINES + +ITEM_PIPELINES +-------------- + +Default: ``[]`` + +The item pipelines to use (a list of classes). + +Example:: + + ITEM_PIPELINES = [ + 'mybot.pipeline.validate.ValidateMyItem', + 'mybot.pipeline.validate.StoreMyItem' + ] + +.. setting:: LOG_ENABLED + +LOG_ENABLED +----------- + +Default: ``True`` + +Enable logging. + +.. setting:: LOG_STDOUT + +LOG_STDOUT +---------- + +Default: ``False`` + +If enabled logging will be sent to standard output, otherwise standard error +will be used. + +.. setting:: LOGFILE + +LOGFILE +------- + +Default: ``None`` + +File name to use for logging output. If None, standard input (or error) will be +used depending on the value of the LOG_STDOUT setting. + +.. setting:: LOGLEVEL + +LOGLEVEL +-------- + +Default: ``'DEBUG'`` + +Minimum level to log. Available levels are: SILENT, CRITICAL, ERROR, WARNING, +INFO, DEBUG, TRACE + +.. setting:: MAIL_FROM + +MAIL_FROM +--------- + +Default: ``'scrapy@localhost'`` + +Email to use as sender address for sending emails using the :ref:`Scrapy e-mail +sending facility `. + +.. setting:: MAIL_HOST + +MAIL_HOST +--------- + +Default: ``'localhost'`` + +Host to use for sending emails using the :ref:`Scrapy e-mail sending facility +`. + +.. setting:: MEMDEBUG_ENABLED + +MEMDEBUG_ENABLED +---------------- + +Default: ``False`` + +Whether to enable memory debugging. + +.. setting:: MEMDEBUG_NOTIFY + +MEMDEBUG_NOTIFY +--------------- + +Default: ``[]`` + +When memory debugging is enabled a memory report will be sent to the specified +addresses if this setting is not empty, otherwise the report will be written to +the log. + +Example:: + + MEMDEBUG_NOTIFY = ['user@example.com'] + +.. setting:: MEMUSAGE_ENABLED + +MEMUSAGE_ENABLED +---------------- + +Default: ``False`` + +Scope: ``scrapy.contrib.memusage`` + +Whether to enable the memory usage extension that will shutdown the Scrapy +process when it exceeds a memory limit, and also notify by email when that +happened. + +See :ref:`topics-extensions-ref-memusage`. + +.. setting:: MEMUSAGE_LIMIT_MB + +MEMUSAGE_LIMIT_MB +----------------- + +Default: ``0`` + +Scope: ``scrapy.contrib.memusage`` + +The maximum amount of memory to allow (in megabytes) before shutting down +Scrapy (if MEMUSAGE_ENABLED is True). If zero, no check will be performed. + +See :ref:`topics-extensions-ref-memusage`. + +.. setting:: MEMUSAGE_NOTIFY_MAIL + +MEMUSAGE_NOTIFY_MAIL +-------------------- + +Default: ``False`` + +Scope: ``scrapy.contrib.memusage`` + +A list of emails to notify if the memory limit has been reached. + +Example:: + + MEMUSAGE_NOTIFY_MAIL = ['user@example.com'] + +See :ref:`topics-extensions-ref-memusage`. + +.. setting:: MEMUSAGE_REPORT + +MEMUSAGE_REPORT +--------------- + +Default: ``False`` + +Scope: ``scrapy.contrib.memusage`` + +Whether to send a memory usage report after each domain has been closed. + +See :ref:`topics-extensions-ref-memusage`. + +.. setting:: MEMUSAGE_WARNING_MB + +MEMUSAGE_WARNING_MB +------------------- + +Default: ``0`` + +Scope: ``scrapy.contrib.memusage`` + +The maximum amount of memory to allow (in megabytes) before sending a warning +email notifying about it. If zero, no warning will be produced. + +.. setting:: MYSQL_CONNECTION_SETTINGS + +MYSQL_CONNECTION_SETTINGS +------------------------- + +Default: ``{}`` + +Scope: ``scrapy.utils.db.mysql_connect`` + +Settings to use for MySQL connections performed through +``scrapy.utils.db.mysql_connect`` + +.. setting:: NEWSPIDER_MODULE + +NEWSPIDER_MODULE +---------------- + +Default: ``''`` + +Module where to create new spiders using the ``genspider`` command. + +Example:: + + NEWSPIDER_MODULE = 'mybot.spiders_dev' + +.. setting:: PROJECT_NAME + +PROJECT_NAME +------------ + +Default: ``Not Defined`` + +The name of the current project. It matches the project module name as created +by ``startproject`` command, and is only defined by project settings file. + +.. setting:: REDIRECT_MAX_TIMES + +REDIRECT_MAX_TIMES +------------------ + +Default: ``20`` + +Defines the maximun times a request can be redirected. After this maximun the +request's response is returned as is. We used Firefox default value for the +same task. + +.. setting:: REDIRECT_MAX_METAREFRESH_DELAY + +REDIRECT_MAX_METAREFRESH_DELAY +------------------------------ + +Default: ``100`` + +Some sites use meta-refresh for redirecting to a session expired page, so we +restrict automatic redirection to a maximum delay (in seconds) + +.. setting:: REDIRECT_PRIORITY_ADJUST + +REDIRECT_PRIORITY_ADJUST +------------------------------ + +Default: ``+2`` + +Adjust redirect request priority relative to original request. +A negative priority adjust means more priority. + +.. setting:: REQUESTS_PER_DOMAIN + +REQUESTS_PER_DOMAIN +------------------- + +Default: ``8`` + +Specifies how many concurrent (ie. simultaneous) requests will be performed per +open spider. + +.. setting:: REQUESTS_QUEUE_SIZE + +REQUESTS_QUEUE_SIZE +------------------- + +Default: ``0`` + +Scope: ``scrapy.contrib.spidermiddleware.limit`` + +If non zero, it will be used as an upper limit for the amount of requests that +can be scheduled per domain. + +.. setting:: ROBOTSTXT_OBEY + +ROBOTSTXT_OBEY +-------------- + +Default: ``False`` + +Scope: ``scrapy.contrib.downloadermiddleware.robotstxt`` + +If enabled, Scrapy will respect robots.txt policies. For more information see +:topic:`robotstxt` + +.. setting:: SCHEDULER + +SCHEDULER +--------- + +Default: ``'scrapy.core.scheduler.Scheduler'`` + +The scheduler to use for crawling. + +.. setting:: SCHEDULER_ORDER + +SCHEDULER_ORDER +--------------- + +Default: ``'BFO'`` + +Scope: ``scrapy.core.scheduler`` + +The order to use for the crawling scheduler. Available orders are: + +* ``'BFO'``: `Breadth-first order`_ - typically consumes more memory but + reaches most relevant pages earlier. + +* ``'DFO'``: `Depth-first order`_ - typically consumes less memory than + but takes longer to reach most relevant pages. + +.. _Breadth-first order: http://en.wikipedia.org/wiki/Breadth-first_search +.. _Depth-first order: http://en.wikipedia.org/wiki/Depth-first_search + +.. setting:: SCHEDULER_MIDDLEWARES + +SCHEDULER_MIDDLEWARES +--------------------- + +Default:: ``{}`` + +A dict containing the scheduler middlewares enabled in your project, and their +orders. + +.. setting:: SCHEDULER_MIDDLEWARES_BASE + +SCHEDULER_MIDDLEWARES_BASE +-------------------------- + +Default:: + + SCHEDULER_MIDDLEWARES_BASE = { + 'scrapy.contrib.schedulermiddleware.duplicatesfilter.DuplicatesFilterMiddleware': 500, + } + +A dict containing the scheduler middlewares enabled by default in Scrapy. You +should never modify this setting in your project, modify +:setting:`SCHEDULER_MIDDLEWARES` instead. + +.. setting:: SPIDERPROFILER_ENABLED + +SPIDERPROFILER_ENABLED +---------------------- + +Default: ``False`` + +Enable the spider profiler. Warning: this could have a big impact in +performance. + +.. setting:: SPIDER_MIDDLEWARES + +SPIDER_MIDDLEWARES +------------------ + +Default:: ``{}`` + +A dict containing the spider middlewares enabled in your project, and their +orders. For more info see :ref:`topics-spider-middleware-setting`. + +.. setting:: SPIDER_MIDDLEWARES_BASE + +SPIDER_MIDDLEWARES_BASE +----------------------- + +Default:: + + { + 'scrapy.contrib.spidermiddleware.httperror.HttpErrorMiddleware': 50, + 'scrapy.contrib.itemsampler.ItemSamplerMiddleware': 100, + 'scrapy.contrib.spidermiddleware.requestlimit.RequestLimitMiddleware': 200, + 'scrapy.contrib.spidermiddleware.restrict.RestrictMiddleware': 300, + 'scrapy.contrib.spidermiddleware.offsite.OffsiteMiddleware': 500, + 'scrapy.contrib.spidermiddleware.referer.RefererMiddleware': 700, + 'scrapy.contrib.spidermiddleware.urllength.UrlLengthMiddleware': 800, + 'scrapy.contrib.spidermiddleware.depth.DepthMiddleware': 900, + } + +A dict containing the spider middlewares enabled by default in Scrapy. You +should never modify this setting in your project, modify +:setting:`SPIDER_MIDDLEWARES` instead. For more info see +:ref:`topics-spider-middleware-setting`. + +.. setting:: SPIDER_MODULES + +SPIDER_MODULES +-------------- + +Default: ``[]`` + +A list of modules where Scrapy will look for spiders. + +Example:: + + SPIDER_MODULES = ['mybot.spiders_prod', 'mybot.spiders_dev'] + +.. setting:: STATS_CLASS + +STATS_CLASS +----------- + +Default: ``'scrapy.stats.collector.MemoryStatsCollector'`` + +The class to use for collecting stats (must implement the Stats Collector API, +or subclass the StatsCollector class). + +.. setting:: STATS_DUMP + +STATS_DUMP +---------- + +Default: ``False`` + +Dump (to log) domain-specific stats collected when a domain is closed, and all +global stats when the Scrapy process finishes (ie. when the engine is +shutdown). + +.. setting:: STATS_ENABLED + +STATS_ENABLED +------------- + +Default: ``True`` + +Enable stats collection. + +.. setting:: STATSMAILER_RCPTS + +STATSMAILER_RCPTS +----------------- + +Default: ``[]`` (empty list) + +Send Scrapy stats after domains finish scrapy. See +:class:`~scrapy.contrib.statsmailer.StatsMailer` for more info. + +.. setting:: TELNETCONSOLE_ENABLED + +TELNETCONSOLE_ENABLED +--------------------- + +Default: ``True`` + +Scope: ``scrapy.management.telnet`` + +A boolean which specifies if the telnet management console will be enabled +(provided its extension is also enabled). + +.. setting:: TELNETCONSOLE_PORT + +TELNETCONSOLE_PORT +------------------ + +Default: ``6023`` + +The port to use for the telnet console. If set to ``None`` or ``0``, a +dynamically assigned port is used. For more info see +:ref:`topics-telnetconsole`. + +.. setting:: TEMPLATES_DIR + +TEMPLATES_DIR +------------- + +Default: ``templates`` dir inside scrapy module + +The directory where to look for template when creating new projects with +scrapy-admin.py newproject. + +.. setting:: URLLENGTH_LIMIT + +URLLENGTH_LIMIT +--------------- + +Default: ``2083`` + +Scope: ``contrib.spidermiddleware.urllength`` + +The maximum URL length to allow for crawled URLs. For more information about +the default value for this setting see: http://www.boutell.com/newfaq/misc/urllength.html + +.. setting:: USER_AGENT + +USER_AGENT +---------- + +Default: ``"%s/%s" % (BOT_NAME, BOT_VERSION)`` + +The default User-Agent to use when crawling, unless overrided. + +.. setting:: WEBCONSOLE_ENABLED + +WEBCONSOLE_ENABLED +------------------ + +Default: True + +A boolean which specifies if the web management console will be enabled +(provided its extension is also enabled). + +.. setting:: WEBCONSOLE_LOGFILE + +WEBCONSOLE_LOGFILE +------------------ + +Default: ``None`` + +A file to use for logging HTTP requests made to the web console. If unset web +the log is sent to standard scrapy log. + +.. setting:: WEBCONSOLE_PORT + +WEBCONSOLE_PORT +--------------- + +Default: ``6080`` + +The port to use for the web console. If set to ``None`` or ``0``, a dynamically +assigned port is used. For more info see :ref:`topics-webconsole`. + diff --git a/docs/ref/signals.rst b/docs/topics/signals.rst similarity index 98% rename from docs/ref/signals.rst rename to docs/topics/signals.rst index 5e1435d6c..5469648e1 100644 --- a/docs/ref/signals.rst +++ b/docs/topics/signals.rst @@ -1,10 +1,11 @@ -.. _signals: +.. _topics-signals: .. module:: scrapy.core.signals :synopsis: Signals definitions -Available Signals -================= +======= +Signals +======= Scrapy uses signals extensively to notify when certain actions occur. You can catch some of those signals in your Scrapy project or extension to perform @@ -19,6 +20,12 @@ For more information about working when see the documentation of .. _pydispatcher: http://pydispatcher.sourceforge.net/ + +.. _topics-signals-ref: + +Built-in signals reference +========================== + Here's a list of signals used in Scrapy and their meaning, in alphabetical order. diff --git a/docs/topics/spider-middleware.rst b/docs/topics/spider-middleware.rst index caddcea7e..16a498664 100644 --- a/docs/topics/spider-middleware.rst +++ b/docs/topics/spider-middleware.rst @@ -105,3 +105,117 @@ no middleware is left and the default exception handling kicks in. If it returns an iterable the ``process_spider_output()`` pipeline kicks in, and no other ``process_spider_exception()`` will be called. + +.. _topics-spider-middleware-ref: + +Built-in spider middleware reference +==================================== + +This page describes all spider middleware components that come with Scrapy. For +information on how to use them and how to write your own spider middleware, see +the :ref:`spider middleware usage guide `. + +For a list of the components enabled by default (and their orders) see the +:setting:`SPIDER_MIDDLEWARES_BASE` setting. + +DepthMiddleware +--------------- + +.. module:: scrapy.contrib.spidermiddleware.depth + +.. class:: DepthMiddleware + + DepthMiddleware is a scrape middleware used for tracking the depth of each + Request inside the site being scraped. It can be used to limit the maximum + depth to scrape or things like that. + + The :class:`DepthMiddleware` can be configured through the following + settings (see the settings documentation for more info): + + * :setting:`DEPTH_LIMIT` - The maximum depth that will be allowed to + crawl for any site. If zero, no limit will be imposed. + * :setting:`DEPTH_STATS` - Whether to collect depth stats. + +HttpErrorMiddleware +------------------- + +.. module:: scrapy.contrib.spidermiddleware.httperror + +.. class:: HttpErrorMiddleware + + Filter out response outside of a range of valid status codes. + + This middleware filters out every response with status outside of the range + 200<=status<300. Spiders can add more exceptions using + ``handle_httpstatus_list`` spider attribute. + +OffsiteMiddleware +----------------- + +.. module:: scrapy.contrib.spidermiddleware.offsite + +.. class:: OffsiteMiddleware + + Filters out Requests for URLs outside the domains covered by the spider. + + This middleware filters out every request whose host names doesn't match + :attr:`~scrapy.spider.BaseSpider.domain_name`, or the spider + :attr:`~scrapy.spider.BaseSpider.domain_name` prefixed by "www.". + Spider can add more domains to exclude using + :attr:`~scrapy.spider.BaseSpider.extra_domain_names` attribute. + +RequestLimitMiddleware +---------------------- + +.. module:: scrapy.contrib.spidermiddleware.requestlimit + +.. class:: RequestLimitMiddleware + + Limits the maximum number of requests in the scheduler for each spider. When + a spider tries to schedule more than the allowed amount of requests, the new + requests (returned by the spider) will be dropped. + + The :class:`RequestLimitMiddleware` can be configured through the following + settings (see the settings documentation for more info): + + * :setting:`REQUESTS_QUEUE_SIZE` - If non zero, it will be used as an + upper limit for the amount of requests that can be scheduled per + domain. Can be set per spider using ``requests_queue_size`` attribute. + +RestrictMiddleware +------------------ + +.. module:: scrapy.contrib.spidermiddleware.restrict + +.. class:: RestrictMiddleware + + Restricts crawling to fixed set of particular URLs. + + The :class:`RestrictMiddleware` can be configured through the following + settings (see the settings documentation for more info): + + * :setting:`RESTRICT_TO_URLS` - Set of URLs allowed to crawl. + +UrlFilterMiddleware +------------------- + +.. module:: scrapy.contrib.spidermiddleware.urlfilter + +.. class:: UrlFilterMiddleware + + Canonicalizes URLs to filter out duplicated ones + +UrlLengthMiddleware +------------------- + +.. module:: scrapy.contrib.spidermiddleware.urllength + +.. class:: UrlLengthMiddleware + + Filters out requests with URLs longer than URLLENGTH_LIMIT + + The :class:`UrlLengthMiddleware` can be configured through the following + settings (see the settings documentation for more info): + + * :setting:`URLLENGTH_LIMIT` - The maximum URL length to allow for crawled URLs. + diff --git a/docs/topics/spiders.rst b/docs/topics/spiders.rst index 63505b4c9..4575d5362 100644 --- a/docs/topics/spiders.rst +++ b/docs/topics/spiders.rst @@ -39,5 +39,379 @@ Even though this cycles applies (more or less) to any kind of spider, there are different kind of default spiders bundled into Scrapy for different purposes. We will talk about those types here. -See :ref:`ref-spiders` for the list of default spiders available in Scrapy. + +.. _topics-spiders-ref: + +Built-in spiders reference +========================== + +.. module:: scrapy.spider + :synopsis: Spiders base class, spider manager and spider middleware + +BaseSpider +---------- + +.. class:: BaseSpider() + +This is the simplest spider, and the one from which every other spider +must inherit from (either the ones that come bundled with Scrapy, or the ones +that you write yourself). It doesn't provide any special functionality. It just +requests the given ``start_urls``/``start_requests``, and calls the spider's +method ``parse`` for each of the resulting responses. + +.. attribute:: BaseSpider.domain_name + + A string which defines the domain name for this spider, which will also be + the unique identifier for this spider (which means you can't have two + spider with the same ``domain_name``). This is the most important spider + attribute and it's required, and it's the name by which Scrapy will known + the spider. + +.. attribute:: BaseSpider.extra_domain_names + + An optional list of strings containing additional domains that this spider + is allowed to crawl. Requests for URLs not belonging to the domain name + specified in :attr:`Spider.domain_name` or this list won't be followed. + +.. attribute:: BaseSpider.start_urls + + Is a list of URLs where the spider will begin to crawl from, when no + particular URLs are specified. So, the first pages downloaded will be those + listed here. The subsequent URLs will be generated successively from data + contained in the start URLs. + +.. method:: BaseSpider.start_requests() + + This method must return an iterable with the first Requests to crawl for + this spider. + + This is the method called by Scrapy when the spider is opened for scraping + when no particular URLs are specified. If particular URLs are specified, + the :meth:`BaseSpider.make_requests_from_url` is used instead to create the + Requests. This method is also called only once from Scrapy, so it's safe to + implement it as a generator. + + The default implementation uses :meth:`BaseSpider.make_requests_from_url` + to generate Requests for each url in :attr:`start_urls`. + + If you want to change the Requests used to start scraping a domain, this is + the method to override. For example, if you need to start by login in using + a POST request, you could do:: + + def start_requests(self): + return [FormRequest("http://www.example.com/login", + formdata={'user': 'john', 'pass': 'secret'}, + callback=self.logged_in)] + + def logged_in(self, response): + # here you would extract links to follow and return Requests for + # each of them, with another callback + pass + +.. method:: BaseSpider.make_requests_from_url(url) + + A method that receives a URL and returns a :class:`~scrapy.http.Request` + object (or a list of :class:`~scrapy.http.Request` objects) to scrape. This + method is used to construct the initial requests in the + :meth:`start_requests` method, and is typically used to convert urls to + requests. + + Unless overridden, this method returns Requests with the :meth:`parse` + method as their callback function, and with dont_filter parameter enabled + (see :class:`~scrapy.http.Request` class for more info). + +.. method:: BaseSpider.parse(response) + + This is the default callback used by the :meth:`start_requests` method, and + will be used to parse the first pages crawled by the spider. + + The ``parse`` method is in charge of processing the response and returning + scraped data and/or more URLs to follow, because of this, the method must + always return a list or at least an empty one. Other Requests callbacks + have the same requirements as the BaseSpider class. + +BaseSpider example +~~~~~~~~~~~~~~~~~~ + +Let's see an example:: + + from scrapy import log # This module is useful for printing out debug information + from scrapy.spider import BaseSpider + + class MySpider(BaseSpider): + domain_name = 'http://www.example.com' + start_urls = [ + 'http://www.example.com/1.html', + 'http://www.example.com/2.html', + 'http://www.example.com/3.html', + ] + + def parse(self, response): + self.log('A response from %s just arrived!' % response.url) + return [] + + SPIDER = MySpider() + +.. module:: scrapy.contrib.spiders + :synopsis: Collection of generic spiders + +CrawlSpider +----------- + +.. class:: CrawlSpider + +This is the most commonly used spider for crawling regular websites, as it +provides a convenient mechanism for following links by defining a set of rules. +It may not be the best suited for your particular web sites or project, but +it's generic enough for several cases, so you can start from it and override it +as need more custom functionality, or just implement your own spider. + +Apart from the attributes inherited from BaseSpider (that you must +specify), this class supports a new attribute: + +.. attribute:: CrawlSpider.rules + + Which is a list of one (or more) :class:`Rule` objects. Each :class:`Rule` + defines a certain behaviour for crawling the site. Rules objects are + described below . + +Crawling rules +~~~~~~~~~~~~~~ + +.. class:: Rule(link_extractor, callback=None, cb_kwargs=None, follow=None, process_links=None) + +``link_extractor`` is a :ref:`Link Extractor ` object which +defines how links will be extracted from each crawled page. + +``callback`` is a callable or a string (in which case a method from the spider +object with that name will be used) to be called for each link extracted with +the specified link_extractor. This callback receives a response as its first +argument and must return a list containing either ScrapedItems and Requests (or +any subclass of them). + +``cb_kwargs`` is a dict containing the keyword arguments to be passed to the +callback function + +``follow`` is a boolean which specified if links should be followed from each +response extracted with this rule. If ``callback`` is None ``follow`` defaults +to ``True``, otherwise it default to ``False``. + +``process_links`` is a callable, or a string (in which case a method from the +spider object with that name will be used) which will be called for each list +of links extracted from each response using the specified ``link_extractor``. +This is mainly used for filtering purposes. + + +CrawlSpider example +------------------- + +Let's now take a look at an example CrawlSpider with rules:: + + from scrapy.contrib.spiders import CrawlSpider, Rule + from scrapy.contrib.linkextractors.sgml import SgmlLinkExtractor + from scrapy.xpath.selector import HtmlXPathSelector + from scrapy.item import ScrapedItem + + class MySpider(CrawlSpider): + domain_name = 'example.com' + start_urls = ['http://www.example.com'] + + rules = ( + # Extract links matching 'category.php' (but not matching 'subsection.php') + # and follow links from them (since no callback means follow=True by default). + Rule(SgmlLinkExtractor(allow=('category\.php', ), deny=('subsection\.php', ))), + + # Extract links matching 'item.php' and parse them with the spider's method parse_item + Rule(SgmlLinkExtractor(allow=('item\.php', )), callback='parse_item'), + ) + + def parse_item(self, response): + self.log('Hi, this is an item page! %s' % response.url) + + hxs = HtmlXPathSelector(response) + item = ScrapedItem() + item.id = hxs.select('//td[@id="item_id"]/text()').re(r'ID: (\d+)') + item.name = hxs.select('//td[@id="item_name"]/text()').extract() + item.description = hxs.select('//td[@id="item_description"]/text()').extract() + return [item] + + SPIDER = MySpider() + + +This spider would start crawling example.com's home page, collecting category +links, and item links, parsing the latter with the +:meth:`XMLFeedSpider.parse_item` method. For each item response, some data will +be extracted from the HTML using XPath, and a ScrapedItem will be filled with +it. + +XMLFeedSpider +------------- + +.. class:: XMLFeedSpider + + XMLFeedSpider is designed for parsing XML feeds by iterating through them by a + certain node name. The iterator can be chosen from: ``iternodes``, ``xml``, + and ``html``. It's recommended to use the ``iternodes`` iterator for + performance reasons, since the ``xml`` and ``html`` iterators generate the + whole DOM at once in order to parse it. However, using ``html`` as the + iterator may be useful when parsing XML with bad markup. + + For setting the iterator and the tag name, you must define the following class + attributes: + + .. attribute:: iterator + + A string which defines the iterator to use. It can be either: + + - ``'iternodes'`` - a fast iterator based on regular expressions + + - ``'html'`` - an iterator which uses HtmlXPathSelector. Keep in mind + this uses DOM parsing and must load all DOM in memory which could be a + problem for big feeds + + - ``'xml'`` - an iterator which uses XmlXPathSelector. Keep in mind + this uses DOM parsing and must load all DOM in memory which could be a + problem for big feeds + + It defaults to: ``'iternodes'``. + + .. attribute:: itertag + + A string with the name of the node (or element) to iterate in. Example:: + + itertag = 'product' + + .. attribute:: namespaces + + A list of ``(prefix, uri)`` tuples which define the namespaces + available in that document that will be processed with this spider. The + ``prefix`` and ``uri`` will be used to automatically register + namespaces using the + :meth:`~scrapy.xpath.XPathSelector.register_namespace` method. + + You can then specify nodes with namespaces in the :attr:`itertag` + attribute. + + Example:: + + class YourSpider(XMLFeedSpider): + + namespaces = [('n', 'http://www.sitemaps.org/schemas/sitemap/0.9')] + itertag = 'n:url' + # ... + + Apart from these new attributes, this spider has the following overrideable + methods too: + + .. method:: adapt_response(response) + + A method that receives the response as soon as it arrives from the spider + middleware and before start parsing it. It can be used used for modifying + the response body before parsing it. This method receives a response and + returns response (it could be the same or another one). + + .. method:: parse_item(response, selector) + + This method is called for the nodes matching the provided tag name + (``itertag``). Receives the response and an XPathSelector for each node. + Overriding this method is mandatory. Otherwise, you spider won't work. + This method must return either a ScrapedItem, a Request, or a list + containing any of them. + + .. warning:: This method will soon change its name to ``parse_node`` + + .. method:: process_results(response, results) + + This method is called for each result (item or request) returned by the + spider, and it's intended to perform any last time processing required + before returning the results to the framework core, for example setting the + item IDs. It receives a list of results and the response which originated + that results. It must return a list of results (Items or Requests).""" + + +XMLFeedSpider example +~~~~~~~~~~~~~~~~~~~~~ + +These spiders are pretty easy to use, let's have at one example:: + + from scrapy import log + from scrapy.contrib.spiders import XMLFeedSpider + from scrapy.item import ScrapedItem + + class MySpider(XMLFeedSpider): + domain_name = 'example.com' + start_urls = ['http://www.example.com/feed.xml'] + iterator = 'iternodes' # This is actually unnecesary, since it's the default value + itertag = 'item' + + def parse_item(self, response, node): + log.msg('Hi, this is a <%s> node!: %s' % (self.itertag, ''.join(node.extract()))) + + item = ScrapedItem() + item.id = node.select('@id').extract() + item.name = node.select('name').extract() + item.description = node.select('description').extract() + return item + + SPIDER = MySpider() + +Basically what we did up there was creating a spider that downloads a feed from +the given ``start_urls``, and then iterates through each of its ``item`` tags, +prints them out, and stores some random data in ScrapedItems. + +CSVFeedSpider +------------- + +.. class:: CSVFeedSpider + +.. warning:: The API of the CSVFeedSpider is not yet stable. Use with caution. + +This spider is very similar to the XMLFeedSpider, although it iterates through +rows, instead of nodes. It also has other two different attributes: + +.. attribute:: CSVFeedSpider.delimiter + + A string with the separator character for each field in the CSV file + Defaults to ``','`` (comma). + +.. attribute:: CSVFeedSpider.headers + + A list of the rows contained in the file CSV feed which will be used for + extracting fields from it. + +In this spider, the method that gets called in each row iteration ``parse_row`` +instead of ``parse_item`` (like in :class:`XMLFeedSpider`). + +.. method:: CSVFeedSpider.parse_row(response, row) + + Receives a response and a dict (representing each row) with a key for each + provided (or detected) header of the CSV file. This spider also gives the + opportunity to override ``adapt_response`` and ``process_results`` methods + for pre and post-processing purposes. + +CSVFeedSpider example +~~~~~~~~~~~~~~~~~~~~~ + +Let's see an example similar to the previous one, but using CSVFeedSpider:: + + from scrapy import log + from scrapy.contrib.spiders import CSVFeedSpider + from scrapy.item import ScrapedItem + + class MySpider(CSVFeedSpider): + domain_name = 'example.com' + start_urls = ['http://www.example.com/feed.csv'] + delimiter = ';' + headers = ['id', 'name', 'description'] + + def parse_row(self, response, row): + log.msg('Hi, this is a row!: %r' % row) + + item = ScrapedItem() + item.id = row['id'] + item.name = row['name'] + item.description = row['description'] + return item + + SPIDER = MySpider() diff --git a/docs/topics/stats.rst b/docs/topics/stats.rst index 2e30ccd17..2afc78397 100644 --- a/docs/topics/stats.rst +++ b/docs/topics/stats.rst @@ -94,7 +94,7 @@ Get all stats from a given domain:: >>> stats.get_stats('pages_crawled', domain='example.com') {'pages_crawled': 1238, 'start_time': datetime.datetime(2009, 7, 14, 21, 47, 28, 977139)} -.. _topics-stats-api: +.. _topics-stats-ref: Stats Collector API =================== diff --git a/docs/topics/telnetconsole.rst b/docs/topics/telnetconsole.rst index e2126b102..af19181ae 100644 --- a/docs/topics/telnetconsole.rst +++ b/docs/topics/telnetconsole.rst @@ -11,9 +11,10 @@ Scrapy comes with a built-in telnet console for inspecting and controlling a Scrapy running process. The telnet console is just a regular python shell running inside the Scrapy process, so you can do literally anything from it. -The telnet console is a :ref:`built-in Scrapy extension ` which -comes enabled by default, but you can also disable it if you want. For more -information about the extension itself see :ref:`ref-extensions-telnetconsole`. +The telnet console is a :ref:`built-in Scrapy extension +` which comes enabled by default, but you can also +disable it if you want. For more information about the extension itself see +:ref:`topics-extensions-ref-telnetconsole`. .. highlight:: none diff --git a/docs/topics/webconsole.rst b/docs/topics/webconsole.rst index 646439122..71037c376 100644 --- a/docs/topics/webconsole.rst +++ b/docs/topics/webconsole.rst @@ -7,11 +7,12 @@ Web Console Scrapy comes with a built-in web server for monitoring and controlling a Scrapy running process. -The web console is :ref:`built-in Scrapy extension ` which -comes enabled by default, but you can also disable it if you're running tight -on memory. +The web console is :ref:`built-in Scrapy extension +` which comes enabled by default, but you can also +disable it if you're running tight on memory. -For more information about this extension see :ref:`ref-extensions-webconsole`. +For more information about this extension see +:ref:`topics-extensions-ref-webconsole`. Writing a web console extension ===============================