diff --git a/docs/_templates/index.html b/docs/_templates/index.html
index 323870d78..a8bf042d5 100644
--- a/docs/_templates/index.html
+++ b/docs/_templates/index.html
@@ -8,7 +8,7 @@
for an overview and tutorial
Using Scrapy
usage guide and key concepts
- API Reference
+
API Reference
all details about Scrapy stable API
Frequently Asked Questions
diff --git a/docs/_templates/layout.html b/docs/_templates/layout.html
index 4d731a7c2..a75908edc 100644
--- a/docs/_templates/layout.html
+++ b/docs/_templates/layout.html
@@ -5,7 +5,7 @@
Home |
Getting Started |
Using Scrapy |
- API reference |
+ API reference |
FAQ |
Search
diff --git a/docs/contents.rst b/docs/contents.rst
index 605176260..6a748f364 100644
--- a/docs/contents.rst
+++ b/docs/contents.rst
@@ -8,6 +8,6 @@ Scrapy documentation contents
intro/index
topics/index
- ref/index
+ reference
faq
experimental/index
diff --git a/docs/faq.rst b/docs/faq.rst
index 8a1fd8667..fb2a7bab0 100644
--- a/docs/faq.rst
+++ b/docs/faq.rst
@@ -69,7 +69,7 @@ You need to install `pywin32`_ because of `this Twisted bug`_.
How can I simulate a user login in my spider?
---------------------------------------------
-See :ref:`ref-request-userlogin`.
+See :ref:`topics-request-response-ref-request-userlogin`.
Can I crawl in depth-first order instead of breadth-first order?
----------------------------------------------------------------
diff --git a/docs/ref/downloader-middleware.rst b/docs/ref/downloader-middleware.rst
deleted file mode 100644
index 976e32c04..000000000
--- a/docs/ref/downloader-middleware.rst
+++ /dev/null
@@ -1,71 +0,0 @@
-.. _ref-downloader-middleware:
-
-========================================
-Built-in downloader middleware reference
-========================================
-
-This page describes all downloader middleware components that come with
-Scrapy. For information on how to use them and how to write your own downloader
-middleware, see the :ref:`downloader middleware usage guide
-`.
-
-For a list of the components enabled by default (and their orders) see the
-:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting.
-
-Available downloader middlewares
-================================
-
-DefaultHeadersMiddleware
-------------------------
-
-.. module:: scrapy.contrib.downloadermiddleware.defaultheaders
- :synopsis: Default Headers Downloader Middleware
-
-.. class:: DefaultHeadersMiddleware
-
- This middleware sets all default requests headers specified in the
- :setting:`DEFAULT_REQUEST_HEADERS` setting.
-
-DebugMiddleware
----------------
-
-.. module:: scrapy.contrib.downloadermiddleware.debug
- :synopsis: Downloader middlewares for debugging
-
-.. class:: DebugMiddleware
-
- This is a convenient middleware to inspect what's passing through the
- downloader middleware. It logs all requests and responses catched by the
- middleware component methods. This middleware does not use any settings and
- does not come enabled by default. Instead, it's meant to be inserted at the
- point of the middleware that you want to inspect.
-
-HttpCacheMiddleware
--------------------
-
-.. module:: scrapy.contrib.downloadermiddleware.httpcache
- :synopsis: HTTP Cache downloader middleware
-
-.. class:: HttpCacheMiddleware
-
- This middleware provides low-level cache to all HTTP requests and responses.
- Every request and its corresponding response are cached and then, when that
- same request is seen again, the response is returned without transferring
- anything from the Internet.
-
- The HTTP cache is useful for testing spiders faster (without having to wait for
- downloads every time) and for trying your spider off-line when you don't have
- an Internet connection.
-
- The :class:`HttpCacheMiddleware` can be configured through the following
- settings (see the settings documentation for more info):
-
- * :setting:`HTTPCACHE_DIR` - this one actually enables the cache besides
- settings the cache dir
- * :setting:`HTTPCACHE_IGNORE_MISSING` - ignoring missing requests instead
- of downloading them
- * :setting:`HTTPCACHE_SECTORIZE` - split HTTP cache in several directories
- (for performance reasons)
- * :setting:`HTTPCACHE_EXPIRATION_SECS` - how many secs until the cache is
- considered out of date
-
diff --git a/docs/ref/extension-manager.rst b/docs/ref/extension-manager.rst
deleted file mode 100644
index 2e0d18dee..000000000
--- a/docs/ref/extension-manager.rst
+++ /dev/null
@@ -1,62 +0,0 @@
-.. _ref-extension-manager:
-
-=================
-Extension Manager
-=================
-
-.. module:: scrapy.extension
- :synopsis: The extension manager
-
-The Extension Manager is responsible for loading and keeping track of installed
-extensions and it's configured through the :setting:`EXTENSIONS` setting which
-contains a dictionary of all available extensions and their order similar to
-how you :ref:`configure the downloader middlewares
-`.
-
-.. class:: ExtensionManager
-
- The extension manager is a singleton object, which is instantiated at module
- loading time and can be accessed like this::
-
- from scrapy.extension import extensions
-
- .. attribute:: loaded
-
- A boolean which is True if extensions are already loaded or False if
- they're not.
-
- .. attribute:: enabled
-
- A dict with the enabled extensions. The keys are the extension class names,
- and the values are the extension objects. Example::
-
- >>> from scrapy.extension import extensions
- >>> extensions.load()
- >>> print extensions.enabled
- {'CoreStats': ,
- 'WebConsoke': ,
- ...
-
- .. attribute:: disabled
-
- A dict with the disabled extensions. The keys are the extension class names,
- and the values are the extension class paths (because objects are never
- instantiated for disabled extensions). Example::
-
- >>> from scrapy.extension import extensions
- >>> extensions.load()
- >>> print extensions.disabled
- {'MemoryDebugger': 'scrapy.contrib.webconsole.stats.MemoryDebugger',
- 'SpiderProfiler': 'scrapy.contrib.spider.profiler.SpiderProfiler',
- ...
-
- .. method:: load()
-
- Load the available extensions configured in the :setting:`EXTENSIONS`
- setting. On a standard run, this method is usually called by the Execution
- Manager, but you may need to call it explicitly if you're dealing with
- code outside Scrapy.
-
- .. method:: reload()
-
- Reload the available extensions. See :meth:`load`.
diff --git a/docs/ref/extensions.rst b/docs/ref/extensions.rst
deleted file mode 100644
index d8c467959..000000000
--- a/docs/ref/extensions.rst
+++ /dev/null
@@ -1,247 +0,0 @@
-.. _ref-extensions:
-
-=============================
-Built-in extensions reference
-=============================
-
-This document explains all extensions that come with Scrapy. For information on
-how to use them and how to write your own extensions, see the :ref:`extensions
-usage guide `.
-
-
-General purpose extensions
-==========================
-
-Core Stats extension
---------------------
-
-.. module:: scrapy.stats.corestats
- :synopsis: Core stats collection
-
-.. class:: scrapy.stats.corestats.CoreStats
-
-Enable the collection of core statistics, provided the stats collection are
-enabled (see :ref:`topics-stats`).
-
-.. _ref-extensions-webconsole:
-
-Web console extension
----------------------
-
-.. module:: scrapy.management.web
- :synopsis: Web management console
-
-.. class:: scrapy.management.web.WebConsole
-
-Provides an extensible web server for managing a Scrapy process. It's enabled
-by the :setting:`WEBCONSOLE_ENABLED` setting. The server will listen in the
-port specified in :setting:`WEBCONSOLE_PORT`, and will log to the file
-specified in :setting:`WEBCONSOLE_LOGFILE`.
-
-The web server is designed to be extended by other extensions which can add
-their own management web interfaces.
-
-See also :ref:`topics-webconsole` for information on how to write your own web
-console extension, and "Web console extensions" below for a list of available
-built-in (web console) extensions.
-
-.. _ref-extensions-telnetconsole:
-
-Telnet console extension
-------------------------
-
-.. module:: scrapy.management.telnet
- :synopsis: Telnet management console
-
-.. class:: scrapy.management.telnet.TelnetConsole
-
-Provides a telnet console for getting into a Python interpreter inside the
-currently running Scrapy process, which can be very useful for debugging.
-
-The telnet console must be enabled by the :setting:`TELNETCONSOLE_ENABLED`
-setting, and the server will listen in the port specified in
-:setting:`WEBCONSOLE_PORT`.
-
-Spider reloader extension
--------------------------
-
-.. module:: scrapy.contrib.spider.reloader
- :synopsis: Spider reloader extension
-
-.. class:: scrapy.contrib.spider.reloader.SpiderReloader
-
-Reload spider objects once they've finished scraping, to release the resources
-and references to other objects they may hold.
-
-.. _ref-extensions-memusage:
-
-Memory usage extension
-----------------------
-
-.. module:: scrapy.contrib.memusage
- :synopsis: Memory usage extension
-
-.. class:: scrapy.contrib.memusage.MemoryUsage
-
-Allows monitoring the memory used by a Scrapy process and:
-
-1, send a notification email when it exceeds a certain value
-2. terminate the Scrapy process when it exceeds a certain value
-
-The notification emails can be triggered when a certain warning value is
-reached (:setting:`MEMUSAGE_WARNING_MB`) and when the maximum value is reached
-(:setting:`MEMUSAGE_LIMIT_MB`) which will also cause the Scrapy process to be
-terminated.
-
-This extension is enabled by the :setting:`MEMUSAGE_ENABLED` setting and
-can be configured with the following settings:
-
-* :setting:`MEMUSAGE_LIMIT_MB`
-* :setting:`MEMUSAGE_WARNING_MB`
-* :setting:`MEMUSAGE_NOTIFY_MAIL`
-* :setting:`MEMUSAGE_REPORT`
-
-Memory debugger extension
--------------------------
-
-.. module:: scrapy.contrib.memdebug
- :synopsis: Memory debugger extension
-
-.. class:: scrapy.contrib.memdebug.MemoryDebugger
-
-A memory debugger which collects some info about objects uncollected by the
-garbage collector and libxml2 memory leaks. To enable this extension turn on
-the :setting:`MEMDEBUG_ENABLED` setting. The report will be printed to standard
-output. If the :setting:`MEMDEBUG_NOTIFY` setting contains a list of emails the
-report will also be sent to those addresses.
-
-Close domain extension
-----------------------
-
-.. module:: scrapy.contrib.closedomain
- :synopsis: Close domain extension
-
-.. class:: scrapy.contrib.closedomain.CloseDomain
-
-Closes a domain/spider automatically when some conditions are met, using a
-specific closing reason for each condition.
-
-The conditions for closing a domain can be configured through the following
-settings. Other conditions will be supported in the future.
-
-.. setting:: CLOSEDOMAIN_TIMEOUT
-
-CLOSEDOMAIN_TIMEOUT
-~~~~~~~~~~~~~~~~~~~
-
-Default: ``0``
-
-An integer which specifies a number of seconds. If the domain remains open for
-more than that number of second, it will be automatically closed with the
-reason ``closedomain_timeout``. If zero (or non set) domains won't be closed by
-timeout.
-
-.. setting:: CLOSEDOMAIN_ITEMPASSED
-
-CLOSEDOMAIN_ITEMPASSED
-~~~~~~~~~~~~~~~~~~~~~~
-
-Default: ``0``
-
-An integer which specifies a number of items. If the spider scrapes more than
-that amount if items and those items are passed by the item pipeline, the
-domain will be closed with the reason ``closedomain_itempassed``. If zero (or
-non set) domains won't be closed by number of passed items.
-
-Stack trace dump extension
----------------------------
-
-.. module:: scrapy.contrib.debug
- :synopsis: Extensions for debugging Scrapy
-
-.. class:: scrapy.contrib.debug.StackTraceDump
-
-Adds a `SIGUSR1`_ signal handler which dumps the stack trace of a runnning
-Scrapy process when a ``SIGUSR1`` signal is catched. After the stack trace is
-dumped, the Scrapy process continues to run normally.
-
-The stack trace is sent to standard output, or to the Scrapy log file if
-:setting:`LOG_STDOUT` is enabled.
-
-This extension only works on POSIX-compliant platforms (ie. not Windows).
-
-.. _SIGUSR1: http://en.wikipedia.org/wiki/SIGUSR1_and_SIGUSR2
-
-StatsMailer extension
----------------------
-
-.. module:: scrapy.contrib.statsmailer
- :synopsis: StatsMailer extension
-
-.. class:: scrapy.contrib.statsmailer.StatsMailer
-
-This simple extension can be used to send a notification email every time a
-domain has finished scraping, including the Scrapy stats collected. The email
-will be sent to all recipients specified in the :setting:`STATSMAILER_RCPTS`
-setting.
-
-Web console extensions
-======================
-
-.. module:: scrapy.contrib.webconsole
- :synopsis: Contains most built-in web console extensions
-
-Here is a list of built-in web console extensions. For clarity "web console
-extension" is abbreviated as "WC extension".
-
-For more information see the see the :ref:`web console documentation
-`.
-
-Scheduler queue WC extension
-----------------------------
-
-.. module:: scrapy.contrib.webconsole.scheduler
- :synopsis: Scheduler queue web console extension
-
-.. class:: scrapy.contrib.webconsole.scheduler.SchedulerQueue
-
-Display a list of all pending Requests in the Scheduler queue, grouped by
-domain/spider.
-
-Spider live stats WC extension
-------------------------------
-
-.. module:: scrapy.contrib.webconsole.livestats
- :synopsis: Spider live stats web console extension
-
-.. class:: scrapy.contrib.webconsole.livestats.LiveStats
-
-Display a table with stats of all spider crawled by the current Scrapy run,
-including:
-
-* Number of items scraped
-* Number of pages crawled
-* Number of pending requests in the scheduler
-* Number of pending requests in the downloader queue
-* Number of requests currently being downloaded
-
-Engine status WC extension
----------------------------
-
-.. module:: scrapy.contrib.webconsole.enginestatus
- :synopsis: Engine status web console extension
-
-.. class:: scrapy.contrib.webconsole.enginestatus.EngineStatus
-
-Display the current status of the Scrapy Engine, which is just the output of
-the Scrapy engine ``getstatus()`` method.
-
-Stats collector dump WC extension
-----------------------------------
-
-.. module:: scrapy.contrib.webconsole.stats
- :synopsis: Stats dump web console extension
-
-.. class:: scrapy.contrib.webconsole.stats.StatsDump
-
-Display the stats collected so far by the stats collector.
diff --git a/docs/ref/index.rst b/docs/ref/index.rst
deleted file mode 100644
index eb0adafde..000000000
--- a/docs/ref/index.rst
+++ /dev/null
@@ -1,26 +0,0 @@
-.. _ref-index:
-
-API Reference
-=============
-
-This section documents the Scrapy |version| API. For more information see :ref:`misc-api-stability`.
-
-.. toctree::
- :maxdepth: 1
-
- request-response
- spiders
- selectors
- settings
- signals
- exceptions
- logging
- email
- extension-manager
- extensions
- downloader-middleware
- spider-middleware
- scheduler-middleware
- link-extractors
-
-* :ref:`topics-stats-api`
diff --git a/docs/ref/link-extractors.rst b/docs/ref/link-extractors.rst
deleted file mode 100644
index 103b1253a..000000000
--- a/docs/ref/link-extractors.rst
+++ /dev/null
@@ -1,121 +0,0 @@
-.. _ref-link-extractors:
-
-=========================
-Available Link Extractors
-=========================
-
-.. module:: scrapy.contrib.linkextractors
- :synopsis: Link extractors classes
-
-All available link extractors classes bundled with Scrapy are provided in the
-:mod:`scrapy.contrib.linkextractors` module.
-
-.. module:: scrapy.contrib.linkextractors.sgml
- :synopsis: SGMLParser-based link extractors
-
-SgmlLinkExtractor
-=================
-
-.. class:: SgmlLinkExtractor(allow=(), deny=(), allow_domains=(), deny_domains=(), restrict_xpaths(), tags=('a', 'area'), attrs=('href'), canonicalize=True, unique=True, process_value=None)
-
- The SgmlLinkExtractor extends the base :class:`BaseSgmlLinkExtractor` by
- providing additional filters that you can specify to extract links,
- including regular expressions patterns that the links must match to be
- extracted. All those filters are configured through these constructor
- parameters:
-
- :param allow: a single regular expression (or list of regular expressions)
- that the (absolute) urls must match in order to be extracted. If not
- given (or empty), it will match all links.
- :type allow: a regular expression (or list of)
-
- :param deny: a single regular expression (or list of regular expressions)
- that the (absolute) urls must match in order to be excluded (ie. not
- extracted). It has precedence over the ``allow`` parameter. If not
- given (or empty) it won't exclude any links.
- :type allow: a regular expression (or list of)
-
- :param allow_domains: is single value or a list of string containing
- domains which will be considered for extracting the links
- :type allow: str or list
-
- :param deny_domains: is single value or a list of strings containing
- domains which which won't be considered for extracting the links
- :type allow: str or list
-
- :param restrict_xpaths: is a XPath (or list of XPath's) which defines
- regions inside the response where links should be extracted from.
- If given, only the text selected by those XPath will be scanned for
- links. See examples below.
- :type restrict_xpaths: str or list
-
- :param tags: a tag or a list of tags to consider when extracting links.
- Defaults to ``('a', 'area')``.
- :type tags: str or list
-
- :param attrs: list of attrbitues which should be considered when looking
- for links to extract (only for those tags specified in the ``tags``
- parameter). Defaults to ``('href',)``
- :type attrs: boolean
-
- :param canonicalize: canonicalize each extracted url (using
- scrapy.utils.url.canonicalize_url). Defaults to ``True``.
- :type canonicalize: boolean
-
- :param unique: whether duplicate filtering should be applied to extracted
- links.
- :type unique: boolean
-
- :param process_value: see ``process_value`` argument of
- :class:`LinkExtractor` class constructor
- :type process_value: boolean
-
-BaseSgmlLinkExtractor
-=====================
-
-.. class:: BaseSgmlLinkExtractor(tag="a", href="href", unique=False, process_value=None)
-
- The purpose of this Link Extractor is only to serve as a base class for the
- :class:`SgmlLinkExtractor`. You should use that one instead.
-
- The constructor arguments are:
-
- :param tag: either a string (with the name of a tag) or a function that
- receives a tag name and returns ``True`` if links should be extracted
- from those tag, or ``False`` if they shouldn't. Defaults to ``'a'``.
- request (once its downloaded) as its first parameter. For more
- information see :ref:`ref-request-callback-arguments` below.
- :type tag: str or callable
-
- :param attr: either string (with the name of a tag attribute), or a
- function that receives a an attribute name and returns ``True`` if
- links should be extracted from it, or ``False`` if the shouldn't.
- Defaults to ``href``.
- :type attr: str or callable
-
- :param unique: is a boolean that specifies if a duplicate filtering should
- be applied to links extracted.
- :type unique: boolean
-
- :param process_value: a function which receives each value extracted from
- the tag and attributes scanned and can modify the value and return a
- new one, or return ``None`` to ignore the link altogether. If not
- given, ``process_value`` defaults to ``lambda x: x``.
-
- .. highlight:: html
-
- For example, to extract links from this code::
-
- Link text
-
- .. highlight:: python
-
- You can use the following function in ``process_value``::
-
- def process_value(value):
- m = re.search("javascript:goToPage\('(.*?)'", value)
- if m:
- return m.group(1)
-
- :type process_value: callable
-
diff --git a/docs/ref/selectors.rst b/docs/ref/selectors.rst
deleted file mode 100644
index 5be263cdc..000000000
--- a/docs/ref/selectors.rst
+++ /dev/null
@@ -1,189 +0,0 @@
-.. _ref-selectors:
-
-===================
-XPath Selectors API
-===================
-
-.. module:: scrapy.xpath
- :synopsis: XPath selectors classes
-
-There are two types of selectors bundled with Scrapy:
-:class:`HtmlXPathSelector` and :class:`XmlXPathSelector`. Both of them
-implement the same :class:`XPathSelector` interface. The only different is that
-one is used to process HTML data and the other XML data.
-
-XPathSelector objects
-=====================
-
-.. class:: XPathSelector(response)
-
- A :class:`XPathSelector` object is a wrapper over response to select
- certain parts of its content.
-
- A :class:`Request` object represents an HTTP request, which is usually
- generated in the Spider and executed by the Downloader, and thus generating
- a :class:`Response`.
-
- ``url`` is a :class:`~scrapy.http.Response` object that will be used for
- selecting and extracting data
-
-
-XPathSelector Methods
----------------------
-
-.. method:: XPathSelector.select(xpath)
-
- Apply the given XPath relative to this XPathSelector and return a list
- of :class:`XPathSelector` objects (ie. a :class:`XPathSelectorList`) with
- the result.
-
- ``xpath`` is a string containing the XPath to apply
-
-.. method:: XPathSelector.re(regex)
-
- Apply the given regex and return a list of unicode strings with the
- matches.
-
- ``regex`` can be either a compiled regular expression or a string which
- will be compiled to a regular expression using ``re.compile(regex)``
-
-.. method:: XPathSelector.extract()
-
- Return a unicode string with the content of this :class:`XPathSelector`
- object.
-
-.. method:: XPathSelector.extract_unquoted()
-
- Return a unicode string with the content of this :class:`XPathSelector`
- without entities or CDATA. This method is intended to be use for text-only
- selectors, like ``//h1/text()`` (but not ``//h1``). If it's used for
- :class:`XPathSelector` objects which don't select a textual content (ie. if
- they contain tags), the output of this method is undefined.
-
-.. method:: XPathSelector.register_namespace(prefix, uri)
-
- Register the given namespace to be used in this :class:`XPathSelector`.
- Without registering namespaces you can't select or extract data from
- non-standard namespaces. See examples below.
-
-.. method:: XPathSelector.__nonzero__()
-
- Returns ``True`` if there is any real content selected by this
- :class:`XPathSelector` or ``False`` otherwise. In other words, the boolean
- value of an XPathSelector is given by the contents it selects.
-
-XPathSelectorList objects
-=========================
-
-.. class:: XPathSelectorList
-
- The :class:`XPathSelectorList` class is subclass of the builtin ``list``
- class, which provides a few additional methods.
-
-
-XPathSelectorList Methods
--------------------------
-
-.. method:: XPathSelectorList.select(xpath)
-
- Call the :meth:`XPathSelector.re` method for all :class:`XPathSelector`
- objects in this list and return their results flattened, as new
- :class:`XPathSelectorList`.
-
- ``xpath`` is the same argument as the one in :meth:`XPathSelector.x`
-
-.. method:: XPathSelector.re(regex)
-
- Call the :meth:`XPathSelector.re` method for all :class:`XPathSelector`
- objects in this list and return their results flattened, as a list of
- unicode strings.
-
- ``regex`` is the same argument as the one in :meth:`XPathSelector.re`
-
-.. method:: XPathSelector.extract()
-
- Call the :meth:`XPathSelector.re` method for all :class:`XPathSelector`
- objects in this list and return their results flattened, as a list of
- unicode strings.
-
-.. method:: XPathSelector.extract_unquoted()
-
- Call the :meth:`XPathSelector.extract_unoquoted` method for all
- :class:`XPathSelector` objects in this list and return their results
- flattened, as a list of unicode strings. This method should not be applied
- to all kinds of XPathSelectors. For more info see
- :meth:`XPathSelector.extract_unoquoted`.
-
-HtmlXPathSelector objects
-=========================
-
-.. class:: HtmlXPathSelector(response)
-
- A subclass of :class:`XPathSelector` for working with HTML content. It uses
- the `libxml2`_ HTML parser. See the :class:`XPathSelector` API for more info.
-
-.. _libxml2: http://xmlsoft.org/
-
-HtmlXPathSelector examples
---------------------------
-
-Here's a couple of :class:`HtmlXPathSelector` examples to illustrate several
-concepts. In all cases we assume there is already a :class:`HtmlPathSelector`
-instanced with a :class:`~scrapy.http.Response` object like this::
-
- x = HtmlXPathSelector(html_response)
-
-1. Select all ```` elements from a HTML response body, returning a list of
- :class:`XPathSelector` objects (ie. a :class:`XPathSelectorList` object)::
-
- x.select("//h1")
-
-2. Extract the text of all ```` elements from a HTML response body,
- returning a list of unicode strings::
-
- x.select("//h1").extract() # this includes the h1 tag
- x.select("//h1/text()").extract() # this excludes the h1 tag
-
-3. Iterate over all ```` tags and print their class attribute::
-
- for node in x.select("//p"):
- ... print node.select("@href")
-
-4. Extract textual data from all `` `` tags without entities, as a list of
- unicode strings::
-
- x.select("//p/text()").extract_unquoted()
-
- # the following line is wrong. extract_unquoted() should only be used
- # with textual XPathSelectors
- x.select("//p").extract_unquoted() # it may work but output is unpredictable
-
-XmlXPathSelector objects
-========================
-
-.. class:: XmlXPathSelector(response)
-
- A subclass of :class:`XPathSelector` for working with XML content. It uses
- the `libxml2`_ XML parser. See the :class:`XPathSelector` API for more info.
-
-XmlXPathSelector examples
--------------------------
-
-Here's a couple of :class:`XmlXPathSelector` examples to illustrate several
-concepts. In all cases we assume there is already a :class:`XmlPathSelector`
-instanced with a :class:`~scrapy.http.Response` object like this::
-
- x = HtmlXPathSelector(xml_response)
-
-1. Select all ```` elements from a XML response body, returning a list of
- :class:`XPathSelector` objects (ie. a :class:`XPathSelectorList` object)::
-
- x.select("//h1")
-
-2. Extract all prices from a `Google Base XML feed`_ which requires registering
- a namespace::
-
- x.register_namespace("g", "http://base.google.com/ns/1.0")
- x.select("//g:price").extract()
-
-.. _Google Base XML feed: http://base.google.com/support/bin/answer.py?hl=en&answer=59461
diff --git a/docs/ref/settings.rst b/docs/ref/settings.rst
deleted file mode 100644
index b3bc77840..000000000
--- a/docs/ref/settings.rst
+++ /dev/null
@@ -1,886 +0,0 @@
-.. _settings:
-
-Available Settings
-==================
-
-Here's a list of all available Scrapy settings, in alphabetical order, along
-with their default values and the scope where they apply.
-
-The scope, where available, shows where the setting is being used, if it's tied
-to any particular component. In that case the module of that component will be
-shown, typically an extension, middleware or pipeline. It also means that the
-component must be enabled in order for the setting to have any effect.
-
-.. setting:: BOT_NAME
-
-BOT_NAME
---------
-
-Default: ``scrapybot``
-
-The name of the bot implemented by this Scrapy project. This will be used to
-construct the User-Agent by default, and also for logging.
-
-.. setting:: BOT_VERSION
-
-BOT_VERSION
------------
-
-Default: ``1.0``
-
-The version of the bot implemented by this Scrapy project. This will be used to
-construct the User-Agent by default.
-
-.. setting:: HTTPCACHE_DIR
-
-HTTPCACHE_DIR
--------------
-
-Default: ``''`` (empty string)
-
-The directory to use for storing the (low-level) HTTP cache. If empty the HTTP
-cache will be disabled.
-
-.. setting:: HTTPCACHE_EXPIRATION_SECS
-
-HTTPCACHE_EXPIRATION_SECS
--------------------------
-
-Default: ``0``
-
-Number of seconds to use for HTTP cache expiration. Requests that were cached
-before this time will be re-downloaded. If zero, cached requests will always
-expire. Negative numbers means requests will never expire.
-
-.. setting:: HTTPCACHE_IGNORE_MISSING
-
-HTTPCACHE_IGNORE_MISSING
-------------------------
-
-Default: ``False``
-
-If enabled, requests not found in the cache will be ignored instead of downloaded.
-
-.. setting:: HTTPCACHE_SECTORIZE
-
-HTTPCACHE_SECTORIZE
--------------------
-
-Default: ``True``
-
-Whether to split HTTP cache storage in several dirs for performance.
-
-.. setting:: COMMANDS_MODULE
-
-COMMANDS_MODULE
----------------
-
-Default: ``''`` (empty string)
-
-A module to use for looking for custom Scrapy commands. This is used to add
-custom command for your Scrapy project.
-
-Example::
-
- COMMANDS_MODULE = 'mybot.commands'
-
-.. setting:: COMMANDS_SETTINGS_MODULE
-
-COMMANDS_SETTINGS_MODULE
-------------------------
-
-Default: ``''`` (empty string)
-
-A module to use for looking for custom Scrapy command settings.
-
-Example::
-
- COMMANDS_SETTINGS_MODULE = 'mybot.conf.commands'
-
-.. setting:: CONCURRENT_DOMAINS
-
-CONCURRENT_DOMAINS
-------------------
-
-Default: ``8``
-
-Maximum number of domains to scrape in parallel.
-
-.. setting:: CONCURRENT_ITEMS
-
-CONCURRENT_ITEMS
-----------------
-
-Default: ``100``
-
-Maximum number of concurrent items (per response) to process in parallel in the
-Item Processor (also known as the Item Pipeline).
-
-.. setting:: COOKIES_DEBUG
-
-COOKIES_DEBUG
--------------
-
-Default: ``False``
-
-Enable debugging message of Cookies Downloader Middleware.
-
-.. setting:: DEFAULT_ITEM_CLASS
-
-DEFAULT_ITEM_CLASS
-------------------
-
-Default: ``'scrapy.item.ScrapedItem'``
-
-The default class that will be used for instantiating items in the :ref:`the
-Scrapy shell `.
-
-.. setting:: DEFAULT_REQUEST_HEADERS
-
-DEFAULT_REQUEST_HEADERS
------------------------
-
-Default::
-
- {
- 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
- 'Accept-Language': 'en',
- }
-
-The default headers used for Scrapy HTTP Requests. They're populated in the
-:class:`~scrapy.contrib.downloadermiddleware.defaultheaders.DefaultHeadersMiddleware`.
-
-.. setting:: DEFAULT_SPIDER
-
-DEFAULT_SPIDER
---------------
-
-Default: ``None``
-
-The default spider class that will be instantiated for URLs for which no
-specific spider is found. This class must have a constructor which receives as
-only parameter the domain name of the given URL.
-
-.. setting:: DEPTH_LIMIT
-
-DEPTH_LIMIT
------------
-
-Default: ``0``
-
-The maximum depth that will be allowed to crawl for any site. If zero, no limit
-will be imposed.
-
-.. setting:: DEPTH_STATS
-
-DEPTH_STATS
------------
-
-Default: ``True``
-
-Whether to collect depth stats.
-
-.. setting:: DOMAIN_SCHEDULER
-
-DOMAIN_SCHEDULER
-----------------
-
-Default: ``'scrapy.contrib.domainsch.FifoDomainScheduler'``
-
-The Domain Scheduler to use. The domain scheduler returns the next domain
-(spider) to scrape.
-
-.. setting:: DOWNLOADER_DEBUG
-
-DOWNLOADER_DEBUG
-----------------
-
-Default: ``False``
-
-Whether to enable the Downloader debugging mode.
-
-.. setting:: DOWNLOADER_MIDDLEWARES
-
-DOWNLOADER_MIDDLEWARES
-----------------------
-
-Default:: ``{}``
-
-A dict containing the downloader middlewares enabled in your project, and their
-orders. For more info see :ref:`topics-downloader-middleware-setting`.
-
-.. setting:: DOWNLOADER_MIDDLEWARES_BASE
-
-DOWNLOADER_MIDDLEWARES_BASE
----------------------------
-
-Default::
-
- {
- 'scrapy.contrib.downloadermiddleware.robotstxt.RobotsTxtMiddleware': 100,
- 'scrapy.contrib.downloadermiddleware.httpauth.HttpAuthMiddleware': 300,
- 'scrapy.contrib.downloadermiddleware.useragent.UserAgentMiddleware': 400,
- 'scrapy.contrib.downloadermiddleware.retry.RetryMiddleware': 500,
- 'scrapy.contrib.downloadermiddleware.defaultheaders.DefaultHeadersMiddleware': 550,
- 'scrapy.contrib.downloadermiddleware.redirect.RedirectMiddleware': 600,
- 'scrapy.contrib.downloadermiddleware.cookies.CookiesMiddleware': 700,
- 'scrapy.contrib.downloadermiddleware.httpcompression.HttpCompressionMiddleware': 800,
- 'scrapy.contrib.downloadermiddleware.stats.DownloaderStats': 850,
- 'scrapy.contrib.downloadermiddleware.httpcache.HttpCacheMiddleware': 900,
- }
-
-A dict containing the downloader middlewares enabled by default in Scrapy. You
-should never modify this setting in your project, modify
-:setting:`DOWNLOADER_MIDDLEWARES` instead. For more info see
-:ref:`topics-downloader-middleware-setting`.
-
-.. setting:: DOWNLOADER_STATS
-
-DOWNLOADER_STATS
-----------------
-
-Default: ``True``
-
-Whether to enable downloader stats collection.
-
-.. setting:: DOWNLOAD_DELAY
-
-DOWNLOAD_DELAY
---------------
-
-Default: ``0``
-
-The amount of time (in secs) that the downloader should wait before downloading
-consecutive pages from the same spider. This can be used to throttle the
-crawling speed to avoid hitting servers too hard. Decimal numbers are
-supported. Example::
-
- DOWNLOAD_DELAY = 0.25 # 250 ms of delay
-
-.. setting:: DOWNLOAD_TIMEOUT
-
-DOWNLOAD_TIMEOUT
-----------------
-
-Default: ``180``
-
-The amount of time (in secs) that the downloader will wait before timing out.
-
-.. setting:: DUPEFILTER_CLASS
-
-DUPEFILTER_CLASS
-----------------
-
-Default: ``'scrapy.contrib.dupefilter.RequestFingerprintDupeFilter'``
-
-The class used to detect and filter duplicate requests.
-
-The default (``RequestFingerprintDupeFilter``) filters based on request fingerprint
-(using ``scrapy.utils.request.request_fingerprint``) and grouping per domain.
-
-.. setting:: EXTENSIONS
-
-EXTENSIONS
-----------
-
-Default:: ``{}``
-
-A dict containing the extensions enabled in your project, and their orders.
-
-.. setting:: EXTENSIONS_BASE
-
-EXTENSIONS_BASE
----------------
-
-Default::
-
- {
- 'scrapy.stats.corestats.CoreStats': 0,
- 'scrapy.management.web.WebConsole': 0,
- 'scrapy.management.telnet.TelnetConsole': 0,
- 'scrapy.contrib.webconsole.scheduler.SchedulerQueue': 0,
- 'scrapy.contrib.webconsole.livestats.LiveStats': 0,
- 'scrapy.contrib.webconsole.spiderctl.Spiderctl': 0,
- 'scrapy.contrib.webconsole.enginestatus.EngineStatus': 0,
- 'scrapy.contrib.webconsole.stats.StatsDump': 0,
- 'scrapy.contrib.spider.reloader.SpiderReloader': 0,
- 'scrapy.contrib.memusage.MemoryUsage': 0,
- 'scrapy.contrib.memdebug.MemoryDebugger': 0,
- 'scrapy.contrib.closedomain.CloseDomain': 0,
- 'scrapy.contrib.debug.StackTraceDump': 0,
- }
-
-The list of available extensions. Keep in mind that some of them need need to
-be enabled through a setting. By default, this setting contains all stable
-built-in extensions.
-
-For more information See the :ref:`extensions user guide `
-and the :ref:`list of available extensions `.
-
-.. setting:: GROUPSETTINGS_ENABLED
-
-GROUPSETTINGS_ENABLED
----------------------
-
-Default: ``False``
-
-Whether to enable group settings where spiders pull their settings from.
-
-.. setting:: GROUPSETTINGS_MODULE
-
-GROUPSETTINGS_MODULE
---------------------
-
-Default: ``''`` (empty string)
-
-The module to use for pulling settings from, if the group settings is enabled.
-
-.. setting:: IMAGES_DIR
-
-IMAGES_DIR
-----------
-
-Default: ``None``
-
-Directory where :class:`ImagesPipeline` will store its images.
-
-For more information see :ref:`topics-images`.
-
-.. setting:: IMAGES_EXPIRES
-
-IMAGES_EXPIRES
---------------
-
-Default = 90
-
-Number of days for an image to be considered `expired` (downloaded again) in
-:class:`ImagesPipeline`.
-
-For more information see :ref:`topics-images`.
-
-.. setting:: IMAGES_MIN_HEIGHT
-
-IMAGES_MIN_HEIGHT
------------------
-
-Default = 0
-
-Minimum height that an image is allowed to have in :class:`ImagesPipeline`.
-
-For more information see :ref:`topics-images-size`.
-
-.. setting:: IMAGES_MIN_WIDTH
-
-IMAGES_MIN_WIDTH
------------------
-
-Default = 0
-
-Minimum width that an image is allowed to have in :class:`ImagesPipeline`.
-
-For more information see :ref:`topics-images-size`.
-
-.. setting:: ITEM_PIPELINES
-
-ITEM_PIPELINES
---------------
-
-Default: ``[]``
-
-The item pipelines to use (a list of classes).
-
-Example::
-
- ITEM_PIPELINES = [
- 'mybot.pipeline.validate.ValidateMyItem',
- 'mybot.pipeline.validate.StoreMyItem'
- ]
-
-.. setting:: LOG_ENABLED
-
-LOG_ENABLED
------------
-
-Default: ``True``
-
-Enable logging.
-
-.. setting:: LOG_STDOUT
-
-LOG_STDOUT
-----------
-
-Default: ``False``
-
-If enabled logging will be sent to standard output, otherwise standard error
-will be used.
-
-.. setting:: LOGFILE
-
-LOGFILE
--------
-
-Default: ``None``
-
-File name to use for logging output. If None, standard input (or error) will be
-used depending on the value of the LOG_STDOUT setting.
-
-.. setting:: LOGLEVEL
-
-LOGLEVEL
---------
-
-Default: ``'DEBUG'``
-
-Minimum level to log. Available levels are: SILENT, CRITICAL, ERROR, WARNING,
-INFO, DEBUG, TRACE
-
-.. setting:: MAIL_FROM
-
-MAIL_FROM
----------
-
-Default: ``'scrapy@localhost'``
-
-Email to use as sender address for sending emails using the :ref:`Scrapy e-mail
-sending facility `.
-
-.. setting:: MAIL_HOST
-
-MAIL_HOST
----------
-
-Default: ``'localhost'``
-
-Host to use for sending emails using the :ref:`Scrapy e-mail sending facility
-`.
-
-.. setting:: MEMDEBUG_ENABLED
-
-MEMDEBUG_ENABLED
-----------------
-
-Default: ``False``
-
-Whether to enable memory debugging.
-
-.. setting:: MEMDEBUG_NOTIFY
-
-MEMDEBUG_NOTIFY
----------------
-
-Default: ``[]``
-
-When memory debugging is enabled a memory report will be sent to the specified
-addresses if this setting is not empty, otherwise the report will be written to
-the log.
-
-Example::
-
- MEMDEBUG_NOTIFY = ['user@example.com']
-
-.. setting:: MEMUSAGE_ENABLED
-
-MEMUSAGE_ENABLED
-----------------
-
-Default: ``False``
-
-Scope: ``scrapy.contrib.memusage``
-
-Whether to enable the memory usage extension that will shutdown the Scrapy
-process when it exceeds a memory limit, and also notify by email when that
-happened.
-
-See :ref:`ref-extensions-memusage`.
-
-.. setting:: MEMUSAGE_LIMIT_MB
-
-MEMUSAGE_LIMIT_MB
------------------
-
-Default: ``0``
-
-Scope: ``scrapy.contrib.memusage``
-
-The maximum amount of memory to allow (in megabytes) before shutting down
-Scrapy (if MEMUSAGE_ENABLED is True). If zero, no check will be performed.
-
-See :ref:`ref-extensions-memusage`.
-
-.. setting:: MEMUSAGE_NOTIFY_MAIL
-
-MEMUSAGE_NOTIFY_MAIL
---------------------
-
-Default: ``False``
-
-Scope: ``scrapy.contrib.memusage``
-
-A list of emails to notify if the memory limit has been reached.
-
-Example::
-
- MEMUSAGE_NOTIFY_MAIL = ['user@example.com']
-
-See :ref:`ref-extensions-memusage`.
-
-.. setting:: MEMUSAGE_REPORT
-
-MEMUSAGE_REPORT
----------------
-
-Default: ``False``
-
-Scope: ``scrapy.contrib.memusage``
-
-Whether to send a memory usage report after each domain has been closed.
-
-See :ref:`ref-extensions-memusage`.
-
-.. setting:: MEMUSAGE_WARNING_MB
-
-MEMUSAGE_WARNING_MB
--------------------
-
-Default: ``0``
-
-Scope: ``scrapy.contrib.memusage``
-
-The maximum amount of memory to allow (in megabytes) before sending a warning
-email notifying about it. If zero, no warning will be produced.
-
-.. setting:: MYSQL_CONNECTION_SETTINGS
-
-MYSQL_CONNECTION_SETTINGS
--------------------------
-
-Default: ``{}``
-
-Scope: ``scrapy.utils.db.mysql_connect``
-
-Settings to use for MySQL connections performed through
-``scrapy.utils.db.mysql_connect``
-
-.. setting:: NEWSPIDER_MODULE
-
-NEWSPIDER_MODULE
-----------------
-
-Default: ``''``
-
-Module where to create new spiders using the ``genspider`` command.
-
-Example::
-
- NEWSPIDER_MODULE = 'mybot.spiders_dev'
-
-.. setting:: PROJECT_NAME
-
-PROJECT_NAME
-------------
-
-Default: ``Not Defined``
-
-The name of the current project. It matches the project module name as created
-by ``startproject`` command, and is only defined by project settings file.
-
-.. setting:: REDIRECT_MAX_TIMES
-
-REDIRECT_MAX_TIMES
-------------------
-
-Default: ``20``
-
-Defines the maximun times a request can be redirected. After this maximun the
-request's response is returned as is. We used Firefox default value for the
-same task.
-
-.. setting:: REDIRECT_MAX_METAREFRESH_DELAY
-
-REDIRECT_MAX_METAREFRESH_DELAY
-------------------------------
-
-Default: ``100``
-
-Some sites use meta-refresh for redirecting to a session expired page, so we
-restrict automatic redirection to a maximum delay (in seconds)
-
-.. setting:: REDIRECT_PRIORITY_ADJUST
-
-REDIRECT_PRIORITY_ADJUST
-------------------------------
-
-Default: ``+2``
-
-Adjust redirect request priority relative to original request.
-A negative priority adjust means more priority.
-
-.. setting:: REQUESTS_PER_DOMAIN
-
-REQUESTS_PER_DOMAIN
--------------------
-
-Default: ``8``
-
-Specifies how many concurrent (ie. simultaneous) requests will be performed per
-open spider.
-
-.. setting:: REQUESTS_QUEUE_SIZE
-
-REQUESTS_QUEUE_SIZE
--------------------
-
-Default: ``0``
-
-Scope: ``scrapy.contrib.spidermiddleware.limit``
-
-If non zero, it will be used as an upper limit for the amount of requests that
-can be scheduled per domain.
-
-.. setting:: ROBOTSTXT_OBEY
-
-ROBOTSTXT_OBEY
---------------
-
-Default: ``False``
-
-Scope: ``scrapy.contrib.downloadermiddleware.robotstxt``
-
-If enabled, Scrapy will respect robots.txt policies. For more information see
-:topic:`robotstxt`
-
-.. setting:: SCHEDULER
-
-SCHEDULER
----------
-
-Default: ``'scrapy.core.scheduler.Scheduler'``
-
-The scheduler to use for crawling.
-
-.. setting:: SCHEDULER_ORDER
-
-SCHEDULER_ORDER
----------------
-
-Default: ``'BFO'``
-
-Scope: ``scrapy.core.scheduler``
-
-The order to use for the crawling scheduler. Available orders are:
-
-* ``'BFO'``: `Breadth-first order`_ - typically consumes more memory but
- reaches most relevant pages earlier.
-
-* ``'DFO'``: `Depth-first order`_ - typically consumes less memory than
- but takes longer to reach most relevant pages.
-
-.. _Breadth-first order: http://en.wikipedia.org/wiki/Breadth-first_search
-.. _Depth-first order: http://en.wikipedia.org/wiki/Depth-first_search
-
-.. setting:: SCHEDULER_MIDDLEWARES
-
-SCHEDULER_MIDDLEWARES
----------------------
-
-Default:: ``{}``
-
-A dict containing the scheduler middlewares enabled in your project, and their
-orders.
-
-.. setting:: SCHEDULER_MIDDLEWARES_BASE
-
-SCHEDULER_MIDDLEWARES_BASE
---------------------------
-
-Default::
-
- SCHEDULER_MIDDLEWARES_BASE = {
- 'scrapy.contrib.schedulermiddleware.duplicatesfilter.DuplicatesFilterMiddleware': 500,
- }
-
-A dict containing the scheduler middlewares enabled by default in Scrapy. You
-should never modify this setting in your project, modify
-:setting:`SCHEDULER_MIDDLEWARES` instead.
-
-.. setting:: SPIDERPROFILER_ENABLED
-
-SPIDERPROFILER_ENABLED
-----------------------
-
-Default: ``False``
-
-Enable the spider profiler. Warning: this could have a big impact in
-performance.
-
-.. setting:: SPIDER_MIDDLEWARES
-
-SPIDER_MIDDLEWARES
-------------------
-
-Default:: ``{}``
-
-A dict containing the spider middlewares enabled in your project, and their
-orders. For more info see :ref:`topics-spider-middleware-setting`.
-
-.. setting:: SPIDER_MIDDLEWARES_BASE
-
-SPIDER_MIDDLEWARES_BASE
------------------------
-
-Default::
-
- {
- 'scrapy.contrib.spidermiddleware.httperror.HttpErrorMiddleware': 50,
- 'scrapy.contrib.itemsampler.ItemSamplerMiddleware': 100,
- 'scrapy.contrib.spidermiddleware.requestlimit.RequestLimitMiddleware': 200,
- 'scrapy.contrib.spidermiddleware.restrict.RestrictMiddleware': 300,
- 'scrapy.contrib.spidermiddleware.offsite.OffsiteMiddleware': 500,
- 'scrapy.contrib.spidermiddleware.referer.RefererMiddleware': 700,
- 'scrapy.contrib.spidermiddleware.urllength.UrlLengthMiddleware': 800,
- 'scrapy.contrib.spidermiddleware.depth.DepthMiddleware': 900,
- }
-
-A dict containing the spider middlewares enabled by default in Scrapy. You
-should never modify this setting in your project, modify
-:setting:`SPIDER_MIDDLEWARES` instead. For more info see
-:ref:`topics-spider-middleware-setting`.
-
-.. setting:: SPIDER_MODULES
-
-SPIDER_MODULES
---------------
-
-Default: ``[]``
-
-A list of modules where Scrapy will look for spiders.
-
-Example::
-
- SPIDER_MODULES = ['mybot.spiders_prod', 'mybot.spiders_dev']
-
-.. setting:: STATS_CLASS
-
-STATS_CLASS
------------
-
-Default: ``'scrapy.stats.collector.MemoryStatsCollector'``
-
-The class to use for collecting stats (must implement the Stats Collector API,
-or subclass the StatsCollector class).
-
-.. setting:: STATS_DUMP
-
-STATS_DUMP
-----------
-
-Default: ``False``
-
-Dump (to log) domain-specific stats collected when a domain is closed, and all
-global stats when the Scrapy process finishes (ie. when the engine is
-shutdown).
-
-.. setting:: STATS_ENABLED
-
-STATS_ENABLED
--------------
-
-Default: ``True``
-
-Enable stats collection.
-
-.. setting:: STATSMAILER_RCPTS
-
-STATSMAILER_RCPTS
------------------
-
-Default: ``[]`` (empty list)
-
-Send Scrapy stats after domains finish scrapy. See
-:class:`~scrapy.contrib.statsmailer.StatsMailer` for more info.
-
-.. setting:: TELNETCONSOLE_ENABLED
-
-TELNETCONSOLE_ENABLED
----------------------
-
-Default: ``True``
-
-Scope: ``scrapy.management.telnet``
-
-A boolean which specifies if the telnet management console will be enabled
-(provided its extension is also enabled).
-
-.. setting:: TELNETCONSOLE_PORT
-
-TELNETCONSOLE_PORT
-------------------
-
-Default: ``6023``
-
-The port to use for the telnet console. If set to ``None`` or ``0``, a
-dynamically assigned port is used. For more info see
-:ref:`topics-telnetconsole`.
-
-.. setting:: TEMPLATES_DIR
-
-TEMPLATES_DIR
--------------
-
-Default: ``templates`` dir inside scrapy module
-
-The directory where to look for template when creating new projects with
-scrapy-admin.py newproject.
-
-.. setting:: URLLENGTH_LIMIT
-
-URLLENGTH_LIMIT
----------------
-
-Default: ``2083``
-
-Scope: ``contrib.spidermiddleware.urllength``
-
-The maximum URL length to allow for crawled URLs. For more information about
-the default value for this setting see: http://www.boutell.com/newfaq/misc/urllength.html
-
-.. setting:: USER_AGENT
-
-USER_AGENT
-----------
-
-Default: ``"%s/%s" % (BOT_NAME, BOT_VERSION)``
-
-The default User-Agent to use when crawling, unless overrided.
-
-.. setting:: WEBCONSOLE_ENABLED
-
-WEBCONSOLE_ENABLED
-------------------
-
-Default: True
-
-A boolean which specifies if the web management console will be enabled
-(provided its extension is also enabled).
-
-.. setting:: WEBCONSOLE_LOGFILE
-
-WEBCONSOLE_LOGFILE
-------------------
-
-Default: ``None``
-
-A file to use for logging HTTP requests made to the web console. If unset web
-the log is sent to standard scrapy log.
-
-.. setting:: WEBCONSOLE_PORT
-
-WEBCONSOLE_PORT
----------------
-
-Default: ``6080``
-
-The port to use for the web console. If set to ``None`` or ``0``, a dynamically
-assigned port is used. For more info see :ref:`topics-webconsole`.
-
diff --git a/docs/ref/spider-middleware.rst b/docs/ref/spider-middleware.rst
deleted file mode 100644
index ca90a28fb..000000000
--- a/docs/ref/spider-middleware.rst
+++ /dev/null
@@ -1,117 +0,0 @@
-.. _ref-spider-middleware:
-
-====================================
-Built-in spider middleware reference
-====================================
-
-This page describes all spider middleware components that come with Scrapy. For
-information on how to use them and how to write your own spider middleware, see
-the :ref:`spider middleware usage guide `.
-
-For a list of the components enabled by default (and their orders) see the
-:setting:`SPIDER_MIDDLEWARES_BASE` setting.
-
-Available spider middlewares
-============================
-
-DepthMiddleware
----------------
-
-.. module:: scrapy.contrib.spidermiddleware.depth
-
-.. class:: DepthMiddleware
-
- DepthMiddleware is a scrape middleware used for tracking the depth of each
- Request inside the site being scraped. It can be used to limit the maximum
- depth to scrape or things like that.
-
- The :class:`DepthMiddleware` can be configured through the following
- settings (see the settings documentation for more info):
-
- * :setting:`DEPTH_LIMIT` - The maximum depth that will be allowed to
- crawl for any site. If zero, no limit will be imposed.
- * :setting:`DEPTH_STATS` - Whether to collect depth stats.
-
-HttpErrorMiddleware
--------------------
-
-.. module:: scrapy.contrib.spidermiddleware.httperror
-
-.. class:: HttpErrorMiddleware
-
- Filter out response outside of a range of valid status codes.
-
- This middleware filters out every response with status outside of the range
- 200<=status<300. Spiders can add more exceptions using
- ``handle_httpstatus_list`` spider attribute.
-
-OffsiteMiddleware
------------------
-
-.. module:: scrapy.contrib.spidermiddleware.offsite
-
-.. class:: OffsiteMiddleware
-
- Filters out Requests for URLs outside the domains covered by the spider.
-
- This middleware filters out every request whose host names doesn't match
- :attr:`~scrapy.spider.BaseSpider.domain_name`, or the spider
- :attr:`~scrapy.spider.BaseSpider.domain_name` prefixed by "www.".
- Spider can add more domains to exclude using
- :attr:`~scrapy.spider.BaseSpider.extra_domain_names` attribute.
-
-RequestLimitMiddleware
-----------------------
-
-.. module:: scrapy.contrib.spidermiddleware.requestlimit
-
-.. class:: RequestLimitMiddleware
-
- Limits the maximum number of requests in the scheduler for each spider. When
- a spider tries to schedule more than the allowed amount of requests, the new
- requests (returned by the spider) will be dropped.
-
- The :class:`RequestLimitMiddleware` can be configured through the following
- settings (see the settings documentation for more info):
-
- * :setting:`REQUESTS_QUEUE_SIZE` - If non zero, it will be used as an
- upper limit for the amount of requests that can be scheduled per
- domain. Can be set per spider using ``requests_queue_size`` attribute.
-
-RestrictMiddleware
-------------------
-
-.. module:: scrapy.contrib.spidermiddleware.restrict
-
-.. class:: RestrictMiddleware
-
- Restricts crawling to fixed set of particular URLs.
-
- The :class:`RestrictMiddleware` can be configured through the following
- settings (see the settings documentation for more info):
-
- * :setting:`RESTRICT_TO_URLS` - Set of URLs allowed to crawl.
-
-UrlFilterMiddleware
--------------------
-
-.. module:: scrapy.contrib.spidermiddleware.urlfilter
-
-.. class:: UrlFilterMiddleware
-
- Canonicalizes URLs to filter out duplicated ones
-
-UrlLengthMiddleware
--------------------
-
-.. module:: scrapy.contrib.spidermiddleware.urllength
-
-.. class:: UrlLengthMiddleware
-
- Filters out requests with URLs longer than URLLENGTH_LIMIT
-
- The :class:`UrlLengthMiddleware` can be configured through the following
- settings (see the settings documentation for more info):
-
- * :setting:`URLLENGTH_LIMIT` - The maximum URL length to allow for crawled URLs.
-
diff --git a/docs/ref/spiders.rst b/docs/ref/spiders.rst
deleted file mode 100644
index 2011562ed..000000000
--- a/docs/ref/spiders.rst
+++ /dev/null
@@ -1,377 +0,0 @@
-.. _ref-spiders:
-
-=========================
-Available Generic Spiders
-=========================
-
-.. module:: scrapy.spider
- :synopsis: Spiders base class, spider manager and spider middleware
-
-BaseSpider
-==========
-
-.. class:: BaseSpider()
-
-This is the simplest spider, and the one from which every other spider
-must inherit from (either the ones that come bundled with Scrapy, or the ones
-that you write yourself). It doesn't provide any special functionality. It just
-requests the given ``start_urls``/``start_requests``, and calls the spider's
-method ``parse`` for each of the resulting responses.
-
-.. attribute:: BaseSpider.domain_name
-
- A string which defines the domain name for this spider, which will also be
- the unique identifier for this spider (which means you can't have two
- spider with the same ``domain_name``). This is the most important spider
- attribute and it's required, and it's the name by which Scrapy will known
- the spider.
-
-.. attribute:: BaseSpider.extra_domain_names
-
- An optional list of strings containing additional domains that this spider
- is allowed to crawl. Requests for URLs not belonging to the domain name
- specified in :attr:`Spider.domain_name` or this list won't be followed.
-
-.. attribute:: BaseSpider.start_urls
-
- Is a list of URLs where the spider will begin to crawl from, when no
- particular URLs are specified. So, the first pages downloaded will be those
- listed here. The subsequent URLs will be generated successively from data
- contained in the start URLs.
-
-.. method:: BaseSpider.start_requests()
-
- This method must return an iterable with the first Requests to crawl for
- this spider.
-
- This is the method called by Scrapy when the spider is opened for scraping
- when no particular URLs are specified. If particular URLs are specified,
- the :meth:`BaseSpider.make_requests_from_url` is used instead to create the
- Requests. This method is also called only once from Scrapy, so it's safe to
- implement it as a generator.
-
- The default implementation uses :meth:`BaseSpider.make_requests_from_url`
- to generate Requests for each url in :attr:`start_urls`.
-
- If you want to change the Requests used to start scraping a domain, this is
- the method to override. For example, if you need to start by login in using
- a POST request, you could do::
-
- def start_requests(self):
- return [FormRequest("http://www.example.com/login",
- formdata={'user': 'john', 'pass': 'secret'},
- callback=self.logged_in)]
-
- def logged_in(self, response):
- # here you would extract links to follow and return Requests for
- # each of them, with another callback
- pass
-
-.. method:: BaseSpider.make_requests_from_url(url)
-
- A method that receives a URL and returns a :class:`~scrapy.http.Request`
- object (or a list of :class:`~scrapy.http.Request` objects) to scrape. This
- method is used to construct the initial requests in the
- :meth:`start_requests` method, and is typically used to convert urls to
- requests.
-
- Unless overridden, this method returns Requests with the :meth:`parse`
- method as their callback function, and with dont_filter parameter enabled
- (see :class:`~scrapy.http.Request` class for more info).
-
-.. method:: BaseSpider.parse(response)
-
- This is the default callback used by the :meth:`start_requests` method, and
- will be used to parse the first pages crawled by the spider.
-
- The ``parse`` method is in charge of processing the response and returning
- scraped data and/or more URLs to follow, because of this, the method must
- always return a list or at least an empty one. Other Requests callbacks
- have the same requirements as the BaseSpider class.
-
-BaseSpider example
-------------------
-
-Let's see an example::
-
- from scrapy import log # This module is useful for printing out debug information
- from scrapy.spider import BaseSpider
-
- class MySpider(BaseSpider):
- domain_name = 'http://www.example.com'
- start_urls = [
- 'http://www.example.com/1.html',
- 'http://www.example.com/2.html',
- 'http://www.example.com/3.html',
- ]
-
- def parse(self, response):
- self.log('A response from %s just arrived!' % response.url)
- return []
-
- SPIDER = MySpider()
-
-.. module:: scrapy.contrib.spiders
- :synopsis: Collection of generic spiders
-
-CrawlSpider
-===========
-
-.. class:: CrawlSpider
-
-This is the most commonly used spider for crawling regular websites, as it
-provides a convenient mechanism for following links by defining a set of rules.
-It may not be the best suited for your particular web sites or project, but
-it's generic enough for several cases, so you can start from it and override it
-as need more custom functionality, or just implement your own spider.
-
-Apart from the attributes inherited from BaseSpider (that you must
-specify), this class supports a new attribute:
-
-.. attribute:: CrawlSpider.rules
-
- Which is a list of one (or more) :class:`Rule` objects. Each :class:`Rule`
- defines a certain behaviour for crawling the site. Rules objects are
- described below .
-
-Crawling rules
---------------
-
-.. class:: Rule(link_extractor, callback=None, cb_kwargs=None, follow=None, process_links=None)
-
-``link_extractor`` is a :ref:`Link Extractor ` object which
-defines how links will be extracted from each crawled page.
-
-``callback`` is a callable or a string (in which case a method from the spider
-object with that name will be used) to be called for each link extracted with
-the specified link_extractor. This callback receives a response as its first
-argument and must return a list containing either ScrapedItems and Requests (or
-any subclass of them).
-
-``cb_kwargs`` is a dict containing the keyword arguments to be passed to the
-callback function
-
-``follow`` is a boolean which specified if links should be followed from each
-response extracted with this rule. If ``callback`` is None ``follow`` defaults
-to ``True``, otherwise it default to ``False``.
-
-``process_links`` is a callable, or a string (in which case a method from the
-spider object with that name will be used) which will be called for each list
-of links extracted from each response using the specified ``link_extractor``.
-This is mainly used for filtering purposes.
-
-
-CrawlSpider example
--------------------
-
-Let's now take a look at an example CrawlSpider with rules::
-
- from scrapy.contrib.spiders import CrawlSpider, Rule
- from scrapy.contrib.linkextractors.sgml import SgmlLinkExtractor
- from scrapy.xpath.selector import HtmlXPathSelector
- from scrapy.item import ScrapedItem
-
- class MySpider(CrawlSpider):
- domain_name = 'example.com'
- start_urls = ['http://www.example.com']
-
- rules = (
- # Extract links matching 'category.php' (but not matching 'subsection.php')
- # and follow links from them (since no callback means follow=True by default).
- Rule(SgmlLinkExtractor(allow=('category\.php', ), deny=('subsection\.php', ))),
-
- # Extract links matching 'item.php' and parse them with the spider's method parse_item
- Rule(SgmlLinkExtractor(allow=('item\.php', )), callback='parse_item'),
- )
-
- def parse_item(self, response):
- self.log('Hi, this is an item page! %s' % response.url)
-
- hxs = HtmlXPathSelector(response)
- item = ScrapedItem()
- item.id = hxs.select('//td[@id="item_id"]/text()').re(r'ID: (\d+)')
- item.name = hxs.select('//td[@id="item_name"]/text()').extract()
- item.description = hxs.select('//td[@id="item_description"]/text()').extract()
- return [item]
-
- SPIDER = MySpider()
-
-
-This spider would start crawling example.com's home page, collecting category
-links, and item links, parsing the latter with the
-:meth:`XMLFeedSpider.parse_item` method. For each item response, some data will
-be extracted from the HTML using XPath, and a ScrapedItem will be filled with
-it.
-
-XMLFeedSpider
-=============
-
-.. class:: XMLFeedSpider
-
- XMLFeedSpider is designed for parsing XML feeds by iterating through them by a
- certain node name. The iterator can be chosen from: ``iternodes``, ``xml``,
- and ``html``. It's recommended to use the ``iternodes`` iterator for
- performance reasons, since the ``xml`` and ``html`` iterators generate the
- whole DOM at once in order to parse it. However, using ``html`` as the
- iterator may be useful when parsing XML with bad markup.
-
- For setting the iterator and the tag name, you must define the following class
- attributes:
-
- .. attribute:: iterator
-
- A string which defines the iterator to use. It can be either:
-
- - ``'iternodes'`` - a fast iterator based on regular expressions
-
- - ``'html'`` - an iterator which uses HtmlXPathSelector. Keep in mind
- this uses DOM parsing and must load all DOM in memory which could be a
- problem for big feeds
-
- - ``'xml'`` - an iterator which uses XmlXPathSelector. Keep in mind
- this uses DOM parsing and must load all DOM in memory which could be a
- problem for big feeds
-
- It defaults to: ``'iternodes'``.
-
- .. attribute:: itertag
-
- A string with the name of the node (or element) to iterate in. Example::
-
- itertag = 'product'
-
- .. attribute:: namespaces
-
- A list of ``(prefix, uri)`` tuples which define the namespaces
- available in that document that will be processed with this spider. The
- ``prefix`` and ``uri`` will be used to automatically register
- namespaces using the
- :meth:`~scrapy.xpath.XPathSelector.register_namespace` method.
-
- You can then specify nodes with namespaces in the :attr:`itertag`
- attribute.
-
- Example::
-
- class YourSpider(XMLFeedSpider):
-
- namespaces = [('n', 'http://www.sitemaps.org/schemas/sitemap/0.9')]
- itertag = 'n:url'
- # ...
-
- Apart from these new attributes, this spider has the following overrideable
- methods too:
-
- .. method:: adapt_response(response)
-
- A method that receives the response as soon as it arrives from the spider
- middleware and before start parsing it. It can be used used for modifying
- the response body before parsing it. This method receives a response and
- returns response (it could be the same or another one).
-
- .. method:: parse_item(response, selector)
-
- This method is called for the nodes matching the provided tag name
- (``itertag``). Receives the response and an XPathSelector for each node.
- Overriding this method is mandatory. Otherwise, you spider won't work.
- This method must return either a ScrapedItem, a Request, or a list
- containing any of them.
-
- .. warning:: This method will soon change its name to ``parse_node``
-
- .. method:: process_results(response, results)
-
- This method is called for each result (item or request) returned by the
- spider, and it's intended to perform any last time processing required
- before returning the results to the framework core, for example setting the
- item IDs. It receives a list of results and the response which originated
- that results. It must return a list of results (Items or Requests)."""
-
-
-XMLFeedSpider example
----------------------
-
-These spiders are pretty easy to use, let's have at one example::
-
- from scrapy import log
- from scrapy.contrib.spiders import XMLFeedSpider
- from scrapy.item import ScrapedItem
-
- class MySpider(XMLFeedSpider):
- domain_name = 'example.com'
- start_urls = ['http://www.example.com/feed.xml']
- iterator = 'iternodes' # This is actually unnecesary, since it's the default value
- itertag = 'item'
-
- def parse_item(self, response, node):
- log.msg('Hi, this is a <%s> node!: %s' % (self.itertag, ''.join(node.extract())))
-
- item = ScrapedItem()
- item.id = node.select('@id').extract()
- item.name = node.select('name').extract()
- item.description = node.select('description').extract()
- return item
-
- SPIDER = MySpider()
-
-Basically what we did up there was creating a spider that downloads a feed from
-the given ``start_urls``, and then iterates through each of its ``item`` tags,
-prints them out, and stores some random data in ScrapedItems.
-
-CSVFeedSpider
-=============
-
-.. class:: CSVFeedSpider
-
-.. warning:: The API of the CSVFeedSpider is not yet stable. Use with caution.
-
-This spider is very similar to the XMLFeedSpider, although it iterates through
-rows, instead of nodes. It also has other two different attributes:
-
-.. attribute:: CSVFeedSpider.delimiter
-
- A string with the separator character for each field in the CSV file
- Defaults to ``','`` (comma).
-
-.. attribute:: CSVFeedSpider.headers
-
- A list of the rows contained in the file CSV feed which will be used for
- extracting fields from it.
-
-In this spider, the method that gets called in each row iteration ``parse_row``
-instead of ``parse_item`` (like in :class:`XMLFeedSpider`).
-
-.. method:: CSVFeedSpider.parse_row(response, row)
-
- Receives a response and a dict (representing each row) with a key for each
- provided (or detected) header of the CSV file. This spider also gives the
- opportunity to override ``adapt_response`` and ``process_results`` methods
- for pre and post-processing purposes.
-
-CSVFeedSpider example
----------------------
-
-Let's see an example similar to the previous one, but using CSVFeedSpider::
-
- from scrapy import log
- from scrapy.contrib.spiders import CSVFeedSpider
- from scrapy.item import ScrapedItem
-
- class MySpider(CSVFeedSpider):
- domain_name = 'example.com'
- start_urls = ['http://www.example.com/feed.csv']
- delimiter = ';'
- headers = ['id', 'name', 'description']
-
- def parse_row(self, response, row):
- log.msg('Hi, this is a row!: %r' % row)
-
- item = ScrapedItem()
- item.id = row['id']
- item.name = row['name']
- item.description = row['description']
- return item
-
- SPIDER = MySpider()
-
-
diff --git a/docs/reference.rst b/docs/reference.rst
new file mode 100644
index 000000000..16897148e
--- /dev/null
+++ b/docs/reference.rst
@@ -0,0 +1,22 @@
+.. _ref:
+
+API Reference
+=============
+
+This section documents the Scrapy |version| API. For more information see :ref:`misc-api-stability`.
+
+* :ref:`topics-request-response`
+* :ref:`topics-spiders-ref`
+* :ref:`topics-selectors-ref`
+* :ref:`topics-settings-ref`
+* :ref:`topics-signals-ref`
+* :ref:`topics-exceptions-ref`
+* :ref:`topics-logging`
+* :ref:`topics-email`
+* :ref:`topics-extensions-ref`
+* :ref:`topics-downloader-middleware-ref`
+* :ref:`topics-spider-middleware-ref`
+* :ref:`topics-scheduler-middleware-ref`
+* :ref:`topics-link-extractors-ref`
+* :ref:`topics-stats-ref`
+
diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst
index 22a4f9e89..9ea0a059f 100644
--- a/docs/topics/downloader-middleware.rst
+++ b/docs/topics/downloader-middleware.rst
@@ -128,3 +128,71 @@ scope, and the original request won't finish until redirected request is
completed. This stop ``process_download_exception()`` middleware as returning Response
would do.
+
+.. _topics-downloader-middleware-ref:
+
+Built-in downloader middleware reference
+========================================
+
+This page describes all downloader middleware components that come with
+Scrapy. For information on how to use them and how to write your own downloader
+middleware, see the :ref:`downloader middleware usage guide
+`.
+
+For a list of the components enabled by default (and their orders) see the
+:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting.
+
+DefaultHeadersMiddleware
+------------------------
+
+.. module:: scrapy.contrib.downloadermiddleware.defaultheaders
+ :synopsis: Default Headers Downloader Middleware
+
+.. class:: DefaultHeadersMiddleware
+
+ This middleware sets all default requests headers specified in the
+ :setting:`DEFAULT_REQUEST_HEADERS` setting.
+
+DebugMiddleware
+---------------
+
+.. module:: scrapy.contrib.downloadermiddleware.debug
+ :synopsis: Downloader middlewares for debugging
+
+.. class:: DebugMiddleware
+
+ This is a convenient middleware to inspect what's passing through the
+ downloader middleware. It logs all requests and responses catched by the
+ middleware component methods. This middleware does not use any settings and
+ does not come enabled by default. Instead, it's meant to be inserted at the
+ point of the middleware that you want to inspect.
+
+HttpCacheMiddleware
+-------------------
+
+.. module:: scrapy.contrib.downloadermiddleware.httpcache
+ :synopsis: HTTP Cache downloader middleware
+
+.. class:: HttpCacheMiddleware
+
+ This middleware provides low-level cache to all HTTP requests and responses.
+ Every request and its corresponding response are cached and then, when that
+ same request is seen again, the response is returned without transferring
+ anything from the Internet.
+
+ The HTTP cache is useful for testing spiders faster (without having to wait for
+ downloads every time) and for trying your spider off-line when you don't have
+ an Internet connection.
+
+ The :class:`HttpCacheMiddleware` can be configured through the following
+ settings (see the settings documentation for more info):
+
+ * :setting:`HTTPCACHE_DIR` - this one actually enables the cache besides
+ settings the cache dir
+ * :setting:`HTTPCACHE_IGNORE_MISSING` - ignoring missing requests instead
+ of downloading them
+ * :setting:`HTTPCACHE_SECTORIZE` - split HTTP cache in several directories
+ (for performance reasons)
+ * :setting:`HTTPCACHE_EXPIRATION_SECS` - how many secs until the cache is
+ considered out of date
+
diff --git a/docs/ref/email.rst b/docs/topics/email.rst
similarity index 99%
rename from docs/ref/email.rst
rename to docs/topics/email.rst
index 76af6e1a1..7fc2f9d60 100644
--- a/docs/ref/email.rst
+++ b/docs/topics/email.rst
@@ -1,4 +1,4 @@
-.. _ref-email:
+.. _topics-email:
=============
Sending email
diff --git a/docs/ref/exceptions.rst b/docs/topics/exceptions.rst
similarity index 90%
rename from docs/ref/exceptions.rst
rename to docs/topics/exceptions.rst
index 9a9288037..0fa6b524a 100644
--- a/docs/ref/exceptions.rst
+++ b/docs/topics/exceptions.rst
@@ -1,10 +1,16 @@
-.. _exceptions:
+.. _topics-exceptions:
+
+==========
+Exceptions
+==========
.. module:: scrapy.core.exceptions
:synopsis: Core exceptions
-Available Exceptions
-====================
+.. _topics-exceptions-ref:
+
+Built-in Exceptions reference
+=============================
Here's a list of all exceptions included in Scrapy and their usage.
diff --git a/docs/topics/extensions.rst b/docs/topics/extensions.rst
index 3172887de..8a8ac01b6 100644
--- a/docs/topics/extensions.rst
+++ b/docs/topics/extensions.rst
@@ -65,21 +65,22 @@ Not all available extensions will be enabled. Some of them usually depend on a
particular setting. For example, the HTTP Cache extension is available by default
but disabled unless the :setting:`HTTPCACHE_DIR` setting is set. Both enabled
and disabled extension can be accessed through the
-:ref:`ref-extension-manager`.
+:ref:`topics-extensions-ref-manager`.
Accessing enabled extensions
============================
Even though it's not usually needed, you can access extension objects through
-the :ref:`ref-extension-manager` which is populated when extensions are loaded.
-For example, to access the ``WebConsole`` extension::
+the :ref:`topics-extensions-ref-manager` which is populated when extensions are
+loaded. For example, to access the ``WebConsole`` extension::
from scrapy.extension import extensions
webconsole_extension = extensions.enabled['WebConsole']
.. seealso::
- :ref:`ref-extension-manager`, for the complete Extension manager reference.
+ :ref:`topics-extensions-ref-manager`, for the complete Extension manager
+ reference.
Writing your own extension
==========================
@@ -110,8 +111,308 @@ everytime a domain/spider is opened and closed::
def domain_closed(self, domain, spider):
log.msg("closed domain %s" % domain)
-Built-in extensions
-===================
-See :ref:`ref-extensions`.
+.. _topics-extensions-ref:
+
+Built-in extensions reference
+=============================
+
+.. _topics-extensions-ref-manager:
+
+Extension manager
+-----------------
+
+.. module:: scrapy.extension
+ :synopsis: The extension manager
+
+The Extension Manager is responsible for loading and keeping track of installed
+extensions and it's configured through the :setting:`EXTENSIONS` setting which
+contains a dictionary of all available extensions and their order similar to
+how you :ref:`configure the downloader middlewares
+`.
+
+.. class:: ExtensionManager
+
+ The extension manager is a singleton object, which is instantiated at module
+ loading time and can be accessed like this::
+
+ from scrapy.extension import extensions
+
+ .. attribute:: loaded
+
+ A boolean which is True if extensions are already loaded or False if
+ they're not.
+
+ .. attribute:: enabled
+
+ A dict with the enabled extensions. The keys are the extension class names,
+ and the values are the extension objects. Example::
+
+ >>> from scrapy.extension import extensions
+ >>> extensions.load()
+ >>> print extensions.enabled
+ {'CoreStats': ,
+ 'WebConsoke': ,
+ ...
+
+ .. attribute:: disabled
+
+ A dict with the disabled extensions. The keys are the extension class names,
+ and the values are the extension class paths (because objects are never
+ instantiated for disabled extensions). Example::
+
+ >>> from scrapy.extension import extensions
+ >>> extensions.load()
+ >>> print extensions.disabled
+ {'MemoryDebugger': 'scrapy.contrib.webconsole.stats.MemoryDebugger',
+ 'SpiderProfiler': 'scrapy.contrib.spider.profiler.SpiderProfiler',
+ ...
+
+ .. method:: load()
+
+ Load the available extensions configured in the :setting:`EXTENSIONS`
+ setting. On a standard run, this method is usually called by the Execution
+ Manager, but you may need to call it explicitly if you're dealing with
+ code outside Scrapy.
+
+ .. method:: reload()
+
+ Reload the available extensions. See :meth:`load`.
+
+General purpose extensions
+--------------------------
+
+Core Stats extension
+~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.stats.corestats
+ :synopsis: Core stats collection
+
+.. class:: scrapy.stats.corestats.CoreStats
+
+Enable the collection of core statistics, provided the stats collection are
+enabled (see :ref:`topics-stats`).
+
+.. _topics-extensions-ref-webconsole:
+
+Web console extension
+~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.management.web
+ :synopsis: Web management console
+
+.. class:: scrapy.management.web.WebConsole
+
+Provides an extensible web server for managing a Scrapy process. It's enabled
+by the :setting:`WEBCONSOLE_ENABLED` setting. The server will listen in the
+port specified in :setting:`WEBCONSOLE_PORT`, and will log to the file
+specified in :setting:`WEBCONSOLE_LOGFILE`.
+
+The web server is designed to be extended by other extensions which can add
+their own management web interfaces.
+
+See also :ref:`topics-webconsole` for information on how to write your own web
+console extension, and "Web console extensions" below for a list of available
+built-in (web console) extensions.
+
+.. _topics-extensions-ref-telnetconsole:
+
+Telnet console extension
+~~~~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.management.telnet
+ :synopsis: Telnet management console
+
+.. class:: scrapy.management.telnet.TelnetConsole
+
+Provides a telnet console for getting into a Python interpreter inside the
+currently running Scrapy process, which can be very useful for debugging.
+
+The telnet console must be enabled by the :setting:`TELNETCONSOLE_ENABLED`
+setting, and the server will listen in the port specified in
+:setting:`WEBCONSOLE_PORT`.
+
+Spider reloader extension
+~~~~~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.contrib.spider.reloader
+ :synopsis: Spider reloader extension
+
+.. class:: scrapy.contrib.spider.reloader.SpiderReloader
+
+Reload spider objects once they've finished scraping, to release the resources
+and references to other objects they may hold.
+
+.. _topics-extensions-ref-memusage:
+
+Memory usage extension
+~~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.contrib.memusage
+ :synopsis: Memory usage extension
+
+.. class:: scrapy.contrib.memusage.MemoryUsage
+
+Allows monitoring the memory used by a Scrapy process and:
+
+1, send a notification email when it exceeds a certain value
+2. terminate the Scrapy process when it exceeds a certain value
+
+The notification emails can be triggered when a certain warning value is
+reached (:setting:`MEMUSAGE_WARNING_MB`) and when the maximum value is reached
+(:setting:`MEMUSAGE_LIMIT_MB`) which will also cause the Scrapy process to be
+terminated.
+
+This extension is enabled by the :setting:`MEMUSAGE_ENABLED` setting and
+can be configured with the following settings:
+
+* :setting:`MEMUSAGE_LIMIT_MB`
+* :setting:`MEMUSAGE_WARNING_MB`
+* :setting:`MEMUSAGE_NOTIFY_MAIL`
+* :setting:`MEMUSAGE_REPORT`
+
+Memory debugger extension
+~~~~~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.contrib.memdebug
+ :synopsis: Memory debugger extension
+
+.. class:: scrapy.contrib.memdebug.MemoryDebugger
+
+A memory debugger which collects some info about objects uncollected by the
+garbage collector and libxml2 memory leaks. To enable this extension turn on
+the :setting:`MEMDEBUG_ENABLED` setting. The report will be printed to standard
+output. If the :setting:`MEMDEBUG_NOTIFY` setting contains a list of emails the
+report will also be sent to those addresses.
+
+Close domain extension
+~~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.contrib.closedomain
+ :synopsis: Close domain extension
+
+.. class:: scrapy.contrib.closedomain.CloseDomain
+
+Closes a domain/spider automatically when some conditions are met, using a
+specific closing reason for each condition.
+
+The conditions for closing a domain can be configured through the following
+settings. Other conditions will be supported in the future.
+
+.. setting:: CLOSEDOMAIN_TIMEOUT
+
+CLOSEDOMAIN_TIMEOUT
+"""""""""""""""""""
+
+Default: ``0``
+
+An integer which specifies a number of seconds. If the domain remains open for
+more than that number of second, it will be automatically closed with the
+reason ``closedomain_timeout``. If zero (or non set) domains won't be closed by
+timeout.
+
+.. setting:: CLOSEDOMAIN_ITEMPASSED
+
+CLOSEDOMAIN_ITEMPASSED
+""""""""""""""""""""""
+
+Default: ``0``
+
+An integer which specifies a number of items. If the spider scrapes more than
+that amount if items and those items are passed by the item pipeline, the
+domain will be closed with the reason ``closedomain_itempassed``. If zero (or
+non set) domains won't be closed by number of passed items.
+
+Stack trace dump extension
+~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.contrib.debug
+ :synopsis: Extensions for debugging Scrapy
+
+.. class:: scrapy.contrib.debug.StackTraceDump
+
+Adds a `SIGUSR1`_ signal handler which dumps the stack trace of a runnning
+Scrapy process when a ``SIGUSR1`` signal is catched. After the stack trace is
+dumped, the Scrapy process continues to run normally.
+
+The stack trace is sent to standard output, or to the Scrapy log file if
+:setting:`LOG_STDOUT` is enabled.
+
+This extension only works on POSIX-compliant platforms (ie. not Windows).
+
+.. _SIGUSR1: http://en.wikipedia.org/wiki/SIGUSR1_and_SIGUSR2
+
+StatsMailer extension
+~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.contrib.statsmailer
+ :synopsis: StatsMailer extension
+
+.. class:: scrapy.contrib.statsmailer.StatsMailer
+
+This simple extension can be used to send a notification email every time a
+domain has finished scraping, including the Scrapy stats collected. The email
+will be sent to all recipients specified in the :setting:`STATSMAILER_RCPTS`
+setting.
+
+Web console extensions
+----------------------
+
+.. module:: scrapy.contrib.webconsole
+ :synopsis: Contains most built-in web console extensions
+
+Here is a list of built-in web console extensions. For clarity "web console
+extension" is abbreviated as "WC extension".
+
+For more information see the see the :ref:`web console documentation
+`.
+
+Scheduler queue WC extension
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.contrib.webconsole.scheduler
+ :synopsis: Scheduler queue web console extension
+
+.. class:: scrapy.contrib.webconsole.scheduler.SchedulerQueue
+
+Display a list of all pending Requests in the Scheduler queue, grouped by
+domain/spider.
+
+Spider live stats WC extension
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.contrib.webconsole.livestats
+ :synopsis: Spider live stats web console extension
+
+.. class:: scrapy.contrib.webconsole.livestats.LiveStats
+
+Display a table with stats of all spider crawled by the current Scrapy run,
+including:
+
+* Number of items scraped
+* Number of pages crawled
+* Number of pending requests in the scheduler
+* Number of pending requests in the downloader queue
+* Number of requests currently being downloaded
+
+Engine status WC extension
+~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.contrib.webconsole.enginestatus
+ :synopsis: Engine status web console extension
+
+.. class:: scrapy.contrib.webconsole.enginestatus.EngineStatus
+
+Display the current status of the Scrapy Engine, which is just the output of
+the Scrapy engine ``getstatus()`` method.
+
+Stats collector dump WC extension
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+.. module:: scrapy.contrib.webconsole.stats
+ :synopsis: Stats dump web console extension
+
+.. class:: scrapy.contrib.webconsole.stats.StatsDump
+
+Display the stats collected so far by the stats collector.
diff --git a/docs/topics/index.rst b/docs/topics/index.rst
index 0615aa748..49a115b96 100644
--- a/docs/topics/index.rst
+++ b/docs/topics/index.rst
@@ -25,3 +25,9 @@ This section describes all key concepts of Scrapy.
robotstxt
firefox
firebug
+ signals
+ logging
+ scheduler-middleware
+ request-response
+ exceptions
+ email
diff --git a/docs/topics/link-extractors.rst b/docs/topics/link-extractors.rst
index 48d87f107..383568952 100644
--- a/docs/topics/link-extractors.rst
+++ b/docs/topics/link-extractors.rst
@@ -24,6 +24,124 @@ your spiders even if you don't subclass from
:class:`~scrapy.contrib.spiders.CrawlSpider`, as its purpose is very simple: to
extract links.
-See :ref:`ref-link-extractors` for the list of available built-in Link
-Extractors, including some examples.
+
+.. _topics-link-extractors-ref:
+
+Built-in link extractors reference
+==================================
+
+.. module:: scrapy.contrib.linkextractors
+ :synopsis: Link extractors classes
+
+All available link extractors classes bundled with Scrapy are provided in the
+:mod:`scrapy.contrib.linkextractors` module.
+
+.. module:: scrapy.contrib.linkextractors.sgml
+ :synopsis: SGMLParser-based link extractors
+
+SgmlLinkExtractor
+-----------------
+
+.. class:: SgmlLinkExtractor(allow=(), deny=(), allow_domains=(), deny_domains=(), restrict_xpaths(), tags=('a', 'area'), attrs=('href'), canonicalize=True, unique=True, process_value=None)
+
+ The SgmlLinkExtractor extends the base :class:`BaseSgmlLinkExtractor` by
+ providing additional filters that you can specify to extract links,
+ including regular expressions patterns that the links must match to be
+ extracted. All those filters are configured through these constructor
+ parameters:
+
+ :param allow: a single regular expression (or list of regular expressions)
+ that the (absolute) urls must match in order to be extracted. If not
+ given (or empty), it will match all links.
+ :type allow: a regular expression (or list of)
+
+ :param deny: a single regular expression (or list of regular expressions)
+ that the (absolute) urls must match in order to be excluded (ie. not
+ extracted). It has precedence over the ``allow`` parameter. If not
+ given (or empty) it won't exclude any links.
+ :type allow: a regular expression (or list of)
+
+ :param allow_domains: is single value or a list of string containing
+ domains which will be considered for extracting the links
+ :type allow: str or list
+
+ :param deny_domains: is single value or a list of strings containing
+ domains which which won't be considered for extracting the links
+ :type allow: str or list
+
+ :param restrict_xpaths: is a XPath (or list of XPath's) which defines
+ regions inside the response where links should be extracted from.
+ If given, only the text selected by those XPath will be scanned for
+ links. See examples below.
+ :type restrict_xpaths: str or list
+
+ :param tags: a tag or a list of tags to consider when extracting links.
+ Defaults to ``('a', 'area')``.
+ :type tags: str or list
+
+ :param attrs: list of attrbitues which should be considered when looking
+ for links to extract (only for those tags specified in the ``tags``
+ parameter). Defaults to ``('href',)``
+ :type attrs: boolean
+
+ :param canonicalize: canonicalize each extracted url (using
+ scrapy.utils.url.canonicalize_url). Defaults to ``True``.
+ :type canonicalize: boolean
+
+ :param unique: whether duplicate filtering should be applied to extracted
+ links.
+ :type unique: boolean
+
+ :param process_value: see ``process_value`` argument of
+ :class:`LinkExtractor` class constructor
+ :type process_value: boolean
+
+BaseSgmlLinkExtractor
+---------------------
+
+.. class:: BaseSgmlLinkExtractor(tag="a", href="href", unique=False, process_value=None)
+
+ The purpose of this Link Extractor is only to serve as a base class for the
+ :class:`SgmlLinkExtractor`. You should use that one instead.
+
+ The constructor arguments are:
+
+ :param tag: either a string (with the name of a tag) or a function that
+ receives a tag name and returns ``True`` if links should be extracted from
+ those tag, or ``False`` if they shouldn't. Defaults to ``'a'``. request
+ (once its downloaded) as its first parameter. For more information see
+ :ref:`topics-request-response-ref-request-callback-arguments`.
+ :type tag: str or callable
+
+ :param attr: either string (with the name of a tag attribute), or a
+ function that receives a an attribute name and returns ``True`` if
+ links should be extracted from it, or ``False`` if the shouldn't.
+ Defaults to ``href``.
+ :type attr: str or callable
+
+ :param unique: is a boolean that specifies if a duplicate filtering should
+ be applied to links extracted.
+ :type unique: boolean
+
+ :param process_value: a function which receives each value extracted from
+ the tag and attributes scanned and can modify the value and return a
+ new one, or return ``None`` to ignore the link altogether. If not
+ given, ``process_value`` defaults to ``lambda x: x``.
+
+ .. highlight:: html
+
+ For example, to extract links from this code::
+
+ Link text
+
+ .. highlight:: python
+
+ You can use the following function in ``process_value``::
+
+ def process_value(value):
+ m = re.search("javascript:goToPage\('(.*?)'", value)
+ if m:
+ return m.group(1)
+
+ :type process_value: callable
diff --git a/docs/ref/logging.rst b/docs/topics/logging.rst
similarity index 99%
rename from docs/ref/logging.rst
rename to docs/topics/logging.rst
index 734401d77..fe677f3c9 100644
--- a/docs/ref/logging.rst
+++ b/docs/topics/logging.rst
@@ -1,4 +1,4 @@
-.. _ref-logging:
+.. _topics-logging:
=======
Logging
diff --git a/docs/ref/request-response.rst b/docs/topics/request-response.rst
similarity index 93%
rename from docs/ref/request-response.rst
rename to docs/topics/request-response.rst
index a94674c3b..49ab5a20f 100644
--- a/docs/ref/request-response.rst
+++ b/docs/topics/request-response.rst
@@ -1,15 +1,12 @@
-.. _ref-request-response:
+.. _topics-request-response:
-============================
-Request and Response objects
-============================
+======================
+Requests and Responses
+======================
.. module:: scrapy.http
:synopsis: Request and Response classes
-Quick overview
-==============
-
Scrapy uses :class:`Request` and :class:`Response` objects for crawling web
sites.
@@ -20,7 +17,9 @@ issued the request.
Both :class:`Request` and :class:`Response` classes have subclasses which adds
additional functionality not required in the base classes. These are described
-below in :ref:`ref-request-subclasses` and :ref:`ref-response-subclasses`.
+below in :ref:`topics-request-response-ref-request-subclasses` and
+:ref:`topics-request-response-ref-response-subclasses`.
+
Request objects
===============
@@ -36,7 +35,7 @@ Request objects
:param callback: the function that will be called with the response of this
request (once its downloaded) as its first parameter. For more information
- see :ref:`ref-request-callback-arguments` below.
+ see :ref:`topics-request-response-ref-request-callback-arguments` below.
:type callback: callable
:param method: the HTTP method of this request. Defaults to ``'GET'``.
@@ -145,18 +144,19 @@ Request objects
.. method:: Request.copy()
Return a new Request which is a copy of this Request. The attribute
- :attr:`Request.meta` is copied, while :attr:`Request.cache` is not. See also
- :ref:`ref-request-callback-arguments`.
+ :attr:`Request.meta` is copied, while :attr:`Request.cache` is not. See
+ also :ref:`topics-request-response-ref-request-callback-arguments`.
.. method:: Request.replace([url, callback, method, headers, body, cookies, meta, encoding, dont_filter])
Return a Request object with the same members, except for those members
- given new values by whichever keyword arguments are specified. The attribute
- :attr:`Request.meta` is copied by default (unless a new value is given
- in the ``meta`` argument). The :attr:`Request.cache` attribute is always
- cleared. See also :ref:`ref-request-callback-arguments`.
+ given new values by whichever keyword arguments are specified. The
+ attribute :attr:`Request.meta` is copied by default (unless a new value
+ is given in the ``meta`` argument). The :attr:`Request.cache` attribute
+ is always cleared. See also
+ :ref:`topics-request-response-ref-request-callback-arguments`.
-.. _ref-request-callback-copy:
+.. _topics-request-response-ref-callback-copy:
Caveats with copying Requests and callbacks
-------------------------------------------
@@ -178,7 +178,7 @@ In the above example, ``request2`` is a copy of ``request`` but it has no
callback, while ``request3`` is a copy of ``request`` and also contains the
callback.
-.. _ref-request-callback-arguments:
+.. _topics-request-response-ref-request-callback-arguments:
Passing arguments to callback functions
---------------------------------------
@@ -230,7 +230,7 @@ Using Request.meta::
referer_url = response.request.meta['referer_url']
self.log("Visited page %s from %s" % (response.url, referer_url))
-.. _ref-request-subclasses:
+.. _topics-request-response-ref-request-subclasses:
Request subclasses
==================
@@ -266,7 +266,8 @@ objects.
Returns a new :class:`FormRequest` object with its form field values
pre-populated with those found in the HTML `` |