Bring back _BASE settings

This commit is contained in:
Jakob de Maeyer 2015-11-06 01:14:49 +01:00
parent 57f87b95d4
commit e66f649894
6 changed files with 169 additions and 88 deletions

View File

@ -23,20 +23,22 @@ Here's an example::
'myproject.middlewares.CustomDownloaderMiddleware': 543,
}
The specified :setting:`DOWNLOADER_MIDDLEWARES` setting is merged with the
default one (i.e. it does not overwrite it) and then sorted by order to get the
final sorted list of enabled middlewares: the first middleware is the one
closer to the engine and the last is the one closer to the downloader.
The :setting:`DOWNLOADER_MIDDLEWARES` setting is merged with the
:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting defined in Scrapy (and not meant
to be overridden) and then sorted by order to get the final sorted list of
enabled middlewares: the first middleware is the one closer to the engine and
the last is the one closer to the downloader.
To decide which order to assign to your middleware see the default
:setting:`DOWNLOADER_MIDDLEWARES` setting and pick a value according to
To decide which order to assign to your middleware see the
:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting and pick a value according to
where you want to insert the middleware. The order does matter because each
middleware performs a different action and your middleware could depend on some
previous (or subsequent) middleware being applied.
If you want to disable a built-in middleware you must define it in your
project's :setting:`DOWNLOADER_MIDDLEWARES` setting and assign ``None`` as its
value. For example, if you want to disable the user-agent middleware::
If you want to disable a built-in middleware (the ones defined in
:setting:`DOWNLOADER_MIDDLEWARES_BASE` and enabled by default) you must define it
in your project's :setting:`DOWNLOADER_MIDDLEWARES` setting and assign `None`
as its value. For example, if you want to disable the user-agent middleware::
DOWNLOADER_MIDDLEWARES = {
'myproject.middlewares.CustomDownloaderMiddleware': 543,
@ -162,7 +164,7 @@ middleware, see the :ref:`downloader middleware usage guide
<topics-downloader-middleware>`.
For a list of the components enabled by default (and their orders) see the
:setting:`DOWNLOADER_MIDDLEWARES` setting.
:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting.
.. _cookies-mw:

View File

@ -42,13 +42,14 @@ by a string: the full Python path to the extension's class name. For example::
As you can see, the :setting:`EXTENSIONS` setting is a dict where the keys are
the extension paths, and their values are the orders, which define the
extension *loading* order. The specified :setting:`EXTENSIONS` setting is merged
with the default one (i.e. it does not overwrite it) and then sorted by order
to get the final sorted list of enabled extensions.
extension *loading* order. The :setting:`EXTENSIONS` setting is merged with the
:setting:`EXTENSIONS_BASE` setting defined in Scrapy (and not meant to be
overridden) and then sorted by order to get the final sorted list of enabled
extensions.
As extensions typically do not depend on each other, their loading order is
irrelevant in most cases. This is why the default :setting:`EXTENSIONS` setting
defines all extensions with the same order (``500``). However, this feature can
irrelevant in most cases. This is why the :setting:`EXTENSIONS_BASE` setting
defines all extensions with the same order (``0``). However, this feature can
be exploited if you need to add an extension which depends on other extensions
already loaded.
@ -63,7 +64,7 @@ Disabling an extension
======================
In order to disable an extension that comes enabled by default (ie. those
included in the default :setting:`EXTENSIONS` setting) you must set its order to
included in the :setting:`EXTENSIONS_BASE` setting) you must set its order to
``None``. For example::
EXTENSIONS = {

View File

@ -265,6 +265,16 @@ Whether to export empty feeds (ie. feeds with no items).
FEED_STORAGES
-------------
Default:: ``{}``
A dict containing additional feed storage backends supported by your project.
The keys are URI schemes and the values are paths to storage classes.
.. setting:: FEED_STORAGES_BASE
FEED_STORAGES_BASE
------------------
Default::
{
@ -275,19 +285,30 @@ Default::
'ftp': 'scrapy.extensions.feedexport.FTPFeedStorage',
}
A dict containing all feed storage backends supported by your project. The keys
are URI schemes and the values are paths to storage classes.
A dict containing the built-in feed storage backends supported by Scrapy. You
can disable any of these backends by assigning ``None`` to their URI scheme in
:setting:`FEED_STORAGES`. E.g., to disable the built-in FTP storage backend
(without replacement), place this in your ``settings.py``::
When you set :setting:`FEED_STORAGES` manually, e.g. in your project's settings
module, it will be merged with the default, not overwrite it. If you want to
disable any of the default feed storage backends, you must assign ``None`` as
their value.
FEED_STORAGES = {
'ftp': None,
}
.. setting:: FEED_EXPORTERS
FEED_EXPORTERS
--------------
Default:: ``{}``
A dict containing additional exporters supported by your project. The keys are
serialization formats and the values are paths to :ref:`Item exporter
<topics-exporters>` classes.
.. setting:: FEED_EXPORTERS_BASE
FEED_EXPORTERS_BASE
-------------------
Default::
{
@ -300,14 +321,14 @@ Default::
'pickle': 'scrapy.exporters.PickleItemExporter',
}
A dict containing all feed exporters supported by your project. The keys are
URI schemes and the values are paths to :ref:`Item exporter <topics-exporters>`
classes.
A dict containing the built-in feed exporters supported by Scrapy. You can
disable any of these exporters by assigning ``None`` to their serialization
format in :setting:`FEED_EXPORTERS`. E.g., to disable the built-in CSV exporter
(without replacement), place this in your ``settings.py``::
When you set :setting:`FEED_EXPORTERS` manually, e.g. in your project's settings
module, it will be merged with the default, not overwrite it. If you want to
disable any of the default feed exporters, you must assign ``None`` as their
value.
FEED_EXPORTERS = {
'csv': None,
}
.. _URI: http://en.wikipedia.org/wiki/Uniform_Resource_Identifier
.. _Amazon S3: http://aws.amazon.com/s3/

View File

@ -269,11 +269,6 @@ Default::
The default headers used for Scrapy HTTP Requests. They're populated in the
:class:`~scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware`.
When you set :setting:`DEFAULT_REQUEST_HEADERS` manually, e.g. in your
project's settings module, it will be merged with the default, not overwrite it.
If you want to disable any of the default request headers (and not replace them)
you must assign ``None`` as their value.
.. setting:: DEPTH_LIMIT
DEPTH_LIMIT
@ -355,6 +350,16 @@ The downloader to use for crawling.
DOWNLOADER_MIDDLEWARES
----------------------
Default:: ``{}``
A dict containing the downloader middlewares enabled in your project, and their
orders. For more info see :ref:`topics-downloader-middleware-setting`.
.. setting:: DOWNLOADER_MIDDLEWARES_BASE
DOWNLOADER_MIDDLEWARES_BASE
---------------------------
Default::
{
@ -375,16 +380,11 @@ Default::
'scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware': 900,
}
A dict containing the downloader middlewares enabled in your project, and their
orders. Low orders are closer to the engine, high orders are closer to the
downloader.
When you set :setting:`DOWNLOADER_MIDDLEWARES` manually, e.g. in your project's
settings module, it will be merged with the default, not overwrite it. If you
want to disable any of the default downloader middlewares you must assign
``None`` as their value.
For more info see :ref:`topics-downloader-middleware-setting`.
A dict containing the downloader middlewares enabled by default in Scrapy. Low
orders are closer to the engine, high orders are closer to the downloader. You
should never modify this setting in your project, modify
:setting:`DOWNLOADER_MIDDLEWARES` instead. For more info see
:ref:`topics-downloader-middleware-setting`.
.. setting:: DOWNLOADER_STATS
@ -425,6 +425,16 @@ spider attribute.
DOWNLOAD_HANDLERS
-----------------
Default: ``{}``
A dict containing the request downloader handlers enabled in your project.
See :setting:`DOWNLOAD_HANDLERS_BASE` for example format.
.. setting:: DOWNLOAD_HANDLERS_BASE
DOWNLOAD_HANDLERS_BASE
----------------------
Default::
{
@ -436,15 +446,16 @@ Default::
}
A dict containing the request downloader handlers enabled in your project.
A dict containing the request download handlers enabled by default in Scrapy.
You should never modify this setting in your project, modify
:setting:`DOWNLOAD_HANDLERS` instead.
When you set :setting:`DOWNLOAD_HANDLERS` manually, e.g. in your project's
settings module, it will be merged with the default, not overwrite it. If you
want to disable any of the default download handlers you must assign ``None``
as their value. For example, if you want to disable the file download handler::
You can disable any of these download handlers by assigning ``None`` to their
URI scheme in :setting:`DOWNLOAD_HANDLERS`. E.g., to disable the built-in FTP
handler (without replacement), place this in your ``settings.py``::
DOWNLOAD_HANDLERS = {
'file': None,
'ftp': None,
}
.. setting:: DOWNLOAD_TIMEOUT
@ -544,6 +555,15 @@ to ``vi`` (on Unix systems) or the IDLE editor (on Windows).
EXTENSIONS
----------
Default:: ``{}``
A dict containing the extensions enabled in your project, and their orders.
.. setting:: EXTENSIONS_BASE
EXTENSIONS_BASE
---------------
Default::
{
@ -558,15 +578,10 @@ Default::
'scrapy.extensions.throttle.AutoThrottle': 0,
}
A dict containing the extensions enabled in your project, and their orders. By
default, this setting contains all stable built-in extensions. Keep in mind that
A dict containing the extensions available by default in Scrapy, and their
orders. This setting contains all stable built-in extensions. Keep in mind that
some of them need to be enabled through a setting.
When you set :setting:`EXTENSIONS` manually, e.g. in your project's settings
module, it will be merged with the default, not overwrite it. If you want to
disable any of the default enabled extensions you must assign ``None`` as their
value.
For more information See the :ref:`extensions user guide <topics-extensions>`
and the :ref:`list of available extensions <topics-extensions-ref>`.
@ -589,6 +604,16 @@ Example::
'mybot.pipelines.validate.StoreMyItem': 800,
}
.. setting:: ITEM_PIPELINES_BASE
ITEM_PIPELINES_BASE
-------------------
Default: ``{}``
A dict containing the pipelines enabled by default in Scrapy. You should never
modify this setting in your project, modify :setting:`ITEM_PIPELINES` instead.
.. setting:: LOG_ENABLED
LOG_ENABLED
@ -878,6 +903,16 @@ The scheduler to use for crawling.
SPIDER_CONTRACTS
----------------
Default:: ``{}``
A dict containing the spider contracts enabled in your project, used for
testing spiders. For more info see :ref:`topics-contracts`.
.. setting:: SPIDER_CONTRACTS_BASE
SPIDER_CONTRACTS_BASE
---------------------
Default::
{
@ -886,13 +921,17 @@ Default::
'scrapy.contracts.default.ScrapesContract': 3,
}
A dict containing the scrapy contracts enabled in your project, used for
testing spiders. For more info see :ref:`topics-contracts`.
A dict containing the scrapy contracts enabled by default in Scrapy. You should
never modify this setting in your project, modify :setting:`SPIDER_CONTRACTS`
instead. For more info see :ref:`topics-contracts`.
When you set :setting:`SPIDER_CONTRACTS` manually, e.g. in your project's
settings module, it will be merged with the default, not overwrite it. If you
want to disable any of the default contracts you must assign ``None`` as their
value.
You can disable any of these contracts by assigning ``None`` to their class
path in :setting:`SPIDER_CONTRACTS`. E.g., to disable the built-in
``ScrapesContract``, place this in your ``settings.py``::
SPIDER_CONTRACTS = {
'scrapy.contracts.default.ScrapesContract': None,
}
.. setting:: SPIDER_LOADER_CLASS
@ -909,6 +948,16 @@ The class that will be used for loading spiders, which must implement the
SPIDER_MIDDLEWARES
------------------
Default:: ``{}``
A dict containing the spider middlewares enabled in your project, and their
orders. For more info see :ref:`topics-spider-middleware-setting`.
.. setting:: SPIDER_MIDDLEWARES_BASE
SPIDER_MIDDLEWARES_BASE
-----------------------
Default::
{
@ -919,14 +968,9 @@ Default::
'scrapy.spidermiddlewares.depth.DepthMiddleware': 900,
}
A dict containing the spider middlewares enabled in your project, and their
orders. Low orders are closer to the engine, high orders are closer to the
spider. For more info see :ref:`topics-spider-middleware-setting`.
When you set :setting:`SPIDER_MIDDLEWARES` manually, e.g. in your project's
settings module, it will be merged with the default, not overwrite it. If you
want to disable any of the default spider middlewares you must assign ``None``
as their value.
A dict containing the spider middlewares enabled by default in Scrapy, and
their orders. Low orders are closer to the engine, high orders are closer to
the spider. For more info see :ref:`topics-spider-middleware-setting`.
.. setting:: SPIDER_MODULES

View File

@ -24,20 +24,22 @@ Here's an example::
'myproject.middlewares.CustomSpiderMiddleware': 543,
}
The specified :setting:`SPIDER_MIDDLEWARES` setting is merged with the default
one (i.e. it does not overwrite it) and then sorted by order to get the final
sorted list of enabled middlewares: the first middleware is the one closer to
the engine and the last is the one closer to the spider.
The :setting:`SPIDER_MIDDLEWARES` setting is merged with the
:setting:`SPIDER_MIDDLEWARES_BASE` setting defined in Scrapy (and not meant to
be overridden) and then sorted by order to get the final sorted list of enabled
middlewares: the first middleware is the one closer to the engine and the last
is the one closer to the spider.
To decide which order to assign to your middleware see the default
:setting:`SPIDER_MIDDLEWARES` setting and pick a value according to where
To decide which order to assign to your middleware see the
:setting:`SPIDER_MIDDLEWARES_BASE` setting and pick a value according to where
you want to insert the middleware. The order does matter because each
middleware performs a different action and your middleware could depend on some
previous (or subsequent) middleware being applied.
If you want to disable a builtin middleware you must define it in your project's
:setting:`SPIDER_MIDDLEWARES` setting and assign ``None`` as its value. For
example, if you want to disable the off-site middleware::
If you want to disable a builtin middleware (the ones defined in
:setting:`SPIDER_MIDDLEWARES_BASE`, and enabled by default) you must define it
in your project :setting:`SPIDER_MIDDLEWARES` setting and assign `None` as its
value. For example, if you want to disable the off-site middleware::
SPIDER_MIDDLEWARES = {
'myproject.middlewares.CustomSpiderMiddleware': 543,
@ -171,7 +173,7 @@ information on how to use them and how to write your own spider middleware, see
the :ref:`spider middleware usage guide <topics-spider-middleware>`.
For a list of the components enabled by default (and their orders) see the
:setting:`SPIDER_MIDDLEWARES` setting.
:setting:`SPIDER_MIDDLEWARES_BASE` setting.
DepthMiddleware
---------------

View File

@ -63,7 +63,8 @@ DNS_TIMEOUT = 60
DOWNLOAD_DELAY = 0
DOWNLOAD_HANDLERS = {
DOWNLOAD_HANDLERS = {}
DOWNLOAD_HANDLERS_BASE = {
'file': 'scrapy.core.downloader.handlers.file.FileDownloadHandler',
'http': 'scrapy.core.downloader.handlers.http.HTTPDownloadHandler',
'https': 'scrapy.core.downloader.handlers.http.HTTPDownloadHandler',
@ -81,7 +82,9 @@ DOWNLOADER = 'scrapy.core.downloader.Downloader'
DOWNLOADER_HTTPCLIENTFACTORY = 'scrapy.core.downloader.webclient.ScrapyHTTPClientFactory'
DOWNLOADER_CLIENTCONTEXTFACTORY = 'scrapy.core.downloader.contextfactory.ScrapyClientContextFactory'
DOWNLOADER_MIDDLEWARES = {
DOWNLOADER_MIDDLEWARES = {}
DOWNLOADER_MIDDLEWARES_BASE = {
# Engine side
'scrapy.downloadermiddlewares.robotstxt.RobotsTxtMiddleware': 100,
'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware': 300,
@ -113,7 +116,9 @@ except KeyError:
else:
EDITOR = 'vi'
EXTENSIONS = {
EXTENSIONS = {}
EXTENSIONS_BASE = {
'scrapy.extensions.corestats.CoreStats': 0,
'scrapy.extensions.telnet.TelnetConsole': 0,
'scrapy.extensions.memusage.MemoryUsage': 0,
@ -130,14 +135,16 @@ FEED_URI_PARAMS = None # a function to extend uri arguments
FEED_FORMAT = 'jsonlines'
FEED_STORE_EMPTY = False
FEED_EXPORT_FIELDS = None
FEED_STORAGES = {
FEED_STORAGES = {}
FEED_STORAGES_BASE = {
'': 'scrapy.extensions.feedexport.FileFeedStorage',
'file': 'scrapy.extensions.feedexport.FileFeedStorage',
'stdout': 'scrapy.extensions.feedexport.StdoutFeedStorage',
's3': 'scrapy.extensions.feedexport.S3FeedStorage',
'ftp': 'scrapy.extensions.feedexport.FTPFeedStorage',
}
FEED_EXPORTERS = {
FEED_EXPORTERS = {}
FEED_EXPORTERS_BASE = {
'json': 'scrapy.exporters.JsonItemExporter',
'jsonlines': 'scrapy.exporters.JsonLinesItemExporter',
'jl': 'scrapy.exporters.JsonLinesItemExporter',
@ -163,6 +170,7 @@ HTTPCACHE_GZIP = False
ITEM_PROCESSOR = 'scrapy.pipelines.ItemPipelineManager'
ITEM_PIPELINES = {}
ITEM_PIPELINES_BASE = {}
LOG_ENABLED = True
LOG_ENCODING = 'utf-8'
@ -221,7 +229,9 @@ SCHEDULER_MEMORY_QUEUE = 'scrapy.squeues.LifoMemoryQueue'
SPIDER_LOADER_CLASS = 'scrapy.spiderloader.SpiderLoader'
SPIDER_MIDDLEWARES = {
SPIDER_MIDDLEWARES = {}
SPIDER_MIDDLEWARES_BASE = {
# Engine side
'scrapy.spidermiddlewares.httperror.HttpErrorMiddleware': 50,
'scrapy.spidermiddlewares.offsite.OffsiteMiddleware': 500,
@ -248,7 +258,8 @@ TELNETCONSOLE_ENABLED = 1
TELNETCONSOLE_PORT = [6023, 6073]
TELNETCONSOLE_HOST = '127.0.0.1'
SPIDER_CONTRACTS = {
SPIDER_CONTRACTS = {}
SPIDER_CONTRACTS_BASE = {
'scrapy.contracts.default.UrlContract': 1,
'scrapy.contracts.default.ReturnsContract': 2,
'scrapy.contracts.default.ScrapesContract': 3,