diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 08d8f3edf..cc0254d29 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -23,20 +23,22 @@ Here's an example:: 'myproject.middlewares.CustomDownloaderMiddleware': 543, } -The specified :setting:`DOWNLOADER_MIDDLEWARES` setting is merged with the -default one (i.e. it does not overwrite it) and then sorted by order to get the -final sorted list of enabled middlewares: the first middleware is the one -closer to the engine and the last is the one closer to the downloader. +The :setting:`DOWNLOADER_MIDDLEWARES` setting is merged with the +:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting defined in Scrapy (and not meant +to be overridden) and then sorted by order to get the final sorted list of +enabled middlewares: the first middleware is the one closer to the engine and +the last is the one closer to the downloader. -To decide which order to assign to your middleware see the default -:setting:`DOWNLOADER_MIDDLEWARES` setting and pick a value according to +To decide which order to assign to your middleware see the +:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting and pick a value according to where you want to insert the middleware. The order does matter because each middleware performs a different action and your middleware could depend on some previous (or subsequent) middleware being applied. -If you want to disable a built-in middleware you must define it in your -project's :setting:`DOWNLOADER_MIDDLEWARES` setting and assign ``None`` as its -value. For example, if you want to disable the user-agent middleware:: +If you want to disable a built-in middleware (the ones defined in +:setting:`DOWNLOADER_MIDDLEWARES_BASE` and enabled by default) you must define it +in your project's :setting:`DOWNLOADER_MIDDLEWARES` setting and assign `None` +as its value. For example, if you want to disable the user-agent middleware:: DOWNLOADER_MIDDLEWARES = { 'myproject.middlewares.CustomDownloaderMiddleware': 543, @@ -162,7 +164,7 @@ middleware, see the :ref:`downloader middleware usage guide `. For a list of the components enabled by default (and their orders) see the -:setting:`DOWNLOADER_MIDDLEWARES` setting. +:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting. .. _cookies-mw: diff --git a/docs/topics/extensions.rst b/docs/topics/extensions.rst index 11c0aadb6..847353868 100644 --- a/docs/topics/extensions.rst +++ b/docs/topics/extensions.rst @@ -42,13 +42,14 @@ by a string: the full Python path to the extension's class name. For example:: As you can see, the :setting:`EXTENSIONS` setting is a dict where the keys are the extension paths, and their values are the orders, which define the -extension *loading* order. The specified :setting:`EXTENSIONS` setting is merged -with the default one (i.e. it does not overwrite it) and then sorted by order -to get the final sorted list of enabled extensions. +extension *loading* order. The :setting:`EXTENSIONS` setting is merged with the +:setting:`EXTENSIONS_BASE` setting defined in Scrapy (and not meant to be +overridden) and then sorted by order to get the final sorted list of enabled +extensions. As extensions typically do not depend on each other, their loading order is -irrelevant in most cases. This is why the default :setting:`EXTENSIONS` setting -defines all extensions with the same order (``500``). However, this feature can +irrelevant in most cases. This is why the :setting:`EXTENSIONS_BASE` setting +defines all extensions with the same order (``0``). However, this feature can be exploited if you need to add an extension which depends on other extensions already loaded. @@ -63,7 +64,7 @@ Disabling an extension ====================== In order to disable an extension that comes enabled by default (ie. those -included in the default :setting:`EXTENSIONS` setting) you must set its order to +included in the :setting:`EXTENSIONS_BASE` setting) you must set its order to ``None``. For example:: EXTENSIONS = { diff --git a/docs/topics/feed-exports.rst b/docs/topics/feed-exports.rst index d8b8da166..03c6fb3fb 100644 --- a/docs/topics/feed-exports.rst +++ b/docs/topics/feed-exports.rst @@ -265,6 +265,16 @@ Whether to export empty feeds (ie. feeds with no items). FEED_STORAGES ------------- +Default:: ``{}`` + +A dict containing additional feed storage backends supported by your project. +The keys are URI schemes and the values are paths to storage classes. + +.. setting:: FEED_STORAGES_BASE + +FEED_STORAGES_BASE +------------------ + Default:: { @@ -275,19 +285,30 @@ Default:: 'ftp': 'scrapy.extensions.feedexport.FTPFeedStorage', } -A dict containing all feed storage backends supported by your project. The keys -are URI schemes and the values are paths to storage classes. +A dict containing the built-in feed storage backends supported by Scrapy. You +can disable any of these backends by assigning ``None`` to their URI scheme in +:setting:`FEED_STORAGES`. E.g., to disable the built-in FTP storage backend +(without replacement), place this in your ``settings.py``:: -When you set :setting:`FEED_STORAGES` manually, e.g. in your project's settings -module, it will be merged with the default, not overwrite it. If you want to -disable any of the default feed storage backends, you must assign ``None`` as -their value. + FEED_STORAGES = { + 'ftp': None, + } .. setting:: FEED_EXPORTERS FEED_EXPORTERS -------------- +Default:: ``{}`` + +A dict containing additional exporters supported by your project. The keys are +serialization formats and the values are paths to :ref:`Item exporter +` classes. + +.. setting:: FEED_EXPORTERS_BASE + +FEED_EXPORTERS_BASE +------------------- Default:: { @@ -300,14 +321,14 @@ Default:: 'pickle': 'scrapy.exporters.PickleItemExporter', } -A dict containing all feed exporters supported by your project. The keys are -URI schemes and the values are paths to :ref:`Item exporter ` -classes. +A dict containing the built-in feed exporters supported by Scrapy. You can +disable any of these exporters by assigning ``None`` to their serialization +format in :setting:`FEED_EXPORTERS`. E.g., to disable the built-in CSV exporter +(without replacement), place this in your ``settings.py``:: -When you set :setting:`FEED_EXPORTERS` manually, e.g. in your project's settings -module, it will be merged with the default, not overwrite it. If you want to -disable any of the default feed exporters, you must assign ``None`` as their -value. + FEED_EXPORTERS = { + 'csv': None, + } .. _URI: http://en.wikipedia.org/wiki/Uniform_Resource_Identifier .. _Amazon S3: http://aws.amazon.com/s3/ diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 8908fae7e..aa0417e1a 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -269,11 +269,6 @@ Default:: The default headers used for Scrapy HTTP Requests. They're populated in the :class:`~scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware`. -When you set :setting:`DEFAULT_REQUEST_HEADERS` manually, e.g. in your -project's settings module, it will be merged with the default, not overwrite it. -If you want to disable any of the default request headers (and not replace them) -you must assign ``None`` as their value. - .. setting:: DEPTH_LIMIT DEPTH_LIMIT @@ -355,6 +350,16 @@ The downloader to use for crawling. DOWNLOADER_MIDDLEWARES ---------------------- +Default:: ``{}`` + +A dict containing the downloader middlewares enabled in your project, and their +orders. For more info see :ref:`topics-downloader-middleware-setting`. + +.. setting:: DOWNLOADER_MIDDLEWARES_BASE + +DOWNLOADER_MIDDLEWARES_BASE +--------------------------- + Default:: { @@ -375,16 +380,11 @@ Default:: 'scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware': 900, } -A dict containing the downloader middlewares enabled in your project, and their -orders. Low orders are closer to the engine, high orders are closer to the -downloader. - -When you set :setting:`DOWNLOADER_MIDDLEWARES` manually, e.g. in your project's -settings module, it will be merged with the default, not overwrite it. If you -want to disable any of the default downloader middlewares you must assign -``None`` as their value. - -For more info see :ref:`topics-downloader-middleware-setting`. +A dict containing the downloader middlewares enabled by default in Scrapy. Low +orders are closer to the engine, high orders are closer to the downloader. You +should never modify this setting in your project, modify +:setting:`DOWNLOADER_MIDDLEWARES` instead. For more info see +:ref:`topics-downloader-middleware-setting`. .. setting:: DOWNLOADER_STATS @@ -425,6 +425,16 @@ spider attribute. DOWNLOAD_HANDLERS ----------------- +Default: ``{}`` + +A dict containing the request downloader handlers enabled in your project. +See :setting:`DOWNLOAD_HANDLERS_BASE` for example format. + +.. setting:: DOWNLOAD_HANDLERS_BASE + +DOWNLOAD_HANDLERS_BASE +---------------------- + Default:: { @@ -436,15 +446,16 @@ Default:: } -A dict containing the request downloader handlers enabled in your project. +A dict containing the request download handlers enabled by default in Scrapy. +You should never modify this setting in your project, modify +:setting:`DOWNLOAD_HANDLERS` instead. -When you set :setting:`DOWNLOAD_HANDLERS` manually, e.g. in your project's -settings module, it will be merged with the default, not overwrite it. If you -want to disable any of the default download handlers you must assign ``None`` -as their value. For example, if you want to disable the file download handler:: +You can disable any of these download handlers by assigning ``None`` to their +URI scheme in :setting:`DOWNLOAD_HANDLERS`. E.g., to disable the built-in FTP +handler (without replacement), place this in your ``settings.py``:: DOWNLOAD_HANDLERS = { - 'file': None, + 'ftp': None, } .. setting:: DOWNLOAD_TIMEOUT @@ -544,6 +555,15 @@ to ``vi`` (on Unix systems) or the IDLE editor (on Windows). EXTENSIONS ---------- +Default:: ``{}`` + +A dict containing the extensions enabled in your project, and their orders. + +.. setting:: EXTENSIONS_BASE + +EXTENSIONS_BASE +--------------- + Default:: { @@ -558,15 +578,10 @@ Default:: 'scrapy.extensions.throttle.AutoThrottle': 0, } -A dict containing the extensions enabled in your project, and their orders. By -default, this setting contains all stable built-in extensions. Keep in mind that +A dict containing the extensions available by default in Scrapy, and their +orders. This setting contains all stable built-in extensions. Keep in mind that some of them need to be enabled through a setting. -When you set :setting:`EXTENSIONS` manually, e.g. in your project's settings -module, it will be merged with the default, not overwrite it. If you want to -disable any of the default enabled extensions you must assign ``None`` as their -value. - For more information See the :ref:`extensions user guide ` and the :ref:`list of available extensions `. @@ -589,6 +604,16 @@ Example:: 'mybot.pipelines.validate.StoreMyItem': 800, } +.. setting:: ITEM_PIPELINES_BASE + +ITEM_PIPELINES_BASE +------------------- + +Default: ``{}`` + +A dict containing the pipelines enabled by default in Scrapy. You should never +modify this setting in your project, modify :setting:`ITEM_PIPELINES` instead. + .. setting:: LOG_ENABLED LOG_ENABLED @@ -878,6 +903,16 @@ The scheduler to use for crawling. SPIDER_CONTRACTS ---------------- +Default:: ``{}`` + +A dict containing the spider contracts enabled in your project, used for +testing spiders. For more info see :ref:`topics-contracts`. + +.. setting:: SPIDER_CONTRACTS_BASE + +SPIDER_CONTRACTS_BASE +--------------------- + Default:: { @@ -886,13 +921,17 @@ Default:: 'scrapy.contracts.default.ScrapesContract': 3, } -A dict containing the scrapy contracts enabled in your project, used for -testing spiders. For more info see :ref:`topics-contracts`. +A dict containing the scrapy contracts enabled by default in Scrapy. You should +never modify this setting in your project, modify :setting:`SPIDER_CONTRACTS` +instead. For more info see :ref:`topics-contracts`. -When you set :setting:`SPIDER_CONTRACTS` manually, e.g. in your project's -settings module, it will be merged with the default, not overwrite it. If you -want to disable any of the default contracts you must assign ``None`` as their -value. +You can disable any of these contracts by assigning ``None`` to their class +path in :setting:`SPIDER_CONTRACTS`. E.g., to disable the built-in +``ScrapesContract``, place this in your ``settings.py``:: + + SPIDER_CONTRACTS = { + 'scrapy.contracts.default.ScrapesContract': None, + } .. setting:: SPIDER_LOADER_CLASS @@ -909,6 +948,16 @@ The class that will be used for loading spiders, which must implement the SPIDER_MIDDLEWARES ------------------ +Default:: ``{}`` + +A dict containing the spider middlewares enabled in your project, and their +orders. For more info see :ref:`topics-spider-middleware-setting`. + +.. setting:: SPIDER_MIDDLEWARES_BASE + +SPIDER_MIDDLEWARES_BASE +----------------------- + Default:: { @@ -919,14 +968,9 @@ Default:: 'scrapy.spidermiddlewares.depth.DepthMiddleware': 900, } -A dict containing the spider middlewares enabled in your project, and their -orders. Low orders are closer to the engine, high orders are closer to the -spider. For more info see :ref:`topics-spider-middleware-setting`. - -When you set :setting:`SPIDER_MIDDLEWARES` manually, e.g. in your project's -settings module, it will be merged with the default, not overwrite it. If you -want to disable any of the default spider middlewares you must assign ``None`` -as their value. +A dict containing the spider middlewares enabled by default in Scrapy, and +their orders. Low orders are closer to the engine, high orders are closer to +the spider. For more info see :ref:`topics-spider-middleware-setting`. .. setting:: SPIDER_MODULES diff --git a/docs/topics/spider-middleware.rst b/docs/topics/spider-middleware.rst index d448801d3..84daaaa55 100644 --- a/docs/topics/spider-middleware.rst +++ b/docs/topics/spider-middleware.rst @@ -24,20 +24,22 @@ Here's an example:: 'myproject.middlewares.CustomSpiderMiddleware': 543, } -The specified :setting:`SPIDER_MIDDLEWARES` setting is merged with the default -one (i.e. it does not overwrite it) and then sorted by order to get the final -sorted list of enabled middlewares: the first middleware is the one closer to -the engine and the last is the one closer to the spider. +The :setting:`SPIDER_MIDDLEWARES` setting is merged with the +:setting:`SPIDER_MIDDLEWARES_BASE` setting defined in Scrapy (and not meant to +be overridden) and then sorted by order to get the final sorted list of enabled +middlewares: the first middleware is the one closer to the engine and the last +is the one closer to the spider. -To decide which order to assign to your middleware see the default -:setting:`SPIDER_MIDDLEWARES` setting and pick a value according to where +To decide which order to assign to your middleware see the +:setting:`SPIDER_MIDDLEWARES_BASE` setting and pick a value according to where you want to insert the middleware. The order does matter because each middleware performs a different action and your middleware could depend on some previous (or subsequent) middleware being applied. -If you want to disable a builtin middleware you must define it in your project's -:setting:`SPIDER_MIDDLEWARES` setting and assign ``None`` as its value. For -example, if you want to disable the off-site middleware:: +If you want to disable a builtin middleware (the ones defined in +:setting:`SPIDER_MIDDLEWARES_BASE`, and enabled by default) you must define it +in your project :setting:`SPIDER_MIDDLEWARES` setting and assign `None` as its +value. For example, if you want to disable the off-site middleware:: SPIDER_MIDDLEWARES = { 'myproject.middlewares.CustomSpiderMiddleware': 543, @@ -171,7 +173,7 @@ information on how to use them and how to write your own spider middleware, see the :ref:`spider middleware usage guide `. For a list of the components enabled by default (and their orders) see the -:setting:`SPIDER_MIDDLEWARES` setting. +:setting:`SPIDER_MIDDLEWARES_BASE` setting. DepthMiddleware --------------- diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 375efcdbb..8435b0354 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -63,7 +63,8 @@ DNS_TIMEOUT = 60 DOWNLOAD_DELAY = 0 -DOWNLOAD_HANDLERS = { +DOWNLOAD_HANDLERS = {} +DOWNLOAD_HANDLERS_BASE = { 'file': 'scrapy.core.downloader.handlers.file.FileDownloadHandler', 'http': 'scrapy.core.downloader.handlers.http.HTTPDownloadHandler', 'https': 'scrapy.core.downloader.handlers.http.HTTPDownloadHandler', @@ -81,7 +82,9 @@ DOWNLOADER = 'scrapy.core.downloader.Downloader' DOWNLOADER_HTTPCLIENTFACTORY = 'scrapy.core.downloader.webclient.ScrapyHTTPClientFactory' DOWNLOADER_CLIENTCONTEXTFACTORY = 'scrapy.core.downloader.contextfactory.ScrapyClientContextFactory' -DOWNLOADER_MIDDLEWARES = { +DOWNLOADER_MIDDLEWARES = {} + +DOWNLOADER_MIDDLEWARES_BASE = { # Engine side 'scrapy.downloadermiddlewares.robotstxt.RobotsTxtMiddleware': 100, 'scrapy.downloadermiddlewares.httpauth.HttpAuthMiddleware': 300, @@ -113,7 +116,9 @@ except KeyError: else: EDITOR = 'vi' -EXTENSIONS = { +EXTENSIONS = {} + +EXTENSIONS_BASE = { 'scrapy.extensions.corestats.CoreStats': 0, 'scrapy.extensions.telnet.TelnetConsole': 0, 'scrapy.extensions.memusage.MemoryUsage': 0, @@ -130,14 +135,16 @@ FEED_URI_PARAMS = None # a function to extend uri arguments FEED_FORMAT = 'jsonlines' FEED_STORE_EMPTY = False FEED_EXPORT_FIELDS = None -FEED_STORAGES = { +FEED_STORAGES = {} +FEED_STORAGES_BASE = { '': 'scrapy.extensions.feedexport.FileFeedStorage', 'file': 'scrapy.extensions.feedexport.FileFeedStorage', 'stdout': 'scrapy.extensions.feedexport.StdoutFeedStorage', 's3': 'scrapy.extensions.feedexport.S3FeedStorage', 'ftp': 'scrapy.extensions.feedexport.FTPFeedStorage', } -FEED_EXPORTERS = { +FEED_EXPORTERS = {} +FEED_EXPORTERS_BASE = { 'json': 'scrapy.exporters.JsonItemExporter', 'jsonlines': 'scrapy.exporters.JsonLinesItemExporter', 'jl': 'scrapy.exporters.JsonLinesItemExporter', @@ -163,6 +170,7 @@ HTTPCACHE_GZIP = False ITEM_PROCESSOR = 'scrapy.pipelines.ItemPipelineManager' ITEM_PIPELINES = {} +ITEM_PIPELINES_BASE = {} LOG_ENABLED = True LOG_ENCODING = 'utf-8' @@ -221,7 +229,9 @@ SCHEDULER_MEMORY_QUEUE = 'scrapy.squeues.LifoMemoryQueue' SPIDER_LOADER_CLASS = 'scrapy.spiderloader.SpiderLoader' -SPIDER_MIDDLEWARES = { +SPIDER_MIDDLEWARES = {} + +SPIDER_MIDDLEWARES_BASE = { # Engine side 'scrapy.spidermiddlewares.httperror.HttpErrorMiddleware': 50, 'scrapy.spidermiddlewares.offsite.OffsiteMiddleware': 500, @@ -248,7 +258,8 @@ TELNETCONSOLE_ENABLED = 1 TELNETCONSOLE_PORT = [6023, 6073] TELNETCONSOLE_HOST = '127.0.0.1' -SPIDER_CONTRACTS = { +SPIDER_CONTRACTS = {} +SPIDER_CONTRACTS_BASE = { 'scrapy.contracts.default.UrlContract': 1, 'scrapy.contracts.default.ReturnsContract': 2, 'scrapy.contracts.default.ScrapesContract': 3,