18 KiB
Extensions
The extensions framework provides a mechanism for inserting your own custom functionality into Scrapy.
Extensions are just regular classes.
Extension settings
Extensions use the :ref:`Scrapy settings <topics-settings>` to manage their settings, just like any other Scrapy code.
System Message: ERROR/3 (<stdin>, line 15); backlink
Unknown interpreted text role "ref".It is customary for extensions to prefix their settings with their own name, to avoid collision with existing (and future) extensions. For example, a hypothetical extension to handle Google Sitemaps would use settings like GOOGLESITEMAP_ENABLED, GOOGLESITEMAP_DEPTH, and so on.
Loading & activating extensions
Extensions are loaded and activated at startup by instantiating a single instance of the extension class per spider being run. All the extension initialization code must be performed in the class __init__ method.
To make an extension available, add it to the :setting:`EXTENSIONS` setting in your Scrapy settings. In :setting:`EXTENSIONS`, each extension is represented by a string: the full Python path to the extension's class name. For example:
System Message: ERROR/3 (<stdin>, line 32); backlink
Unknown interpreted text role "setting".System Message: ERROR/3 (<stdin>, line 32); backlink
Unknown interpreted text role "setting".System Message: WARNING/2 (<stdin>, line 36)
Cannot analyze code. Pygments package not found.
.. code-block:: python
EXTENSIONS = {
"scrapy.extensions.corestats.CoreStats": 500,
"scrapy.extensions.telnet.TelnetConsole": 500,
}
As you can see, the :setting:`EXTENSIONS` setting is a dict where the keys are the extension paths, and their values are the orders, which define the extension loading order. The :setting:`EXTENSIONS` setting is merged with the :setting:`EXTENSIONS_BASE` setting defined in Scrapy (and not meant to be overridden) and then sorted by order to get the final sorted list of enabled extensions.
System Message: ERROR/3 (<stdin>, line 44); backlink
Unknown interpreted text role "setting".System Message: ERROR/3 (<stdin>, line 44); backlink
Unknown interpreted text role "setting".System Message: ERROR/3 (<stdin>, line 44); backlink
Unknown interpreted text role "setting".As extensions typically do not depend on each other, their loading order is irrelevant in most cases. This is why the :setting:`EXTENSIONS_BASE` setting defines all extensions with the same order (0). However, this feature can be exploited if you need to add an extension which depends on other extensions already loaded.
System Message: ERROR/3 (<stdin>, line 51); backlink
Unknown interpreted text role "setting".Available, enabled and disabled extensions
Not all available extensions will be enabled. Some of them usually depend on a particular setting. For example, the HTTP Cache extension is available by default but disabled unless the :setting:`HTTPCACHE_ENABLED` setting is set.
System Message: ERROR/3 (<stdin>, line 60); backlink
Unknown interpreted text role "setting".Disabling an extension
In order to disable an extension that comes enabled by default (i.e. those included in the :setting:`EXTENSIONS_BASE` setting) you must set its order to None. For example:
System Message: ERROR/3 (<stdin>, line 67); backlink
Unknown interpreted text role "setting".System Message: WARNING/2 (<stdin>, line 71)
Cannot analyze code. Pygments package not found.
.. code-block:: python
EXTENSIONS = {
"scrapy.extensions.corestats.CoreStats": None,
}
Writing your own extension
Each extension is a Python class. The main entry point for a Scrapy extension (this also includes middlewares and pipelines) is the from_crawler class method which receives a Crawler instance. Through the Crawler object you can access settings, signals, stats, and also control the crawling behaviour.
Typically, extensions connect to :ref:`signals <topics-signals>` and perform tasks triggered by them.
System Message: ERROR/3 (<stdin>, line 85); backlink
Unknown interpreted text role "ref".Finally, if the from_crawler method raises the :exc:`~scrapy.exceptions.NotConfigured` exception, the extension will be disabled. Otherwise, the extension will be enabled.
System Message: ERROR/3 (<stdin>, line 88); backlink
Unknown interpreted text role "exc".Sample extension
Here we will implement a simple extension to illustrate the concepts described in the previous section. This extension will log a message every time:
- a spider is opened
- a spider is closed
- a specific number of items are scraped
The extension will be enabled through the MYEXT_ENABLED setting and the number of items will be specified through the MYEXT_ITEMCOUNT setting.
Here is the code of such extension:
System Message: WARNING/2 (<stdin>, line 107)
Cannot analyze code. Pygments package not found.
.. code-block:: python
import logging
from scrapy import signals
from scrapy.exceptions import NotConfigured
logger = logging.getLogger(__name__)
class SpiderOpenCloseLogging:
def __init__(self, item_count):
self.item_count = item_count
self.items_scraped = 0
@classmethod
def from_crawler(cls, crawler):
# first check if the extension should be enabled and raise
# NotConfigured otherwise
if not crawler.settings.getbool("MYEXT_ENABLED"):
raise NotConfigured
# get the number of items from settings
item_count = crawler.settings.getint("MYEXT_ITEMCOUNT", 1000)
# instantiate the extension object
ext = cls(item_count)
# connect the extension object to signals
crawler.signals.connect(ext.spider_opened, signal=signals.spider_opened)
crawler.signals.connect(ext.spider_closed, signal=signals.spider_closed)
crawler.signals.connect(ext.item_scraped, signal=signals.item_scraped)
# return the extension object
return ext
def spider_opened(self, spider):
logger.info("opened spider %s", spider.name)
def spider_closed(self, spider):
logger.info("closed spider %s", spider.name)
def item_scraped(self, item, spider):
self.items_scraped += 1
if self.items_scraped % self.item_count == 0:
logger.info("scraped %d items", self.items_scraped)
Built-in extensions reference
General purpose extensions
Log Stats extension
System Message: ERROR/3 (<stdin>, line 165)
Unknown directive type "module".
.. module:: scrapy.extensions.logstats :synopsis: Basic stats logging
Log basic stats like crawled pages and scraped items.
Core Stats extension
System Message: ERROR/3 (<stdin>, line 175)
Unknown directive type "module".
.. module:: scrapy.extensions.corestats :synopsis: Core stats collection
Enable the collection of core statistics, provided the stats collection is enabled (see :ref:`topics-stats`).
System Message: ERROR/3 (<stdin>, line 180); backlink
Unknown interpreted text role "ref".Telnet console extension
System Message: ERROR/3 (<stdin>, line 188)
Unknown directive type "module".
.. module:: scrapy.extensions.telnet :synopsis: Telnet console
Provides a telnet console for getting into a Python interpreter inside the currently running Scrapy process, which can be very useful for debugging.
The telnet console must be enabled by the :setting:`TELNETCONSOLE_ENABLED` setting, and the server will listen in the port specified in :setting:`TELNETCONSOLE_PORT`.
System Message: ERROR/3 (<stdin>, line 196); backlink
Unknown interpreted text role "setting".System Message: ERROR/3 (<stdin>, line 196); backlink
Unknown interpreted text role "setting".Memory usage extension
System Message: ERROR/3 (<stdin>, line 205)
Unknown directive type "module".
.. module:: scrapy.extensions.memusage :synopsis: Memory usage extension
Note
This extension does not work in Windows.
Monitors the memory used by the Scrapy process that runs the spider and:
- sends a notification e-mail when it exceeds a certain value
- closes the spider when it exceeds a certain value
The notification e-mails can be triggered when a certain warning value is reached (:setting:`MEMUSAGE_WARNING_MB`) and when the maximum value is reached (:setting:`MEMUSAGE_LIMIT_MB`) which will also cause the spider to be closed and the Scrapy process to be terminated.
System Message: ERROR/3 (<stdin>, line 217); backlink
Unknown interpreted text role "setting".System Message: ERROR/3 (<stdin>, line 217); backlink
Unknown interpreted text role "setting".This extension is enabled by the :setting:`MEMUSAGE_ENABLED` setting and can be configured with the following settings:
System Message: ERROR/3 (<stdin>, line 222); backlink
Unknown interpreted text role "setting".-
System Message: ERROR/3 (<stdin>, line 225); backlink
Unknown interpreted text role "setting".
:setting:`MEMUSAGE_WARNING_MB`
System Message: ERROR/3 (<stdin>, line 226); backlink
Unknown interpreted text role "setting".
:setting:`MEMUSAGE_NOTIFY_MAIL`
System Message: ERROR/3 (<stdin>, line 227); backlink
Unknown interpreted text role "setting".
:setting:`MEMUSAGE_CHECK_INTERVAL_SECONDS`
System Message: ERROR/3 (<stdin>, line 228); backlink
Unknown interpreted text role "setting".
Memory debugger extension
System Message: ERROR/3 (<stdin>, line 233)
Unknown directive type "module".
.. module:: scrapy.extensions.memdebug :synopsis: Memory debugger extension
An extension for debugging memory usage. It collects information about:
objects uncollected by the Python garbage collector
objects left alive that shouldn't. For more info, see :ref:`topics-leaks-trackrefs`
System Message: ERROR/3 (<stdin>, line 241); backlink
Unknown interpreted text role "ref".
To enable this extension, turn on the :setting:`MEMDEBUG_ENABLED` setting. The info will be stored in the stats.
System Message: ERROR/3 (<stdin>, line 243); backlink
Unknown interpreted text role "setting".Spider state extension
System Message: ERROR/3 (<stdin>, line 251)
Unknown directive type "module".
.. module:: scrapy.extensions.spiderstate :synopsis: Spider state extension
Manages spider state data by loading it before a crawl and saving it after.
Give a value to the :setting:`JOBDIR` setting to enable this extension. When enabled, this extension manages the :attr:`~scrapy.Spider.state` attribute of your :class:`~scrapy.Spider` instance:
System Message: ERROR/3 (<stdin>, line 258); backlink
Unknown interpreted text role "setting".System Message: ERROR/3 (<stdin>, line 258); backlink
Unknown interpreted text role "attr".System Message: ERROR/3 (<stdin>, line 258); backlink
Unknown interpreted text role "class".When your spider closes (:signal:`spider_closed`), the contents of its :attr:`~scrapy.Spider.state` attribute are serialized into a file named spider.state in the :setting:`JOBDIR` folder.
System Message: ERROR/3 (<stdin>, line 262); backlink
Unknown interpreted text role "signal".
System Message: ERROR/3 (<stdin>, line 262); backlink
Unknown interpreted text role "attr".
System Message: ERROR/3 (<stdin>, line 262); backlink
Unknown interpreted text role "setting".
When your spider opens (:signal:`spider_opened`), if a previously-generated spider.state file exists in the :setting:`JOBDIR` folder, it is loaded into the :attr:`~scrapy.Spider.state` attribute.
System Message: ERROR/3 (<stdin>, line 265); backlink
Unknown interpreted text role "signal".
System Message: ERROR/3 (<stdin>, line 265); backlink
Unknown interpreted text role "setting".
System Message: ERROR/3 (<stdin>, line 265); backlink
Unknown interpreted text role "attr".
For an example, see :ref:`topics-keeping-persistent-state-between-batches`.
System Message: ERROR/3 (<stdin>, line 270); backlink
Unknown interpreted text role "ref".Close spider extension
System Message: ERROR/3 (<stdin>, line 275)
Unknown directive type "module".
.. module:: scrapy.extensions.closespider :synopsis: Close spider extension
Closes a spider automatically when some conditions are met, using a specific closing reason for each condition.
The conditions for closing a spider can be configured through the following settings:
:setting:`CLOSESPIDER_TIMEOUT`
System Message: ERROR/3 (<stdin>, line 286); backlink
Unknown interpreted text role "setting".
:setting:`CLOSESPIDER_TIMEOUT_NO_ITEM`
System Message: ERROR/3 (<stdin>, line 287); backlink
Unknown interpreted text role "setting".
:setting:`CLOSESPIDER_ITEMCOUNT`
System Message: ERROR/3 (<stdin>, line 288); backlink
Unknown interpreted text role "setting".
:setting:`CLOSESPIDER_PAGECOUNT`
System Message: ERROR/3 (<stdin>, line 289); backlink
Unknown interpreted text role "setting".
:setting:`CLOSESPIDER_ERRORCOUNT`
System Message: ERROR/3 (<stdin>, line 290); backlink
Unknown interpreted text role "setting".
Note
When a certain closing condition is met, requests which are currently in the downloader queue (up to :setting:`CONCURRENT_REQUESTS` requests) are still processed.
System Message: ERROR/3 (<stdin>, line 294); backlink
Unknown interpreted text role "setting".System Message: ERROR/3 (<stdin>, line 298)
Unknown directive type "setting".
.. setting:: CLOSESPIDER_TIMEOUT
CLOSESPIDER_TIMEOUT
Default: 0
An integer which specifies a number of seconds. If the spider remains open for more than that number of second, it will be automatically closed with the reason closespider_timeout. If zero (or non set), spiders won't be closed by timeout.
System Message: ERROR/3 (<stdin>, line 310)
Unknown directive type "setting".
.. setting:: CLOSESPIDER_TIMEOUT_NO_ITEM
CLOSESPIDER_TIMEOUT_NO_ITEM
Default: 0
An integer which specifies a number of seconds. If the spider has not produced any items in the last number of seconds, it will be closed with the reason closespider_timeout_no_item. If zero (or non set), spiders won't be closed regardless if it hasn't produced any items.
System Message: ERROR/3 (<stdin>, line 322)
Unknown directive type "setting".
.. setting:: CLOSESPIDER_ITEMCOUNT
CLOSESPIDER_ITEMCOUNT
Default: 0
An integer which specifies a number of items. If the spider scrapes more than that amount and those items are passed by the item pipeline, the spider will be closed with the reason closespider_itemcount. If zero (or non set), spiders won't be closed by number of passed items.
System Message: ERROR/3 (<stdin>, line 334)
Unknown directive type "setting".
.. setting:: CLOSESPIDER_PAGECOUNT
CLOSESPIDER_PAGECOUNT
Default: 0
An integer which specifies the maximum number of responses to crawl. If the spider crawls more than that, the spider will be closed with the reason closespider_pagecount. If zero (or non set), spiders won't be closed by number of crawled responses.
System Message: ERROR/3 (<stdin>, line 346)
Unknown directive type "setting".
.. setting:: CLOSESPIDER_PAGECOUNT_NO_ITEM
CLOSESPIDER_PAGECOUNT_NO_ITEM
Default: 0
An integer which specifies the maximum number of consecutive responses to crawl without items scraped. If the spider crawls more consecutive responses than that and no items are scraped in the meantime, the spider will be closed with the reason closespider_pagecount_no_item. If zero (or not set), spiders won't be closed by number of crawled responses with no items.
System Message: ERROR/3 (<stdin>, line 359)
Unknown directive type "setting".
.. setting:: CLOSESPIDER_ERRORCOUNT
CLOSESPIDER_ERRORCOUNT
Default: 0
An integer which specifies the maximum number of errors to receive before closing the spider. If the spider generates more than that number of errors, it will be closed with the reason closespider_errorcount. If zero (or non set), spiders won't be closed by number of errors.
StatsMailer extension
System Message: ERROR/3 (<stdin>, line 374)
Unknown directive type "module".
.. module:: scrapy.extensions.statsmailer :synopsis: StatsMailer extension
This simple extension can be used to send a notification e-mail every time a domain has finished scraping, including the Scrapy stats collected. The email will be sent to all recipients specified in the :setting:`STATSMAILER_RCPTS` setting.
System Message: ERROR/3 (<stdin>, line 379); backlink
Unknown interpreted text role "setting".Emails can be sent using the :class:`~scrapy.mail.MailSender` class. To see a full list of parameters, including examples on how to instantiate :class:`~scrapy.mail.MailSender` and use mail settings, see :ref:`topics-email`.
System Message: ERROR/3 (<stdin>, line 384); backlink
Unknown interpreted text role "class".System Message: ERROR/3 (<stdin>, line 384); backlink
Unknown interpreted text role "class".System Message: ERROR/3 (<stdin>, line 384); backlink
Unknown interpreted text role "ref".System Message: ERROR/3 (<stdin>, line 389)
Unknown directive type "module".
.. module:: scrapy.extensions.debug :synopsis: Extensions for debugging Scrapy
System Message: ERROR/3 (<stdin>, line 392)
Unknown directive type "module".
.. module:: scrapy.extensions.periodic_log :synopsis: Periodic stats logging
Periodic log extension
This extension periodically logs rich stat data as a JSON object:
2023-08-04 02:30:57 [scrapy.extensions.logstats] INFO: Crawled 976 pages (at 162 pages/min), scraped 925 items (at 161 items/min)
2023-08-04 02:30:57 [scrapy.extensions.periodic_log] INFO: {
"delta": {
"downloader/request_bytes": 55582,
"downloader/request_count": 162,
"downloader/request_method_count/GET": 162,
"downloader/response_bytes": 618133,
"downloader/response_count": 162,
"downloader/response_status_count/200": 162,
"item_scraped_count": 161
},
"stats": {
"downloader/request_bytes": 338243,
"downloader/request_count": 992,
"downloader/request_method_count/GET": 992,
"downloader/response_bytes": 3836736,
"downloader/response_count": 976,
"downloader/response_status_count/200": 976,
"item_scraped_count": 925,
"log_count/INFO": 21,
"log_count/WARNING": 1,
"scheduler/dequeued": 992,
"scheduler/dequeued/memory": 992,
"scheduler/enqueued": 1050,
"scheduler/enqueued/memory": 1050
},
"time": {
"elapsed": 360.008903,
"log_interval": 60.0,
"log_interval_real": 60.006694,
"start_time": "2023-08-03 23:24:57",
"utcnow": "2023-08-03 23:30:57"
}
}
This extension logs the following configurable sections:
"delta" shows how some numeric stats have changed since the last stats log message.
The :setting:`PERIODIC_LOG_DELTA` setting determines the target stats. They must have int or float values.
System Message: ERROR/3 (<stdin>, line 442); backlink
Unknown interpreted text role "setting".
"stats" shows the current value of some stats.
The :setting:`PERIODIC_LOG_STATS` setting determines the target stats.
System Message: ERROR/3 (<stdin>, line 447); backlink
Unknown interpreted text role "setting".
"time" shows detailed timing data.
The :setting:`PERIODIC_LOG_TIMING_ENABLED` setting determines whether or not to show this section.
System Message: ERROR/3 (<stdin>, line 451); backlink
Unknown interpreted text role "setting".
This extension logs data at the start, then on a fixed time interval configurable through the :setting:`LOGSTATS_INTERVAL` setting, and finally right before the crawl ends.
System Message: ERROR/3 (<stdin>, line 454); backlink
Unknown interpreted text role "setting".Example extension configuration:
System Message: WARNING/2 (<stdin>, line 461)
Cannot analyze code. Pygments package not found.
.. code-block:: python
custom_settings = {
"LOG_LEVEL": "INFO",
"PERIODIC_LOG_STATS": {
"include": ["downloader/", "scheduler/", "log_count/", "item_scraped_count/"],
},
"PERIODIC_LOG_DELTA": {"include": ["downloader/"]},
"PERIODIC_LOG_TIMING_ENABLED": True,
"EXTENSIONS": {
"scrapy.extensions.periodic_log.PeriodicLog": 0,
},
}
System Message: ERROR/3 (<stdin>, line 475)
Unknown directive type "setting".
.. setting:: PERIODIC_LOG_DELTA
PERIODIC_LOG_DELTA
Default: None
- "PERIODIC_LOG_DELTA": True - show deltas for all int and float stat values.
- "PERIODIC_LOG_DELTA": {"include": ["downloader/", "scheduler/"]} - show deltas for stats with names containing any configured substring.
- "PERIODIC_LOG_DELTA": {"exclude": ["downloader/"]} - show deltas for all stats with names not containing any configured substring.
System Message: ERROR/3 (<stdin>, line 486)
Unknown directive type "setting".
.. setting:: PERIODIC_LOG_STATS
PERIODIC_LOG_STATS
Default: None
- "PERIODIC_LOG_STATS": True - show the current value of all stats.
- "PERIODIC_LOG_STATS": {"include": ["downloader/", "scheduler/"]} - show current values for stats with names containing any configured substring.
- "PERIODIC_LOG_STATS": {"exclude": ["downloader/"]} - show current values for all stats with names not containing any configured substring.
System Message: ERROR/3 (<stdin>, line 498)
Unknown directive type "setting".
.. setting:: PERIODIC_LOG_TIMING_ENABLED
PERIODIC_LOG_TIMING_ENABLED
Default: False
True enables logging of timing data (i.e. the "time" section).
Debugging extensions
Stack trace dump extension
Dumps information about the running process when a SIGQUIT or SIGUSR2 signal is received. The information dumped is the following:
engine status (using scrapy.utils.engine.get_engine_status())
live references (see :ref:`topics-leaks-trackrefs`)
System Message: ERROR/3 (<stdin>, line 520); backlink
Unknown interpreted text role "ref".
stack trace of all threads
After the stack trace and engine status is dumped, the Scrapy process continues running normally.
This extension only works on POSIX-compliant platforms (i.e. not Windows), because the SIGQUIT and SIGUSR2 signals are not available on Windows.
There are at least two ways to send Scrapy the SIGQUIT signal:
By pressing Ctrl-while a Scrapy process is running (Linux only?)
By running this command (assuming <pid> is the process id of the Scrapy process):
kill -QUIT <pid>
Debugger extension
Invokes a :doc:`Python debugger <library/pdb>` inside a running Scrapy process when a SIGUSR2 signal is received. After the debugger is exited, the Scrapy process continues running normally.
System Message: ERROR/3 (<stdin>, line 545); backlink
Unknown interpreted text role "doc".This extension only works on POSIX-compliant platforms (i.e. not Windows).