diff --git a/docs/index.rst b/docs/index.rst index 1a9cf636c..f4417c445 100644 --- a/docs/index.rst +++ b/docs/index.rst @@ -152,10 +152,10 @@ Solving specific problems :hidden: faq + topics/optimize topics/debug topics/contracts topics/practices - topics/broad-crawls topics/developer-tools topics/dynamic-content topics/leaks @@ -170,6 +170,9 @@ Solving specific problems :doc:`faq` Get answers to most frequently asked questions. +:doc:`topics/optimize` + Tune Scrapy for specific use cases and to save on specific resources. + :doc:`topics/debug` Learn how to debug common problems of your Scrapy spider. @@ -179,9 +182,6 @@ Solving specific problems :doc:`topics/practices` Get familiar with some Scrapy common practices. -:doc:`topics/broad-crawls` - Tune Scrapy for crawling a lot domains in parallel. - :doc:`topics/developer-tools` Learn how to scrape with your browser's developer tools. diff --git a/docs/news.rst b/docs/news.rst index 34d0b0880..890b8eab0 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -723,7 +723,7 @@ Bug fixes - During :ref:`feed export `, do not close the underlying file from :ref:`built-in post-processing plugins - `. + `. (:issue:`5932`, :issue:`6178`, :issue:`6239`) - :class:`LinkExtractor ` @@ -2042,8 +2042,9 @@ New features :issue:`5161`, :issue:`5203`) - You can now apply :ref:`post-processing ` to feeds, and - :ref:`built-in post-processing plugins ` are provided for - output file compression. (:issue:`2174`, :issue:`5168`, :issue:`5190`) + :ref:`built-in post-processing plugins ` are + provided for output file compression. + (:issue:`2174`, :issue:`5168`, :issue:`5190`) - The :setting:`FEEDS` setting now supports :class:`pathlib.Path` objects as keys. (:issue:`5383`, :issue:`5384`) diff --git a/docs/topics/feed-exports.rst b/docs/topics/feed-exports.rst index 7f401f0c7..aff3cb29f 100644 --- a/docs/topics/feed-exports.rst +++ b/docs/topics/feed-exports.rst @@ -341,14 +341,15 @@ ItemFilter .. _post-processing: -Post-Processing +Post-processing =============== .. versionadded:: 2.6.0 -Scrapy provides an option to activate plugins to post-process feeds before they are exported -to feed storages. In addition to using :ref:`builtin plugins `, you -can create your own :ref:`plugins `. +Scrapy provides an option to activate plugins to post-process feeds before they +are exported to feed storages. In addition to using :ref:`built-in plugins +`, you can create your own :ref:`plugins +`. These plugins can be activated through the ``postprocessing`` option of a feed. The option must be passed a list of post-processing plugins in the order you want @@ -356,10 +357,10 @@ the feed to be processed. These plugins can be declared either as an import stri or with the imported class of the plugin. Parameters to plugins can be passed through the feed options. See :ref:`feed options ` for examples. -.. _builtin-plugins: +.. _post-processing-plugins: -Built-in Plugins ----------------- +Built-in post-processing plugins +-------------------------------- .. autoclass:: scrapy.extensions.postprocessing.GzipPlugin @@ -367,39 +368,16 @@ Built-in Plugins .. autoclass:: scrapy.extensions.postprocessing.Bz2Plugin -.. _custom-plugins: -Custom Plugins --------------- +.. _custom-post-processing-plugin: -Each plugin is a class that must implement the following methods: +Writing a post-processing plugin +-------------------------------- -.. method:: __init__(self, file, feed_options) - - Initialize the plugin. - - :param file: file-like object having at least the `write`, `tell` and `close` methods implemented - - :param feed_options: feed-specific :ref:`options ` - :type feed_options: :class:`dict` - -.. method:: write(self, data) - - Process and write `data` (:class:`bytes` or :class:`memoryview`) into the plugin's target file. - It must return number of bytes written. - -.. method:: close(self) - - Clean up the plugin. - - For example, you might want to close a file wrapper that you might have - used to compress data written into the file received in the ``__init__`` - method. - - .. warning:: Do not close the file from the ``__init__`` method. - -To pass a parameter to your plugin, use :ref:`feed options `. You -can then access those parameters from the ``__init__`` method of your plugin. +.. autoclass:: scrapy.extensions.postprocessing.PostprocessingPluginProtocol() + :members: + :special-members: __init__ + :member-order: bysource Settings diff --git a/docs/topics/broad-crawls.rst b/docs/topics/optimize.rst similarity index 88% rename from docs/topics/broad-crawls.rst rename to docs/topics/optimize.rst index b50b64cd6..655415cd4 100644 --- a/docs/topics/broad-crawls.rst +++ b/docs/topics/optimize.rst @@ -1,8 +1,30 @@ +.. _optimize: + +============= +Optimizations +============= + +Scrapy offers different ways to optimize crawls based on :ref:`resource +constraints ` and :ref:`use cases `. + +.. _optimize-resources: + +Lowering resource usage +======================= + +… + +.. + TODO: + Network input and output + optional compression packages + + +.. _broad-crawls: .. _topics-broad-crawls: -============ -Broad crawls -============ +Optimizing broad crawls +======================= While Scrapy is well suited for **broad crawls**, i.e. crawls that target many websites, the default :ref:`settings ` are optimized for @@ -18,7 +40,7 @@ For broad crawls, consider these adjustments: .. _broad-crawls-concurrency: Increase concurrency -==================== +-------------------- Concurrency is the number of requests that are processed in parallel. There is a global limit (:setting:`CONCURRENT_REQUESTS`) and an additional limit that @@ -49,7 +71,7 @@ concern, you might need to lower your global concurrency limit accordingly. Increase Twisted IO thread pool maximum size -============================================ +-------------------------------------------- Currently Scrapy does DNS resolution in a blocking way with usage of thread pool. With higher concurrency levels the crawling could be slow or even fail @@ -64,7 +86,7 @@ To increase maximum thread pool size use: REACTOR_THREADPOOL_MAXSIZE = 20 Setup your own DNS -================== +------------------ If you have multiple crawling processes and single central DNS, it can act like DoS attack on the DNS server resulting to slow down of entire network or @@ -72,7 +94,7 @@ even blocking your machines. To avoid this setup your own DNS server with local cache and upstream to some large DNS like OpenDNS or Verizon. Reduce log level -================ +---------------- When doing broad crawls you are often only interested in the crawl rates you get and any errors found. These stats are reported by Scrapy when using the @@ -88,7 +110,7 @@ To set the log level use: LOG_LEVEL = "INFO" Disable cookies -=============== +--------------- Disable cookies unless you *really* need. Cookies are often not needed when doing broad crawls (search engine crawlers ignore them), and they improve @@ -102,7 +124,7 @@ To disable cookies use: COOKIES_ENABLED = False Disable retries -=============== +--------------- Retrying failed HTTP requests can slow down the crawls substantially, specially when sites causes are very slow (or fail) to respond, thus causing a timeout @@ -116,7 +138,7 @@ To disable retries use: RETRY_ENABLED = False Reduce download timeout -======================= +----------------------- Unless you are crawling from a very slow connection (which shouldn't be the case for broad crawls) reduce the download timeout so that stuck requests are @@ -129,7 +151,7 @@ To reduce the download timeout use: DOWNLOAD_TIMEOUT = 15 Disable redirects -================= +----------------- Consider disabling redirects, unless you are interested in following them. When doing broad crawls it's common to save redirects and resolve them when @@ -146,7 +168,7 @@ To disable redirects use: .. _broad-crawls-bfo: Crawl in BFO order -================== +------------------ :ref:`Scrapy crawls in DFO order by default `. @@ -158,7 +180,7 @@ final depth is reached, which can significantly increase memory usage. Be mindful of memory leaks -========================== +-------------------------- If your broad crawl shows a high memory usage, in addition to :ref:`crawling in BFO order ` and :ref:`lowering concurrency @@ -167,7 +189,7 @@ BFO order ` and :ref:`lowering concurrency Install a specific Twisted reactor -================================== +---------------------------------- If the crawl is exceeding the system's capabilities, you might want to try installing a specific Twisted reactor, via the :setting:`TWISTED_REACTOR` setting. diff --git a/docs/topics/scheduler.rst b/docs/topics/scheduler.rst index 30ddda492..38adee1f7 100644 --- a/docs/topics/scheduler.rst +++ b/docs/topics/scheduler.rst @@ -6,50 +6,122 @@ Scheduler .. module:: scrapy.core.scheduler -The **scheduler** is a :ref:`component ` that stores pending -requests, drops unwanted requests, and determines in which order pending -requests are sent. +The **scheduler** is a :ref:`component ` that :ref:`stores +pending requests, drops unwanted requests, and determines in which order +pending requests are sent `. -It is set in the :setting:`SCHEDULER` setting. +It is set in the :setting:`SCHEDULER` setting. There is only 1 built-in +scheduler, :class:`~scrapy.core.scheduler.Scheduler`, but you can +:ref:`implement your own `. -Pending requests may come from seeding (see :setting:`SEEDING_POLICY`), -spider callbacks (:attr:`Request.callback `), -:ref:`spider middlewares ` or :ref:`downloader -middlewares `. +.. _scheduler-responsibilities: -How requests are **stored** depends on the scheduler. The built-in scheduler, -:class:`~scrapy.core.scheduler.Scheduler`, can store requests in memory or -disk. Other schedulers may rely, for example, on frontier, queue, database or -storage services. +Scheduler responsibilities +========================== -Which requests are **dropped** also depends on the scheduler. It is recommended -for schedulers to use the configured :setting:`DUPEFILTER_CLASS` and take into -account :attr:`Request.dont_filter `, but -schedulers are free to follow their own criteria for dropping requests. +A scheduler must: -How requests are **prioritized**, i.e. in which order they are sent, depends on -the scheduler as well. Schedulers may take into account :attr:`Request.priority -` and applicable built-in settings (e.g. -:setting:`SCHEDULER_PRIORITY_QUEUE`, :setting:`SCHEDULER_MEMORY_QUEUE`, -:setting:`SCHEDULER_DISK_QUEUE`), but schedulers may also ignore any of those -parameters at will. +- Store pending requests. + + The built-in scheduler stores requests in memory or disk. Other schedulers + may rely, for example, on frontier, queue, database or storage services. + + Pending requests may come from seeding (see :setting:`SEEDING_POLICY`), + spider callbacks (:attr:`Request.callback `), + :ref:`spider middlewares ` or :ref:`downloader + middlewares `. + +- Drop unwanted requests. + + It is recommended for schedulers to use the configured + :setting:`DUPEFILTER_CLASS` and take into account + :attr:`Request.dont_filter `, but schedulers + may follow different criteria for dropping requests. + +- Return requests in the order they should be sent. + + To determine the right order, schedulers may take into account + :attr:`Request.priority ` and applicable built-in + settings (e.g. :setting:`SCHEDULER_PRIORITY_QUEUE`, + :setting:`SCHEDULER_MEMORY_QUEUE`, :setting:`SCHEDULER_DISK_QUEUE`), but + schedulers may also ignore any of those parameters at will. + +Built-in components +=================== Built-in scheduler -================== +------------------ .. autoclass:: Scheduler() -Writing a scheduler -=================== +.. _priority-queues: -.. tip:: Before writing a custom scheduler, see - :class:`~scrapy.core.scheduler.Scheduler` to learn how to customize the - default scheduler. +Built-in priority queues +------------------------ + +Components for :setting:`SCHEDULER_PRIORITY_QUEUE`: + +.. autoclass:: scrapy.pqueues.ScrapyPriorityQueue() +.. autoclass:: scrapy.pqueues.DownloaderAwarePriorityQueue() + + +.. _memory-queues: + +Built-in memory queues +---------------------- + +Components for :setting:`SCHEDULER_MEMORY_QUEUE`: + +.. autoclass:: scrapy.squeues.FifoMemoryQueue() +.. autoclass:: scrapy.squeues.LifoMemoryQueue() + + +.. _disk-queues: + +Built-in disk queues +-------------------- + +Components for :setting:`SCHEDULER_DISK_QUEUE`: + +.. autoclass:: scrapy.squeues.PickleFifoDiskQueue() +.. autoclass:: scrapy.squeues.PickleLifoDiskQueue() +.. autoclass:: scrapy.squeues.MarshalFifoDiskQueue() +.. autoclass:: scrapy.squeues.MarshalLifoDiskQueue() + + +Writing custom components +========================= + +.. _custom-scheduler: + +Writing a scheduler +------------------- Schedulers should subclass :class:`BaseScheduler` and implement its abstract methods: .. autoclass:: BaseScheduler - :members: - :member-order: bysource + :members: + :member-order: bysource + + +.. _custom-priority-queue: + +Writing a priority queue +------------------------ + +.. autoclass:: scrapy.pqueues.PriorityQueueProtocol() + :members: + :special-members: __init__, __len__ + :member-order: bysource + +.. _custom-internal-queue: + +Writing an internal queue +------------------------- + +.. autoclass:: scrapy.pqueues.QueueProtocol() + :members: + :special-members: __len__ + :member-order: bysource diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index a1a4ab710..4d573a9f6 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -520,6 +520,9 @@ This setting also affects :setting:`DOWNLOAD_DELAY` and :ref:`topics-autothrottle`: if :setting:`CONCURRENT_REQUESTS_PER_IP` is non-zero, download delay is enforced per IP, not per domain. +Cannot work with :class:`~scrapy.pqueues.DownloaderAwarePriorityQueue`. + + .. setting:: DEFAULT_DROPITEM_LOG_LEVEL DEFAULT_DROPITEM_LOG_LEVEL @@ -1673,11 +1676,10 @@ SCHEDULER_DEBUG Default: ``False`` -Setting to ``True`` will log debug information about the requests scheduler. -This currently logs (only once) if the requests cannot be serialized to disk. -Stats counter (``scheduler/unserializable``) tracks the number of times this happens. +Asks the :ref:`scheduler ` to log debug messages. -Example entry in logs:: +Causes the default scheduler, :class:`~scrapy.core.scheduler.Scheduler`, to log +a message about the first unserializable request. For example:: 1956-01-31 00:00:00+0800 [scrapy.core.scheduler] ERROR: Unable to serialize request: - reason: cannot serialize @@ -1690,20 +1692,35 @@ Example entry in logs:: SCHEDULER_DISK_QUEUE -------------------- -Default: ``'scrapy.squeues.PickleLifoDiskQueue'`` +Default: :class:`~scrapy.squeues.PickleLifoDiskQueue` + +Queue that :ref:`scheduling components ` should use to store +scheduled requests on disk. The queue also determines the order in which those +requests are returned. See :ref:`disk-queues` and :ref:`custom-internal-queue`. + +When :setting:`JOBDIR` is set, :ref:`built-in priority queues +` use this component for each set of +same-:attr:`~scrapy.Request.priority` requests. -Type of disk queue that will be used by scheduler. Other available types are -``scrapy.squeues.PickleFifoDiskQueue``, ``scrapy.squeues.MarshalFifoDiskQueue``, -``scrapy.squeues.MarshalLifoDiskQueue``. .. setting:: SCHEDULER_MEMORY_QUEUE SCHEDULER_MEMORY_QUEUE ---------------------- -Default: ``'scrapy.squeues.LifoMemoryQueue'`` -Type of in-memory queue used by scheduler. Other available type is: -``scrapy.squeues.FifoMemoryQueue``. +Default: :class:`~scrapy.squeues.LifoMemoryQueue` + +Queue that :ref:`scheduling components ` should use to store +scheduled requests in memory. The queue also determines the order in which +those requests are returned. See :ref:`memory-queues` and +:ref:`custom-internal-queue`. + +When :setting:`JOBDIR` is *not* set, :ref:`built-in priority queues +` use this component for sets of +same-:attr:`~scrapy.Request.priority` requests. When :setting:`JOBDIR` *is* +set, they are also used, as a fallback for requests that +:setting:`SCHEDULER_DISK_QUEUE` cannot serialize. + .. setting:: SCHEDULER_PRIORITY_QUEUE @@ -1712,16 +1729,13 @@ SCHEDULER_PRIORITY_QUEUE Default: :class:`~scrapy.pqueues.ScrapyPriorityQueue` -Queue used by the :ref:`scheduler ` to sort scheduled -requests by :attr:`Request.priority `. +Queue that :ref:`scheduling components ` should use for +request prioritization. See :ref:`priority-queues` and +:ref:`custom-priority-queue`. -Scheduled requests with the same priority are stored in nested queues, either -:setting:`SCHEDULER_MEMORY_QUEUE` or :setting:`SCHEDULER_DISK_QUEUE`. +The default scheduler, :class:`~scrapy.core.scheduler.Scheduler`, uses this +component. -The following built-in priority queues are available: - -.. autoclass:: scrapy.pqueues.ScrapyPriorityQueue -.. autoclass:: scrapy.pqueues.DownloaderAwarePriorityQueue .. setting:: SCRAPER_SLOT_MAX_ACTIVE_SIZE diff --git a/scrapy/core/scheduler.py b/scrapy/core/scheduler.py index efdbce5aa..c89d0b2f2 100644 --- a/scrapy/core/scheduler.py +++ b/scrapy/core/scheduler.py @@ -23,7 +23,7 @@ if TYPE_CHECKING: from scrapy.crawler import Crawler from scrapy.dupefilters import BaseDupeFilter from scrapy.http.request import Request - from scrapy.pqueues import ScrapyPriorityQueue + from scrapy.pqueues import PriorityQueueProtocol from scrapy.statscollectors import StatsCollector @@ -50,7 +50,8 @@ class BaseSchedulerMeta(type): class BaseScheduler(metaclass=BaseSchedulerMeta): - """Base class for :ref:`schedulers `.""" + """Base class for :ref:`scheduler ` :ref:`components + `.""" @abstractmethod def enqueue_request(self, request: Request) -> bool: @@ -122,14 +123,7 @@ class Scheduler(BaseScheduler): Requests are dropped if :attr:`~scrapy.Request.dont_filter` is ``False`` and :setting:`DUPEFILTER_CLASS` flags them as duplicate requests. - :setting:`SCHEDULER_PRIORITY_QUEUE` handles request prioritization. For - same-priority requests, their prioritization depends on - :setting:`SCHEDULER_MEMORY_QUEUE`, and also on - :setting:`SCHEDULER_DISK_QUEUE` if :setting:`JOBDIR` is set. - - If :setting:`JOBDIR` is set, :setting:`SCHEDULER_MEMORY_QUEUE` is used for - requests that cannot be serialized to disk. Memory requests always take - priority over disk requests. + :setting:`SCHEDULER_PRIORITY_QUEUE` handles request prioritization. The following stats are generated: @@ -144,7 +138,7 @@ class Scheduler(BaseScheduler): scheduler/unserializable If the value of the ``scheduler/unserializable`` stat is non-zero, consider - enabling :setting:`SCHEDULER_DEBUG` to log a warning messages with details + enabling :setting:`SCHEDULER_DEBUG` to log a warning message with details about the first unserializable request, to try and figure out how to make it serializable. @@ -159,12 +153,12 @@ class Scheduler(BaseScheduler): mqclass: type[BaseQueue] | None = None, logunser: bool = False, stats: StatsCollector | None = None, - pqclass: type[ScrapyPriorityQueue] | None = None, + pqclass: type[PriorityQueueProtocol] | None = None, crawler: Crawler | None = None, ): self.df: BaseDupeFilter = dupefilter self.dqdir: str | None = self._dqdir(jobdir) - self.pqclass: type[ScrapyPriorityQueue] | None = pqclass + self.pqclass: type[PriorityQueueProtocol] | None = pqclass self.dqclass: type[BaseQueue] | None = dqclass self.mqclass: type[BaseQueue] | None = mqclass self.logunser: bool = logunser @@ -190,8 +184,8 @@ class Scheduler(BaseScheduler): def open(self, spider: Spider) -> Deferred[None] | None: self.spider: Spider = spider - self.mqs: ScrapyPriorityQueue = self._mq() - self.dqs: ScrapyPriorityQueue | None = self._dq() if self.dqdir else None + self.mqs: PriorityQueueProtocol = self._mq() + self.dqs: PriorityQueueProtocol | None = self._dq() if self.dqdir else None return self.df.open() def close(self, reason: str) -> Deferred[None] | None: @@ -263,7 +257,7 @@ class Scheduler(BaseScheduler): return self.dqs.pop() return None - def _mq(self) -> ScrapyPriorityQueue: + def _mq(self) -> PriorityQueueProtocol: """Create a new priority queue instance, with in-memory storage""" assert self.crawler assert self.pqclass @@ -274,7 +268,7 @@ class Scheduler(BaseScheduler): key="", ) - def _dq(self) -> ScrapyPriorityQueue: + def _dq(self) -> PriorityQueueProtocol: """Create a new priority queue instance, with disk storage""" assert self.crawler assert self.dqdir @@ -304,13 +298,13 @@ class Scheduler(BaseScheduler): return str(dqdir) return None - def _read_dqs_state(self, dqdir: str) -> list[int]: + def _read_dqs_state(self, dqdir: str) -> Any: path = Path(dqdir, "active.json") if not path.exists(): return [] with path.open(encoding="utf-8") as f: - return cast(list[int], json.load(f)) + return cast(Any, json.load(f)) - def _write_dqs_state(self, dqdir: str, state: list[int]) -> None: + def _write_dqs_state(self, dqdir: str, state: Any) -> None: with Path(dqdir, "active.json").open("w", encoding="utf-8") as f: json.dump(state, f) diff --git a/scrapy/extensions/postprocessing.py b/scrapy/extensions/postprocessing.py index b1fa160c8..b8f5d7071 100644 --- a/scrapy/extensions/postprocessing.py +++ b/scrapy/extensions/postprocessing.py @@ -1,16 +1,47 @@ -""" -Extension for processing data before they are exported to feeds. -""" +"""Extension for processing data before they are exported to feeds.""" + +from __future__ import annotations from bz2 import BZ2File from gzip import GzipFile from io import IOBase from lzma import LZMAFile -from typing import IO, Any, BinaryIO, cast +from typing import IO, Any, BinaryIO, Protocol, cast from scrapy.utils.misc import load_object +class PostprocessingPluginProtocol(Protocol): + """:class:`~typing.Protocol` of :ref:`post-processing plugins + `.""" + + def __init__(self, file: BinaryIO, feed_options: dict[str, Any]) -> None: + """Initialize the plugin. + + *file* is a file-like object that implements at least the + :meth:`~typing.BinaryIO.write`, :meth:`~typing.BinaryIO.tell` and + :meth:`~typing.BinaryIO.close` methods. + + *feed_options* are the :ref:`options of the feed ` that + uses the plugin, which may be used to modify the behavior of the + plugin. + """ + + def write(self, data: bytes | memoryview) -> int: + """Process and write *data* into the target file of the plugin, and + return the number of bytes written.""" + + def close(self) -> None: + """Run clean-up code. + + For example, you might want to close a file wrapper that you might have + used to compress data written into the file received in the + :meth:`__init__` method. + + .. warning:: Do *not* close the file from the :meth:`__init__` method. + """ + + class GzipPlugin: """ Compresses received data using `gzip `_. @@ -82,9 +113,6 @@ class LZMAPlugin: - `lzma_preset` - `lzma_filters` - .. note:: - ``lzma_filters`` cannot be used in pypy version 7.3.1 and older. - See :py:class:`lzma.LZMAFile` for more info about parameters. """ diff --git a/scrapy/pqueues.py b/scrapy/pqueues.py index 5aba0bf25..ac754819c 100644 --- a/scrapy/pqueues.py +++ b/scrapy/pqueues.py @@ -2,7 +2,10 @@ from __future__ import annotations import hashlib import logging -from typing import TYPE_CHECKING, Protocol, cast +from typing import TYPE_CHECKING, Any, Protocol, cast + +# typing.Self requires Python 3.11 +from typing_extensions import Self from scrapy import Request from scrapy.utils.misc import build_from_crawler @@ -10,9 +13,6 @@ from scrapy.utils.misc import build_from_crawler if TYPE_CHECKING: from collections.abc import Iterable - # typing.Self requires Python 3.11 - from typing_extensions import Self - from scrapy.core.downloader import Downloader from scrapy.crawler import Crawler @@ -38,41 +38,119 @@ def _path_safe(text: str) -> str: class QueueProtocol(Protocol): - """Protocol for downstream queues of ``ScrapyPriorityQueue``.""" + """:class:`~typing.Protocol` of queues for the + :setting:`SCHEDULER_MEMORY_QUEUE` and :setting:`SCHEDULER_DISK_QUEUE` + settings. - def push(self, request: Request) -> None: ... + Queues may also define a ``peek()`` method, identical to :meth:`pop` except + that the returned request is not removed from the queue. + """ - def pop(self) -> Request | None: ... + def push(self, request: Request) -> None: + """Add *request* to the queue. - def close(self) -> None: ... + Raise :exc:`ValueError` if *request* cannot be stored, e.g. if + the queue stores requests on disk but it cannot serialize *request*. + """ - def __len__(self) -> int: ... + def pop(self) -> Request | None: + """Remove the next request from the queue and return it, or return + ``None`` if there are no requests.""" + + def close(self) -> None: + """Called when the queue is closed. May be used for cleanup code.""" + + def __len__(self) -> int: + """Return the number of requests that are currently in the queue.""" + + +class PriorityQueueProtocol(Protocol): + """:class:`~typing.Protocol` of queues for the + :setting:`SCHEDULER_PRIORITY_QUEUE` setting.""" + + @classmethod + def from_crawler( + cls, + crawler: Crawler, + downstream_queue_cls: type[QueueProtocol], + key: str, + startprios: Any = None, + ) -> Self: + """Create an instance of the queue. + + See :meth:`__init__` for details. + """ + + def __init__( + self, + crawler: Crawler, + downstream_queue_cls: type[QueueProtocol], + key: str, + startprios: Any = None, + ): + """Initializes the queue. + + *crawler* is the running crawler. + + *downstream_queue_cls* is an :ref:`internal queue + `, e.g. a :ref:`memory queue ` or + a :ref:`disk queue `. Each set of same-priority requests + should be stored in an instance of this queue. + + *key* is the path where the queue should serialize requests, if the + queue uses disk storage. Queues that do not use disk storage may ignore + this parameter. + + If :ref:`resuming a job `, *startprios* is the return + value of the previous call to :meth:`close`. Otherwise, it is a falsy + value (e.g. ``None`` or an empty :class:`list`). + """ + + def push(self, request: Request) -> None: + """Add *request* to the queue. + + Raise :exc:`ValueError` if *request* cannot be stored, e.g. if + the queue stores requests on disk but it cannot serialize *request*. + """ + + def pop(self) -> Request | None: + """Remove the next request from the queue and return it, or return + ``None`` if there are no requests.""" + + def close(self) -> Any: + """Called when closing the queue if the priority queue was initialized + with a disk-based internal queue. + + It must return JSON-serializable data representing the internal state + of the queue. The returned data will be passed back to :meth:`__init__` + as *startprio* when :ref:`resuming a job `. + """ + + def __len__(self) -> int: + """Return the number of requests that are currently in the queue.""" class ScrapyPriorityQueue: """Default scheduler priority queue (:setting:`SCHEDULER_PRIORITY_QUEUE`). - The internal queue must implement the following methods: + Sorts requests based solely on :attr:`Request.priority + `. - * push(obj) - * pop() - * close() - * __len__() + Requests with the same :attr:`Request.priority ` + value are sorted by the corresponding internal queue, + :setting:`SCHEDULER_MEMORY_QUEUE` or :setting:`SCHEDULER_DISK_QUEUE`. - Optionally, the queue could provide a ``peek`` method, that should return the - next object to be returned by ``pop``, but without removing it from the queue. + Which internal queue is used depends on the value of :setting:`JOBDIR`: - ``__init__`` method of ScrapyPriorityQueue receives a downstream_queue_cls - argument, which is a class used to instantiate a new (internal) queue when - a new priority is allocated. + - If :setting:`JOBDIR` is not set, :setting:`SCHEDULER_MEMORY_QUEUE` is + always used. - Only integer priorities should be used. Lower numbers are higher - priorities. - - startprios is a sequence of priorities to start with. If the queue was - previously closed leaving some priority buckets non-empty, those priorities - should be passed in startprios. + - If :setting:`JOBDIR` is set, :setting:`SCHEDULER_DISK_QUEUE` is used by + default, while :setting:`SCHEDULER_MEMORY_QUEUE` is used as a fallback + for requests that :setting:`SCHEDULER_DISK_QUEUE` cannot serialize. + When returning a request, memory requests always take precedence over + disk requests. """ @classmethod @@ -182,16 +260,18 @@ class DownloaderInterface: class DownloaderAwarePriorityQueue: - """PriorityQueue which takes Downloader activity into account: - domains (slots) with the least amount of active downloads are dequeued - first. + """Scheduler priority queue (:setting:`SCHEDULER_PRIORITY_QUEUE`) that + accounts for the download slot (usually the domain name) of active requests + (i.e. requests being downloaded). - Another available type is - ``scrapy.pqueues.DownloaderAwarePriorityQueue``. - ``scrapy.pqueues.DownloaderAwarePriorityQueue`` works better than - ``scrapy.pqueues.ScrapyPriorityQueue`` when you crawl many different - domains in parallel. But currently ``scrapy.pqueues.DownloaderAwarePriorityQueue`` - does not work together with :setting:`CONCURRENT_REQUESTS_PER_IP`. + It prioritizes requests that have their download slot in common with + *fewer* active requests. Otherwise, it works like + :class:`ScrapyPriorityQueue`. + + It works better than :class:`ScrapyPriorityQueue` for :ref:`broad crawls + `. + + Cannot work with :setting:`CONCURRENT_REQUESTS_PER_IP`. """ @classmethod