mirror of https://github.com/scrapy/scrapy.git
Add Jobdir documentation (#5260)
* Add Jobdir documentation * Fix link format * remove doubtful comments * some more edits * Keep the information in jobs.rst an example, and provide details in the reference docs of the corresponding components * Minor edits --------- Co-authored-by: Adrian Chaves <adrian@zyte.com>
This commit is contained in:
parent
16929c0991
commit
66fe5de139
|
|
@ -17,15 +17,21 @@ facilities:
|
|||
* an extension that keeps some spider state (key/value pairs) persistent
|
||||
between batches
|
||||
|
||||
.. _job-dir:
|
||||
|
||||
Job directory
|
||||
=============
|
||||
|
||||
To enable persistence support you just need to define a *job directory* through
|
||||
the ``JOBDIR`` setting. This directory will be for storing all required data to
|
||||
keep the state of a single job (i.e. a spider run). It's important to note that
|
||||
this directory must not be shared by different spiders, or even different
|
||||
jobs/runs of the same spider, as it's meant to be used for storing the state of
|
||||
a *single* job.
|
||||
To enable persistence support, define a *job directory* through the
|
||||
:setting:`JOBDIR` setting.
|
||||
|
||||
The job directory will store all required data to keep the state of a *single*
|
||||
job (i.e. a spider run), so that if stopped cleanly, it can be resumed later.
|
||||
|
||||
.. warning:: This directory must *not* be shared by different spiders, or even
|
||||
different jobs of the same spider.
|
||||
|
||||
See also :ref:`job-dir-contents`.
|
||||
|
||||
How to use it
|
||||
=============
|
||||
|
|
@ -65,6 +71,14 @@ Persistence gotchas
|
|||
There are a few things to keep in mind if you want to be able to use the Scrapy
|
||||
persistence support:
|
||||
|
||||
Pause limitations
|
||||
-----------------
|
||||
|
||||
Job pausing and resuming is only supported when the spider is paused by
|
||||
stopping it cleanly. Forced, sudden or otherwise unclean shutdown can lead to
|
||||
data corruption in the job directory, which may prevent the spider from
|
||||
resuming correctly.
|
||||
|
||||
Cookies expiration
|
||||
------------------
|
||||
|
||||
|
|
@ -72,7 +86,6 @@ Cookies may expire. So, if you don't resume your spider quickly the requests
|
|||
scheduled may no longer work. This won't be an issue if your spider doesn't rely
|
||||
on cookies.
|
||||
|
||||
|
||||
.. _request-serialization:
|
||||
|
||||
Request serialization
|
||||
|
|
@ -86,3 +99,52 @@ running :class:`~scrapy.Spider` class.
|
|||
If you wish to log the requests that couldn't be serialized, you can set the
|
||||
:setting:`SCHEDULER_DEBUG` setting to ``True`` in the project's settings page.
|
||||
It is ``False`` by default.
|
||||
|
||||
.. _job-dir-contents:
|
||||
|
||||
Job directory contents
|
||||
======================
|
||||
|
||||
The contents of a job directory depend on the components used during the job.
|
||||
Components known to write in the job directory include the :ref:`scheduler
|
||||
<topics-scheduler>` and the :class:`~scrapy.extensions.spiderstate.SpiderState`
|
||||
extension. See the reference documentation of the corresponding components for
|
||||
details.
|
||||
|
||||
For example, with default settings, the job directory may look like this:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
├── requests.queue
|
||||
| ├── active.json
|
||||
| └── {hostname}-{hash}
|
||||
| └── {priority}{s?}
|
||||
| ├── q{00000}
|
||||
| └── info.json
|
||||
├── requests.seen
|
||||
└── spider.state
|
||||
|
||||
Where:
|
||||
|
||||
- :class:`~scrapy.core.scheduler.Scheduler` creates the ``requests.queue/``
|
||||
directory and the ``active.json`` file, the latter containing the state
|
||||
data returned by :meth:`DownloaderAwarePriorityQueue.close()
|
||||
<scrapy.pqueues.DownloaderAwarePriorityQueue.close>` the last time the job
|
||||
was paused.
|
||||
|
||||
- :class:`~scrapy.pqueues.DownloaderAwarePriorityQueue` creates the
|
||||
``{hostname}-{hash}`` directories.
|
||||
|
||||
- :class:`~scrapy.pqueues.ScrapyPriorityQueue` creates the ``{priority}{s?}``
|
||||
directories.
|
||||
|
||||
- :class:`scrapy.squeues.PickleLifoDiskQueue`, a subclass of
|
||||
:class:`queuelib.LifoDiskQueue` that uses :mod:`pickle` to serialize
|
||||
:class:`dict` representations of :class:`scrapy.Request` objects, creates
|
||||
the ``info.json`` and ``q{00000}`` files.
|
||||
|
||||
- :class:`~scrapy.dupefilters.RFPDupeFilter` creates the ``requests.seen``
|
||||
file.
|
||||
|
||||
- :class:`~scrapy.extensions.spiderstate.SpiderState` creates the
|
||||
``spider.state`` file.
|
||||
|
|
|
|||
|
|
@ -32,3 +32,10 @@ Default scheduler
|
|||
.. autoclass:: Scheduler()
|
||||
:members:
|
||||
:special-members: __init__, __len__
|
||||
|
||||
|
||||
Priority queues
|
||||
===============
|
||||
|
||||
.. autoclass:: scrapy.pqueues.DownloaderAwarePriorityQueue
|
||||
.. autoclass:: scrapy.pqueues.ScrapyPriorityQueue
|
||||
|
|
|
|||
|
|
@ -1702,10 +1702,10 @@ the user agent to use in the robots.txt file.
|
|||
SCHEDULER
|
||||
---------
|
||||
|
||||
Default: ``'scrapy.core.scheduler.Scheduler'``
|
||||
Default: :class:`~scrapy.core.scheduler.Scheduler`
|
||||
|
||||
The scheduler class to be used for crawling.
|
||||
See the :ref:`topics-scheduler` topic for details.
|
||||
The scheduler class to be used for crawling. See :ref:`topics-scheduler` for
|
||||
details.
|
||||
|
||||
.. setting:: SCHEDULER_DEBUG
|
||||
|
||||
|
|
@ -1755,12 +1755,14 @@ Type of in-memory queue used by the scheduler. Other available type is:
|
|||
SCHEDULER_PRIORITY_QUEUE
|
||||
------------------------
|
||||
|
||||
Default: ``'scrapy.pqueues.DownloaderAwarePriorityQueue'``
|
||||
Default: :class:`~scrapy.pqueues.DownloaderAwarePriorityQueue`
|
||||
|
||||
Type of priority queue used by the scheduler. Another available type is
|
||||
``scrapy.pqueues.ScrapyPriorityQueue``.
|
||||
``scrapy.pqueues.DownloaderAwarePriorityQueue`` works better than
|
||||
``scrapy.pqueues.ScrapyPriorityQueue`` when you crawl many different
|
||||
Type of priority queue used by the scheduler.
|
||||
|
||||
Another available type is :class:`~scrapy.pqueues.ScrapyPriorityQueue`.
|
||||
|
||||
:class:`~scrapy.pqueues.DownloaderAwarePriorityQueue` works better than
|
||||
:class:`~scrapy.pqueues.ScrapyPriorityQueue` when you crawl many different
|
||||
domains in parallel.
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -4,7 +4,7 @@ import json
|
|||
import logging
|
||||
from abc import abstractmethod
|
||||
from pathlib import Path
|
||||
from typing import TYPE_CHECKING, Any, cast
|
||||
from typing import TYPE_CHECKING, Any
|
||||
from warnings import warn
|
||||
|
||||
# working around https://github.com/sphinx-doc/sphinx/issues/10400
|
||||
|
|
@ -128,7 +128,7 @@ class BaseScheduler(metaclass=BaseSchedulerMeta):
|
|||
|
||||
|
||||
class Scheduler(BaseScheduler):
|
||||
"""Default scheduler.
|
||||
r"""Default scheduler.
|
||||
|
||||
Requests are stored into priority queues
|
||||
(:setting:`SCHEDULER_PRIORITY_QUEUE`) that sort requests by
|
||||
|
|
@ -190,7 +190,8 @@ class Scheduler(BaseScheduler):
|
|||
following :ref:`settings <topics-settings>`:
|
||||
|
||||
| :setting:`DEPTH_PRIORITY` = ``1``
|
||||
| :setting:`SCHEDULER_DISK_QUEUE` = ``"scrapy.squeues.PickleFifoDiskQueue"``
|
||||
| :setting:`SCHEDULER_DISK_QUEUE` =
|
||||
``"scrapy.squeues.PickleFifoDiskQueue"``
|
||||
| :setting:`SCHEDULER_MEMORY_QUEUE` = ``"scrapy.squeues.FifoMemoryQueue"``
|
||||
|
||||
.. _BFO order: https://en.wikipedia.org/wiki/Breadth-first_search
|
||||
|
|
@ -219,6 +220,37 @@ class Scheduler(BaseScheduler):
|
|||
order. Lowering those settings to ``1`` enforces the desired order except
|
||||
for the very first request, but it significantly slows down the crawl as a
|
||||
whole.
|
||||
|
||||
Job directory contents
|
||||
======================
|
||||
|
||||
.. warning:: The files that this class generates in the :ref:`job directory
|
||||
<job-dir>` are an implementation detail, and may change without a
|
||||
warning in a future version of Scrapy. Do not rely on the following
|
||||
information for anything other than debugging purposes.
|
||||
|
||||
When using :setting:`JOBDIR`, this scheduler class:
|
||||
|
||||
- Creates a directory named ``requests.queue`` inside the :ref:`job
|
||||
directory <job-dir>`, meant to keep track of all requests stored in
|
||||
the scheduler (i.e. not downloaded yet).
|
||||
|
||||
- Generates inside that directory an ``active.json`` file with a JSON
|
||||
representation of the state (``startprios``) of
|
||||
:setting:`SCHEDULER_PRIORITY_QUEUE`.
|
||||
|
||||
The file is generated whenever the job stops (cleanly) and is loaded
|
||||
when resuming the job.
|
||||
|
||||
- Instantiates the configured :setting:`SCHEDULER_PRIORITY_QUEUE` with
|
||||
``requests.queue/`` as persistence directory (*key*) and
|
||||
:setting:`SCHEDULER_DISK_QUEUE` as *downstream_queue_cls*. The priority
|
||||
queue may create additional files and directories inside that
|
||||
directory, directly or though instances of
|
||||
:setting:`SCHEDULER_DISK_QUEUE`.
|
||||
|
||||
This scheduler class also uses the configured :setting:`DUPEFILTER_CLASS`,
|
||||
which may also write data inside the job directory.
|
||||
"""
|
||||
|
||||
@classmethod
|
||||
|
|
@ -486,13 +518,13 @@ class Scheduler(BaseScheduler):
|
|||
return str(dqdir)
|
||||
return None
|
||||
|
||||
def _read_dqs_state(self, dqdir: str) -> list[int]:
|
||||
def _read_dqs_state(self, dqdir: str) -> Any:
|
||||
path = Path(dqdir, "active.json")
|
||||
if not path.exists():
|
||||
return []
|
||||
with path.open(encoding="utf-8") as f:
|
||||
return cast("list[int]", json.load(f))
|
||||
return json.load(f)
|
||||
|
||||
def _write_dqs_state(self, dqdir: str, state: list[int]) -> None:
|
||||
def _write_dqs_state(self, dqdir: str, state: Any) -> None:
|
||||
with Path(dqdir, "active.json").open("w", encoding="utf-8") as f:
|
||||
json.dump(state, f)
|
||||
|
|
|
|||
|
|
@ -55,6 +55,18 @@ class RFPDupeFilter(BaseDupeFilter):
|
|||
filters out requests with the canonical
|
||||
(:func:`w3lib.url.canonicalize_url`) :attr:`~scrapy.http.Request.url`,
|
||||
:attr:`~scrapy.http.Request.method` and :attr:`~scrapy.http.Request.body`.
|
||||
|
||||
Job directory contents
|
||||
======================
|
||||
|
||||
.. warning:: The files that this class generates in the :ref:`job directory
|
||||
<job-dir>` are an implementation detail, and may change without a
|
||||
warning in a future version of Scrapy. Do not rely on the following
|
||||
information for anything other than debugging purposes.
|
||||
|
||||
When using :setting:`JOBDIR`, seen fingerprints are tracked in a file named
|
||||
``requests.seen`` in the :ref:`job directory <job-dir>`, which contains 1
|
||||
request fingerprint per line.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
|
|
|
|||
|
|
@ -51,16 +51,17 @@ class QueueProtocol(Protocol):
|
|||
|
||||
class ScrapyPriorityQueue:
|
||||
"""A priority queue implemented using multiple internal queues (typically,
|
||||
FIFO queues). It uses one internal queue for each priority value. The internal
|
||||
queue must implement the following methods:
|
||||
FIFO queues). It uses one internal queue for each priority value. The
|
||||
internal queue must implement the following methods:
|
||||
|
||||
* push(obj)
|
||||
* pop()
|
||||
* close()
|
||||
* __len__()
|
||||
|
||||
Optionally, the queue could provide a ``peek`` method, that should return the
|
||||
next object to be returned by ``pop``, but without removing it from the queue.
|
||||
Optionally, the queue could provide a ``peek`` method, that should return
|
||||
the next object to be returned by ``pop``, but without removing it from the
|
||||
queue.
|
||||
|
||||
``__init__`` method of ScrapyPriorityQueue receives a downstream_queue_cls
|
||||
argument, which is a class used to instantiate a new (internal) queue when
|
||||
|
|
@ -72,6 +73,28 @@ class ScrapyPriorityQueue:
|
|||
startprios is a sequence of priorities to start with. If the queue was
|
||||
previously closed leaving some priority buckets non-empty, those priorities
|
||||
should be passed in startprios.
|
||||
|
||||
Disk persistence
|
||||
================
|
||||
|
||||
.. warning:: The files that this class generates on disk are an
|
||||
implementation detail, and may change without a warning in a future
|
||||
version of Scrapy. Do not rely on the following information for
|
||||
anything other than debugging purposes.
|
||||
|
||||
When a component instantiates this class with a non-empty *key* argument,
|
||||
*key* is used as a persistence directory.
|
||||
|
||||
For every request enqueued, this class checks:
|
||||
|
||||
- Whether the request is a :ref:`start request <start-requests>` or not.
|
||||
|
||||
- The :data:`~scrapy.Request.priority` of the request.
|
||||
|
||||
For each combination of the above seen, this class creates an instance of
|
||||
*downstream_queue_cls* with *key* set to a subdirectory of the persistence
|
||||
directory, named as the request priority (e.g. ``1``), with an ``s`` suffix
|
||||
in case of a start request (e.g. ``1s``).
|
||||
"""
|
||||
|
||||
@classmethod
|
||||
|
|
@ -255,6 +278,26 @@ class DownloaderAwarePriorityQueue:
|
|||
"""PriorityQueue which takes Downloader activity into account:
|
||||
domains (slots) with the least amount of active downloads are dequeued
|
||||
first.
|
||||
|
||||
Disk persistence
|
||||
================
|
||||
|
||||
.. warning:: The files that this class generates on disk are an
|
||||
implementation detail, and may change without a warning in a future
|
||||
version of Scrapy. Do not rely on the following information for
|
||||
anything other than debugging purposes.
|
||||
|
||||
When a component instantiates this class with a non-empty *key* argument,
|
||||
*key* is used as a persistence directory, and inside that directory this
|
||||
class creates a subdirectory per download slot (domain).
|
||||
|
||||
Those subdirectories are named after the corresponding download slot, with
|
||||
path-unsafe characters replaced by underscores and an MD5 hash suffix to
|
||||
avoid collisions.
|
||||
|
||||
For each download slot, this class creates an instance of
|
||||
:class:`ScrapyPriorityQueue` with the download slot subdirectory as *key*
|
||||
and its own *downstream_queue_cls*.
|
||||
"""
|
||||
|
||||
@classmethod
|
||||
|
|
|
|||
Loading…
Reference in New Issue