Docs for the reactorless mode (#7385)

* Reactorless mode docs.

* Unset the ftp handler in the reactorless mode.

* Cleanup reactorless subprocess tests.

* Better handling of incompatible TWISTED_ENABLED cases.

* Improvements.

* Improve early error logging in test scripts.

* Address feedback.

* Address feedback.
This commit is contained in:
Andrey Rakhmatullin 2026-04-01 20:40:34 +05:00 committed by GitHub
parent fa76ca52e9
commit 0c6ccf50b3
No known key found for this signature in database
GPG Key ID: B5690EEEBB952194
23 changed files with 542 additions and 124 deletions

View File

@ -663,7 +663,7 @@ New features
(:issue:`4463`, :issue:`6804`)
- Added :func:`scrapy.utils.asyncio.is_asyncio_available` as an alternative
to :func:`scrapy.utils.defer.is_asyncio_reactor_installed` with a
to :func:`scrapy.utils.reactor.is_asyncio_reactor_installed` with a
future-proof name and semantics.
(:issue:`6827`)

View File

@ -129,6 +129,173 @@ example:
.. autofunction:: scrapy.utils.reactor.is_asyncio_reactor_installed
.. _asyncio-without-reactor:
Using Scrapy without a Twisted reactor
======================================
.. versionadded:: 2.15.0
.. warning::
This is currently experimental and may not be suitable for production use.
It's possible to use Scrapy without installing a Twisted reactor at all, by
setting the :setting:`TWISTED_ENABLED` setting to ``False``. In this mode
Scrapy will use the asyncio event loop directly, and most of the Scrapy
functionality will work in the same way.
Doing this provides several benefits in certain use cases:
* A Twisted reactor, once stopped, cannot be started again. This prevents, for
example, using several instances of
:class:`~scrapy.crawler.AsyncCrawlerProcess` in the same process when they
use a reactor, but with ``TWISTED_ENABLED=False`` it becomes possible.
* There may be limitations imposed by
:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor` and related
Twisted code, such as the requirement of using
:class:`~asyncio.SelectorEventLoop` on Windows (see :ref:`asyncio-windows`),
that do not apply if the reactor is not used.
* :class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor` manages the
underlying event loop, and while :class:`~scrapy.crawler.AsyncCrawlerRunner`
can use a pre-existing reactor which, in turn, can use a pre-existing event
loop, it's easier to use :class:`~scrapy.crawler.AsyncCrawlerRunner` with a
pre-existing loop directly.
* Omitting the reactor machinery may improve performance and reliability.
Limitations
-----------
As some Scrapy features and components require a reactor, they don't work and
are disabled without it. Replacements that don't require a reactor may be added
in future Scrapy versions. The following features are not available:
* The default HTTP(S) download handler,
:class:`~scrapy.core.downloader.handlers.http11.HTTP11DownloadHandler` (this
is likely the biggest difference; Scrapy provides an HTTP(S) download handler
that doesn't require a reactor and will be used instead of it:
:class:`~scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler`)
* :class:`~scrapy.core.downloader.handlers.ftp.FTPDownloadHandler`
* :class:`~scrapy.core.downloader.handlers.http2.H2DownloadHandler`
* :ref:`topics-shell`
* :ref:`topics-telnetconsole`
* :class:`~scrapy.crawler.CrawlerRunner` and
:class:`~scrapy.crawler.CrawlerProcess`
(:class:`~scrapy.crawler.AsyncCrawlerProcess` and
:class:`~scrapy.crawler.AsyncCrawlerRunner` are available)
* Twisted-specific DNS resolvers (the :setting:`DNS_RESOLVER` setting)
* User and 3rd-party code that requires a reactor (see :ref:`below
<asyncio-without-reactor-migrate>` for examples)
Note that importing Twisted modules and, among other things, creating and using
:class:`~twisted.internet.defer.Deferred` objects doesn't require a reactor, so
code that uses :class:`~twisted.internet.defer.Deferred`,
:class:`~twisted.python.failure.Failure` and some other Twisted APIs will not
necessarily stop working.
Other differences
-----------------
When :setting:`TWISTED_ENABLED` is set to ``False``, Scrapy will change the
defaults of some other settings:
* :setting:`TELNETCONSOLE_ENABLED` is set to ``False``.
* The ``"http"`` and ``"https"`` keys in :setting:`DOWNLOAD_HANDLERS_BASE` are
set to ``"scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler"``.
* The ``"ftp"`` key in :setting:`DOWNLOAD_HANDLERS_BASE` is set to ``None``.
Thus, :class:`~scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler` is
used by default for making HTTP(S) requests. Please refer to its documentation
for its differences and limitations compared to
:class:`~scrapy.core.downloader.handlers.http11.HTTP11DownloadHandler`.
Additionally, :class:`~scrapy.crawler.AsyncCrawlerProcess` will install a
:term:`meta path finder` that prevents :mod:`twisted.internet.reactor` from
being imported.
.. _asyncio-without-reactor-migrate:
Adding support to existing code
-------------------------------
Code that doesn't directly use Twisted APIs or APIs that depend on Twisted ones
doesn't need special support for running without a reactor.
Here are some examples of APIs and patterns that need a replacement:
* Using :meth:`reactor.callLater()
<twisted.internet.base.ReactorBase.callLater>` for sleeping or delayed calls.
You can use :meth:`asyncio.loop.call_later` instead.
* Using :func:`twisted.internet.threads.deferToThread`,
:meth:`reactor.callFromThread()
<twisted.internet.base.ReactorBase.callFromThread>` and related APIs to
execute code in other threads. You can use :func:`asyncio.to_thread`,
:meth:`asyncio.loop.call_soon_threadsafe` and related APIs instead.
* Using :class:`twisted.internet.task.LoopingCall` for scheduling repeated
tasks. As there is no direct replacement in the standard library, you may
need to write your own one using :func:`asyncio.sleep` in a task.
* Using Twisted network client and server APIs (:meth:`reactor.connectTCP()
<twisted.internet.interfaces.IReactorTCP.connectTCP>`,
:meth:`reactor.listenTCP()
<twisted.internet.interfaces.IReactorTCP.listenTCP>`,
:mod:`twisted.web.client`, :mod:`twisted.mail.smtp` etc.). You can use other
built-in or 3rd-party libraries for this.
* Using :class:`~scrapy.crawler.CrawlerProcess` or
:class:`~scrapy.crawler.CrawlerRunner`. You should use
:class:`~scrapy.crawler.AsyncCrawlerProcess` or
:class:`~scrapy.crawler.AsyncCrawlerRunner` respectively instead.
* Checking whether ``asyncio`` support is available with
:func:`scrapy.utils.reactor.is_asyncio_reactor_installed`. You should use
:func:`scrapy.utils.asyncio.is_asyncio_available` instead.
Scrapy provides unified helpers for some of these examples:
.. autofunction:: scrapy.utils.asyncio.call_later
.. autofunction:: scrapy.utils.asyncio.create_looping_call
.. autoclass:: scrapy.utils.asyncio.AsyncioLoopingCall
.. autofunction:: scrapy.utils.asyncio.run_in_thread
If your code needs to know whether the reactor is available, you can either
check for the value of the :setting:`TWISTED_ENABLED` setting (you need access
to the :class:`~scrapy.crawler.Crawler` instance to do this) or use the
following function:
.. autofunction:: scrapy.utils.reactorless.is_reactorless
In general, code that doesn't use the reactor (directly or indirectly) can be
used unmodified both with the asyncio reactor and without a reactor. This
includes code that converts Deferreds to futures and vice versa as described in
:ref:`asyncio-await-dfd`.
Troubleshooting
---------------
**ImportError: Import of twisted.internet.reactor is forbidden when running
without a Twisted reactor [...]:** Scrapy is configured to run without a
reactor, but some code imported :mod:`twisted.internet.reactor`, most likely
because that code needs a reactor to be used. You need to stop using this code
or set :setting:`TWISTED_ENABLED` back to ``True``. It's also possible that the
reactor isn't really needed but was installed due to the problem described in
:ref:`asyncio-preinstalled-reactor`, in which case it should be enough to fix
the problematic imports.
**RuntimeError: TWISTED_ENABLED is False but a Twisted reactor is installed:**
Scrapy is configured to run without a reactor, but a reactor is already
installed before the Scrapy code is executed. If you are trying to set
:setting:`TWISTED_ENABLED` via :ref:`per-spider settings <spider-settings>`,
it's currently unsupported.
**RuntimeError: We expected a Twisted reactor to be installed but it isn't:**
Scrapy is configured to run with a reactor and not to install one, but a
reactor wasn't installed before the Scrapy code is executed. If you are trying
to set :setting:`TWISTED_ENABLED` via :ref:`per-spider settings
<spider-settings>`, it's currently unsupported.
**RuntimeError: <class> doesn't support TWISTED_ENABLED=False:** The listed
class cannot be used with :setting:`TWISTED_ENABLED` set to ``False``. There
may be a replacement in the :ref:`documentation above
<asyncio-without-reactor>` or the documentation of the affected class.
.. _asyncio-windows:
Windows-specific notes
@ -149,6 +316,9 @@ automatically when you change the :setting:`TWISTED_REACTOR` setting or call
them together with Scrapy on Windows (but you should be able to use
them on WSL or native Linux).
.. note:: This problem doesn't apply when not using the reactor, see
:ref:`asyncio-without-reactor`.
.. _playwright: https://github.com/microsoft/playwright-python

View File

@ -140,6 +140,9 @@ This handler supports ``ftp://host/path`` FTP URIs.
It's implemented using :mod:`twisted.protocols.ftp`.
.. note::
This handler is not supported when :setting:`TWISTED_ENABLED` is ``False``.
.. _twisted-http2-handler:
H2DownloadHandler
@ -193,6 +196,9 @@ If you want to use this handler you need to replace the default one for the
.. _http2 faq: https://http2.github.io/faq/#does-http2-require-encryption
.. _server pushes: https://datatracker.ietf.org/doc/html/rfc7540#section-8.2
.. note::
This handler is not supported when :setting:`TWISTED_ENABLED` is ``False``.
HTTP11DownloadHandler
---------------------
@ -206,6 +212,9 @@ uses the HTTP/1.1 protocol for them.
It's implemented using :mod:`twisted.web.client`.
.. note::
This handler is not supported when :setting:`TWISTED_ENABLED` is ``False``.
HttpxDownloadHandler
--------------------

View File

@ -166,6 +166,86 @@ with :class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor`):
.. seealso:: :doc:`twisted:core/howto/reactor-basics`
And here are examples of using these classes with :setting:`TWISTED_ENABLED`
set to ``False``.
Simple usage of :class:`~scrapy.crawler.AsyncCrawlerProcess`:
.. code-block:: python
import scrapy
from scrapy.crawler import AsyncCrawlerProcess
class MySpider(scrapy.Spider):
# Your spider definition
...
process = AsyncCrawlerProcess(
settings={
"TWISTED_ENABLED": False,
}
)
process.crawl(MySpider)
process.start() # the script will block here until the crawling is finished
With ``TWISTED_ENABLED=False`` you can use several instances of
:class:`~scrapy.crawler.AsyncCrawlerProcess` in the same process:
.. code-block:: python
import scrapy
from scrapy.crawler import AsyncCrawlerProcess
class MySpider(scrapy.Spider):
# Your spider definition
...
process1 = AsyncCrawlerProcess(
settings={
"TWISTED_ENABLED": False,
}
)
process1.crawl(MySpider)
process1.start()
process2 = AsyncCrawlerProcess(
settings={
"TWISTED_ENABLED": False,
}
)
process2.crawl(MySpider)
process2.start()
Using :func:`asyncio.run` with :class:`~scrapy.crawler.AsyncCrawlerRunner`:
.. code-block:: python
import asyncio
import scrapy
from scrapy.crawler import AsyncCrawlerRunner
from scrapy.utils.log import configure_logging
class MySpider(scrapy.Spider):
# Your spider definition
...
async def main():
configure_logging({"LOG_FORMAT": "%(levelname)s: %(message)s"})
runner = AsyncCrawlerRunner(settings={"TWISTED_ENABLED": False})
await runner.crawl(MySpider) # completes when the spider finishes
asyncio.run(main())
.. _run-multiple-spiders:
Running multiple spiders in the same process

View File

@ -303,11 +303,12 @@ Pre-crawler settings
These settings cannot be :ref:`set from a spider <spider-settings>`.
These settings are :setting:`SPIDER_LOADER_CLASS` and settings used by the
corresponding :ref:`component <topics-components>`, e.g.
:setting:`SPIDER_MODULES` and :setting:`SPIDER_LOADER_WARN_ONLY` for the
default component.
These settings are:
- :setting:`TWISTED_ENABLED`
- :setting:`SPIDER_LOADER_CLASS` and settings used by the corresponding
spider loader class, e.g. :setting:`SPIDER_MODULES` and
:setting:`SPIDER_LOADER_WARN_ONLY` for the default spider loader class.
.. _reactor-settings:
@ -356,6 +357,9 @@ ignoring the value of :setting:`TWISTED_REACTOR` and using the value of
e.g. in :ref:`per-spider settings <spider-settings>`, an exception will be
raised.
All of these settings, except for :setting:`ASYNCIO_EVENT_LOOP`, are only used
when the Twisted reactor is used, i.e. when :setting:`TWISTED_ENABLED` is
``True``.
.. _topics-settings-ref:
@ -651,6 +655,13 @@ Default: ``True``
Whether to enable DNS in-memory cache.
.. note::
This setting is only used by
:class:`~scrapy.resolver.CachingThreadedResolver` and
:class:`~scrapy.resolver.CachingHostnameResolver`. It has no effect when
:setting:`TWISTED_ENABLED` is ``False``, and may have no effect either when
:setting:`DNS_RESOLVER` is set to a different resolver.
.. setting:: DNSCACHE_SIZE
DNSCACHE_SIZE
@ -658,7 +669,7 @@ DNSCACHE_SIZE
Default: ``10000``
DNS in-memory cache size.
DNS in-memory cache size, see :setting:`DNSCACHE_ENABLED`.
.. setting:: DNS_RESOLVER
@ -667,12 +678,16 @@ DNS_RESOLVER
Default: ``'scrapy.resolver.CachingThreadedResolver'``
The class to be used to resolve DNS names. The default ``scrapy.resolver.CachingThreadedResolver``
supports specifying a timeout for DNS requests via the :setting:`DNS_TIMEOUT` setting,
but works only with IPv4 addresses. Scrapy provides an alternative resolver,
The class to be used by Twisted to resolve DNS names. The default
``scrapy.resolver.CachingThreadedResolver`` supports specifying a timeout for
DNS requests via the :setting:`DNS_TIMEOUT` setting, but works only with IPv4
addresses. Scrapy provides an alternative resolver,
``scrapy.resolver.CachingHostnameResolver``, which supports IPv4/IPv6 addresses but does not
take the :setting:`DNS_TIMEOUT` setting into account.
.. note::
This setting has no effect when :setting:`TWISTED_ENABLED` is ``False``.
.. setting:: DNS_TIMEOUT
DNS_TIMEOUT
@ -682,6 +697,12 @@ Default: ``60``
Timeout for processing of DNS queries in seconds. Float is supported.
.. note::
This setting is only used by
:class:`~scrapy.resolver.CachingThreadedResolver`. It has no effect when
:setting:`TWISTED_ENABLED` is ``False``, and may have no effect either when
:setting:`DNS_RESOLVER` is set to a different resolver.
.. setting:: DOWNLOADER
DOWNLOADER
@ -922,6 +943,20 @@ Default:
"ftp": "scrapy.core.downloader.handlers.ftp.FTPDownloadHandler",
}
(when :setting:`TWISTED_ENABLED` is ``True``)
.. code-block:: python
{
"data": "scrapy.core.downloader.handlers.datauri.DataURIDownloadHandler",
"file": "scrapy.core.downloader.handlers.file.FileDownloadHandler",
"http": "scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler",
"https": "scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler",
"s3": "scrapy.core.downloader.handlers.s3.S3DownloadHandler",
"ftp": None,
}
(when :setting:`TWISTED_ENABLED` is ``False``)
A dict containing the :ref:`download handlers <topics-download-handlers>`
enabled by default in Scrapy. You should never modify this setting in your
@ -1954,7 +1989,7 @@ For more info see: :ref:`topics-stats`.
TELNETCONSOLE_ENABLED
---------------------
Default: ``True``
Default: ``True`` (``False`` when :setting:`TWISTED_ENABLED` is ``False``)
A boolean which specifies if the :ref:`telnet console <topics-telnetconsole>`
will be enabled (provided its extension is also enabled).
@ -1973,6 +2008,35 @@ command.
The project name must not conflict with the name of custom files or directories
in the ``project`` subdirectory.
.. setting:: TWISTED_ENABLED
TWISTED_ENABLED
---------------
Default: ``True``
Whether to install and use the Twisted reactor.
If this is set to ``True``, Scrapy will use the Twisted reactor and will
install one according to the :setting:`TWISTED_REACTOR` setting value when
appropriate (e.g. when running via :ref:`the command-line tool
<topics-commands>`). This is the traditional mode of using Scrapy.
If this is set to ``False``, Scrapy will use the asyncio event loop directly
and will not attempt to install or use a reactor. Features that require a
reactor won't be available, but Twisted APIs that don't require a reactor,
including :class:`~twisted.internet.defer.Deferred` and
:class:`~twisted.python.failure.Failure`, will still be available. On the other
hand, limitations related to Twisted reactors (such as not being able to start
a reactor in the same process where a reactor was previously started and
stopped) will not apply. This mode is currently experimental and may not be
suitable for production use. It may also not be supported by 3rd-party code.
See :ref:`asyncio-without-reactor` for more information about this mode.
.. note:: This setting can't be set :ref:`per-spider <spider-settings>`.
.. versionadded:: 2.15.0
.. setting:: TWISTED_REACTOR
TWISTED_REACTOR

View File

@ -17,6 +17,9 @@ spider, without having to run the spider to test every change.
Once you get familiarized with the Scrapy shell, you'll see that it's an
invaluable tool for developing and debugging your spiders.
.. note::
This feature is not supported when :setting:`TWISTED_ENABLED` is ``False``.
Configuring the shell
=====================

View File

@ -26,6 +26,9 @@ disable it if you want. For more information about the extension itself see
Please avoid using telnet console over insecure connections,
or disable it completely using :setting:`TELNETCONSOLE_ENABLED` option.
.. note::
This feature is not supported when :setting:`TWISTED_ENABLED` is ``False``.
.. highlight:: none
How to access the telnet console

View File

@ -107,16 +107,26 @@ class Crawler:
use_reactor = self.settings.getbool("TWISTED_ENABLED")
if use_reactor:
# We either install a reactor or expect one to be installed.
reactor_class: str = self.settings["TWISTED_REACTOR"]
event_loop: str = self.settings["ASYNCIO_EVENT_LOOP"]
if self._init_reactor:
# this needs to be done after the spider settings are merged,
# but before something imports twisted.internet.reactor
# We need to install a reactor.
# This needs to be done after the spider settings are merged,
# but before something imports twisted.internet.reactor.
if reactor_class:
# Install a specific reactor.
install_reactor(reactor_class, event_loop)
else:
# Install the default one.
from twisted.internet import reactor # noqa: F401
elif not is_reactor_installed():
# We need a reactor to be already installed.
raise RuntimeError(
"We expected a Twisted reactor to be installed but it isn't."
)
if reactor_class:
# We need to check that the correct reactor is installed.
verify_installed_reactor(reactor_class)
if is_asyncio_reactor_installed() and event_loop:
verify_installed_asyncio_event_loop(event_loop)
@ -124,6 +134,11 @@ class Crawler:
if self._init_reactor or reactor_class:
log_reactor_info()
else:
# We expect a reactor to not be installed.
if is_reactor_installed():
raise RuntimeError(
"TWISTED_ENABLED is False but a Twisted reactor is installed."
)
logger.debug("Not using a Twisted reactor")
self._apply_reactorless_default_settings()
@ -147,6 +162,7 @@ class Crawler:
self.settings["DOWNLOAD_HANDLERS_BASE"][scheme] = (
"scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler"
)
self.settings["DOWNLOAD_HANDLERS_BASE"]["ftp"] = None
# Cannot use @deferred_f_from_coro_f because that relies on the reactor
# being installed already, which is done within _apply_settings(), inside
@ -477,17 +493,24 @@ class CrawlerRunner(CrawlerRunnerBase):
class AsyncCrawlerRunner(CrawlerRunnerBase):
"""
This is a convenient helper class that keeps track of, manages and runs
crawlers inside an already setup :mod:`~twisted.internet.reactor`.
crawlers inside an already setup :mod:`~twisted.internet.reactor` or
asyncio event loop.
The AsyncCrawlerRunner object must be instantiated with a
:class:`~scrapy.settings.Settings` object.
When the :setting:`TWISTED_ENABLED` setting is set to ``True``, this class
requires a reactor to be installed and uses it, otherwise it requires a
reactor to not be installed but requires an asyncio event loop to be
installed and uses it.
This class shouldn't be needed (since Scrapy is responsible of using it
accordingly) unless writing scripts that manually handle the crawling
process. See :ref:`run-from-script` for an example.
This class provides coroutine APIs. It requires
:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor`.
:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor` when used
with a reactor.
"""
def __init__(self, settings: dict[str, Any] | Settings | None = None):
@ -528,6 +551,10 @@ class AsyncCrawlerRunner(CrawlerRunnerBase):
"it must be a spider class (or a Crawler object)"
)
if self.settings.getbool("TWISTED_ENABLED"):
if not is_reactor_installed():
raise RuntimeError(
"We expected a Twisted reactor to be installed but it isn't."
)
if not is_asyncio_reactor_installed():
raise RuntimeError(
f"When TWISTED_ENABLED is True, {type(self).__name__} "
@ -716,8 +743,8 @@ class CrawlerProcess(CrawlerProcessBase, CrawlerRunner):
) -> None:
"""
This method starts a :mod:`~twisted.internet.reactor`, adjusts its pool
size to :setting:`REACTOR_THREADPOOL_MAXSIZE`, and installs a DNS cache
based on :setting:`DNSCACHE_ENABLED` and :setting:`DNSCACHE_SIZE`.
size to :setting:`REACTOR_THREADPOOL_MAXSIZE`, and installs a DNS
resolver based on :setting:`DNSCACHE_ENABLED`.
If ``stop_after_crawl`` is True, the reactor will be stopped after all
crawlers have finished, using :meth:`join`.
@ -757,6 +784,10 @@ class AsyncCrawlerProcess(CrawlerProcessBase, AsyncCrawlerRunner):
The AsyncCrawlerProcess object must be instantiated with a
:class:`~scrapy.settings.Settings` object.
When the :setting:`TWISTED_ENABLED` setting is set to ``True``, this class
installs a reactor and uses it, otherwise it requires a reactor to not be
installed but installs an asyncio event loop and uses it.
:param install_root_handler: whether to install root logging handler
(default: True)
@ -765,7 +796,8 @@ class AsyncCrawlerProcess(CrawlerProcessBase, AsyncCrawlerRunner):
process. See :ref:`run-from-script` for an example.
This class provides coroutine APIs. It requires
:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor`.
:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor` when used
with a reactor.
"""
def __init__(
@ -808,9 +840,12 @@ class AsyncCrawlerProcess(CrawlerProcessBase, AsyncCrawlerRunner):
self, stop_after_crawl: bool = True, install_signal_handlers: bool = True
) -> None:
"""
This method starts a :mod:`~twisted.internet.reactor`, adjusts its pool
size to :setting:`REACTOR_THREADPOOL_MAXSIZE`, and installs a DNS cache
based on :setting:`DNSCACHE_ENABLED` and :setting:`DNSCACHE_SIZE`.
This method starts a :mod:`~twisted.internet.reactor`/asyncio event
loop, depending on the value of the :setting:`TWISTED_ENABLED` setting.
When using a reactor it adjusts its pool size to
:setting:`REACTOR_THREADPOOL_MAXSIZE` and installs a DNS resolver based
on :setting:`DNSCACHE_ENABLED`.
If ``stop_after_crawl`` is True, the reactor will be stopped after all
crawlers have finished, using :meth:`join`.

View File

@ -132,13 +132,16 @@ async def _parallel_asyncio(
class AsyncioLoopingCall:
"""A simple implementation of a periodic call using asyncio, keeping
some API and behavior compatibility with the Twisted ``LoopingCall``.
some API and behavior compatibility with
:class:`~twisted.internet.task.LoopingCall`.
The function is called every *interval* seconds, independent of the finish
time of the previous call. If the function is still running when it's time
to call it again, calls are skipped until the function finishes.
The function must not return a coroutine or a ``Deferred``.
.. versionadded:: 2.14.0
"""
def __init__(self, func: Callable[_P, _T], *args: _P.args, **kwargs: _P.kwargs):
@ -216,8 +219,12 @@ def create_looping_call(
) -> AsyncioLoopingCall | LoopingCall:
"""Create an instance of a looping call class.
This creates an instance of :class:`AsyncioLoopingCall` or
:class:`LoopingCall`, depending on whether asyncio support is available.
This creates an instance of
:class:`~scrapy.utils.asyncio.AsyncioLoopingCall` or
:class:`~twisted.internet.task.LoopingCall`, depending on whether asyncio
support is available.
.. versionadded:: 2.14.0
"""
if is_asyncio_available():
return AsyncioLoopingCall(func, *args, **kwargs)
@ -229,8 +236,11 @@ def call_later(
) -> CallLaterResult:
"""Schedule a function to be called after a delay.
This uses either ``loop.call_later()`` or ``reactor.callLater()``, depending
on whether asyncio support is available.
This uses either :meth:`asyncio.loop.call_later` or
:meth:`reactor.callLater() <twisted.internet.base.ReactorBase.callLater>`,
depending on whether asyncio support is available.
.. versionadded:: 2.14.0
"""
if is_asyncio_available():
loop = asyncio.get_event_loop()
@ -249,6 +259,8 @@ class CallLaterResult:
no ``active()`` (as there is no such public API in
:class:`asyncio.TimerHandle`) but ``cancel()`` can be called on already
called or cancelled instances.
.. versionadded:: 2.14.0
"""
_timer_handle: asyncio.TimerHandle | None = None

View File

@ -0,0 +1,36 @@
from __future__ import annotations
import logging
from typing import TYPE_CHECKING
import scrapy
from scrapy.crawler import AsyncCrawlerProcess
from scrapy.utils.reactorless import is_reactorless
if TYPE_CHECKING:
from asyncio import Task
class NoRequestsSpider(scrapy.Spider):
name = "no_request"
custom_settings = {
"TWISTED_ENABLED": False,
}
async def start(self):
self.logger.info(f"is_reactorless(): {is_reactorless()}")
return
yield
def log_task_exception(task: Task) -> None:
try:
task.result()
except Exception:
logging.exception("Crawl task failed") # noqa: LOG015
process = AsyncCrawlerProcess()
task = process.crawl(NoRequestsSpider)
task.add_done_callback(log_task_exception)
process.start()

View File

@ -12,16 +12,7 @@ class NoRequestsSpider(scrapy.Spider):
yield
process = AsyncCrawlerProcess(
settings={
"TWISTED_ENABLED": False,
"DOWNLOAD_HANDLERS": {
"http": None,
"https": None,
"ftp": None,
},
}
)
process = AsyncCrawlerProcess(settings={"TWISTED_ENABLED": False})
process.crawl(NoRequestsSpider)
process.start()

View File

@ -3,13 +3,4 @@ from scrapy.utils.reactor import install_reactor
install_reactor("twisted.internet.asyncioreactor.AsyncioSelectorReactor")
AsyncCrawlerProcess(
settings={
"TWISTED_ENABLED": False,
"DOWNLOAD_HANDLERS": {
"http": None,
"https": None,
"ftp": None,
},
}
)
AsyncCrawlerProcess(settings={"TWISTED_ENABLED": False})

View File

@ -12,16 +12,7 @@ class NoRequestsSpider(scrapy.Spider):
yield
process = AsyncCrawlerProcess(
settings={
"TWISTED_ENABLED": False,
"DOWNLOAD_HANDLERS": {
"http": None,
"https": None,
"ftp": None,
},
}
)
process = AsyncCrawlerProcess(settings={"TWISTED_ENABLED": False})
process.crawl(NoRequestsSpider)
process.start()

View File

@ -1,25 +0,0 @@
import scrapy
from scrapy.crawler import AsyncCrawlerProcess
class NoRequestsSpider(scrapy.Spider):
name = "no_request"
async def start(self):
return
yield
process = AsyncCrawlerProcess(
settings={
"TWISTED_ENABLED": False,
"DOWNLOAD_HANDLERS": {
"http": None,
"https": None,
"ftp": None,
},
}
)
process.crawl(NoRequestsSpider)
process.start()

View File

@ -13,11 +13,6 @@ class NoRequestsSpider(scrapy.Spider):
process = AsyncCrawlerProcess(
settings={
"TWISTED_ENABLED": False,
"DOWNLOAD_HANDLERS": {
"http": None,
"https": None,
"ftp": None,
},
"TELNETCONSOLE_ENABLED": False,
}
)

View File

@ -13,11 +13,6 @@ class NoRequestsSpider(scrapy.Spider):
process = AsyncCrawlerProcess(
settings={
"TWISTED_ENABLED": False,
"DOWNLOAD_HANDLERS": {
"http": None,
"https": None,
"ftp": None,
},
"TELNETCONSOLE_ENABLED": True,
}
)

View File

@ -0,0 +1,22 @@
import asyncio
from scrapy import Spider
from scrapy.crawler import AsyncCrawlerRunner
from scrapy.utils.log import configure_logging
class NoRequestsSpider(Spider):
name = "no_request"
async def start(self):
return
yield
async def main() -> None:
configure_logging()
runner = AsyncCrawlerRunner()
await runner.crawl(NoRequestsSpider)
asyncio.run(main())

View File

@ -0,0 +1,27 @@
import asyncio
from scrapy import Spider
from scrapy.crawler import AsyncCrawlerRunner
from scrapy.utils.log import configure_logging
from scrapy.utils.reactorless import is_reactorless
class NoRequestsSpider(Spider):
name = "no_request"
custom_settings = {
"TWISTED_ENABLED": False,
}
async def start(self):
self.logger.info(f"is_reactorless(): {is_reactorless()}")
return
yield
async def main() -> None:
configure_logging()
runner = AsyncCrawlerRunner()
await runner.crawl(NoRequestsSpider)
asyncio.run(main())

View File

@ -17,16 +17,7 @@ class DataSpider(Spider):
async def main() -> None:
configure_logging()
runner = AsyncCrawlerRunner(
settings={
"TWISTED_ENABLED": False,
"DOWNLOAD_HANDLERS": {
"http": None,
"https": None,
"ftp": None,
},
}
)
runner = AsyncCrawlerRunner(settings={"TWISTED_ENABLED": False})
await runner.crawl(DataSpider)

View File

@ -16,16 +16,7 @@ class NoRequestsSpider(Spider):
async def main() -> None:
configure_logging()
runner = AsyncCrawlerRunner(
settings={
"TWISTED_ENABLED": False,
"DOWNLOAD_HANDLERS": {
"http": None,
"https": None,
"ftp": None,
},
}
)
runner = AsyncCrawlerRunner(settings={"TWISTED_ENABLED": False})
await runner.crawl(NoRequestsSpider)

View File

@ -17,16 +17,7 @@ class NoRequestsSpider(Spider):
async def main() -> None:
configure_logging()
runner = AsyncCrawlerRunner(
settings={
"TWISTED_ENABLED": False,
"DOWNLOAD_HANDLERS": {
"http": None,
"https": None,
"ftp": None,
},
}
)
runner = AsyncCrawlerRunner(settings={"TWISTED_ENABLED": False})
await runner.crawl(NoRequestsSpider)

View File

@ -0,0 +1,19 @@
from twisted.python import log
from scrapy import Spider
from scrapy.crawler import CrawlerRunner
from scrapy.utils.log import configure_logging
class NoRequestsSpider(Spider):
name = "no_request"
async def start(self):
return
yield
configure_logging()
runner = CrawlerRunner()
d = runner.crawl(NoRequestsSpider)
d.addErrback(log.err)

View File

@ -352,7 +352,16 @@ class TestAsyncCrawlerProcessSubprocess(TestCrawlerProcessSubprocessBase):
assert "Spider closed (finished)" in log
assert "is_reactorless(): True" in log
assert "ERROR: " not in log
assert "WARNING: " not in log
assert log.count("WARNING: HttpxDownloadHandler is experimental") == 2
assert log.count("WARNING: ") == 2
def test_reactorless_custom_settings(self):
"""Setting TWISTED_ENABLED=False in spider settings is not currently supported,
AsyncCrawlerProcess will install a reactor in this case.
"""
log = self.run_script("reactorless_custom_settings.py")
assert "Spider closed (finished)" not in log
assert "TWISTED_ENABLED is False but a Twisted reactor is installed." in log
def test_reactorless_datauri(self):
log = self.run_script("reactorless_datauri.py")
@ -371,7 +380,7 @@ class TestAsyncCrawlerProcessSubprocess(TestCrawlerProcessSubprocessBase):
def test_reactorless_telnetconsole_default(self):
"""By default TWISTED_ENABLED=False silently sets TELNETCONSOLE_ENABLED=False."""
log = self.run_script("reactorless_telnetconsole_default.py")
log = self.run_script("reactorless_simple.py") # no need for a separate script
assert "Not using a Twisted reactor" in log
assert "Spider closed (finished)" in log
assert "The TelnetConsole extension requires a Twisted reactor" not in log
@ -466,6 +475,14 @@ class TestCrawlerRunnerSubprocessBase(ScriptRunnerMixin):
"setting (uvloop.Loop)"
) in log
def test_no_reactor(self):
log = self.run_script("no_reactor.py")
assert "Spider closed (finished)" not in log
assert (
"RuntimeError: We expected a Twisted reactor to be installed but it isn't."
in log
)
class TestCrawlerRunnerSubprocess(TestCrawlerRunnerSubprocessBase):
@property
@ -521,7 +538,16 @@ class TestAsyncCrawlerRunnerSubprocess(TestCrawlerRunnerSubprocessBase):
assert "Spider closed (finished)" in log
assert "is_reactorless(): True" in log
assert "ERROR: " not in log
assert "WARNING: " not in log
assert log.count("WARNING: HttpxDownloadHandler is experimental") == 2
assert log.count("WARNING: ") == 2
def test_reactorless_custom_settings(self):
"""Setting TWISTED_ENABLED=False in spider settings is not currently supported,
AsyncCrawlerRunner will expect a reactor installed by the user.
"""
log = self.run_script("reactorless_custom_settings.py")
assert "Spider closed (finished)" not in log
assert "We expected a Twisted reactor to be installed but it isn't." in log
def test_reactorless_datauri(self):
log = self.run_script("reactorless_datauri.py")
@ -530,7 +556,8 @@ class TestAsyncCrawlerRunnerSubprocess(TestCrawlerRunnerSubprocessBase):
assert "{'data': 'foo'}" in log
assert "'item_scraped_count': 1" in log
assert "ERROR: " not in log
assert "WARNING: " not in log
assert log.count("WARNING: HttpxDownloadHandler is experimental") == 2
assert log.count("WARNING: ") == 2
def test_reactorless_reactor(self):
log = self.run_script("reactorless_reactor.py")