mirror of https://github.com/scrapy/scrapy.git
Docs for the reactorless mode (#7385)
* Reactorless mode docs. * Unset the ftp handler in the reactorless mode. * Cleanup reactorless subprocess tests. * Better handling of incompatible TWISTED_ENABLED cases. * Improvements. * Improve early error logging in test scripts. * Address feedback. * Address feedback.
This commit is contained in:
parent
fa76ca52e9
commit
0c6ccf50b3
|
|
@ -663,7 +663,7 @@ New features
|
||||||
(:issue:`4463`, :issue:`6804`)
|
(:issue:`4463`, :issue:`6804`)
|
||||||
|
|
||||||
- Added :func:`scrapy.utils.asyncio.is_asyncio_available` as an alternative
|
- Added :func:`scrapy.utils.asyncio.is_asyncio_available` as an alternative
|
||||||
to :func:`scrapy.utils.defer.is_asyncio_reactor_installed` with a
|
to :func:`scrapy.utils.reactor.is_asyncio_reactor_installed` with a
|
||||||
future-proof name and semantics.
|
future-proof name and semantics.
|
||||||
(:issue:`6827`)
|
(:issue:`6827`)
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -129,6 +129,173 @@ example:
|
||||||
.. autofunction:: scrapy.utils.reactor.is_asyncio_reactor_installed
|
.. autofunction:: scrapy.utils.reactor.is_asyncio_reactor_installed
|
||||||
|
|
||||||
|
|
||||||
|
.. _asyncio-without-reactor:
|
||||||
|
|
||||||
|
Using Scrapy without a Twisted reactor
|
||||||
|
======================================
|
||||||
|
|
||||||
|
.. versionadded:: 2.15.0
|
||||||
|
|
||||||
|
.. warning::
|
||||||
|
This is currently experimental and may not be suitable for production use.
|
||||||
|
|
||||||
|
It's possible to use Scrapy without installing a Twisted reactor at all, by
|
||||||
|
setting the :setting:`TWISTED_ENABLED` setting to ``False``. In this mode
|
||||||
|
Scrapy will use the asyncio event loop directly, and most of the Scrapy
|
||||||
|
functionality will work in the same way.
|
||||||
|
|
||||||
|
Doing this provides several benefits in certain use cases:
|
||||||
|
|
||||||
|
* A Twisted reactor, once stopped, cannot be started again. This prevents, for
|
||||||
|
example, using several instances of
|
||||||
|
:class:`~scrapy.crawler.AsyncCrawlerProcess` in the same process when they
|
||||||
|
use a reactor, but with ``TWISTED_ENABLED=False`` it becomes possible.
|
||||||
|
* There may be limitations imposed by
|
||||||
|
:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor` and related
|
||||||
|
Twisted code, such as the requirement of using
|
||||||
|
:class:`~asyncio.SelectorEventLoop` on Windows (see :ref:`asyncio-windows`),
|
||||||
|
that do not apply if the reactor is not used.
|
||||||
|
* :class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor` manages the
|
||||||
|
underlying event loop, and while :class:`~scrapy.crawler.AsyncCrawlerRunner`
|
||||||
|
can use a pre-existing reactor which, in turn, can use a pre-existing event
|
||||||
|
loop, it's easier to use :class:`~scrapy.crawler.AsyncCrawlerRunner` with a
|
||||||
|
pre-existing loop directly.
|
||||||
|
* Omitting the reactor machinery may improve performance and reliability.
|
||||||
|
|
||||||
|
Limitations
|
||||||
|
-----------
|
||||||
|
|
||||||
|
As some Scrapy features and components require a reactor, they don't work and
|
||||||
|
are disabled without it. Replacements that don't require a reactor may be added
|
||||||
|
in future Scrapy versions. The following features are not available:
|
||||||
|
|
||||||
|
* The default HTTP(S) download handler,
|
||||||
|
:class:`~scrapy.core.downloader.handlers.http11.HTTP11DownloadHandler` (this
|
||||||
|
is likely the biggest difference; Scrapy provides an HTTP(S) download handler
|
||||||
|
that doesn't require a reactor and will be used instead of it:
|
||||||
|
:class:`~scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler`)
|
||||||
|
* :class:`~scrapy.core.downloader.handlers.ftp.FTPDownloadHandler`
|
||||||
|
* :class:`~scrapy.core.downloader.handlers.http2.H2DownloadHandler`
|
||||||
|
* :ref:`topics-shell`
|
||||||
|
* :ref:`topics-telnetconsole`
|
||||||
|
* :class:`~scrapy.crawler.CrawlerRunner` and
|
||||||
|
:class:`~scrapy.crawler.CrawlerProcess`
|
||||||
|
(:class:`~scrapy.crawler.AsyncCrawlerProcess` and
|
||||||
|
:class:`~scrapy.crawler.AsyncCrawlerRunner` are available)
|
||||||
|
* Twisted-specific DNS resolvers (the :setting:`DNS_RESOLVER` setting)
|
||||||
|
* User and 3rd-party code that requires a reactor (see :ref:`below
|
||||||
|
<asyncio-without-reactor-migrate>` for examples)
|
||||||
|
|
||||||
|
Note that importing Twisted modules and, among other things, creating and using
|
||||||
|
:class:`~twisted.internet.defer.Deferred` objects doesn't require a reactor, so
|
||||||
|
code that uses :class:`~twisted.internet.defer.Deferred`,
|
||||||
|
:class:`~twisted.python.failure.Failure` and some other Twisted APIs will not
|
||||||
|
necessarily stop working.
|
||||||
|
|
||||||
|
Other differences
|
||||||
|
-----------------
|
||||||
|
|
||||||
|
When :setting:`TWISTED_ENABLED` is set to ``False``, Scrapy will change the
|
||||||
|
defaults of some other settings:
|
||||||
|
|
||||||
|
* :setting:`TELNETCONSOLE_ENABLED` is set to ``False``.
|
||||||
|
* The ``"http"`` and ``"https"`` keys in :setting:`DOWNLOAD_HANDLERS_BASE` are
|
||||||
|
set to ``"scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler"``.
|
||||||
|
* The ``"ftp"`` key in :setting:`DOWNLOAD_HANDLERS_BASE` is set to ``None``.
|
||||||
|
|
||||||
|
Thus, :class:`~scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler` is
|
||||||
|
used by default for making HTTP(S) requests. Please refer to its documentation
|
||||||
|
for its differences and limitations compared to
|
||||||
|
:class:`~scrapy.core.downloader.handlers.http11.HTTP11DownloadHandler`.
|
||||||
|
|
||||||
|
Additionally, :class:`~scrapy.crawler.AsyncCrawlerProcess` will install a
|
||||||
|
:term:`meta path finder` that prevents :mod:`twisted.internet.reactor` from
|
||||||
|
being imported.
|
||||||
|
|
||||||
|
.. _asyncio-without-reactor-migrate:
|
||||||
|
|
||||||
|
Adding support to existing code
|
||||||
|
-------------------------------
|
||||||
|
|
||||||
|
Code that doesn't directly use Twisted APIs or APIs that depend on Twisted ones
|
||||||
|
doesn't need special support for running without a reactor.
|
||||||
|
|
||||||
|
Here are some examples of APIs and patterns that need a replacement:
|
||||||
|
|
||||||
|
* Using :meth:`reactor.callLater()
|
||||||
|
<twisted.internet.base.ReactorBase.callLater>` for sleeping or delayed calls.
|
||||||
|
You can use :meth:`asyncio.loop.call_later` instead.
|
||||||
|
* Using :func:`twisted.internet.threads.deferToThread`,
|
||||||
|
:meth:`reactor.callFromThread()
|
||||||
|
<twisted.internet.base.ReactorBase.callFromThread>` and related APIs to
|
||||||
|
execute code in other threads. You can use :func:`asyncio.to_thread`,
|
||||||
|
:meth:`asyncio.loop.call_soon_threadsafe` and related APIs instead.
|
||||||
|
* Using :class:`twisted.internet.task.LoopingCall` for scheduling repeated
|
||||||
|
tasks. As there is no direct replacement in the standard library, you may
|
||||||
|
need to write your own one using :func:`asyncio.sleep` in a task.
|
||||||
|
* Using Twisted network client and server APIs (:meth:`reactor.connectTCP()
|
||||||
|
<twisted.internet.interfaces.IReactorTCP.connectTCP>`,
|
||||||
|
:meth:`reactor.listenTCP()
|
||||||
|
<twisted.internet.interfaces.IReactorTCP.listenTCP>`,
|
||||||
|
:mod:`twisted.web.client`, :mod:`twisted.mail.smtp` etc.). You can use other
|
||||||
|
built-in or 3rd-party libraries for this.
|
||||||
|
* Using :class:`~scrapy.crawler.CrawlerProcess` or
|
||||||
|
:class:`~scrapy.crawler.CrawlerRunner`. You should use
|
||||||
|
:class:`~scrapy.crawler.AsyncCrawlerProcess` or
|
||||||
|
:class:`~scrapy.crawler.AsyncCrawlerRunner` respectively instead.
|
||||||
|
* Checking whether ``asyncio`` support is available with
|
||||||
|
:func:`scrapy.utils.reactor.is_asyncio_reactor_installed`. You should use
|
||||||
|
:func:`scrapy.utils.asyncio.is_asyncio_available` instead.
|
||||||
|
|
||||||
|
Scrapy provides unified helpers for some of these examples:
|
||||||
|
|
||||||
|
.. autofunction:: scrapy.utils.asyncio.call_later
|
||||||
|
.. autofunction:: scrapy.utils.asyncio.create_looping_call
|
||||||
|
.. autoclass:: scrapy.utils.asyncio.AsyncioLoopingCall
|
||||||
|
.. autofunction:: scrapy.utils.asyncio.run_in_thread
|
||||||
|
|
||||||
|
If your code needs to know whether the reactor is available, you can either
|
||||||
|
check for the value of the :setting:`TWISTED_ENABLED` setting (you need access
|
||||||
|
to the :class:`~scrapy.crawler.Crawler` instance to do this) or use the
|
||||||
|
following function:
|
||||||
|
|
||||||
|
.. autofunction:: scrapy.utils.reactorless.is_reactorless
|
||||||
|
|
||||||
|
In general, code that doesn't use the reactor (directly or indirectly) can be
|
||||||
|
used unmodified both with the asyncio reactor and without a reactor. This
|
||||||
|
includes code that converts Deferreds to futures and vice versa as described in
|
||||||
|
:ref:`asyncio-await-dfd`.
|
||||||
|
|
||||||
|
Troubleshooting
|
||||||
|
---------------
|
||||||
|
|
||||||
|
**ImportError: Import of twisted.internet.reactor is forbidden when running
|
||||||
|
without a Twisted reactor [...]:** Scrapy is configured to run without a
|
||||||
|
reactor, but some code imported :mod:`twisted.internet.reactor`, most likely
|
||||||
|
because that code needs a reactor to be used. You need to stop using this code
|
||||||
|
or set :setting:`TWISTED_ENABLED` back to ``True``. It's also possible that the
|
||||||
|
reactor isn't really needed but was installed due to the problem described in
|
||||||
|
:ref:`asyncio-preinstalled-reactor`, in which case it should be enough to fix
|
||||||
|
the problematic imports.
|
||||||
|
|
||||||
|
**RuntimeError: TWISTED_ENABLED is False but a Twisted reactor is installed:**
|
||||||
|
Scrapy is configured to run without a reactor, but a reactor is already
|
||||||
|
installed before the Scrapy code is executed. If you are trying to set
|
||||||
|
:setting:`TWISTED_ENABLED` via :ref:`per-spider settings <spider-settings>`,
|
||||||
|
it's currently unsupported.
|
||||||
|
|
||||||
|
**RuntimeError: We expected a Twisted reactor to be installed but it isn't:**
|
||||||
|
Scrapy is configured to run with a reactor and not to install one, but a
|
||||||
|
reactor wasn't installed before the Scrapy code is executed. If you are trying
|
||||||
|
to set :setting:`TWISTED_ENABLED` via :ref:`per-spider settings
|
||||||
|
<spider-settings>`, it's currently unsupported.
|
||||||
|
|
||||||
|
**RuntimeError: <class> doesn't support TWISTED_ENABLED=False:** The listed
|
||||||
|
class cannot be used with :setting:`TWISTED_ENABLED` set to ``False``. There
|
||||||
|
may be a replacement in the :ref:`documentation above
|
||||||
|
<asyncio-without-reactor>` or the documentation of the affected class.
|
||||||
|
|
||||||
|
|
||||||
.. _asyncio-windows:
|
.. _asyncio-windows:
|
||||||
|
|
||||||
Windows-specific notes
|
Windows-specific notes
|
||||||
|
|
@ -149,6 +316,9 @@ automatically when you change the :setting:`TWISTED_REACTOR` setting or call
|
||||||
them together with Scrapy on Windows (but you should be able to use
|
them together with Scrapy on Windows (but you should be able to use
|
||||||
them on WSL or native Linux).
|
them on WSL or native Linux).
|
||||||
|
|
||||||
|
.. note:: This problem doesn't apply when not using the reactor, see
|
||||||
|
:ref:`asyncio-without-reactor`.
|
||||||
|
|
||||||
.. _playwright: https://github.com/microsoft/playwright-python
|
.. _playwright: https://github.com/microsoft/playwright-python
|
||||||
|
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -140,6 +140,9 @@ This handler supports ``ftp://host/path`` FTP URIs.
|
||||||
|
|
||||||
It's implemented using :mod:`twisted.protocols.ftp`.
|
It's implemented using :mod:`twisted.protocols.ftp`.
|
||||||
|
|
||||||
|
.. note::
|
||||||
|
This handler is not supported when :setting:`TWISTED_ENABLED` is ``False``.
|
||||||
|
|
||||||
.. _twisted-http2-handler:
|
.. _twisted-http2-handler:
|
||||||
|
|
||||||
H2DownloadHandler
|
H2DownloadHandler
|
||||||
|
|
@ -193,6 +196,9 @@ If you want to use this handler you need to replace the default one for the
|
||||||
.. _http2 faq: https://http2.github.io/faq/#does-http2-require-encryption
|
.. _http2 faq: https://http2.github.io/faq/#does-http2-require-encryption
|
||||||
.. _server pushes: https://datatracker.ietf.org/doc/html/rfc7540#section-8.2
|
.. _server pushes: https://datatracker.ietf.org/doc/html/rfc7540#section-8.2
|
||||||
|
|
||||||
|
.. note::
|
||||||
|
This handler is not supported when :setting:`TWISTED_ENABLED` is ``False``.
|
||||||
|
|
||||||
HTTP11DownloadHandler
|
HTTP11DownloadHandler
|
||||||
---------------------
|
---------------------
|
||||||
|
|
||||||
|
|
@ -206,6 +212,9 @@ uses the HTTP/1.1 protocol for them.
|
||||||
|
|
||||||
It's implemented using :mod:`twisted.web.client`.
|
It's implemented using :mod:`twisted.web.client`.
|
||||||
|
|
||||||
|
.. note::
|
||||||
|
This handler is not supported when :setting:`TWISTED_ENABLED` is ``False``.
|
||||||
|
|
||||||
HttpxDownloadHandler
|
HttpxDownloadHandler
|
||||||
--------------------
|
--------------------
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -166,6 +166,86 @@ with :class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor`):
|
||||||
|
|
||||||
.. seealso:: :doc:`twisted:core/howto/reactor-basics`
|
.. seealso:: :doc:`twisted:core/howto/reactor-basics`
|
||||||
|
|
||||||
|
And here are examples of using these classes with :setting:`TWISTED_ENABLED`
|
||||||
|
set to ``False``.
|
||||||
|
|
||||||
|
Simple usage of :class:`~scrapy.crawler.AsyncCrawlerProcess`:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
import scrapy
|
||||||
|
from scrapy.crawler import AsyncCrawlerProcess
|
||||||
|
|
||||||
|
|
||||||
|
class MySpider(scrapy.Spider):
|
||||||
|
# Your spider definition
|
||||||
|
...
|
||||||
|
|
||||||
|
|
||||||
|
process = AsyncCrawlerProcess(
|
||||||
|
settings={
|
||||||
|
"TWISTED_ENABLED": False,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
process.crawl(MySpider)
|
||||||
|
process.start() # the script will block here until the crawling is finished
|
||||||
|
|
||||||
|
With ``TWISTED_ENABLED=False`` you can use several instances of
|
||||||
|
:class:`~scrapy.crawler.AsyncCrawlerProcess` in the same process:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
import scrapy
|
||||||
|
from scrapy.crawler import AsyncCrawlerProcess
|
||||||
|
|
||||||
|
|
||||||
|
class MySpider(scrapy.Spider):
|
||||||
|
# Your spider definition
|
||||||
|
...
|
||||||
|
|
||||||
|
|
||||||
|
process1 = AsyncCrawlerProcess(
|
||||||
|
settings={
|
||||||
|
"TWISTED_ENABLED": False,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
process1.crawl(MySpider)
|
||||||
|
process1.start()
|
||||||
|
|
||||||
|
process2 = AsyncCrawlerProcess(
|
||||||
|
settings={
|
||||||
|
"TWISTED_ENABLED": False,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
process2.crawl(MySpider)
|
||||||
|
process2.start()
|
||||||
|
|
||||||
|
Using :func:`asyncio.run` with :class:`~scrapy.crawler.AsyncCrawlerRunner`:
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
|
||||||
|
import scrapy
|
||||||
|
from scrapy.crawler import AsyncCrawlerRunner
|
||||||
|
from scrapy.utils.log import configure_logging
|
||||||
|
|
||||||
|
|
||||||
|
class MySpider(scrapy.Spider):
|
||||||
|
# Your spider definition
|
||||||
|
...
|
||||||
|
|
||||||
|
|
||||||
|
async def main():
|
||||||
|
configure_logging({"LOG_FORMAT": "%(levelname)s: %(message)s"})
|
||||||
|
runner = AsyncCrawlerRunner(settings={"TWISTED_ENABLED": False})
|
||||||
|
await runner.crawl(MySpider) # completes when the spider finishes
|
||||||
|
|
||||||
|
|
||||||
|
asyncio.run(main())
|
||||||
|
|
||||||
|
|
||||||
.. _run-multiple-spiders:
|
.. _run-multiple-spiders:
|
||||||
|
|
||||||
Running multiple spiders in the same process
|
Running multiple spiders in the same process
|
||||||
|
|
|
||||||
|
|
@ -303,11 +303,12 @@ Pre-crawler settings
|
||||||
|
|
||||||
These settings cannot be :ref:`set from a spider <spider-settings>`.
|
These settings cannot be :ref:`set from a spider <spider-settings>`.
|
||||||
|
|
||||||
These settings are :setting:`SPIDER_LOADER_CLASS` and settings used by the
|
These settings are:
|
||||||
corresponding :ref:`component <topics-components>`, e.g.
|
|
||||||
:setting:`SPIDER_MODULES` and :setting:`SPIDER_LOADER_WARN_ONLY` for the
|
|
||||||
default component.
|
|
||||||
|
|
||||||
|
- :setting:`TWISTED_ENABLED`
|
||||||
|
- :setting:`SPIDER_LOADER_CLASS` and settings used by the corresponding
|
||||||
|
spider loader class, e.g. :setting:`SPIDER_MODULES` and
|
||||||
|
:setting:`SPIDER_LOADER_WARN_ONLY` for the default spider loader class.
|
||||||
|
|
||||||
.. _reactor-settings:
|
.. _reactor-settings:
|
||||||
|
|
||||||
|
|
@ -356,6 +357,9 @@ ignoring the value of :setting:`TWISTED_REACTOR` and using the value of
|
||||||
e.g. in :ref:`per-spider settings <spider-settings>`, an exception will be
|
e.g. in :ref:`per-spider settings <spider-settings>`, an exception will be
|
||||||
raised.
|
raised.
|
||||||
|
|
||||||
|
All of these settings, except for :setting:`ASYNCIO_EVENT_LOOP`, are only used
|
||||||
|
when the Twisted reactor is used, i.e. when :setting:`TWISTED_ENABLED` is
|
||||||
|
``True``.
|
||||||
|
|
||||||
.. _topics-settings-ref:
|
.. _topics-settings-ref:
|
||||||
|
|
||||||
|
|
@ -651,6 +655,13 @@ Default: ``True``
|
||||||
|
|
||||||
Whether to enable DNS in-memory cache.
|
Whether to enable DNS in-memory cache.
|
||||||
|
|
||||||
|
.. note::
|
||||||
|
This setting is only used by
|
||||||
|
:class:`~scrapy.resolver.CachingThreadedResolver` and
|
||||||
|
:class:`~scrapy.resolver.CachingHostnameResolver`. It has no effect when
|
||||||
|
:setting:`TWISTED_ENABLED` is ``False``, and may have no effect either when
|
||||||
|
:setting:`DNS_RESOLVER` is set to a different resolver.
|
||||||
|
|
||||||
.. setting:: DNSCACHE_SIZE
|
.. setting:: DNSCACHE_SIZE
|
||||||
|
|
||||||
DNSCACHE_SIZE
|
DNSCACHE_SIZE
|
||||||
|
|
@ -658,7 +669,7 @@ DNSCACHE_SIZE
|
||||||
|
|
||||||
Default: ``10000``
|
Default: ``10000``
|
||||||
|
|
||||||
DNS in-memory cache size.
|
DNS in-memory cache size, see :setting:`DNSCACHE_ENABLED`.
|
||||||
|
|
||||||
.. setting:: DNS_RESOLVER
|
.. setting:: DNS_RESOLVER
|
||||||
|
|
||||||
|
|
@ -667,12 +678,16 @@ DNS_RESOLVER
|
||||||
|
|
||||||
Default: ``'scrapy.resolver.CachingThreadedResolver'``
|
Default: ``'scrapy.resolver.CachingThreadedResolver'``
|
||||||
|
|
||||||
The class to be used to resolve DNS names. The default ``scrapy.resolver.CachingThreadedResolver``
|
The class to be used by Twisted to resolve DNS names. The default
|
||||||
supports specifying a timeout for DNS requests via the :setting:`DNS_TIMEOUT` setting,
|
``scrapy.resolver.CachingThreadedResolver`` supports specifying a timeout for
|
||||||
but works only with IPv4 addresses. Scrapy provides an alternative resolver,
|
DNS requests via the :setting:`DNS_TIMEOUT` setting, but works only with IPv4
|
||||||
|
addresses. Scrapy provides an alternative resolver,
|
||||||
``scrapy.resolver.CachingHostnameResolver``, which supports IPv4/IPv6 addresses but does not
|
``scrapy.resolver.CachingHostnameResolver``, which supports IPv4/IPv6 addresses but does not
|
||||||
take the :setting:`DNS_TIMEOUT` setting into account.
|
take the :setting:`DNS_TIMEOUT` setting into account.
|
||||||
|
|
||||||
|
.. note::
|
||||||
|
This setting has no effect when :setting:`TWISTED_ENABLED` is ``False``.
|
||||||
|
|
||||||
.. setting:: DNS_TIMEOUT
|
.. setting:: DNS_TIMEOUT
|
||||||
|
|
||||||
DNS_TIMEOUT
|
DNS_TIMEOUT
|
||||||
|
|
@ -682,6 +697,12 @@ Default: ``60``
|
||||||
|
|
||||||
Timeout for processing of DNS queries in seconds. Float is supported.
|
Timeout for processing of DNS queries in seconds. Float is supported.
|
||||||
|
|
||||||
|
.. note::
|
||||||
|
This setting is only used by
|
||||||
|
:class:`~scrapy.resolver.CachingThreadedResolver`. It has no effect when
|
||||||
|
:setting:`TWISTED_ENABLED` is ``False``, and may have no effect either when
|
||||||
|
:setting:`DNS_RESOLVER` is set to a different resolver.
|
||||||
|
|
||||||
.. setting:: DOWNLOADER
|
.. setting:: DOWNLOADER
|
||||||
|
|
||||||
DOWNLOADER
|
DOWNLOADER
|
||||||
|
|
@ -922,6 +943,20 @@ Default:
|
||||||
"ftp": "scrapy.core.downloader.handlers.ftp.FTPDownloadHandler",
|
"ftp": "scrapy.core.downloader.handlers.ftp.FTPDownloadHandler",
|
||||||
}
|
}
|
||||||
|
|
||||||
|
(when :setting:`TWISTED_ENABLED` is ``True``)
|
||||||
|
|
||||||
|
.. code-block:: python
|
||||||
|
|
||||||
|
{
|
||||||
|
"data": "scrapy.core.downloader.handlers.datauri.DataURIDownloadHandler",
|
||||||
|
"file": "scrapy.core.downloader.handlers.file.FileDownloadHandler",
|
||||||
|
"http": "scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler",
|
||||||
|
"https": "scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler",
|
||||||
|
"s3": "scrapy.core.downloader.handlers.s3.S3DownloadHandler",
|
||||||
|
"ftp": None,
|
||||||
|
}
|
||||||
|
|
||||||
|
(when :setting:`TWISTED_ENABLED` is ``False``)
|
||||||
|
|
||||||
A dict containing the :ref:`download handlers <topics-download-handlers>`
|
A dict containing the :ref:`download handlers <topics-download-handlers>`
|
||||||
enabled by default in Scrapy. You should never modify this setting in your
|
enabled by default in Scrapy. You should never modify this setting in your
|
||||||
|
|
@ -1954,7 +1989,7 @@ For more info see: :ref:`topics-stats`.
|
||||||
TELNETCONSOLE_ENABLED
|
TELNETCONSOLE_ENABLED
|
||||||
---------------------
|
---------------------
|
||||||
|
|
||||||
Default: ``True``
|
Default: ``True`` (``False`` when :setting:`TWISTED_ENABLED` is ``False``)
|
||||||
|
|
||||||
A boolean which specifies if the :ref:`telnet console <topics-telnetconsole>`
|
A boolean which specifies if the :ref:`telnet console <topics-telnetconsole>`
|
||||||
will be enabled (provided its extension is also enabled).
|
will be enabled (provided its extension is also enabled).
|
||||||
|
|
@ -1973,6 +2008,35 @@ command.
|
||||||
The project name must not conflict with the name of custom files or directories
|
The project name must not conflict with the name of custom files or directories
|
||||||
in the ``project`` subdirectory.
|
in the ``project`` subdirectory.
|
||||||
|
|
||||||
|
.. setting:: TWISTED_ENABLED
|
||||||
|
|
||||||
|
TWISTED_ENABLED
|
||||||
|
---------------
|
||||||
|
|
||||||
|
Default: ``True``
|
||||||
|
|
||||||
|
Whether to install and use the Twisted reactor.
|
||||||
|
|
||||||
|
If this is set to ``True``, Scrapy will use the Twisted reactor and will
|
||||||
|
install one according to the :setting:`TWISTED_REACTOR` setting value when
|
||||||
|
appropriate (e.g. when running via :ref:`the command-line tool
|
||||||
|
<topics-commands>`). This is the traditional mode of using Scrapy.
|
||||||
|
|
||||||
|
If this is set to ``False``, Scrapy will use the asyncio event loop directly
|
||||||
|
and will not attempt to install or use a reactor. Features that require a
|
||||||
|
reactor won't be available, but Twisted APIs that don't require a reactor,
|
||||||
|
including :class:`~twisted.internet.defer.Deferred` and
|
||||||
|
:class:`~twisted.python.failure.Failure`, will still be available. On the other
|
||||||
|
hand, limitations related to Twisted reactors (such as not being able to start
|
||||||
|
a reactor in the same process where a reactor was previously started and
|
||||||
|
stopped) will not apply. This mode is currently experimental and may not be
|
||||||
|
suitable for production use. It may also not be supported by 3rd-party code.
|
||||||
|
See :ref:`asyncio-without-reactor` for more information about this mode.
|
||||||
|
|
||||||
|
.. note:: This setting can't be set :ref:`per-spider <spider-settings>`.
|
||||||
|
|
||||||
|
.. versionadded:: 2.15.0
|
||||||
|
|
||||||
.. setting:: TWISTED_REACTOR
|
.. setting:: TWISTED_REACTOR
|
||||||
|
|
||||||
TWISTED_REACTOR
|
TWISTED_REACTOR
|
||||||
|
|
|
||||||
|
|
@ -17,6 +17,9 @@ spider, without having to run the spider to test every change.
|
||||||
Once you get familiarized with the Scrapy shell, you'll see that it's an
|
Once you get familiarized with the Scrapy shell, you'll see that it's an
|
||||||
invaluable tool for developing and debugging your spiders.
|
invaluable tool for developing and debugging your spiders.
|
||||||
|
|
||||||
|
.. note::
|
||||||
|
This feature is not supported when :setting:`TWISTED_ENABLED` is ``False``.
|
||||||
|
|
||||||
Configuring the shell
|
Configuring the shell
|
||||||
=====================
|
=====================
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -26,6 +26,9 @@ disable it if you want. For more information about the extension itself see
|
||||||
Please avoid using telnet console over insecure connections,
|
Please avoid using telnet console over insecure connections,
|
||||||
or disable it completely using :setting:`TELNETCONSOLE_ENABLED` option.
|
or disable it completely using :setting:`TELNETCONSOLE_ENABLED` option.
|
||||||
|
|
||||||
|
.. note::
|
||||||
|
This feature is not supported when :setting:`TWISTED_ENABLED` is ``False``.
|
||||||
|
|
||||||
.. highlight:: none
|
.. highlight:: none
|
||||||
|
|
||||||
How to access the telnet console
|
How to access the telnet console
|
||||||
|
|
|
||||||
|
|
@ -107,16 +107,26 @@ class Crawler:
|
||||||
|
|
||||||
use_reactor = self.settings.getbool("TWISTED_ENABLED")
|
use_reactor = self.settings.getbool("TWISTED_ENABLED")
|
||||||
if use_reactor:
|
if use_reactor:
|
||||||
|
# We either install a reactor or expect one to be installed.
|
||||||
reactor_class: str = self.settings["TWISTED_REACTOR"]
|
reactor_class: str = self.settings["TWISTED_REACTOR"]
|
||||||
event_loop: str = self.settings["ASYNCIO_EVENT_LOOP"]
|
event_loop: str = self.settings["ASYNCIO_EVENT_LOOP"]
|
||||||
if self._init_reactor:
|
if self._init_reactor:
|
||||||
# this needs to be done after the spider settings are merged,
|
# We need to install a reactor.
|
||||||
# but before something imports twisted.internet.reactor
|
# This needs to be done after the spider settings are merged,
|
||||||
|
# but before something imports twisted.internet.reactor.
|
||||||
if reactor_class:
|
if reactor_class:
|
||||||
|
# Install a specific reactor.
|
||||||
install_reactor(reactor_class, event_loop)
|
install_reactor(reactor_class, event_loop)
|
||||||
else:
|
else:
|
||||||
|
# Install the default one.
|
||||||
from twisted.internet import reactor # noqa: F401
|
from twisted.internet import reactor # noqa: F401
|
||||||
|
elif not is_reactor_installed():
|
||||||
|
# We need a reactor to be already installed.
|
||||||
|
raise RuntimeError(
|
||||||
|
"We expected a Twisted reactor to be installed but it isn't."
|
||||||
|
)
|
||||||
if reactor_class:
|
if reactor_class:
|
||||||
|
# We need to check that the correct reactor is installed.
|
||||||
verify_installed_reactor(reactor_class)
|
verify_installed_reactor(reactor_class)
|
||||||
if is_asyncio_reactor_installed() and event_loop:
|
if is_asyncio_reactor_installed() and event_loop:
|
||||||
verify_installed_asyncio_event_loop(event_loop)
|
verify_installed_asyncio_event_loop(event_loop)
|
||||||
|
|
@ -124,6 +134,11 @@ class Crawler:
|
||||||
if self._init_reactor or reactor_class:
|
if self._init_reactor or reactor_class:
|
||||||
log_reactor_info()
|
log_reactor_info()
|
||||||
else:
|
else:
|
||||||
|
# We expect a reactor to not be installed.
|
||||||
|
if is_reactor_installed():
|
||||||
|
raise RuntimeError(
|
||||||
|
"TWISTED_ENABLED is False but a Twisted reactor is installed."
|
||||||
|
)
|
||||||
logger.debug("Not using a Twisted reactor")
|
logger.debug("Not using a Twisted reactor")
|
||||||
self._apply_reactorless_default_settings()
|
self._apply_reactorless_default_settings()
|
||||||
|
|
||||||
|
|
@ -147,6 +162,7 @@ class Crawler:
|
||||||
self.settings["DOWNLOAD_HANDLERS_BASE"][scheme] = (
|
self.settings["DOWNLOAD_HANDLERS_BASE"][scheme] = (
|
||||||
"scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler"
|
"scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler"
|
||||||
)
|
)
|
||||||
|
self.settings["DOWNLOAD_HANDLERS_BASE"]["ftp"] = None
|
||||||
|
|
||||||
# Cannot use @deferred_f_from_coro_f because that relies on the reactor
|
# Cannot use @deferred_f_from_coro_f because that relies on the reactor
|
||||||
# being installed already, which is done within _apply_settings(), inside
|
# being installed already, which is done within _apply_settings(), inside
|
||||||
|
|
@ -477,17 +493,24 @@ class CrawlerRunner(CrawlerRunnerBase):
|
||||||
class AsyncCrawlerRunner(CrawlerRunnerBase):
|
class AsyncCrawlerRunner(CrawlerRunnerBase):
|
||||||
"""
|
"""
|
||||||
This is a convenient helper class that keeps track of, manages and runs
|
This is a convenient helper class that keeps track of, manages and runs
|
||||||
crawlers inside an already setup :mod:`~twisted.internet.reactor`.
|
crawlers inside an already setup :mod:`~twisted.internet.reactor` or
|
||||||
|
asyncio event loop.
|
||||||
|
|
||||||
The AsyncCrawlerRunner object must be instantiated with a
|
The AsyncCrawlerRunner object must be instantiated with a
|
||||||
:class:`~scrapy.settings.Settings` object.
|
:class:`~scrapy.settings.Settings` object.
|
||||||
|
|
||||||
|
When the :setting:`TWISTED_ENABLED` setting is set to ``True``, this class
|
||||||
|
requires a reactor to be installed and uses it, otherwise it requires a
|
||||||
|
reactor to not be installed but requires an asyncio event loop to be
|
||||||
|
installed and uses it.
|
||||||
|
|
||||||
This class shouldn't be needed (since Scrapy is responsible of using it
|
This class shouldn't be needed (since Scrapy is responsible of using it
|
||||||
accordingly) unless writing scripts that manually handle the crawling
|
accordingly) unless writing scripts that manually handle the crawling
|
||||||
process. See :ref:`run-from-script` for an example.
|
process. See :ref:`run-from-script` for an example.
|
||||||
|
|
||||||
This class provides coroutine APIs. It requires
|
This class provides coroutine APIs. It requires
|
||||||
:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor`.
|
:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor` when used
|
||||||
|
with a reactor.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
def __init__(self, settings: dict[str, Any] | Settings | None = None):
|
def __init__(self, settings: dict[str, Any] | Settings | None = None):
|
||||||
|
|
@ -528,6 +551,10 @@ class AsyncCrawlerRunner(CrawlerRunnerBase):
|
||||||
"it must be a spider class (or a Crawler object)"
|
"it must be a spider class (or a Crawler object)"
|
||||||
)
|
)
|
||||||
if self.settings.getbool("TWISTED_ENABLED"):
|
if self.settings.getbool("TWISTED_ENABLED"):
|
||||||
|
if not is_reactor_installed():
|
||||||
|
raise RuntimeError(
|
||||||
|
"We expected a Twisted reactor to be installed but it isn't."
|
||||||
|
)
|
||||||
if not is_asyncio_reactor_installed():
|
if not is_asyncio_reactor_installed():
|
||||||
raise RuntimeError(
|
raise RuntimeError(
|
||||||
f"When TWISTED_ENABLED is True, {type(self).__name__} "
|
f"When TWISTED_ENABLED is True, {type(self).__name__} "
|
||||||
|
|
@ -716,8 +743,8 @@ class CrawlerProcess(CrawlerProcessBase, CrawlerRunner):
|
||||||
) -> None:
|
) -> None:
|
||||||
"""
|
"""
|
||||||
This method starts a :mod:`~twisted.internet.reactor`, adjusts its pool
|
This method starts a :mod:`~twisted.internet.reactor`, adjusts its pool
|
||||||
size to :setting:`REACTOR_THREADPOOL_MAXSIZE`, and installs a DNS cache
|
size to :setting:`REACTOR_THREADPOOL_MAXSIZE`, and installs a DNS
|
||||||
based on :setting:`DNSCACHE_ENABLED` and :setting:`DNSCACHE_SIZE`.
|
resolver based on :setting:`DNSCACHE_ENABLED`.
|
||||||
|
|
||||||
If ``stop_after_crawl`` is True, the reactor will be stopped after all
|
If ``stop_after_crawl`` is True, the reactor will be stopped after all
|
||||||
crawlers have finished, using :meth:`join`.
|
crawlers have finished, using :meth:`join`.
|
||||||
|
|
@ -757,6 +784,10 @@ class AsyncCrawlerProcess(CrawlerProcessBase, AsyncCrawlerRunner):
|
||||||
The AsyncCrawlerProcess object must be instantiated with a
|
The AsyncCrawlerProcess object must be instantiated with a
|
||||||
:class:`~scrapy.settings.Settings` object.
|
:class:`~scrapy.settings.Settings` object.
|
||||||
|
|
||||||
|
When the :setting:`TWISTED_ENABLED` setting is set to ``True``, this class
|
||||||
|
installs a reactor and uses it, otherwise it requires a reactor to not be
|
||||||
|
installed but installs an asyncio event loop and uses it.
|
||||||
|
|
||||||
:param install_root_handler: whether to install root logging handler
|
:param install_root_handler: whether to install root logging handler
|
||||||
(default: True)
|
(default: True)
|
||||||
|
|
||||||
|
|
@ -765,7 +796,8 @@ class AsyncCrawlerProcess(CrawlerProcessBase, AsyncCrawlerRunner):
|
||||||
process. See :ref:`run-from-script` for an example.
|
process. See :ref:`run-from-script` for an example.
|
||||||
|
|
||||||
This class provides coroutine APIs. It requires
|
This class provides coroutine APIs. It requires
|
||||||
:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor`.
|
:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor` when used
|
||||||
|
with a reactor.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
def __init__(
|
def __init__(
|
||||||
|
|
@ -808,9 +840,12 @@ class AsyncCrawlerProcess(CrawlerProcessBase, AsyncCrawlerRunner):
|
||||||
self, stop_after_crawl: bool = True, install_signal_handlers: bool = True
|
self, stop_after_crawl: bool = True, install_signal_handlers: bool = True
|
||||||
) -> None:
|
) -> None:
|
||||||
"""
|
"""
|
||||||
This method starts a :mod:`~twisted.internet.reactor`, adjusts its pool
|
This method starts a :mod:`~twisted.internet.reactor`/asyncio event
|
||||||
size to :setting:`REACTOR_THREADPOOL_MAXSIZE`, and installs a DNS cache
|
loop, depending on the value of the :setting:`TWISTED_ENABLED` setting.
|
||||||
based on :setting:`DNSCACHE_ENABLED` and :setting:`DNSCACHE_SIZE`.
|
|
||||||
|
When using a reactor it adjusts its pool size to
|
||||||
|
:setting:`REACTOR_THREADPOOL_MAXSIZE` and installs a DNS resolver based
|
||||||
|
on :setting:`DNSCACHE_ENABLED`.
|
||||||
|
|
||||||
If ``stop_after_crawl`` is True, the reactor will be stopped after all
|
If ``stop_after_crawl`` is True, the reactor will be stopped after all
|
||||||
crawlers have finished, using :meth:`join`.
|
crawlers have finished, using :meth:`join`.
|
||||||
|
|
|
||||||
|
|
@ -132,13 +132,16 @@ async def _parallel_asyncio(
|
||||||
|
|
||||||
class AsyncioLoopingCall:
|
class AsyncioLoopingCall:
|
||||||
"""A simple implementation of a periodic call using asyncio, keeping
|
"""A simple implementation of a periodic call using asyncio, keeping
|
||||||
some API and behavior compatibility with the Twisted ``LoopingCall``.
|
some API and behavior compatibility with
|
||||||
|
:class:`~twisted.internet.task.LoopingCall`.
|
||||||
|
|
||||||
The function is called every *interval* seconds, independent of the finish
|
The function is called every *interval* seconds, independent of the finish
|
||||||
time of the previous call. If the function is still running when it's time
|
time of the previous call. If the function is still running when it's time
|
||||||
to call it again, calls are skipped until the function finishes.
|
to call it again, calls are skipped until the function finishes.
|
||||||
|
|
||||||
The function must not return a coroutine or a ``Deferred``.
|
The function must not return a coroutine or a ``Deferred``.
|
||||||
|
|
||||||
|
.. versionadded:: 2.14.0
|
||||||
"""
|
"""
|
||||||
|
|
||||||
def __init__(self, func: Callable[_P, _T], *args: _P.args, **kwargs: _P.kwargs):
|
def __init__(self, func: Callable[_P, _T], *args: _P.args, **kwargs: _P.kwargs):
|
||||||
|
|
@ -216,8 +219,12 @@ def create_looping_call(
|
||||||
) -> AsyncioLoopingCall | LoopingCall:
|
) -> AsyncioLoopingCall | LoopingCall:
|
||||||
"""Create an instance of a looping call class.
|
"""Create an instance of a looping call class.
|
||||||
|
|
||||||
This creates an instance of :class:`AsyncioLoopingCall` or
|
This creates an instance of
|
||||||
:class:`LoopingCall`, depending on whether asyncio support is available.
|
:class:`~scrapy.utils.asyncio.AsyncioLoopingCall` or
|
||||||
|
:class:`~twisted.internet.task.LoopingCall`, depending on whether asyncio
|
||||||
|
support is available.
|
||||||
|
|
||||||
|
.. versionadded:: 2.14.0
|
||||||
"""
|
"""
|
||||||
if is_asyncio_available():
|
if is_asyncio_available():
|
||||||
return AsyncioLoopingCall(func, *args, **kwargs)
|
return AsyncioLoopingCall(func, *args, **kwargs)
|
||||||
|
|
@ -229,8 +236,11 @@ def call_later(
|
||||||
) -> CallLaterResult:
|
) -> CallLaterResult:
|
||||||
"""Schedule a function to be called after a delay.
|
"""Schedule a function to be called after a delay.
|
||||||
|
|
||||||
This uses either ``loop.call_later()`` or ``reactor.callLater()``, depending
|
This uses either :meth:`asyncio.loop.call_later` or
|
||||||
on whether asyncio support is available.
|
:meth:`reactor.callLater() <twisted.internet.base.ReactorBase.callLater>`,
|
||||||
|
depending on whether asyncio support is available.
|
||||||
|
|
||||||
|
.. versionadded:: 2.14.0
|
||||||
"""
|
"""
|
||||||
if is_asyncio_available():
|
if is_asyncio_available():
|
||||||
loop = asyncio.get_event_loop()
|
loop = asyncio.get_event_loop()
|
||||||
|
|
@ -249,6 +259,8 @@ class CallLaterResult:
|
||||||
no ``active()`` (as there is no such public API in
|
no ``active()`` (as there is no such public API in
|
||||||
:class:`asyncio.TimerHandle`) but ``cancel()`` can be called on already
|
:class:`asyncio.TimerHandle`) but ``cancel()`` can be called on already
|
||||||
called or cancelled instances.
|
called or cancelled instances.
|
||||||
|
|
||||||
|
.. versionadded:: 2.14.0
|
||||||
"""
|
"""
|
||||||
|
|
||||||
_timer_handle: asyncio.TimerHandle | None = None
|
_timer_handle: asyncio.TimerHandle | None = None
|
||||||
|
|
|
||||||
|
|
@ -0,0 +1,36 @@
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
from typing import TYPE_CHECKING
|
||||||
|
|
||||||
|
import scrapy
|
||||||
|
from scrapy.crawler import AsyncCrawlerProcess
|
||||||
|
from scrapy.utils.reactorless import is_reactorless
|
||||||
|
|
||||||
|
if TYPE_CHECKING:
|
||||||
|
from asyncio import Task
|
||||||
|
|
||||||
|
|
||||||
|
class NoRequestsSpider(scrapy.Spider):
|
||||||
|
name = "no_request"
|
||||||
|
custom_settings = {
|
||||||
|
"TWISTED_ENABLED": False,
|
||||||
|
}
|
||||||
|
|
||||||
|
async def start(self):
|
||||||
|
self.logger.info(f"is_reactorless(): {is_reactorless()}")
|
||||||
|
return
|
||||||
|
yield
|
||||||
|
|
||||||
|
|
||||||
|
def log_task_exception(task: Task) -> None:
|
||||||
|
try:
|
||||||
|
task.result()
|
||||||
|
except Exception:
|
||||||
|
logging.exception("Crawl task failed") # noqa: LOG015
|
||||||
|
|
||||||
|
|
||||||
|
process = AsyncCrawlerProcess()
|
||||||
|
task = process.crawl(NoRequestsSpider)
|
||||||
|
task.add_done_callback(log_task_exception)
|
||||||
|
process.start()
|
||||||
|
|
@ -12,16 +12,7 @@ class NoRequestsSpider(scrapy.Spider):
|
||||||
yield
|
yield
|
||||||
|
|
||||||
|
|
||||||
process = AsyncCrawlerProcess(
|
process = AsyncCrawlerProcess(settings={"TWISTED_ENABLED": False})
|
||||||
settings={
|
|
||||||
"TWISTED_ENABLED": False,
|
|
||||||
"DOWNLOAD_HANDLERS": {
|
|
||||||
"http": None,
|
|
||||||
"https": None,
|
|
||||||
"ftp": None,
|
|
||||||
},
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
process.crawl(NoRequestsSpider)
|
process.crawl(NoRequestsSpider)
|
||||||
process.start()
|
process.start()
|
||||||
|
|
|
||||||
|
|
@ -3,13 +3,4 @@ from scrapy.utils.reactor import install_reactor
|
||||||
|
|
||||||
install_reactor("twisted.internet.asyncioreactor.AsyncioSelectorReactor")
|
install_reactor("twisted.internet.asyncioreactor.AsyncioSelectorReactor")
|
||||||
|
|
||||||
AsyncCrawlerProcess(
|
AsyncCrawlerProcess(settings={"TWISTED_ENABLED": False})
|
||||||
settings={
|
|
||||||
"TWISTED_ENABLED": False,
|
|
||||||
"DOWNLOAD_HANDLERS": {
|
|
||||||
"http": None,
|
|
||||||
"https": None,
|
|
||||||
"ftp": None,
|
|
||||||
},
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
|
||||||
|
|
@ -12,16 +12,7 @@ class NoRequestsSpider(scrapy.Spider):
|
||||||
yield
|
yield
|
||||||
|
|
||||||
|
|
||||||
process = AsyncCrawlerProcess(
|
process = AsyncCrawlerProcess(settings={"TWISTED_ENABLED": False})
|
||||||
settings={
|
|
||||||
"TWISTED_ENABLED": False,
|
|
||||||
"DOWNLOAD_HANDLERS": {
|
|
||||||
"http": None,
|
|
||||||
"https": None,
|
|
||||||
"ftp": None,
|
|
||||||
},
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
process.crawl(NoRequestsSpider)
|
process.crawl(NoRequestsSpider)
|
||||||
process.start()
|
process.start()
|
||||||
|
|
|
||||||
|
|
@ -1,25 +0,0 @@
|
||||||
import scrapy
|
|
||||||
from scrapy.crawler import AsyncCrawlerProcess
|
|
||||||
|
|
||||||
|
|
||||||
class NoRequestsSpider(scrapy.Spider):
|
|
||||||
name = "no_request"
|
|
||||||
|
|
||||||
async def start(self):
|
|
||||||
return
|
|
||||||
yield
|
|
||||||
|
|
||||||
|
|
||||||
process = AsyncCrawlerProcess(
|
|
||||||
settings={
|
|
||||||
"TWISTED_ENABLED": False,
|
|
||||||
"DOWNLOAD_HANDLERS": {
|
|
||||||
"http": None,
|
|
||||||
"https": None,
|
|
||||||
"ftp": None,
|
|
||||||
},
|
|
||||||
}
|
|
||||||
)
|
|
||||||
|
|
||||||
process.crawl(NoRequestsSpider)
|
|
||||||
process.start()
|
|
||||||
|
|
@ -13,11 +13,6 @@ class NoRequestsSpider(scrapy.Spider):
|
||||||
process = AsyncCrawlerProcess(
|
process = AsyncCrawlerProcess(
|
||||||
settings={
|
settings={
|
||||||
"TWISTED_ENABLED": False,
|
"TWISTED_ENABLED": False,
|
||||||
"DOWNLOAD_HANDLERS": {
|
|
||||||
"http": None,
|
|
||||||
"https": None,
|
|
||||||
"ftp": None,
|
|
||||||
},
|
|
||||||
"TELNETCONSOLE_ENABLED": False,
|
"TELNETCONSOLE_ENABLED": False,
|
||||||
}
|
}
|
||||||
)
|
)
|
||||||
|
|
|
||||||
|
|
@ -13,11 +13,6 @@ class NoRequestsSpider(scrapy.Spider):
|
||||||
process = AsyncCrawlerProcess(
|
process = AsyncCrawlerProcess(
|
||||||
settings={
|
settings={
|
||||||
"TWISTED_ENABLED": False,
|
"TWISTED_ENABLED": False,
|
||||||
"DOWNLOAD_HANDLERS": {
|
|
||||||
"http": None,
|
|
||||||
"https": None,
|
|
||||||
"ftp": None,
|
|
||||||
},
|
|
||||||
"TELNETCONSOLE_ENABLED": True,
|
"TELNETCONSOLE_ENABLED": True,
|
||||||
}
|
}
|
||||||
)
|
)
|
||||||
|
|
|
||||||
|
|
@ -0,0 +1,22 @@
|
||||||
|
import asyncio
|
||||||
|
|
||||||
|
from scrapy import Spider
|
||||||
|
from scrapy.crawler import AsyncCrawlerRunner
|
||||||
|
from scrapy.utils.log import configure_logging
|
||||||
|
|
||||||
|
|
||||||
|
class NoRequestsSpider(Spider):
|
||||||
|
name = "no_request"
|
||||||
|
|
||||||
|
async def start(self):
|
||||||
|
return
|
||||||
|
yield
|
||||||
|
|
||||||
|
|
||||||
|
async def main() -> None:
|
||||||
|
configure_logging()
|
||||||
|
runner = AsyncCrawlerRunner()
|
||||||
|
await runner.crawl(NoRequestsSpider)
|
||||||
|
|
||||||
|
|
||||||
|
asyncio.run(main())
|
||||||
|
|
@ -0,0 +1,27 @@
|
||||||
|
import asyncio
|
||||||
|
|
||||||
|
from scrapy import Spider
|
||||||
|
from scrapy.crawler import AsyncCrawlerRunner
|
||||||
|
from scrapy.utils.log import configure_logging
|
||||||
|
from scrapy.utils.reactorless import is_reactorless
|
||||||
|
|
||||||
|
|
||||||
|
class NoRequestsSpider(Spider):
|
||||||
|
name = "no_request"
|
||||||
|
custom_settings = {
|
||||||
|
"TWISTED_ENABLED": False,
|
||||||
|
}
|
||||||
|
|
||||||
|
async def start(self):
|
||||||
|
self.logger.info(f"is_reactorless(): {is_reactorless()}")
|
||||||
|
return
|
||||||
|
yield
|
||||||
|
|
||||||
|
|
||||||
|
async def main() -> None:
|
||||||
|
configure_logging()
|
||||||
|
runner = AsyncCrawlerRunner()
|
||||||
|
await runner.crawl(NoRequestsSpider)
|
||||||
|
|
||||||
|
|
||||||
|
asyncio.run(main())
|
||||||
|
|
@ -17,16 +17,7 @@ class DataSpider(Spider):
|
||||||
|
|
||||||
async def main() -> None:
|
async def main() -> None:
|
||||||
configure_logging()
|
configure_logging()
|
||||||
runner = AsyncCrawlerRunner(
|
runner = AsyncCrawlerRunner(settings={"TWISTED_ENABLED": False})
|
||||||
settings={
|
|
||||||
"TWISTED_ENABLED": False,
|
|
||||||
"DOWNLOAD_HANDLERS": {
|
|
||||||
"http": None,
|
|
||||||
"https": None,
|
|
||||||
"ftp": None,
|
|
||||||
},
|
|
||||||
}
|
|
||||||
)
|
|
||||||
await runner.crawl(DataSpider)
|
await runner.crawl(DataSpider)
|
||||||
|
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -16,16 +16,7 @@ class NoRequestsSpider(Spider):
|
||||||
|
|
||||||
async def main() -> None:
|
async def main() -> None:
|
||||||
configure_logging()
|
configure_logging()
|
||||||
runner = AsyncCrawlerRunner(
|
runner = AsyncCrawlerRunner(settings={"TWISTED_ENABLED": False})
|
||||||
settings={
|
|
||||||
"TWISTED_ENABLED": False,
|
|
||||||
"DOWNLOAD_HANDLERS": {
|
|
||||||
"http": None,
|
|
||||||
"https": None,
|
|
||||||
"ftp": None,
|
|
||||||
},
|
|
||||||
}
|
|
||||||
)
|
|
||||||
await runner.crawl(NoRequestsSpider)
|
await runner.crawl(NoRequestsSpider)
|
||||||
|
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -17,16 +17,7 @@ class NoRequestsSpider(Spider):
|
||||||
|
|
||||||
async def main() -> None:
|
async def main() -> None:
|
||||||
configure_logging()
|
configure_logging()
|
||||||
runner = AsyncCrawlerRunner(
|
runner = AsyncCrawlerRunner(settings={"TWISTED_ENABLED": False})
|
||||||
settings={
|
|
||||||
"TWISTED_ENABLED": False,
|
|
||||||
"DOWNLOAD_HANDLERS": {
|
|
||||||
"http": None,
|
|
||||||
"https": None,
|
|
||||||
"ftp": None,
|
|
||||||
},
|
|
||||||
}
|
|
||||||
)
|
|
||||||
await runner.crawl(NoRequestsSpider)
|
await runner.crawl(NoRequestsSpider)
|
||||||
|
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -0,0 +1,19 @@
|
||||||
|
from twisted.python import log
|
||||||
|
|
||||||
|
from scrapy import Spider
|
||||||
|
from scrapy.crawler import CrawlerRunner
|
||||||
|
from scrapy.utils.log import configure_logging
|
||||||
|
|
||||||
|
|
||||||
|
class NoRequestsSpider(Spider):
|
||||||
|
name = "no_request"
|
||||||
|
|
||||||
|
async def start(self):
|
||||||
|
return
|
||||||
|
yield
|
||||||
|
|
||||||
|
|
||||||
|
configure_logging()
|
||||||
|
runner = CrawlerRunner()
|
||||||
|
d = runner.crawl(NoRequestsSpider)
|
||||||
|
d.addErrback(log.err)
|
||||||
|
|
@ -352,7 +352,16 @@ class TestAsyncCrawlerProcessSubprocess(TestCrawlerProcessSubprocessBase):
|
||||||
assert "Spider closed (finished)" in log
|
assert "Spider closed (finished)" in log
|
||||||
assert "is_reactorless(): True" in log
|
assert "is_reactorless(): True" in log
|
||||||
assert "ERROR: " not in log
|
assert "ERROR: " not in log
|
||||||
assert "WARNING: " not in log
|
assert log.count("WARNING: HttpxDownloadHandler is experimental") == 2
|
||||||
|
assert log.count("WARNING: ") == 2
|
||||||
|
|
||||||
|
def test_reactorless_custom_settings(self):
|
||||||
|
"""Setting TWISTED_ENABLED=False in spider settings is not currently supported,
|
||||||
|
AsyncCrawlerProcess will install a reactor in this case.
|
||||||
|
"""
|
||||||
|
log = self.run_script("reactorless_custom_settings.py")
|
||||||
|
assert "Spider closed (finished)" not in log
|
||||||
|
assert "TWISTED_ENABLED is False but a Twisted reactor is installed." in log
|
||||||
|
|
||||||
def test_reactorless_datauri(self):
|
def test_reactorless_datauri(self):
|
||||||
log = self.run_script("reactorless_datauri.py")
|
log = self.run_script("reactorless_datauri.py")
|
||||||
|
|
@ -371,7 +380,7 @@ class TestAsyncCrawlerProcessSubprocess(TestCrawlerProcessSubprocessBase):
|
||||||
|
|
||||||
def test_reactorless_telnetconsole_default(self):
|
def test_reactorless_telnetconsole_default(self):
|
||||||
"""By default TWISTED_ENABLED=False silently sets TELNETCONSOLE_ENABLED=False."""
|
"""By default TWISTED_ENABLED=False silently sets TELNETCONSOLE_ENABLED=False."""
|
||||||
log = self.run_script("reactorless_telnetconsole_default.py")
|
log = self.run_script("reactorless_simple.py") # no need for a separate script
|
||||||
assert "Not using a Twisted reactor" in log
|
assert "Not using a Twisted reactor" in log
|
||||||
assert "Spider closed (finished)" in log
|
assert "Spider closed (finished)" in log
|
||||||
assert "The TelnetConsole extension requires a Twisted reactor" not in log
|
assert "The TelnetConsole extension requires a Twisted reactor" not in log
|
||||||
|
|
@ -466,6 +475,14 @@ class TestCrawlerRunnerSubprocessBase(ScriptRunnerMixin):
|
||||||
"setting (uvloop.Loop)"
|
"setting (uvloop.Loop)"
|
||||||
) in log
|
) in log
|
||||||
|
|
||||||
|
def test_no_reactor(self):
|
||||||
|
log = self.run_script("no_reactor.py")
|
||||||
|
assert "Spider closed (finished)" not in log
|
||||||
|
assert (
|
||||||
|
"RuntimeError: We expected a Twisted reactor to be installed but it isn't."
|
||||||
|
in log
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
class TestCrawlerRunnerSubprocess(TestCrawlerRunnerSubprocessBase):
|
class TestCrawlerRunnerSubprocess(TestCrawlerRunnerSubprocessBase):
|
||||||
@property
|
@property
|
||||||
|
|
@ -521,7 +538,16 @@ class TestAsyncCrawlerRunnerSubprocess(TestCrawlerRunnerSubprocessBase):
|
||||||
assert "Spider closed (finished)" in log
|
assert "Spider closed (finished)" in log
|
||||||
assert "is_reactorless(): True" in log
|
assert "is_reactorless(): True" in log
|
||||||
assert "ERROR: " not in log
|
assert "ERROR: " not in log
|
||||||
assert "WARNING: " not in log
|
assert log.count("WARNING: HttpxDownloadHandler is experimental") == 2
|
||||||
|
assert log.count("WARNING: ") == 2
|
||||||
|
|
||||||
|
def test_reactorless_custom_settings(self):
|
||||||
|
"""Setting TWISTED_ENABLED=False in spider settings is not currently supported,
|
||||||
|
AsyncCrawlerRunner will expect a reactor installed by the user.
|
||||||
|
"""
|
||||||
|
log = self.run_script("reactorless_custom_settings.py")
|
||||||
|
assert "Spider closed (finished)" not in log
|
||||||
|
assert "We expected a Twisted reactor to be installed but it isn't." in log
|
||||||
|
|
||||||
def test_reactorless_datauri(self):
|
def test_reactorless_datauri(self):
|
||||||
log = self.run_script("reactorless_datauri.py")
|
log = self.run_script("reactorless_datauri.py")
|
||||||
|
|
@ -530,7 +556,8 @@ class TestAsyncCrawlerRunnerSubprocess(TestCrawlerRunnerSubprocessBase):
|
||||||
assert "{'data': 'foo'}" in log
|
assert "{'data': 'foo'}" in log
|
||||||
assert "'item_scraped_count': 1" in log
|
assert "'item_scraped_count': 1" in log
|
||||||
assert "ERROR: " not in log
|
assert "ERROR: " not in log
|
||||||
assert "WARNING: " not in log
|
assert log.count("WARNING: HttpxDownloadHandler is experimental") == 2
|
||||||
|
assert log.count("WARNING: ") == 2
|
||||||
|
|
||||||
def test_reactorless_reactor(self):
|
def test_reactorless_reactor(self):
|
||||||
log = self.run_script("reactorless_reactor.py")
|
log = self.run_script("reactorless_reactor.py")
|
||||||
|
|
|
||||||
Loading…
Reference in New Issue