diff --git a/docs/topics/asyncio.rst b/docs/topics/asyncio.rst index 3d4ebe088..d0b99c38e 100644 --- a/docs/topics/asyncio.rst +++ b/docs/topics/asyncio.rst @@ -150,6 +150,12 @@ Using Scrapy without a Twisted reactor .. warning:: This is currently experimental and may not be suitable for production use. +.. note:: As the Twisted download handlers cannot be used without a reactor, + the default download handler in this mode is + :class:`~scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler`. You + will need to additionally install the ``httpx`` library to use it, unless + you switch to some different handler. + It's possible to use Scrapy without installing a Twisted reactor at all, by setting the :setting:`TWISTED_REACTOR_ENABLED` setting to ``False``. In this mode Scrapy will use the asyncio event loop directly, and most of the Scrapy diff --git a/docs/topics/practices.rst b/docs/topics/practices.rst index 216b7e64f..23738c98c 100644 --- a/docs/topics/practices.rst +++ b/docs/topics/practices.rst @@ -247,6 +247,110 @@ Using :func:`asyncio.run` with :class:`~scrapy.crawler.AsyncCrawlerRunner`: asyncio.run(main()) +.. _run-spiders-in-apps: + +Running spiders inside existing applications +============================================ + +You may want to run Scrapy spiders inside an existing application. In simple +cases (e.g. task queues that spawn a process for every task, or applications +that can execute tasks synchronously in the same process) you can use the same +approach as for standalone scripts (see :ref:`run-from-script`). More complex +cases, e.g. asynchronous web applications, have additional caveats and +limitations. + +If the application runs its own Twisted reactor, you can use +:class:`~scrapy.crawler.AsyncCrawlerRunner` or +:class:`~scrapy.crawler.CrawlerRunner` to run spiders using this reactor, see +:ref:`run-from-script` for examples. + +If the application doesn't run a Twisted reactor or an asyncio event loop (for +example, a Django web app deployed with a WSGI server such as uWSGI), you can +use :class:`~scrapy.crawler.AsyncCrawlerProcess` with +:setting:`TWISTED_REACTOR_ENABLED` set to ``False``, so that Scrapy starts and +stops an asyncio event loop for every spider run: + +.. code-block:: python + + import scrapy + from django.http import HttpResponse + from scrapy.crawler import AsyncCrawlerProcess + + + class MySpider(scrapy.Spider): + # Your spider definition + ... + + + def crawl_view(request): + process = AsyncCrawlerProcess(settings={"TWISTED_REACTOR_ENABLED": False}) + process.crawl(MySpider) + process.start() # returns when the spider finishes + return HttpResponse("Crawling finished") + +If the application runs its own asyncio event loop (for example, a Django web +app deployed with an ASGI server such as uvicorn), you can use +:class:`~scrapy.crawler.AsyncCrawlerRunner` with +:setting:`TWISTED_REACTOR_ENABLED` set to ``False``, so that Scrapy uses the +existing event loop: + +.. code-block:: python + + import scrapy + from django.http import HttpResponse + from scrapy.crawler import AsyncCrawlerRunner + + + class MySpider(scrapy.Spider): + # Your spider definition + ... + + + async def crawl_view(request): + runner = AsyncCrawlerRunner(settings={"TWISTED_REACTOR_ENABLED": False}) + await runner.crawl(MySpider) # completes when the spider finishes + return HttpResponse("Crawling finished") + +.. note:: Running Scrapy without a Twisted reactor is experimental and has + some limitations, described in :ref:`asyncio-without-reactor`. + +.. _run-in-notebook: + +Running spiders in Jupyter notebooks +==================================== + +You can run Scrapy spiders in Jupyter notebooks. You need to use +:class:`~scrapy.crawler.AsyncCrawlerRunner` with +:setting:`TWISTED_REACTOR_ENABLED` set to ``False`` for this, so that Scrapy +uses the event loop provided by the notebook kernel. As +:class:`~scrapy.crawler.AsyncCrawlerRunner` doesn't configure logging, and you +most likely want to see the spider log in the notebook, you should call +:func:`scrapy.utils.log.configure_logging`. Here is a full example, which +supports rerunning both as a single cell and as separate cells: + +.. code-block:: python + + from scrapy import Spider + from scrapy.crawler import AsyncCrawlerRunner + from scrapy.utils.log import configure_logging + + configure_logging() + + + class BooksSpider(Spider): + name = "books" + start_urls = ["https://books.toscrape.com"] + + def parse(self, response): + for book in response.css("h3"): + yield {"title": book.css("a::attr(title)").get()} + + + runner = AsyncCrawlerRunner({"TWISTED_REACTOR_ENABLED": False}) + await runner.crawl(BooksSpider) + +.. note:: Running Scrapy without a Twisted reactor is experimental and has + some limitations, described in :ref:`asyncio-without-reactor`. .. _run-multiple-spiders: