mirror of https://github.com/scrapy/scrapy.git
Add docs about running Scrapy from apps and notebooks. (#7751)
This commit is contained in:
parent
4d4a04f318
commit
1c5404dce5
|
|
@ -150,6 +150,12 @@ Using Scrapy without a Twisted reactor
|
|||
.. warning::
|
||||
This is currently experimental and may not be suitable for production use.
|
||||
|
||||
.. note:: As the Twisted download handlers cannot be used without a reactor,
|
||||
the default download handler in this mode is
|
||||
:class:`~scrapy.core.downloader.handlers._httpx.HttpxDownloadHandler`. You
|
||||
will need to additionally install the ``httpx`` library to use it, unless
|
||||
you switch to some different handler.
|
||||
|
||||
It's possible to use Scrapy without installing a Twisted reactor at all, by
|
||||
setting the :setting:`TWISTED_REACTOR_ENABLED` setting to ``False``. In this
|
||||
mode Scrapy will use the asyncio event loop directly, and most of the Scrapy
|
||||
|
|
|
|||
|
|
@ -247,6 +247,110 @@ Using :func:`asyncio.run` with :class:`~scrapy.crawler.AsyncCrawlerRunner`:
|
|||
|
||||
asyncio.run(main())
|
||||
|
||||
.. _run-spiders-in-apps:
|
||||
|
||||
Running spiders inside existing applications
|
||||
============================================
|
||||
|
||||
You may want to run Scrapy spiders inside an existing application. In simple
|
||||
cases (e.g. task queues that spawn a process for every task, or applications
|
||||
that can execute tasks synchronously in the same process) you can use the same
|
||||
approach as for standalone scripts (see :ref:`run-from-script`). More complex
|
||||
cases, e.g. asynchronous web applications, have additional caveats and
|
||||
limitations.
|
||||
|
||||
If the application runs its own Twisted reactor, you can use
|
||||
:class:`~scrapy.crawler.AsyncCrawlerRunner` or
|
||||
:class:`~scrapy.crawler.CrawlerRunner` to run spiders using this reactor, see
|
||||
:ref:`run-from-script` for examples.
|
||||
|
||||
If the application doesn't run a Twisted reactor or an asyncio event loop (for
|
||||
example, a Django web app deployed with a WSGI server such as uWSGI), you can
|
||||
use :class:`~scrapy.crawler.AsyncCrawlerProcess` with
|
||||
:setting:`TWISTED_REACTOR_ENABLED` set to ``False``, so that Scrapy starts and
|
||||
stops an asyncio event loop for every spider run:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
import scrapy
|
||||
from django.http import HttpResponse
|
||||
from scrapy.crawler import AsyncCrawlerProcess
|
||||
|
||||
|
||||
class MySpider(scrapy.Spider):
|
||||
# Your spider definition
|
||||
...
|
||||
|
||||
|
||||
def crawl_view(request):
|
||||
process = AsyncCrawlerProcess(settings={"TWISTED_REACTOR_ENABLED": False})
|
||||
process.crawl(MySpider)
|
||||
process.start() # returns when the spider finishes
|
||||
return HttpResponse("Crawling finished")
|
||||
|
||||
If the application runs its own asyncio event loop (for example, a Django web
|
||||
app deployed with an ASGI server such as uvicorn), you can use
|
||||
:class:`~scrapy.crawler.AsyncCrawlerRunner` with
|
||||
:setting:`TWISTED_REACTOR_ENABLED` set to ``False``, so that Scrapy uses the
|
||||
existing event loop:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
import scrapy
|
||||
from django.http import HttpResponse
|
||||
from scrapy.crawler import AsyncCrawlerRunner
|
||||
|
||||
|
||||
class MySpider(scrapy.Spider):
|
||||
# Your spider definition
|
||||
...
|
||||
|
||||
|
||||
async def crawl_view(request):
|
||||
runner = AsyncCrawlerRunner(settings={"TWISTED_REACTOR_ENABLED": False})
|
||||
await runner.crawl(MySpider) # completes when the spider finishes
|
||||
return HttpResponse("Crawling finished")
|
||||
|
||||
.. note:: Running Scrapy without a Twisted reactor is experimental and has
|
||||
some limitations, described in :ref:`asyncio-without-reactor`.
|
||||
|
||||
.. _run-in-notebook:
|
||||
|
||||
Running spiders in Jupyter notebooks
|
||||
====================================
|
||||
|
||||
You can run Scrapy spiders in Jupyter notebooks. You need to use
|
||||
:class:`~scrapy.crawler.AsyncCrawlerRunner` with
|
||||
:setting:`TWISTED_REACTOR_ENABLED` set to ``False`` for this, so that Scrapy
|
||||
uses the event loop provided by the notebook kernel. As
|
||||
:class:`~scrapy.crawler.AsyncCrawlerRunner` doesn't configure logging, and you
|
||||
most likely want to see the spider log in the notebook, you should call
|
||||
:func:`scrapy.utils.log.configure_logging`. Here is a full example, which
|
||||
supports rerunning both as a single cell and as separate cells:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
from scrapy import Spider
|
||||
from scrapy.crawler import AsyncCrawlerRunner
|
||||
from scrapy.utils.log import configure_logging
|
||||
|
||||
configure_logging()
|
||||
|
||||
|
||||
class BooksSpider(Spider):
|
||||
name = "books"
|
||||
start_urls = ["https://books.toscrape.com"]
|
||||
|
||||
def parse(self, response):
|
||||
for book in response.css("h3"):
|
||||
yield {"title": book.css("a::attr(title)").get()}
|
||||
|
||||
|
||||
runner = AsyncCrawlerRunner({"TWISTED_REACTOR_ENABLED": False})
|
||||
await runner.crawl(BooksSpider)
|
||||
|
||||
.. note:: Running Scrapy without a Twisted reactor is experimental and has
|
||||
some limitations, described in :ref:`asyncio-without-reactor`.
|
||||
|
||||
.. _run-multiple-spiders:
|
||||
|
||||
|
|
|
|||
Loading…
Reference in New Issue