scrapy/docs/topics/spiders.rst

17 KiB

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>

Spiders

Spiders are classes that define how a site, or a group of sites, is scraped: which requests to send, and how to parse their responses to extract data and to send additional requests.

A crawl goes as follows:

  1. Scrapy iterates the :meth:`~scrapy.Spider.start` method of the spider to get the initial requests. By default, that method yields a :class:`~scrapy.Request` object for each URL in :attr:`~scrapy.Spider.start_urls`, with :meth:`~scrapy.Spider.parse` as :ref:`callback <callbacks>`.

    System Message: ERROR/3 (<stdin>, line 13); backlink

    Unknown interpreted text role "meth".

    System Message: ERROR/3 (<stdin>, line 13); backlink

    Unknown interpreted text role "class".

    System Message: ERROR/3 (<stdin>, line 13); backlink

    Unknown interpreted text role "attr".

    System Message: ERROR/3 (<stdin>, line 13); backlink

    Unknown interpreted text role "meth".

    System Message: ERROR/3 (<stdin>, line 13); backlink

    Unknown interpreted text role "ref".

  2. Scrapy downloads each request and calls its callback with the resulting :class:`~scrapy.http.Response`.

    System Message: ERROR/3 (<stdin>, line 19); backlink

    Unknown interpreted text role "class".

  3. Callbacks parse the response, typically using :ref:`topics-selectors`, and return or yield :ref:`item objects <topics-items>` with the extracted data and :class:`~scrapy.Request` objects to continue the crawl, which go back to step 2. See :ref:`callback-output`.

    System Message: ERROR/3 (<stdin>, line 22); backlink

    Unknown interpreted text role "ref".

    System Message: ERROR/3 (<stdin>, line 22); backlink

    Unknown interpreted text role "ref".

    System Message: ERROR/3 (<stdin>, line 22); backlink

    Unknown interpreted text role "class".

    System Message: ERROR/3 (<stdin>, line 22); backlink

    Unknown interpreted text role "ref".

  4. Items go through :ref:`item pipelines <topics-item-pipeline>`, and are usually stored through :ref:`topics-feed-exports`.

    System Message: ERROR/3 (<stdin>, line 27); backlink

    Unknown interpreted text role "ref".

    System Message: ERROR/3 (<stdin>, line 27); backlink

    Unknown interpreted text role "ref".

Scrapy includes different spider classes for different purposes, described below.

scrapy.Spider

System Message: ERROR/3 (<stdin>, line 39)

Unknown directive type "autoclass".

.. autoclass:: scrapy.Spider

    .. autoattribute:: name

    .. attribute:: allowed_domains
        :type: list[str]

        The domains that this spider is allowed to crawl, if any. Requests for
        URLs not belonging to the domain names specified in this list (or their
        subdomains) won't be followed if
        :class:`~scrapy.downloadermiddlewares.offsite.OffsiteMiddleware` is
        enabled.

        .. versionchanged:: VERSION
           Changes to this attribute during a crawl are now taken into account.

        Let's say your target url is ``https://www.example.com/1.html``,
        then add ``'example.com'`` to the list.

        You may modify this attribute while the spider runs, e.g. to allow
        domains that you only learn about from an earlier response. The change
        affects requests scheduled after it.

    .. autoattribute:: start_urls

    .. autoattribute:: custom_settings

    .. autoattribute:: crawler

    .. autoattribute:: settings

    .. autoattribute:: logger

    .. attribute:: state
        :type: dict[str, Any]

        Spider state to persist between batches.
        See :ref:`topics-keeping-persistent-state-between-batches` for details.

    .. automethod:: update_settings

    .. automethod:: from_crawler

    .. automethod:: start

    .. automethod:: parse

    .. method:: closed(reason)

        Called when the spider closes. This method provides a shortcut to
        signals.connect() for the :signal:`spider_closed` signal.

Let's see an example:

System Message: WARNING/2 (<stdin>, line 93)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy


    class MySpider(scrapy.Spider):
        name = "example.com"
        allowed_domains = ["example.com"]
        start_urls = [
            "http://www.example.com/1.html",
            "http://www.example.com/2.html",
            "http://www.example.com/3.html",
        ]

        def parse(self, response):
            self.logger.info("A response from %s just arrived!", response.url)

Return multiple Requests and items from a single callback:

System Message: WARNING/2 (<stdin>, line 112)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy


    class MySpider(scrapy.Spider):
        name = "example.com"
        allowed_domains = ["example.com"]
        start_urls = [
            "http://www.example.com/1.html",
            "http://www.example.com/2.html",
            "http://www.example.com/3.html",
        ]

        def parse(self, response):
            for h3 in response.xpath("//h3").getall():
                yield {"title": h3}

            for href in response.xpath("//a/@href").getall():
                yield scrapy.Request(response.urljoin(href), self.parse)

Instead of :attr:`~.start_urls` you can use :meth:`~scrapy.Spider.start` directly; to give data more structure you can use :class:`~scrapy.Item` objects:

System Message: ERROR/3 (<stdin>, line 133); backlink

Unknown interpreted text role "attr".

System Message: ERROR/3 (<stdin>, line 133); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 133); backlink

Unknown interpreted text role "class".

System Message: WARNING/2 (<stdin>, line 138)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy
    from myproject.items import MyItem


    class MySpider(scrapy.Spider):
        name = "example.com"
        allowed_domains = ["example.com"]

        async def start(self):
            yield scrapy.Request("http://www.example.com/1.html", self.parse)
            yield scrapy.Request("http://www.example.com/2.html", self.parse)
            yield scrapy.Request("http://www.example.com/3.html", self.parse)

        def parse(self, response):
            for h3 in response.xpath("//h3").getall():
                yield MyItem(title=h3)

            for href in response.xpath("//a/@href").getall():
                yield scrapy.Request(response.urljoin(href), self.parse)

Spider arguments

Spiders can receive arguments that modify their behaviour. Some common uses for spider arguments are to define the start URLs or to restrict the crawl to certain sections of the site, but they can be used to configure any functionality of the spider.

Spider arguments are passed through the :command:`crawl` command using the -a option. For example:

System Message: ERROR/3 (<stdin>, line 170); backlink

Unknown interpreted text role "command".
scrapy crawl myspider -a category=electronics

Spiders can access arguments in their __init__ methods:

System Message: WARNING/2 (<stdin>, line 177)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy


    class MySpider(scrapy.Spider):
        name = "myspider"

        def __init__(self, category=None, *args, **kwargs):
            super().__init__(*args, **kwargs)
            self.start_urls = [f"http://www.example.com/categories/{category}"]
            # ...

The default __init__ method will take any spider arguments and copy them to the spider as attributes. The above example can also be written as follows:

System Message: WARNING/2 (<stdin>, line 194)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy


    class MySpider(scrapy.Spider):
        name = "myspider"

        async def start(self):
            yield scrapy.Request(f"http://www.example.com/categories/{self.category}")

If you are :ref:`running Scrapy from a script <run-from-script>`, you can specify spider arguments when calling :meth:`CrawlerProcess.crawl <scrapy.crawler.CrawlerProcess.crawl>` or :meth:`CrawlerRunner.crawl <scrapy.crawler.CrawlerRunner.crawl>`:

System Message: ERROR/3 (<stdin>, line 205); backlink

Unknown interpreted text role "ref".

System Message: ERROR/3 (<stdin>, line 205); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 205); backlink

Unknown interpreted text role "meth".

System Message: WARNING/2 (<stdin>, line 211)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    process = CrawlerProcess()
    process.crawl(MySpider, category="electronics")

Keep in mind that spider arguments are only strings. The spider will not do any parsing on its own. If you were to set the start_urls attribute from the command line, you would have to parse it on your own into a list using something like :func:`ast.literal_eval` or :func:`json.loads` and then set it as an attribute. Otherwise, you would cause iteration over a start_urls string (a very common python pitfall) resulting in each character being seen as a separate url.

System Message: ERROR/3 (<stdin>, line 216); backlink

Unknown interpreted text role "func".

System Message: ERROR/3 (<stdin>, line 216); backlink

Unknown interpreted text role "func".

Spider arguments can also be passed through the Scrapyd schedule.json API. See Scrapyd documentation.

scrapy-spider-metadata parameters

Another alternative to pass spider arguments is the library scrapy-spider-metadata.

This allows for Scrapy spiders to define, validate, document and pre-process their arguments as Pydantic models.

The example shows how to define typed parameters where a string argument is automatically converted to an integer:

System Message: WARNING/2 (<stdin>, line 242)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy
    from pydantic import BaseModel
    from scrapy_spider_metadata import Args


    class MyParams(BaseModel):
        pages: int


    class BookSpider(Args[MyParams], scrapy.Spider):
        name = "bookspider"
        start_urls = ["http://books.toscrape.com/catalogue"]

        async def start(self):
            for start_url in self.start_urls:
                for index in range(1, self.args.pages + 1):
                    yield scrapy.Request(f"{start_url}/page-{index}.html")

        def parse(self, response):
            book_links = response.css("article.product_pod h3 a::attr(href)").getall()
            for book_link in book_links:
                yield response.follow(book_link, self.parse_book)

        def parse_book(self, response):
            yield {
                "title": response.css("h1::text").get(),
                "price": response.css("p.price_color::text").get(),
            }

This spider can be called from the command line:

scrapy crawl bookspider -a pages=2

Start requests

Start requests are :class:`~scrapy.Request` objects yielded from the :meth:`~scrapy.Spider.start` method of a spider or from the :meth:`~scrapy.spidermiddlewares.SpiderMiddleware.process_start` method of a :ref:`spider middleware <topics-spider-middleware>`.

System Message: ERROR/3 (<stdin>, line 282); backlink

Unknown interpreted text role "class".

System Message: ERROR/3 (<stdin>, line 282); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 282); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 282); backlink

Unknown interpreted text role "ref".

System Message: ERROR/3 (<stdin>, line 287)

Unknown directive type "seealso".

.. seealso:: :ref:`start-request-order`

Delaying start request iteration

Scrapy iterates :meth:`~scrapy.Spider.start` as fast as it yields, so all start requests reach the scheduler early in the crawl, however many they are. To minimize the number of requests in the scheduler at any given time, and with it resource usage (memory, or disk when using :setting:`JOBDIR`), override :meth:`~scrapy.Spider.start` to pause its iteration whenever there are scheduled requests:

System Message: ERROR/3 (<stdin>, line 294); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 294); backlink

Unknown interpreted text role "setting".

System Message: ERROR/3 (<stdin>, line 294); backlink

Unknown interpreted text role "meth".

System Message: WARNING/2 (<stdin>, line 301)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    async def start(self):
        async for item_or_request in super().start():
            if self.crawler.engine.needs_backout():
                await self.crawler.signals.wait_for(signals.scheduler_empty)
            yield item_or_request

Generic Spiders

Scrapy comes with some useful generic spiders that you can use to subclass your spiders from. Their aim is to provide convenient functionality for a few common scraping cases, like following all links on a site based on certain rules, crawling from Sitemaps, or parsing an XML/CSV feed.

For the examples used in the following spiders, we'll assume you have a project with a TestItem declared in a myproject.items module:

System Message: WARNING/2 (<stdin>, line 322)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from dataclasses import dataclass


    @dataclass
    class TestItem:
        id: str | None = None
        name: str | None = None
        description: str | None = None


System Message: ERROR/3 (<stdin>, line 334)

Unknown directive type "currentmodule".

.. currentmodule:: scrapy.spiders

CrawlSpider

System Message: ERROR/3 (<stdin>, line 339)

Unknown directive type "autoclass".

.. autoclass:: CrawlSpider

    .. autoattribute:: rules

    .. automethod:: parse_start_url

Crawling rules

System Message: ERROR/3 (<stdin>, line 348)

Unknown directive type "autoclass".

.. autoclass:: Rule

CrawlSpider example

Let's now take a look at an example CrawlSpider with rules:

System Message: WARNING/2 (<stdin>, line 355)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import CrawlSpider, Rule
    from scrapy.linkextractors import LinkExtractor


    class MySpider(CrawlSpider):
        name = "example.com"
        allowed_domains = ["example.com"]
        start_urls = ["http://www.example.com"]

        rules = (
            # Extract links matching 'category.php' (but not matching 'subsection.php')
            # and follow links from them (since no callback means follow=True by default).
            Rule(LinkExtractor(allow=(r"category\.php",), deny=(r"subsection\.php",))),
            # Extract links matching 'item.php' and parse them with the spider's method parse_item
            Rule(LinkExtractor(allow=(r"item\.php",)), callback="parse_item"),
        )

        def parse_item(self, response):
            self.logger.info("Hi, this is an item page! %s", response.url)
            item = {}
            item["id"] = response.xpath('//td[@id="item_id"]/text()').re(r"ID: (\d+)")
            item["name"] = response.xpath('//td[@id="item_name"]/text()').get()
            item["description"] = response.xpath(
                '//td[@id="item_description"]/text()'
            ).get()
            item["link_text"] = response.meta["link_text"]
            url = response.xpath('//td[@id="additional_data"]/@href').get()
            return response.follow(
                url, self.parse_additional_page, cb_kwargs=dict(item=item)
            )

        def parse_additional_page(self, response, item):
            item["additional_data"] = response.xpath(
                '//p[@id="additional_data"]/text()'
            ).get()
            return item


This spider would start crawling example.com's home page, collecting category links, and item links, parsing the latter with the parse_item method. For each item response, some data will be extracted from the HTML using XPath, and a dictionary will be filled with it.

XMLFeedSpider

System Message: ERROR/3 (<stdin>, line 403)

Unknown directive type "autoclass".

.. autoclass:: XMLFeedSpider

    .. autoattribute:: iterator

    .. autoattribute:: itertag

    .. autoattribute:: namespaces

    .. automethod:: adapt_response

    .. automethod:: parse_node

    .. automethod:: process_results


XMLFeedSpider example

These spiders are pretty easy to use, let's have a look at one example:

System Message: WARNING/2 (<stdin>, line 424)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import XMLFeedSpider
    from myproject.items import TestItem


    class MySpider(XMLFeedSpider):
        name = "example.com"
        allowed_domains = ["example.com"]
        start_urls = ["http://www.example.com/feed.xml"]
        iterator = "iternodes"  # This is actually unnecessary, since it's the default value
        itertag = "item"

        def parse_node(self, response, node):
            self.logger.info(
                "Hi, this is a <%s> node!: %s", self.itertag, "".join(node.getall())
            )

            item = TestItem()
            item.id = node.xpath("@id").get()
            item.name = node.xpath("name").get()
            item.description = node.xpath("description").get()
            return item

Basically what we did up there was to create a spider that downloads a feed from the given start_urls, and then iterates through each of its item tags, prints them out, and stores some random data in an :class:`~scrapy.Item`.

System Message: ERROR/3 (<stdin>, line 448); backlink

Unknown interpreted text role "class".

CSVFeedSpider

System Message: ERROR/3 (<stdin>, line 455)

Unknown directive type "autoclass".

.. autoclass:: CSVFeedSpider

    .. autoattribute:: delimiter

    .. autoattribute:: quotechar

    .. autoattribute:: headers

    .. automethod:: adapt_response

    .. automethod:: parse_row

    .. automethod:: process_results

CSVFeedSpider example

Let's see an example similar to the previous one, but using a :class:`CSVFeedSpider`:

System Message: ERROR/3 (<stdin>, line 472); backlink

Unknown interpreted text role "class".

System Message: WARNING/2 (<stdin>, line 476)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import CSVFeedSpider
    from myproject.items import TestItem


    class MySpider(CSVFeedSpider):
        name = "example.com"
        allowed_domains = ["example.com"]
        start_urls = ["http://www.example.com/feed.csv"]
        delimiter = ";"
        quotechar = "'"
        headers = ["id", "name", "description"]

        def parse_row(self, response, row):
            self.logger.info("Hi, this is a row!: %r", row)

            item = TestItem()
            item.id = row["id"]
            item.name = row["name"]
            item.description = row["description"]
            return item


SitemapSpider

System Message: ERROR/3 (<stdin>, line 503)

Unknown directive type "autoclass".

.. autoclass:: SitemapSpider

    .. autoattribute:: sitemap_urls

    .. autoattribute:: sitemap_rules

    .. autoattribute:: sitemap_follow

    .. autoattribute:: sitemap_alternate_links

    .. automethod:: sitemap_filter


SitemapSpider examples

Simplest example: process all urls discovered through sitemaps using the parse callback:

System Message: WARNING/2 (<stdin>, line 522)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import SitemapSpider


    class MySpider(SitemapSpider):
        sitemap_urls = ["http://www.example.com/sitemap.xml"]

        def parse(self, response):
            pass  # ... scrape item here ...

Process some urls with certain callback and other urls with a different callback:

System Message: WARNING/2 (<stdin>, line 536)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import SitemapSpider


    class MySpider(SitemapSpider):
        sitemap_urls = ["http://www.example.com/sitemap.xml"]
        sitemap_rules = [
            ("/product/", "parse_product"),
            ("/category/", "parse_category"),
        ]

        def parse_product(self, response):
            pass  # ... scrape product ...

        def parse_category(self, response):
            pass  # ... scrape category ...

Follow sitemaps defined in the robots.txt file and only follow sitemaps whose url contains /sitemap_shop:

System Message: WARNING/2 (<stdin>, line 557)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import SitemapSpider


    class MySpider(SitemapSpider):
        sitemap_urls = ["http://www.example.com/robots.txt"]
        sitemap_rules = [
            ("/shop/", "parse_shop"),
        ]
        sitemap_follow = ["/sitemap_shops"]

        def parse_shop(self, response):
            pass  # ... scrape shop here ...

Combine SitemapSpider with other sources of urls:

System Message: WARNING/2 (<stdin>, line 574)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy import Request
    from scrapy.spiders import SitemapSpider


    class MySpider(SitemapSpider):
        sitemap_urls = ["http://www.example.com/robots.txt"]
        sitemap_rules = [
            ("/shop/", "parse_shop"),
        ]

        other_urls = ["http://www.example.com/about"]

        async def start(self):
            async for item_or_request in super().start():
                yield item_or_request
            for url in self.other_urls:
                yield Request(url, self.parse_other)

        def parse_shop(self, response):
            pass  # ... scrape shop here ...

        def parse_other(self, response):
            pass  # ... scrape other here ...

</html>