scrapy/docs/topics/spiders.rst

17 KiB

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>

Spiders

Spiders are classes that define how a site, or a group of sites, is scraped: which requests to send, and how to parse their responses to extract data and to send additional requests.

A crawl goes as follows:

  1. Scrapy iterates the :meth:`~scrapy.Spider.start` method of the spider to get the initial requests. By default, that method yields a :class:`~scrapy.Request` object for each URL in :attr:`~scrapy.Spider.start_urls`, with :meth:`~scrapy.Spider.parse` as :ref:`callback <callbacks>`.

    System Message: ERROR/3 (<stdin>, line 13); backlink

    Unknown interpreted text role "meth".

    System Message: ERROR/3 (<stdin>, line 13); backlink

    Unknown interpreted text role "class".

    System Message: ERROR/3 (<stdin>, line 13); backlink

    Unknown interpreted text role "attr".

    System Message: ERROR/3 (<stdin>, line 13); backlink

    Unknown interpreted text role "meth".

    System Message: ERROR/3 (<stdin>, line 13); backlink

    Unknown interpreted text role "ref".

  2. Scrapy downloads each request and calls its callback with the resulting :class:`~scrapy.http.Response`.

    System Message: ERROR/3 (<stdin>, line 19); backlink

    Unknown interpreted text role "class".

  3. Callbacks parse the response, typically using :ref:`topics-selectors`, and return or yield :ref:`item objects <topics-items>` with the extracted data and :class:`~scrapy.Request` objects to continue the crawl, which go back to step 2. See :ref:`callback-output`.

    System Message: ERROR/3 (<stdin>, line 22); backlink

    Unknown interpreted text role "ref".

    System Message: ERROR/3 (<stdin>, line 22); backlink

    Unknown interpreted text role "ref".

    System Message: ERROR/3 (<stdin>, line 22); backlink

    Unknown interpreted text role "class".

    System Message: ERROR/3 (<stdin>, line 22); backlink

    Unknown interpreted text role "ref".

  4. Items go through :ref:`item pipelines <topics-item-pipeline>`, and are usually stored through :ref:`topics-feed-exports`.

    System Message: ERROR/3 (<stdin>, line 27); backlink

    Unknown interpreted text role "ref".

    System Message: ERROR/3 (<stdin>, line 27); backlink

    Unknown interpreted text role "ref".

Scrapy includes different spider classes for different purposes, described below.

scrapy.Spider

System Message: ERROR/3 (<stdin>, line 39)

Unknown directive type "autoclass".

.. autoclass:: scrapy.Spider

    .. autoattribute:: name

    .. attribute:: allowed_domains
        :type: list[str]

        The domains that this spider is allowed to crawl, if any. Requests for
        URLs not belonging to the domain names specified in this list (or their
        subdomains) won't be followed if
        :class:`~scrapy.downloadermiddlewares.offsite.OffsiteMiddleware` is
        enabled.

        Let's say your target url is ``https://www.example.com/1.html``,
        then add ``'example.com'`` to the list.

    .. autoattribute:: start_urls

    .. autoattribute:: custom_settings

    .. autoattribute:: crawler

    .. autoattribute:: settings

    .. autoattribute:: logger

    .. attribute:: state
        :type: dict[str, Any]

        Spider state to persist between batches.
        See :ref:`topics-keeping-persistent-state-between-batches` for details.

    .. automethod:: update_settings

    .. automethod:: from_crawler

    .. automethod:: start

    .. automethod:: parse

    .. method:: closed(reason)

        Called when the spider closes. This method provides a shortcut to
        signals.connect() for the :signal:`spider_closed` signal.

Let's see an example:

System Message: WARNING/2 (<stdin>, line 86)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy


    class MySpider(scrapy.Spider):
        name = "example.com"
        allowed_domains = ["example.com"]
        start_urls = [
            "http://www.example.com/1.html",
            "http://www.example.com/2.html",
            "http://www.example.com/3.html",
        ]

        def parse(self, response):
            self.logger.info("A response from %s just arrived!", response.url)

Return multiple Requests and items from a single callback:

System Message: WARNING/2 (<stdin>, line 105)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy


    class MySpider(scrapy.Spider):
        name = "example.com"
        allowed_domains = ["example.com"]
        start_urls = [
            "http://www.example.com/1.html",
            "http://www.example.com/2.html",
            "http://www.example.com/3.html",
        ]

        def parse(self, response):
            for h3 in response.xpath("//h3").getall():
                yield {"title": h3}

            for href in response.xpath("//a/@href").getall():
                yield scrapy.Request(response.urljoin(href), self.parse)

Instead of :attr:`~.start_urls` you can use :meth:`~scrapy.Spider.start` directly; to give data more structure you can use :class:`~scrapy.Item` objects:

System Message: ERROR/3 (<stdin>, line 126); backlink

Unknown interpreted text role "attr".

System Message: ERROR/3 (<stdin>, line 126); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 126); backlink

Unknown interpreted text role "class".

System Message: WARNING/2 (<stdin>, line 131)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy
    from myproject.items import MyItem


    class MySpider(scrapy.Spider):
        name = "example.com"
        allowed_domains = ["example.com"]

        async def start(self):
            yield scrapy.Request("http://www.example.com/1.html", self.parse)
            yield scrapy.Request("http://www.example.com/2.html", self.parse)
            yield scrapy.Request("http://www.example.com/3.html", self.parse)

        def parse(self, response):
            for h3 in response.xpath("//h3").getall():
                yield MyItem(title=h3)

            for href in response.xpath("//a/@href").getall():
                yield scrapy.Request(response.urljoin(href), self.parse)

Spider arguments

Spiders can receive arguments that modify their behaviour. Some common uses for spider arguments are to define the start URLs or to restrict the crawl to certain sections of the site, but they can be used to configure any functionality of the spider.

Spider arguments are passed through the :command:`crawl` command using the -a option. For example:

System Message: ERROR/3 (<stdin>, line 163); backlink

Unknown interpreted text role "command".
scrapy crawl myspider -a category=electronics

Spiders can access arguments in their __init__ methods:

System Message: WARNING/2 (<stdin>, line 170)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy


    class MySpider(scrapy.Spider):
        name = "myspider"

        def __init__(self, category=None, *args, **kwargs):
            super().__init__(*args, **kwargs)
            self.start_urls = [f"http://www.example.com/categories/{category}"]
            # ...

The default __init__ method will take any spider arguments and copy them to the spider as attributes. The above example can also be written as follows:

System Message: WARNING/2 (<stdin>, line 187)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy


    class MySpider(scrapy.Spider):
        name = "myspider"

        async def start(self):
            yield scrapy.Request(f"http://www.example.com/categories/{self.category}")

If you are :ref:`running Scrapy from a script <run-from-script>`, you can specify spider arguments when calling :meth:`CrawlerProcess.crawl <scrapy.crawler.CrawlerProcess.crawl>` or :meth:`CrawlerRunner.crawl <scrapy.crawler.CrawlerRunner.crawl>`:

System Message: ERROR/3 (<stdin>, line 198); backlink

Unknown interpreted text role "ref".

System Message: ERROR/3 (<stdin>, line 198); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 198); backlink

Unknown interpreted text role "meth".

System Message: WARNING/2 (<stdin>, line 204)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    process = CrawlerProcess()
    process.crawl(MySpider, category="electronics")

Keep in mind that spider arguments are only strings. The spider will not do any parsing on its own. If you were to set the start_urls attribute from the command line, you would have to parse it on your own into a list using something like :func:`ast.literal_eval` or :func:`json.loads` and then set it as an attribute. Otherwise, you would cause iteration over a start_urls string (a very common python pitfall) resulting in each character being seen as a separate url.

System Message: ERROR/3 (<stdin>, line 209); backlink

Unknown interpreted text role "func".

System Message: ERROR/3 (<stdin>, line 209); backlink

Unknown interpreted text role "func".

Spider arguments can also be passed through the Scrapyd schedule.json API. See Scrapyd documentation.

scrapy-spider-metadata parameters

Another alternative to pass spider arguments is the library scrapy-spider-metadata.

This allows for Scrapy spiders to define, validate, document and pre-process their arguments as Pydantic models.

The example shows how to define typed parameters where a string argument is automatically converted to an integer:

System Message: WARNING/2 (<stdin>, line 235)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy
    from pydantic import BaseModel
    from scrapy_spider_metadata import Args


    class MyParams(BaseModel):
        pages: int


    class BookSpider(Args[MyParams], scrapy.Spider):
        name = "bookspider"
        start_urls = ["http://books.toscrape.com/catalogue"]

        async def start(self):
            for start_url in self.start_urls:
                for index in range(1, self.args.pages + 1):
                    yield scrapy.Request(f"{start_url}/page-{index}.html")

        def parse(self, response):
            book_links = response.css("article.product_pod h3 a::attr(href)").getall()
            for book_link in book_links:
                yield response.follow(book_link, self.parse_book)

        def parse_book(self, response):
            yield {
                "title": response.css("h1::text").get(),
                "price": response.css("p.price_color::text").get(),
            }

This spider can be called from the command line:

scrapy crawl bookspider -a pages=2

Start requests

Start requests are :class:`~scrapy.Request` objects yielded from the :meth:`~scrapy.Spider.start` method of a spider or from the :meth:`~scrapy.spidermiddlewares.SpiderMiddleware.process_start` method of a :ref:`spider middleware <topics-spider-middleware>`.

System Message: ERROR/3 (<stdin>, line 275); backlink

Unknown interpreted text role "class".

System Message: ERROR/3 (<stdin>, line 275); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 275); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 275); backlink

Unknown interpreted text role "ref".

System Message: ERROR/3 (<stdin>, line 280)

Unknown directive type "seealso".

.. seealso:: :ref:`start-request-order`

Delaying start request iteration

You can override the :meth:`~scrapy.Spider.start` method as follows to pause its iteration whenever there are scheduled requests:

System Message: ERROR/3 (<stdin>, line 287); backlink

Unknown interpreted text role "meth".

System Message: WARNING/2 (<stdin>, line 290)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    async def start(self):
        async for item_or_request in super().start():
            if self.crawler.engine.needs_backout():
                await self.crawler.signals.wait_for(signals.scheduler_empty)
            yield item_or_request

This can help minimize the number of requests in the scheduler at any given time, to minimize resource usage (memory or disk, depending on :setting:`JOBDIR`).

System Message: ERROR/3 (<stdin>, line 298); backlink

Unknown interpreted text role "setting".

Generic Spiders

Scrapy comes with some useful generic spiders that you can use to subclass your spiders from. Their aim is to provide convenient functionality for a few common scraping cases, like following all links on a site based on certain rules, crawling from Sitemaps, or parsing an XML/CSV feed.

For the examples used in the following spiders, we'll assume you have a project with a TestItem declared in a myproject.items module:

System Message: WARNING/2 (<stdin>, line 315)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from dataclasses import dataclass


    @dataclass
    class TestItem:
        id: str | None = None
        name: str | None = None
        description: str | None = None


System Message: ERROR/3 (<stdin>, line 327)

Unknown directive type "currentmodule".

.. currentmodule:: scrapy.spiders

CrawlSpider

System Message: ERROR/3 (<stdin>, line 332)

Unknown directive type "autoclass".

.. autoclass:: CrawlSpider

    .. autoattribute:: rules

    .. automethod:: parse_start_url

Crawling rules

System Message: ERROR/3 (<stdin>, line 341)

Unknown directive type "autoclass".

.. autoclass:: Rule

CrawlSpider example

Let's now take a look at an example CrawlSpider with rules:

System Message: WARNING/2 (<stdin>, line 348)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import CrawlSpider, Rule
    from scrapy.linkextractors import LinkExtractor


    class MySpider(CrawlSpider):
        name = "example.com"
        allowed_domains = ["example.com"]
        start_urls = ["http://www.example.com"]

        rules = (
            # Extract links matching 'category.php' (but not matching 'subsection.php')
            # and follow links from them (since no callback means follow=True by default).
            Rule(LinkExtractor(allow=(r"category\.php",), deny=(r"subsection\.php",))),
            # Extract links matching 'item.php' and parse them with the spider's method parse_item
            Rule(LinkExtractor(allow=(r"item\.php",)), callback="parse_item"),
        )

        def parse_item(self, response):
            self.logger.info("Hi, this is an item page! %s", response.url)
            item = {}
            item["id"] = response.xpath('//td[@id="item_id"]/text()').re(r"ID: (\d+)")
            item["name"] = response.xpath('//td[@id="item_name"]/text()').get()
            item["description"] = response.xpath(
                '//td[@id="item_description"]/text()'
            ).get()
            item["link_text"] = response.meta["link_text"]
            url = response.xpath('//td[@id="additional_data"]/@href').get()
            return response.follow(
                url, self.parse_additional_page, cb_kwargs=dict(item=item)
            )

        def parse_additional_page(self, response, item):
            item["additional_data"] = response.xpath(
                '//p[@id="additional_data"]/text()'
            ).get()
            return item


This spider would start crawling example.com's home page, collecting category links, and item links, parsing the latter with the parse_item method. For each item response, some data will be extracted from the HTML using XPath, and a dictionary will be filled with it.

XMLFeedSpider

System Message: ERROR/3 (<stdin>, line 396)

Unknown directive type "autoclass".

.. autoclass:: XMLFeedSpider

    .. autoattribute:: iterator

    .. autoattribute:: itertag

    .. autoattribute:: namespaces

    .. automethod:: adapt_response

    .. automethod:: parse_node

    .. automethod:: process_results


XMLFeedSpider example

These spiders are pretty easy to use, let's have a look at one example:

System Message: WARNING/2 (<stdin>, line 417)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import XMLFeedSpider
    from myproject.items import TestItem


    class MySpider(XMLFeedSpider):
        name = "example.com"
        allowed_domains = ["example.com"]
        start_urls = ["http://www.example.com/feed.xml"]
        iterator = "iternodes"  # This is actually unnecessary, since it's the default value
        itertag = "item"

        def parse_node(self, response, node):
            self.logger.info(
                "Hi, this is a <%s> node!: %s", self.itertag, "".join(node.getall())
            )

            item = TestItem()
            item.id = node.xpath("@id").get()
            item.name = node.xpath("name").get()
            item.description = node.xpath("description").get()
            return item

Basically what we did up there was to create a spider that downloads a feed from the given start_urls, and then iterates through each of its item tags, prints them out, and stores some random data in an :class:`~scrapy.Item`.

System Message: ERROR/3 (<stdin>, line 441); backlink

Unknown interpreted text role "class".

CSVFeedSpider

System Message: ERROR/3 (<stdin>, line 448)

Unknown directive type "autoclass".

.. autoclass:: CSVFeedSpider

    .. autoattribute:: delimiter

    .. autoattribute:: quotechar

    .. autoattribute:: headers

    .. automethod:: adapt_response

    .. automethod:: parse_row

    .. automethod:: process_results

CSVFeedSpider example

Let's see an example similar to the previous one, but using a :class:`CSVFeedSpider`:

System Message: ERROR/3 (<stdin>, line 465); backlink

Unknown interpreted text role "class".

System Message: WARNING/2 (<stdin>, line 469)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import CSVFeedSpider
    from myproject.items import TestItem


    class MySpider(CSVFeedSpider):
        name = "example.com"
        allowed_domains = ["example.com"]
        start_urls = ["http://www.example.com/feed.csv"]
        delimiter = ";"
        quotechar = "'"
        headers = ["id", "name", "description"]

        def parse_row(self, response, row):
            self.logger.info("Hi, this is a row!: %r", row)

            item = TestItem()
            item.id = row["id"]
            item.name = row["name"]
            item.description = row["description"]
            return item


SitemapSpider

System Message: ERROR/3 (<stdin>, line 496)

Unknown directive type "autoclass".

.. autoclass:: SitemapSpider

    .. autoattribute:: sitemap_urls

    .. autoattribute:: sitemap_rules

    .. autoattribute:: sitemap_follow

    .. autoattribute:: sitemap_alternate_links

    .. automethod:: sitemap_filter


SitemapSpider examples

Simplest example: process all urls discovered through sitemaps using the parse callback:

System Message: WARNING/2 (<stdin>, line 515)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import SitemapSpider


    class MySpider(SitemapSpider):
        sitemap_urls = ["http://www.example.com/sitemap.xml"]

        def parse(self, response):
            pass  # ... scrape item here ...

Process some urls with certain callback and other urls with a different callback:

System Message: WARNING/2 (<stdin>, line 529)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import SitemapSpider


    class MySpider(SitemapSpider):
        sitemap_urls = ["http://www.example.com/sitemap.xml"]
        sitemap_rules = [
            ("/product/", "parse_product"),
            ("/category/", "parse_category"),
        ]

        def parse_product(self, response):
            pass  # ... scrape product ...

        def parse_category(self, response):
            pass  # ... scrape category ...

Follow sitemaps defined in the robots.txt file and only follow sitemaps whose url contains /sitemap_shop:

System Message: WARNING/2 (<stdin>, line 550)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy.spiders import SitemapSpider


    class MySpider(SitemapSpider):
        sitemap_urls = ["http://www.example.com/robots.txt"]
        sitemap_rules = [
            ("/shop/", "parse_shop"),
        ]
        sitemap_follow = ["/sitemap_shops"]

        def parse_shop(self, response):
            pass  # ... scrape shop here ...

Combine SitemapSpider with other sources of urls:

System Message: WARNING/2 (<stdin>, line 567)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from scrapy import Request
    from scrapy.spiders import SitemapSpider


    class MySpider(SitemapSpider):
        sitemap_urls = ["http://www.example.com/robots.txt"]
        sitemap_rules = [
            ("/shop/", "parse_shop"),
        ]

        other_urls = ["http://www.example.com/about"]

        async def start(self):
            async for item_or_request in super().start():
                yield item_or_request
            for url in self.other_urls:
                yield Request(url, self.parse_other)

        def parse_shop(self, response):
            pass  # ... scrape shop here ...

        def parse_other(self, response):
            pass  # ... scrape other here ...

</html>