scrapy/docs/intro/tutorial.rst

30 KiB

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>

Scrapy Tutorial

In this tutorial, we'll assume that Scrapy is already installed on your system. If that's not the case, see :ref:`intro-install`.

System Message: ERROR/3 (<stdin>, line 7); backlink

Unknown interpreted text role "ref".

We are going to scrape quotes.toscrape.com, a website that lists quotes from famous authors.

This tutorial will walk you through these tasks:

  1. Creating a new Scrapy project

  2. Writing a :ref:`spider <topics-spiders>` to crawl a site and extract data

    System Message: ERROR/3 (<stdin>, line 16); backlink

    Unknown interpreted text role "ref".

  3. Exporting the scraped data using the command line

  4. Changing spider to recursively follow links

  5. Using spider arguments

Scrapy is written in Python. The more you learn about Python, the more you can get out of Scrapy.

If you're already familiar with other languages and want to learn Python quickly, the Python Tutorial is a good resource.

If you're new to programming and want to start with Python, the following books may be useful to you:

You can also take a look at this list of Python resources for non-programmers, as well as the suggested resources in the learnpython-subreddit.

Creating a project

Before you start scraping, you will have to set up a new Scrapy project. Enter a directory where you'd like to store your code and run:

scrapy startproject tutorial

This will create a tutorial directory with the following contents:

tutorial/
    scrapy.cfg            # deploy configuration file

    tutorial/             # project's Python module, you'll import your code from here
        __init__.py

        items.py          # project items definition file

        middlewares.py    # project middlewares file

        pipelines.py      # project pipelines file

        settings.py       # project settings file

        spiders/          # a directory where you'll later put your spiders
            __init__.py

Our first Spider

Spiders are classes that you define and that Scrapy uses to scrape information from a website (or a group of websites). They must subclass :class:`~scrapy.Spider` and define the initial requests to be made, and optionally, how to follow links in pages and parse the downloaded page content to extract data.

System Message: ERROR/3 (<stdin>, line 79); backlink

Unknown interpreted text role "class".

This is the code for our first Spider. Save it in a file named quotes_spider.py under the tutorial/spiders directory in your project:

System Message: WARNING/2 (<stdin>, line 87)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from pathlib import Path

    import scrapy


    class QuotesSpider(scrapy.Spider):
        name = "quotes"

        def start_requests(self):
            urls = [
                "https://quotes.toscrape.com/page/1/",
                "https://quotes.toscrape.com/page/2/",
            ]
            for url in urls:
                yield scrapy.Request(url=url, callback=self.parse)

        def parse(self, response):
            page = response.url.split("/")[-2]
            filename = f"quotes-{page}.html"
            Path(filename).write_bytes(response.body)
            self.log(f"Saved file {filename}")


As you can see, our Spider subclasses :class:`scrapy.Spider <scrapy.Spider>` and defines some attributes and methods:

System Message: ERROR/3 (<stdin>, line 112); backlink

Unknown interpreted text role "class".
  • :attr:`~scrapy.Spider.name`: identifies the Spider. It must be unique within a project, that is, you can't set the same name for different Spiders.

    System Message: ERROR/3 (<stdin>, line 115); backlink

    Unknown interpreted text role "attr".

  • :meth:`~scrapy.Spider.start_requests`: must return an iterable of Requests (you can return a list of requests or write a generator function) which the Spider will begin to crawl from. Subsequent requests will be generated successively from these initial requests.

    System Message: ERROR/3 (<stdin>, line 119); backlink

    Unknown interpreted text role "meth".

  • :meth:`~scrapy.Spider.parse`: a method that will be called to handle the response downloaded for each of the requests made. The response parameter is an instance of :class:`~scrapy.http.TextResponse` that holds the page content and has further helpful methods to handle it.

    System Message: ERROR/3 (<stdin>, line 124); backlink

    Unknown interpreted text role "meth".

    System Message: ERROR/3 (<stdin>, line 124); backlink

    Unknown interpreted text role "class".

    The :meth:`~scrapy.Spider.parse` method usually parses the response, extracting the scraped data as dicts and also finding new URLs to follow and creating new requests (:class:`~scrapy.Request`) from them.

    System Message: ERROR/3 (<stdin>, line 129); backlink

    Unknown interpreted text role "meth".

    System Message: ERROR/3 (<stdin>, line 129); backlink

    Unknown interpreted text role "class".

How to run our spider

To put our spider to work, go to the project's top level directory and run:

scrapy crawl quotes

This command runs the spider named quotes that we've just added, that will send some requests for the quotes.toscrape.com domain. You will get an output similar to this:

... (omitted for brevity)
2016-12-16 21:24:05 [scrapy.core.engine] INFO: Spider opened
2016-12-16 21:24:05 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
2016-12-16 21:24:05 [scrapy.extensions.telnet] DEBUG: Telnet console listening on 127.0.0.1:6023
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (404) <GET https://quotes.toscrape.com/robots.txt> (referer: None)
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/1/> (referer: None)
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/2/> (referer: None)
2016-12-16 21:24:05 [quotes] DEBUG: Saved file quotes-1.html
2016-12-16 21:24:05 [quotes] DEBUG: Saved file quotes-2.html
2016-12-16 21:24:05 [scrapy.core.engine] INFO: Closing spider (finished)
...

Now, check the files in the current directory. You should notice that two new files have been created: quotes-1.html and quotes-2.html, with the content for the respective URLs, as our parse method instructs.

Note

If you are wondering why we haven't parsed the HTML yet, hold on, we will cover that soon.

What just happened under the hood?

Scrapy schedules the :class:`scrapy.Request <scrapy.Request>` objects returned by the start_requests method of the Spider. Upon receiving a response for each one, it instantiates :class:`~scrapy.http.Response` objects and calls the callback method associated with the request (in this case, the parse method) passing the response as an argument.

System Message: ERROR/3 (<stdin>, line 167); backlink

Unknown interpreted text role "class".

System Message: ERROR/3 (<stdin>, line 167); backlink

Unknown interpreted text role "class".

A shortcut to the start_requests method

Instead of implementing a :meth:`~scrapy.Spider.start_requests` method that generates :class:`scrapy.Request <scrapy.Request>` objects from URLs, you can just define a :attr:`~scrapy.Spider.start_urls` class attribute with a list of URLs. This list will then be used by the default implementation of :meth:`~scrapy.Spider.start_requests` to create the initial requests for your spider.

System Message: ERROR/3 (<stdin>, line 176); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 176); backlink

Unknown interpreted text role "class".

System Message: ERROR/3 (<stdin>, line 176); backlink

Unknown interpreted text role "attr".

System Message: ERROR/3 (<stdin>, line 176); backlink

Unknown interpreted text role "meth".

System Message: WARNING/2 (<stdin>, line 183)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    from pathlib import Path

    import scrapy


    class QuotesSpider(scrapy.Spider):
        name = "quotes"
        start_urls = [
            "https://quotes.toscrape.com/page/1/",
            "https://quotes.toscrape.com/page/2/",
        ]

        def parse(self, response):
            page = response.url.split("/")[-2]
            filename = f"quotes-{page}.html"
            Path(filename).write_bytes(response.body)

The :meth:`~scrapy.Spider.parse` method will be called to handle each of the requests for those URLs, even though we haven't explicitly told Scrapy to do so. This happens because :meth:`~scrapy.Spider.parse` is Scrapy's default callback method, which is called for requests without an explicitly assigned callback.

System Message: ERROR/3 (<stdin>, line 202); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 202); backlink

Unknown interpreted text role "meth".

Extracting data

The best way to learn how to extract data with Scrapy is trying selectors using the :ref:`Scrapy shell <topics-shell>`. Run:

System Message: ERROR/3 (<stdin>, line 212); backlink

Unknown interpreted text role "ref".
scrapy shell 'https://quotes.toscrape.com/page/1/'

Note

Remember to always enclose URLs in quotes when running Scrapy shell from the command line, otherwise URLs containing arguments (i.e. & character) will not work.

On Windows, use double quotes instead:

scrapy shell "https://quotes.toscrape.com/page/1/"

You will see something like:

[ ... Scrapy log here ... ]
2016-09-19 12:09:27 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/1/> (referer: None)
[s] Available Scrapy objects:
[s]   scrapy     scrapy module (contains scrapy.Request, scrapy.Selector, etc)
[s]   crawler    <scrapy.crawler.Crawler object at 0x7fa91d888c90>
[s]   item       {}
[s]   request    <GET https://quotes.toscrape.com/page/1/>
[s]   response   <200 https://quotes.toscrape.com/page/1/>
[s]   settings   <scrapy.settings.Settings object at 0x7fa91d888c10>
[s]   spider     <DefaultSpider 'default' at 0x7fa91c8af990>
[s] Useful shortcuts:
[s]   shelp()           Shell help (print this help)
[s]   fetch(req_or_url) Fetch request (or URL) and update local objects
[s]   view(response)    View response in a browser

Using the shell, you can try selecting elements using CSS with the response object:

System Message: WARNING/2 (<stdin>, line 251)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> response.css("title")
    [<Selector query='descendant-or-self::title' data='<title>Quotes to Scrape</title>'>]

The result of running response.css('title') is a list-like object called :class:`~scrapy.selector.SelectorList`, which represents a list of :class:`~scrapy.Selector` objects that wrap around XML/HTML elements and allow you to run further queries to refine the selection or extract the data.

System Message: ERROR/3 (<stdin>, line 256); backlink

Unknown interpreted text role "class".

System Message: ERROR/3 (<stdin>, line 256); backlink

Unknown interpreted text role "class".

To extract the text from the title above, you can do:

System Message: WARNING/2 (<stdin>, line 264)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> response.css("title::text").getall()
    ['Quotes to Scrape']

There are two things to note here: one is that we've added ::text to the CSS query, to mean we want to select only the text elements directly inside <title> element. If we don't specify ::text, we'd get the full title element, including its tags:

System Message: WARNING/2 (<stdin>, line 274)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> response.css("title").getall()
    ['<title>Quotes to Scrape</title>']

The other thing is that the result of calling .getall() is a list: it is possible that a selector returns more than one result, so we extract them all. When you know you just want the first result, as in this case, you can do:

System Message: WARNING/2 (<stdin>, line 283)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> response.css("title::text").get()
    'Quotes to Scrape'

As an alternative, you could've written:

System Message: WARNING/2 (<stdin>, line 290)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> response.css("title::text")[0].get()
    'Quotes to Scrape'

Accessing an index on a :class:`~scrapy.selector.SelectorList` instance will raise an :exc:`IndexError` exception if there are no results:

System Message: ERROR/3 (<stdin>, line 295); backlink

Unknown interpreted text role "class".

System Message: ERROR/3 (<stdin>, line 295); backlink

Unknown interpreted text role "exc".

System Message: WARNING/2 (<stdin>, line 298)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> response.css("noelement")[0].get()
    Traceback (most recent call last):
    ...
    IndexError: list index out of range

You might want to use .get() directly on the :class:`~scrapy.selector.SelectorList` instance instead, which returns None if there are no results:

System Message: ERROR/3 (<stdin>, line 305); backlink

Unknown interpreted text role "class".

System Message: WARNING/2 (<stdin>, line 309)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> response.css("noelement").get()

There's a lesson here: for most scraping code, you want it to be resilient to errors due to things not being found on a page, so that even if some parts fail to be scraped, you can at least get some data.

Besides the :meth:`~scrapy.selector.SelectorList.getall` and :meth:`~scrapy.selector.SelectorList.get` methods, you can also use the :meth:`~scrapy.selector.SelectorList.re` method to extract using :doc:`regular expressions <library/re>`:

System Message: ERROR/3 (<stdin>, line 317); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 317); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 317); backlink

Unknown interpreted text role "meth".

System Message: ERROR/3 (<stdin>, line 317); backlink

Unknown interpreted text role "doc".

System Message: WARNING/2 (<stdin>, line 322)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> response.css("title::text").re(r"Quotes.*")
    ['Quotes to Scrape']
    >>> response.css("title::text").re(r"Q\w+")
    ['Quotes']
    >>> response.css("title::text").re(r"(\w+) to (\w+)")
    ['Quotes', 'Scrape']

In order to find the proper CSS selectors to use, you might find it useful to open the response page from the shell in your web browser using view(response). You can use your browser's developer tools to inspect the HTML and come up with a selector (see :ref:`topics-developer-tools`).

System Message: ERROR/3 (<stdin>, line 331); backlink

Unknown interpreted text role "ref".

Selector Gadget is also a nice tool to quickly find CSS selector for visually selected elements, which works in many browsers.

XPath: a brief intro

Besides CSS, Scrapy selectors also support using XPath expressions:

System Message: WARNING/2 (<stdin>, line 347)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> response.xpath("//title")
    [<Selector query='//title' data='<title>Quotes to Scrape</title>'>]
    >>> response.xpath("//title/text()").get()
    'Quotes to Scrape'

XPath expressions are very powerful, and are the foundation of Scrapy Selectors. In fact, CSS selectors are converted to XPath under-the-hood. You can see that if you read the text representation of the selector objects in the shell closely.

While perhaps not as popular as CSS selectors, XPath expressions offer more power because besides navigating the structure, it can also look at the content. Using XPath, you're able to select things like: the link that contains the text "Next Page". This makes XPath very fitting to the task of scraping, and we encourage you to learn XPath even if you already know how to construct CSS selectors, it will make scraping much easier.

We won't cover much of XPath here, but you can read more about :ref:`using XPath with Scrapy Selectors here <topics-selectors>`. To learn more about XPath, we recommend this tutorial to learn XPath through examples, and this tutorial to learn "how to think in XPath".

System Message: ERROR/3 (<stdin>, line 366); backlink

Unknown interpreted text role "ref".

Extracting quotes and authors

Now that you know a bit about selection and extraction, let's complete our spider by writing the code to extract the quotes from the web page.

Each quote in https://quotes.toscrape.com is represented by HTML elements that look like this:

System Message: WARNING/2 (<stdin>, line 384)

Cannot analyze code. Pygments package not found.

.. code-block:: html

    <div class="quote">
        <span class="text">“The world as we have created it is a process of our
        thinking. It cannot be changed without changing our thinking.”</span>
        <span>
            by <small class="author">Albert Einstein</small>
            <a href="/author/Albert-Einstein">(about)</a>
        </span>
        <div class="tags">
            Tags:
            <a class="tag" href="/tag/change/page/1/">change</a>
            <a class="tag" href="/tag/deep-thoughts/page/1/">deep-thoughts</a>
            <a class="tag" href="/tag/thinking/page/1/">thinking</a>
            <a class="tag" href="/tag/world/page/1/">world</a>
        </div>
    </div>

Let's open up scrapy shell and play a bit to find out how to extract the data we want:

scrapy shell 'https://quotes.toscrape.com'

We get a list of selectors for the quote HTML elements with:

System Message: WARNING/2 (<stdin>, line 409)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> response.css("div.quote")
    [<Selector query="descendant-or-self::div[@class and contains(concat(' ', normalize-space(@class), ' '), ' quote ')]" data='<div class="quote" itemscope itemtype...'>,
    <Selector query="descendant-or-self::div[@class and contains(concat(' ', normalize-space(@class), ' '), ' quote ')]" data='<div class="quote" itemscope itemtype...'>,
    ...]

Each of the selectors returned by the query above allows us to run further queries over their sub-elements. Let's assign the first selector to a variable, so that we can run our CSS selectors directly on a particular quote:

System Message: WARNING/2 (<stdin>, line 420)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> quote = response.css("div.quote")[0]

Now, let's extract the text, author and tags from that quote using the quote object we just created:

System Message: WARNING/2 (<stdin>, line 427)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> text = quote.css("span.text::text").get()
    >>> text
    '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”'
    >>> author = quote.css("small.author::text").get()
    >>> author
    'Albert Einstein'

Given that the tags are a list of strings, we can use the .getall() method to get all of them:

System Message: WARNING/2 (<stdin>, line 439)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> tags = quote.css("div.tags a.tag::text").getall()
    >>> tags
    ['change', 'deep-thoughts', 'thinking', 'world']

Having figured out how to extract each bit, we can now iterate over all the quote elements and put them together into a Python dictionary:

System Message: WARNING/2 (<stdin>, line 452)

Cannot analyze code. Pygments package not found.

.. code-block:: pycon

    >>> for quote in response.css("div.quote"):
    ...     text = quote.css("span.text::text").get()
    ...     author = quote.css("small.author::text").get()
    ...     tags = quote.css("div.tags a.tag::text").getall()
    ...     print(dict(text=text, author=author, tags=tags))
    ...
    {'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein', 'tags': ['change', 'deep-thoughts', 'thinking', 'world']}
    {'text': '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', 'author': 'J.K. Rowling', 'tags': ['abilities', 'choices']}
    ...

Extracting data in our spider

Let's get back to our spider. Until now, it hasn't extracted any data in particular, just saving the whole HTML page to a local file. Let's integrate the extraction logic above into our spider.

A Scrapy spider typically generates many dictionaries containing the data extracted from the page. To do that, we use the yield Python keyword in the callback, as you can see below:

System Message: WARNING/2 (<stdin>, line 475)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy


    class QuotesSpider(scrapy.Spider):
        name = "quotes"
        start_urls = [
            "https://quotes.toscrape.com/page/1/",
            "https://quotes.toscrape.com/page/2/",
        ]

        def parse(self, response):
            for quote in response.css("div.quote"):
                yield {
                    "text": quote.css("span.text::text").get(),
                    "author": quote.css("small.author::text").get(),
                    "tags": quote.css("div.tags a.tag::text").getall(),
                }

To run this spider, exit the scrapy shell by entering:

quit()

Then, run:

scrapy crawl quotes

Now, it should output the extracted data with the log:

2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 https://quotes.toscrape.com/page/1/>
{'tags': ['life', 'love'], 'author': 'André Gide', 'text': '“It is better to be hated for what you are than to be loved for what you are not.”'}
2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 https://quotes.toscrape.com/page/1/>
{'tags': ['edison', 'failure', 'inspirational', 'paraphrased'], 'author': 'Thomas A. Edison', 'text': "“I have not failed. I've just found 10,000 ways that won't work.”"}

Storing the scraped data

The simplest way to store the scraped data is by using :ref:`Feed exports <topics-feed-exports>`, with the following command:

System Message: ERROR/3 (<stdin>, line 516); backlink

Unknown interpreted text role "ref".
scrapy crawl quotes -O quotes.json

That will generate a quotes.json file containing all scraped items, serialized in JSON.

The -O command-line switch overwrites any existing file; use -o instead to append new content to any existing file. However, appending to a JSON file makes the file contents invalid JSON. When appending to a file, consider using a different serialization format, such as JSON Lines:

scrapy crawl quotes -o quotes.jsonl

The JSON Lines format is useful because it's stream-like, so you can easily append new records to it. It doesn't have the same problem as JSON when you run twice. Also, as each record is a separate line, you can process big files without having to fit everything in memory, there are tools like JQ to help do that at the command-line.

In small projects (like the one in this tutorial), that should be enough. However, if you want to perform more complex things with the scraped items, you can write an :ref:`Item Pipeline <topics-item-pipeline>`. A placeholder file for Item Pipelines has been set up for you when the project is created, in tutorial/pipelines.py. Though you don't need to implement any item pipelines if you just want to store the scraped items.

System Message: ERROR/3 (<stdin>, line 537); backlink

Unknown interpreted text role "ref".

Using spider arguments

You can provide command line arguments to your spiders by using the -a option when running them:

scrapy crawl quotes -O quotes-humor.json -a tag=humor

These arguments are passed to the Spider's __init__ method and become spider attributes by default.

In this example, the value provided for the tag argument will be available via self.tag. You can use this to make your spider fetch only quotes with a specific tag, building the URL based on the argument:

System Message: WARNING/2 (<stdin>, line 789)

Cannot analyze code. Pygments package not found.

.. code-block:: python

    import scrapy


    class QuotesSpider(scrapy.Spider):
        name = "quotes"

        def start_requests(self):
            url = "https://quotes.toscrape.com/"
            tag = getattr(self, "tag", None)
            if tag is not None:
                url = url + "tag/" + tag
            yield scrapy.Request(url, self.parse)

        def parse(self, response):
            for quote in response.css("div.quote"):
                yield {
                    "text": quote.css("span.text::text").get(),
                    "author": quote.css("small.author::text").get(),
                }

            next_page = response.css("li.next a::attr(href)").get()
            if next_page is not None:
                yield response.follow(next_page, self.parse)


If you pass the tag=humor argument to this spider, you'll notice that it will only visit URLs from the humor tag, such as https://quotes.toscrape.com/tag/humor.

You can :ref:`learn more about handling spider arguments here <spiderargs>`.

System Message: ERROR/3 (<stdin>, line 820); backlink

Unknown interpreted text role "ref".

Next steps

This tutorial covered only the basics of Scrapy, but there's a lot of other features not mentioned here. Check the :ref:`topics-whatelse` section in the :ref:`intro-overview` chapter for a quick overview of the most important ones.

System Message: ERROR/3 (<stdin>, line 825); backlink

Unknown interpreted text role "ref".

System Message: ERROR/3 (<stdin>, line 825); backlink

Unknown interpreted text role "ref".

You can continue from the section :ref:`section-basics` to know more about the command-line tool, spiders, selectors and other things the tutorial hasn't covered like modeling the scraped data. If you'd prefer to play with an example project, check the :ref:`intro-examples` section.

System Message: ERROR/3 (<stdin>, line 829); backlink

Unknown interpreted text role "ref".

System Message: ERROR/3 (<stdin>, line 829); backlink

Unknown interpreted text role "ref".
</html>