23 KiB
Scrapy Tutorial
In this tutorial, we'll assume that Scrapy is already installed on your system. If that's not the case, see :ref:`intro-install`.
System Message: ERROR/3 (<stdin>, line 7); backlink
Unknown interpreted text role "ref".We are going to scrape quotes.toscrape.com, a website that lists quotes from famous authors.
This tutorial will walk you through these tasks:
Creating a new Scrapy project
Writing a :ref:`spider <topics-spiders>` to crawl a site and extract data
System Message: ERROR/3 (<stdin>, line 16); backlink
Unknown interpreted text role "ref".
Exporting the scraped data using command line
Scrapy is written in Python. If you're new to the language you might want to start by getting an idea of what the language is like, to get the most out of Scrapy. If you're already familiar with other languages, and want to learn Python quickly, we recommend Learn Python The Hard Way. If you're new to programming and want to start with Python, take a look at this list of Python resources for non-programmers.
Creating a project
Before you start scraping, you will have to set up a new Scrapy project. Enter a directory where you'd like to store your code and run:
scrapy startproject tutorial
This will create a tutorial directory with the following contents:
tutorial/
scrapy.cfg # deploy configuration file
tutorial/ # project's Python module, you'll import your code from here
__init__.py
items.py # project items file
pipelines.py # project pipelines file
settings.py # project settings file
spiders/ # a directory where you'll later put your spiders
__init__.py
Our first Spider
Spiders are classes that you define and that Scrapy uses to scrape information from a website (or group of websites). They must subclass :class:`scrapy.Spider` and define the initial requests to make, how to follow links in the pages, and how to parse the downloaded page content to extract data.
System Message: ERROR/3 (<stdin>, line 59); backlink
Unknown interpreted text role "class".This is the code for our first Spider. Save it in a file named quotes_spider.py under the tutorial/spiders directory in your project:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
def start_requests(self):
urls = [
'http://quotes.toscrape.com/page/1/',
'http://quotes.toscrape.com/page/2/',
]
for url in urls:
yield scrapy.Request(url=url, callback=self.parse)
def parse(self, response):
page = response.url.split("/")[-2]
filename = 'quotes-%s.html' % page
with open(filename, 'wb') as f:
f.write(response.body)
As you can see, our Spider subclasses :class:`scrapy.Spider <scrapy.spiders.Spider>` and defines some attributes and methods:
System Message: ERROR/3 (<stdin>, line 89); backlink
Unknown interpreted text role "class".:attr:`~scrapy.spiders.Spider.name`: identifies the Spider. It must be unique within a project, that is, you can't set the same name for different Spiders.
System Message: ERROR/3 (<stdin>, line 92); backlink
Unknown interpreted text role "attr".
:meth:`~scrapy.spiders.Spider.start_requests`: must return a list of requests where the Spider will begin to crawl from. Subsequent requests will be generated successively from these initial requests.
System Message: ERROR/3 (<stdin>, line 96); backlink
Unknown interpreted text role "meth".
:meth:`~scrapy.spiders.Spider.parse`: a method that will be called to handle the response downloaded for each of the requests made. The response parameter is an instance of :class:`~scrapy.http.Response` that holds the page content and has further helpful methods to handle it.
System Message: ERROR/3 (<stdin>, line 100); backlink
Unknown interpreted text role "meth".
System Message: ERROR/3 (<stdin>, line 100); backlink
Unknown interpreted text role "class".
The :meth:`~scrapy.spiders.Spider.parse` method usually parses the response, extracting the scraped data as dicts and also finding new URLs to follow and creating new requests (:class:`~scrapy.http.Request`) from them.
System Message: ERROR/3 (<stdin>, line 105); backlink
Unknown interpreted text role "meth".
System Message: ERROR/3 (<stdin>, line 105); backlink
Unknown interpreted text role "class".
How to run our spider
To put our spider to work, go to the project's top level directory and run:
scrapy crawl quotes
This command runs the spider with name quotes that we've just added, that will send some requests for the quotes.toscrape.com domain. You will get an output similar to this:
2016-09-01 16:51:27 [scrapy] INFO: Scrapy started (bot: tutorial)
2016-09-01 16:51:27 [scrapy] INFO: Overridden settings: {...}
2016-09-01 16:51:27 [scrapy] INFO: Enabled extensions: ...
2016-09-01 16:51:27 [scrapy] INFO: Enabled downloader middlewares: ...
2016-09-01 16:51:27 [scrapy] INFO: Enabled spider middlewares: ...
2016-09-01 16:51:27 [scrapy] INFO: Enabled item pipelines: ...
2016-09-01 16:51:27 [scrapy] INFO: Spider opened
2016-09-01 16:51:27 [scrapy] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
2016-09-01 16:51:28 [scrapy] DEBUG: Crawled (404) <GET http://quotes.toscrape.com/robots.txt> (referer: None)
2016-09-01 16:51:28 [scrapy] DEBUG: Crawled (200) <GET http://quotes.toscrape.com/page/1/> (referer: None)
2016-09-01 16:51:29 [scrapy] DEBUG: Crawled (200) <GET http://quotes.toscrape.com/page/2/> (referer: None)
2016-09-01 16:51:29 [scrapy] INFO: Closing spider (finished)
Now, check the files in the current directory. You should notice that two new files have been created: quotes-1.html and quotes-2.html, with the content for the respective URLs, as our parse method instructs.
Note
If you are wondering why we haven't parsed the HTML yet, hold on, we will cover that soon.
What just happened under the hood?
Scrapy schedules the :class:`scrapy.Request <scrapy.http.Request>` objects returned by the start_requests method of the Spider. Upon receiving a response for each one, it instantiates :class:`scrapy.http.Response` objects and calls the parse callback method passing the response as argument.
System Message: ERROR/3 (<stdin>, line 144); backlink
Unknown interpreted text role "class".System Message: ERROR/3 (<stdin>, line 144); backlink
Unknown interpreted text role "class".A shortcut to the start_requests method
Instead of implementing a :meth:`~scrapy.spiders.Spider.start_requests` method that generates :class:`scrapy.Request <scrapy.http.Request>` objects from URLs, you can just define a :attr:`~scrapy.spiders.Spider.start_urls` class attribute with a list of URLs. This list will then be used by the default implementation of :meth:`~scrapy.spiders.Spider.start_requests` to create the initial requests for your spider:
System Message: ERROR/3 (<stdin>, line 153); backlink
Unknown interpreted text role "meth".System Message: ERROR/3 (<stdin>, line 153); backlink
Unknown interpreted text role "class".System Message: ERROR/3 (<stdin>, line 153); backlink
Unknown interpreted text role "attr".System Message: ERROR/3 (<stdin>, line 153); backlink
Unknown interpreted text role "meth".import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
'http://quotes.toscrape.com/page/1/',
'http://quotes.toscrape.com/page/2/',
]
def parse(self, response):
page = response.url.split("/")[-2]
filename = 'quotes-%s.html' % page
with open(filename, 'wb') as f:
f.write(response.body)
The :meth:`~scrapy.spiders.Spider.parse` method will be called to handle each of the requests for those URLs, even though we haven't explicitely told Scrapy to do so. This happens because :meth:`~scrapy.spiders.Spider.parse` is Scrapy's default callback method that is called for any request that have been generated with no callback explicitely assigned to handle it.
System Message: ERROR/3 (<stdin>, line 176); backlink
Unknown interpreted text role "meth".System Message: ERROR/3 (<stdin>, line 176); backlink
Unknown interpreted text role "meth".Extracting data
The best way to learn how to extract data with Scrapy is trying selectors using the shell :ref:`Scrapy shell <topics-shell>`. Run:
System Message: ERROR/3 (<stdin>, line 186); backlink
Unknown interpreted text role "ref".scrapy crawl http://quotes.toscrape.com/page/1/
You will see something like:
[ ... Scrapy log here ... ]
2016-09-19 12:09:27 [scrapy] DEBUG: Crawled (200) <GET http://quotes.toscrape.com/page/1/> (referer: None)
[s] Available Scrapy objects:
[s] crawler <scrapy.crawler.Crawler object at 0x7fa91d888c90>
[s] item {}
[s] request <GET http://quotes.toscrape.com/page/1/>
[s] response <200 http://quotes.toscrape.com/page/1/>
[s] settings <scrapy.settings.Settings object at 0x7fa91d888c10>
[s] spider <DefaultSpider 'default' at 0x7fa91c8af990>
[s] Useful shortcuts:
[s] shelp() Shell help (print this help)
[s] fetch(req_or_url) Fetch request (or URL) and update local objects
[s] view(response) View response in a browser
>>>
Using the shell, you can try selecting elements using CSS with the response object:
>>> response.css('title')
[<Selector xpath=u'descendant-or-self::title' data=u'<title>Quotes to Scrape</title>'>]
The result of running response.css('title') is a list-like object called :class:`~scrapy.selector.SelectorList`, which represents a list of :class:`~scrapy.selector.Selector` objects that wrap around XML/HTML elements and allow you to run further queries to fine-grain the selection or extract the data.
System Message: ERROR/3 (<stdin>, line 214); backlink
Unknown interpreted text role "class".System Message: ERROR/3 (<stdin>, line 214); backlink
Unknown interpreted text role "class".To extract the text from the title above, you can do:
>>> response.css('title::text').extract()
[u'Quotes to Scrape']
There are two things to note here: one is that we've added ::text to the CSS query, to mean that we want to select the text from inside the title element.
The other is that the result of calling .extract() is a list, because we're dealing with an instance :class:`~scrapy.selector.SelectorList`. When you know you just want the first result, as in this case, you can do:
System Message: ERROR/3 (<stdin>, line 228); backlink
Unknown interpreted text role "class".>>> response.css('title::text').extract_first()
u'Quotes to Scrape'
As an alternative, you could've written:
>>> response.css('title::text')[0].extract()
u'Quotes to Scrape'
However, using .extract_first() avoids an IndexError and returns None when it doesn't find any element matching the selection.
There's a lesson here: for most scraping code, you want it to be resilient to errors due to things not being found on a page, so that even if some parts fail to be scraped, you can at least get some data.
Besides the :meth:`~scrapy.selector.Selector.extract` and :meth:`~scrapy.selector.SelectorList.extract_first` methods, you can also use the :meth:`~scrapy.selector.Selector.re` method to extract using a regular expression:
System Message: ERROR/3 (<stdin>, line 247); backlink
Unknown interpreted text role "meth".System Message: ERROR/3 (<stdin>, line 247); backlink
Unknown interpreted text role "meth".System Message: ERROR/3 (<stdin>, line 247); backlink
Unknown interpreted text role "meth".>>> response.css('title::text').re('Quotes.*')
[u'Quotes to Scrape']
>>> response.css('title::text').re('Q\w+')
[u'Quotes']
>>> response.css('title::text').re('(\w+) to (\w+)')
[u'Quotes', u'Scrape']
In order to find the proper CSS selectors to use, you might find useful opening the response page from the shell in your web browser using view(response). You can use your browser developer tools or extensions like Firebug. For more information see :ref:`topics-firebug` and :ref:`topics-firefox`.
System Message: ERROR/3 (<stdin>, line 259); backlink
Unknown interpreted text role "ref".System Message: ERROR/3 (<stdin>, line 259); backlink
Unknown interpreted text role "ref".XPath: a brief intro
Besides CSS, Scrapy selectors also support using XPath expressions:
>>> response.xpath('//title')
[<Selector xpath='//title' data=u'<title>Quotes to Scrape</title>'>]
>>> response.xpath('//title/text()').extract_first()
u'Quotes to Scrape'
XPath expressions are very powerful, and are the foundation of Scrapy Selectors. In fact, CSS selectors are converted to XPath under-the-hood. You can see that if you read closely the text representation of the selector objects in the shell.
While perhaps not as popular as CSS selectors, XPath expressions offer more power because besides navigating the structure, it can also look at the content. Using XPath, you're able to select things like: select the link that contains the text "Next Page". This makes XPath very fitting to the task of scraping, and we encourage you to learn XPath even if you already know how to construct CSS selectors, it will make scraping much easier.
We won't cover much of XPath here. To learn more about XPath, we recommend this tutorial to learn XPath through examples, and this tutorial to learn "how to think in XPath".
Extraction wrap-up
Now that you know a bit about selection and extraction, let's complete our spider by writing the code to extract the quotes from the webpage.
Each quote in http://quotes.toscrape.com is represented by HTML code that looks like this:
<div class="quote">
<span class="text">“The world as we have created it is a process of our
thinking. It cannot be changed without changing our thinking.”</span>
<span>
by <small class="author">Albert Einstein</small>
<a href="/author/Albert-Einstein">(about)</a>
</span>
<div class="tags">
Tags:
<a class="tag" href="/tag/change/page/1/">change</a>
<a class="tag" href="/tag/deep-thoughts/page/1/">deep-thoughts</a>
<a class="tag" href="/tag/thinking/page/1/">thinking</a>
<a class="tag" href="/tag/world/page/1/">world</a>
</div>
</div>
Let's open up scrapy shell and play a bit to find out how to extract the data we want:
$ scrapy shell http://quotes.toscrape.com
We get a list of selectors to the quotes using:
>>> response.css("div.quote")
Each of the selectors returned by the query above allows us to run further queries over the quotes itselves. Let's assign the first selector to a variable, so that we can run our CSS selectors directly on a particular quote:
>>> quote = response.css("div.quote")[0]
Now, let's extract title, author and the tags from that quote using the quote object we just created:
>>> title = quote.css("span.text ::text").extract_first()
>>> title
'“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”'
>>> author = quote.css("small.author ::text").extract_first()
>>> author
'Albert Einstein'
Given that the tags is a list of strings, we can use the .extract() method to get all of them:
>>> tags = quote.css("div.tags a.tag ::text").extract()
>>> tags
['change', 'deep-thoughts', 'thinking', 'world']
Now, we can iterate over all the quotes in the page and use the CSS selectors we defined to extract data:
>>> for quote in response.css("div.quote"):
... text = quote.css("span.text ::text").extract_first()
... author = quote.css("small.author ::text").extract_first()
... tags = quote.css("div.tags a.tag ::text").extract()
... print("{} - {} - {}".format(text, author, tags))
Extracting data in our spider
Until now, the spider we built doesn't extract any data in particular. I just saves the whole HTML page to a local file. Now, let's integrate the extraction logic above in our spider.
A Scrapy spider typically generates many dictionaries containing the data extracted from the page. To do that, we use the yield Python keyword, as you can see below:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
'http://quotes.toscrape.com/page/1/',
'http://quotes.toscrape.com/page/2/',
]
def parse(self, response):
for quote in response.css('div.quote'):
yield {
'text': quote.css('span.text::text').extract_first(),
'author': quote.css('span small::text').extract_first(),
'tags': quote.css("div.tags a.tag ::text").extract(),
}
If you run this spider, it will output the extracted data with the log:
2016-09-19 18:57:19 [scrapy] DEBUG: Scraped from <200 http://quotes.toscrape.com/page/1/>
{'tags': ['life', 'love'], 'author': 'André Gide', 'text': '“It is better to be hated for what you are than to be loved for what you are not.”'}
2016-09-19 18:57:19 [scrapy] DEBUG: Scraped from <200 http://quotes.toscrape.com/page/1/>
{'tags': ['edison', 'failure', 'inspirational', 'paraphrased'], 'author': 'Thomas A. Edison', 'text': "“I have not failed. I've just found 10,000 ways that won't work.”"}
:ref:`Later in the tutorial <storing-data>`, we will see how to save this data to a file.
System Message: ERROR/3 (<stdin>, line 398); backlink
Unknown interpreted text role "ref".Following links
Let's say, instead of just scraping the stuff from the first two pages from http://quotes.toscrape.com, you want quotes from all the pages in the website.
Now that you know how to extract data from pages, let's see how to follow links from them.
Here is a modification of our spider that recursively follows the link to the next page, extracting data from it:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
'http://quotes.toscrape.com/page/1/',
]
def parse(self, response):
for quote in response.css('div.quote'):
yield {
'text': quote.css('span.text::text').extract_first(),
'author': quote.css('span small::text').extract_first(),
}
next_page = response.css('li.next a::attr("href")').extract_first()
if next_page is not None:
next_page = response.urljoin(next_page)
yield scrapy.Request(next_page, callback=self.parse)
Now, after extracting the data, the parse() method looks for the link to the next page, builds a full absolute URL using the response.urljoin method (since the links can be relative) and yields a new request to the next page, registering itself as callback to handle the data extraction for the next page and to keep the crawling going through all the pages.
What you see here is Scrapy's mechanism of following links: when you yield a Request in a callback method, Scrapy will schedule that request to be sent and register a callback method to be executed when that request finishes.
Using this, you can build complex crawlers that follow links according to rules you define, and extract different kinds of data depending on the page it's visiting.
In our example, it creates a sort of loop, following all the links to the next page until it doesn't find one -- handy for crawling blogs, forums and other sites with pagination.
Another common pattern is to build an item with data from more than one page, using a :ref:`trick to pass additional data to the callbacks <topics-request-response-ref-request-callback-arguments>`.
System Message: ERROR/3 (<stdin>, line 453); backlink
Unknown interpreted text role "ref".Adding a spider argument
You can provide command line arguments to your spiders by using the -a option when running them:
scrapy crawl quotes -o items.json -a tag=humor
In this example, the value provided for the tag argument will be available via a spider attribute. Using this, you could make your spider get only quotes tagged with a specific tag, building the URL based on the argument:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
def start_requests(self):
url = 'http://quotes.toscrape.com/'
tag = getattr(self, 'tag', None)
if tag is not None:
url = url + 'tag/' + tag
yield scrapy.Request(url)
def parse(self, response):
for quote in response.css('div.quote'):
yield {
'text': quote.css('span.text::text').extract_first(),
'author': quote.css('span small a::text').extract_first(),
}
next_page = response.css('li.next a::attr("href")').extract_first()
if next_page is not None:
next_page = response.urljoin(next_page)
yield scrapy.Request(next_page, callback=self.parse)
If you pass the tag=humor argument to this spider, you'll notice that it will only visit URLs from the humor tag, such as http://quotes.toscrape.com/tag/humor.
Storing the scraped data
The simplest way to store the scraped data is by using :ref:`Feed exports <topics-feed-exports>`, with the following command:
System Message: ERROR/3 (<stdin>, line 561); backlink
Unknown interpreted text role "ref".scrapy crawl quotes -o items.json
That will generate an items.json file containing all scraped items, serialized in JSON.
In small projects (like the one in this tutorial), that should be enough. However, if you want to perform more complex things with the scraped items, you can write an :ref:`Item Pipeline <topics-item-pipeline>`. As with Items, a placeholder file for Item Pipelines has been set up for you when the project is created, in tutorial/pipelines.py. Though you don't need to implement any item pipelines if you just want to store the scraped items.
System Message: ERROR/3 (<stdin>, line 569); backlink
Unknown interpreted text role "ref".Next steps
This tutorial covered only the basics of Scrapy, but there's a lot of other features not mentioned here. Check the :ref:`topics-whatelse` section in :ref:`intro-overview` chapter for a quick overview of the most important ones.
System Message: ERROR/3 (<stdin>, line 579); backlink
Unknown interpreted text role "ref".System Message: ERROR/3 (<stdin>, line 579); backlink
Unknown interpreted text role "ref".Then, we recommend you continue by playing with an example project (see :ref:`intro-examples`), and then continue with the section :ref:`section-basics`.
System Message: ERROR/3 (<stdin>, line 583); backlink
Unknown interpreted text role "ref".System Message: ERROR/3 (<stdin>, line 583); backlink
Unknown interpreted text role "ref".