From 83b8046fdcae660276634a03b99d51c629fa8b7c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Fri, 8 Nov 2019 16:26:46 +0100 Subject: [PATCH] Do not indent doctests from the documentation unnecessarily --- docs/intro/tutorial.rst | 120 +++---- docs/news.rst | 12 +- docs/topics/developer-tools.rst | 12 +- docs/topics/dynamic-content.rst | 28 +- docs/topics/items.rst | 106 +++--- docs/topics/leaks.rst | 140 ++++---- docs/topics/loaders.rst | 114 ++++--- docs/topics/selectors.rst | 565 ++++++++++++++++---------------- docs/topics/shell.rst | 77 +++-- docs/topics/stats.rst | 12 +- 10 files changed, 589 insertions(+), 597 deletions(-) diff --git a/docs/intro/tutorial.rst b/docs/intro/tutorial.rst index 6b15a5fbd..33b1a969b 100644 --- a/docs/intro/tutorial.rst +++ b/docs/intro/tutorial.rst @@ -252,30 +252,30 @@ The result of running ``response.css('title')`` is a list-like object called and allow you to run further queries to fine-grain the selection or extract the data. -To extract the text from the title above, you can do:: +To extract the text from the title above, you can do: - >>> response.css('title::text').getall() - ['Quotes to Scrape'] +>>> response.css('title::text').getall() +['Quotes to Scrape'] There are two things to note here: one is that we've added ``::text`` to the CSS query, to mean we want to select only the text elements directly inside ```` element. If we don't specify ``::text``, we'd get the full title -element, including its tags:: +element, including its tags: - >>> response.css('title').getall() - ['<title>Quotes to Scrape'] +>>> response.css('title').getall() +['Quotes to Scrape'] The other thing is that the result of calling ``.getall()`` is a list: it is possible that a selector returns more than one result, so we extract them all. -When you know you just want the first result, as in this case, you can do:: +When you know you just want the first result, as in this case, you can do: - >>> response.css('title::text').get() - 'Quotes to Scrape' +>>> response.css('title::text').get() +'Quotes to Scrape' -As an alternative, you could've written:: +As an alternative, you could've written: - >>> response.css('title::text')[0].get() - 'Quotes to Scrape' +>>> response.css('title::text')[0].get() +'Quotes to Scrape' However, using ``.get()`` directly on a :class:`~scrapy.selector.SelectorList` instance avoids an ``IndexError`` and returns ``None`` when it doesn't @@ -288,14 +288,14 @@ to be scraped, you can at least get **some** data. Besides the :meth:`~scrapy.selector.SelectorList.getall` and :meth:`~scrapy.selector.SelectorList.get` methods, you can also use the :meth:`~scrapy.selector.SelectorList.re` method to extract using `regular -expressions`_:: +expressions`_: - >>> response.css('title::text').re(r'Quotes.*') - ['Quotes to Scrape'] - >>> response.css('title::text').re(r'Q\w+') - ['Quotes'] - >>> response.css('title::text').re(r'(\w+) to (\w+)') - ['Quotes', 'Scrape'] +>>> response.css('title::text').re(r'Quotes.*') +['Quotes to Scrape'] +>>> response.css('title::text').re(r'Q\w+') +['Quotes'] +>>> response.css('title::text').re(r'(\w+) to (\w+)') +['Quotes', 'Scrape'] In order to find the proper CSS selectors to use, you might find useful opening the response page from the shell in your web browser using ``view(response)``. @@ -312,12 +312,12 @@ visually selected elements, which works in many browsers. XPath: a brief intro ^^^^^^^^^^^^^^^^^^^^ -Besides `CSS`_, Scrapy selectors also support using `XPath`_ expressions:: +Besides `CSS`_, Scrapy selectors also support using `XPath`_ expressions: - >>> response.xpath('//title') - [] - >>> response.xpath('//title/text()').get() - 'Quotes to Scrape' +>>> response.xpath('//title') +[] +>>> response.xpath('//title/text()').get() +'Quotes to Scrape' XPath expressions are very powerful, and are the foundation of Scrapy Selectors. In fact, CSS selectors are converted to XPath under-the-hood. You @@ -372,35 +372,35 @@ we want:: $ scrapy shell 'http://quotes.toscrape.com' -We get a list of selectors for the quote HTML elements with:: +We get a list of selectors for the quote HTML elements with: - >>> response.css("div.quote") - [, - , - ...] +>>> response.css("div.quote") +[, + , + ...] Each of the selectors returned by the query above allows us to run further queries over their sub-elements. Let's assign the first selector to a -variable, so that we can run our CSS selectors directly on a particular quote:: +variable, so that we can run our CSS selectors directly on a particular quote: - >>> quote = response.css("div.quote")[0] +>>> quote = response.css("div.quote")[0] Now, let's extract ``text``, ``author`` and the ``tags`` from that quote -using the ``quote`` object we just created:: +using the ``quote`` object we just created: - >>> text = quote.css("span.text::text").get() - >>> text - '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”' - >>> author = quote.css("small.author::text").get() - >>> author - 'Albert Einstein' +>>> text = quote.css("span.text::text").get() +>>> text +'“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”' +>>> author = quote.css("small.author::text").get() +>>> author +'Albert Einstein' Given that the tags are a list of strings, we can use the ``.getall()`` method -to get all of them:: +to get all of them: - >>> tags = quote.css("div.tags a.tag::text").getall() - >>> tags - ['change', 'deep-thoughts', 'thinking', 'world'] +>>> tags = quote.css("div.tags a.tag::text").getall() +>>> tags +['change', 'deep-thoughts', 'thinking', 'world'] .. invisible-code-block: python @@ -409,16 +409,16 @@ to get all of them:: .. skip: next if(version_info < (3, 6), reason="Only Python 3.6+ dictionaries match the output") Having figured out how to extract each bit, we can now iterate over all the -quotes elements and put them together into a Python dictionary:: +quotes elements and put them together into a Python dictionary: - >>> for quote in response.css("div.quote"): - ... text = quote.css("span.text::text").get() - ... author = quote.css("small.author::text").get() - ... tags = quote.css("div.tags a.tag::text").getall() - ... print(dict(text=text, author=author, tags=tags)) - {'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein', 'tags': ['change', 'deep-thoughts', 'thinking', 'world']} - {'text': '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', 'author': 'J.K. Rowling', 'tags': ['abilities', 'choices']} - ... +>>> for quote in response.css("div.quote"): +... text = quote.css("span.text::text").get() +... author = quote.css("small.author::text").get() +... tags = quote.css("div.tags a.tag::text").getall() +... print(dict(text=text, author=author, tags=tags)) +{'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein', 'tags': ['change', 'deep-thoughts', 'thinking', 'world']} +{'text': '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', 'author': 'J.K. Rowling', 'tags': ['abilities', 'choices']} +... Extracting data in our spider ----------------------------- @@ -516,23 +516,23 @@ markup: -We can try extracting it in the shell:: +We can try extracting it in the shell: - >>> response.css('li.next a').get() - 'Next ' +>>> response.css('li.next a').get() +'Next ' This gets the anchor element, but we want the attribute ``href``. For that, Scrapy supports a CSS extension that lets you select the attribute contents, -like this:: +like this: - >>> response.css('li.next a::attr(href)').get() - '/page/2/' +>>> response.css('li.next a::attr(href)').get() +'/page/2/' There is also an ``attrib`` property available -(see :ref:`selecting-attributes` for more):: +(see :ref:`selecting-attributes` for more): - >>> response.css('li.next a').attrib['href'] - '/page/2/' +>>> response.css('li.next a').attrib['href'] +'/page/2/' Let's see now our spider modified to recursively follow the link to the next page, extracting data from it:: diff --git a/docs/news.rst b/docs/news.rst index 9dfd28508..217382c57 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -47,13 +47,13 @@ Backward-incompatible changes (:issue:`2111`, :issue:`3392`, :issue:`3442`, :issue:`3450`) * :class:`~scrapy.loader.ItemLoader` now turns the values of its input item - into lists:: + into lists: - >>> item = MyItem() - >>> item['field'] = 'value1' - >>> loader = ItemLoader(item=item) - >>> item['field'] - ['value1'] + >>> item = MyItem() + >>> item['field'] = 'value1' + >>> loader = ItemLoader(item=item) + >>> item['field'] + ['value1'] This is needed to allow adding values to existing fields (``loader.add_value('field', 'value2')``). diff --git a/docs/topics/developer-tools.rst b/docs/topics/developer-tools.rst index bf14643be..1c9315cd8 100644 --- a/docs/topics/developer-tools.rst +++ b/docs/topics/developer-tools.rst @@ -112,13 +112,13 @@ see each quote: With this knowledge we can refine our XPath: Instead of a path to follow, we'll simply select all ``span`` tags with the ``class="text"`` by using -the `has-class-extension`_:: +the `has-class-extension`_: - >>> response.xpath('//span[has-class("text")]/text()').getall() - ['"The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”, - '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', - '“There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.”', - (...)] +>>> response.xpath('//span[has-class("text")]/text()').getall() +['"The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', + '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', + '“There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.”', + ...] And with one simple, cleverer XPath we are able to extract all quotes from the page. We could have constructed a loop over our first XPath to increase diff --git a/docs/topics/dynamic-content.rst b/docs/topics/dynamic-content.rst index 8334ddcec..1c3607860 100644 --- a/docs/topics/dynamic-content.rst +++ b/docs/topics/dynamic-content.rst @@ -172,27 +172,27 @@ data from it: data in JSON format, which you can then parse with `json.loads`_. For example, if the JavaScript code contains a separate line like - ``var data = {"field": "value"};`` you can extract that data as follows:: + ``var data = {"field": "value"};`` you can extract that data as follows: - >>> pattern = r'\bvar\s+data\s*=\s*(\{.*?\})\s*;\s*\n' - >>> json_data = response.css('script::text').re_first(pattern) - >>> json.loads(json_data) - {'field': 'value'} + >>> pattern = r'\bvar\s+data\s*=\s*(\{.*?\})\s*;\s*\n' + >>> json_data = response.css('script::text').re_first(pattern) + >>> json.loads(json_data) + {'field': 'value'} - Otherwise, use js2xml_ to convert the JavaScript code into an XML document that you can parse using :ref:`selectors `. For example, if the JavaScript code contains - ``var data = {field: "value"};`` you can extract that data as follows:: + ``var data = {field: "value"};`` you can extract that data as follows: - >>> import js2xml - >>> import lxml.etree - >>> from parsel import Selector - >>> javascript = response.css('script::text').get() - >>> xml = lxml.etree.tostring(js2xml.parse(javascript), encoding='unicode') - >>> selector = Selector(text=xml) - >>> selector.css('var[name="data"]').get() - 'value' + >>> import js2xml + >>> import lxml.etree + >>> from parsel import Selector + >>> javascript = response.css('script::text').get() + >>> xml = lxml.etree.tostring(js2xml.parse(javascript), encoding='unicode') + >>> selector = Selector(text=xml) + >>> selector.css('var[name="data"]').get() + 'value' .. _topics-javascript-rendering: diff --git a/docs/topics/items.rst b/docs/topics/items.rst index cdf60208e..15313775b 100644 --- a/docs/topics/items.rst +++ b/docs/topics/items.rst @@ -84,77 +84,74 @@ notice the API is very similar to the `dict API`_. Creating items -------------- -:: +>>> product = Product(name='Desktop PC', price=1000) +>>> print(product) +Product(name='Desktop PC', price=1000) - >>> product = Product(name='Desktop PC', price=1000) - >>> print(product) - Product(name='Desktop PC', price=1000) Getting field values -------------------- -:: +>>> product['name'] +Desktop PC +>>> product.get('name') +Desktop PC - >>> product['name'] - Desktop PC - >>> product.get('name') - Desktop PC +>>> product['price'] +1000 - >>> product['price'] - 1000 +>>> product['last_updated'] +Traceback (most recent call last): + ... +KeyError: 'last_updated' - >>> product['last_updated'] - Traceback (most recent call last): - ... - KeyError: 'last_updated' +>>> product.get('last_updated', 'not set') +not set - >>> product.get('last_updated', 'not set') - not set +>>> product['lala'] # getting unknown field +Traceback (most recent call last): + ... +KeyError: 'lala' - >>> product['lala'] # getting unknown field - Traceback (most recent call last): - ... - KeyError: 'lala' +>>> product.get('lala', 'unknown field') +'unknown field' - >>> product.get('lala', 'unknown field') - 'unknown field' +>>> 'name' in product # is name field populated? +True - >>> 'name' in product # is name field populated? - True +>>> 'last_updated' in product # is last_updated populated? +False - >>> 'last_updated' in product # is last_updated populated? - False +>>> 'last_updated' in product.fields # is last_updated a declared field? +True - >>> 'last_updated' in product.fields # is last_updated a declared field? - True +>>> 'lala' in product.fields # is lala a declared field? +False - >>> 'lala' in product.fields # is lala a declared field? - False Setting field values -------------------- -:: +>>> product['last_updated'] = 'today' +>>> product['last_updated'] +today - >>> product['last_updated'] = 'today' - >>> product['last_updated'] - today +>>> product['lala'] = 'test' # setting unknown field +Traceback (most recent call last): + ... +KeyError: 'Product does not support field: lala' - >>> product['lala'] = 'test' # setting unknown field - Traceback (most recent call last): - ... - KeyError: 'Product does not support field: lala' Accessing all populated values ------------------------------ -To access all populated values, just use the typical `dict API`_:: +To access all populated values, just use the typical `dict API`_: - >>> product.keys() - ['price', 'name'] +>>> product.keys() +['price', 'name'] - >>> product.items() - [('price', 1000), ('name', 'Desktop PC')] +>>> product.items() +[('price', 1000), ('name', 'Desktop PC')] .. _copying-items: @@ -194,20 +191,21 @@ To create a deep copy, call :meth:`~scrapy.item.Item.deepcopy` instead Other common tasks ------------------ -Creating dicts from items:: +Creating dicts from items: - >>> dict(product) # create a dict from all populated values - {'price': 1000, 'name': 'Desktop PC'} +>>> dict(product) # create a dict from all populated values +{'price': 1000, 'name': 'Desktop PC'} -Creating items from dicts:: +Creating items from dicts: - >>> Product({'name': 'Laptop PC', 'price': 1500}) - Product(price=1500, name='Laptop PC') +>>> Product({'name': 'Laptop PC', 'price': 1500}) +Product(price=1500, name='Laptop PC') + +>>> Product({'name': 'Laptop PC', 'lala': 1500}) # warning: unknown field in dict +Traceback (most recent call last): + ... +KeyError: 'Product does not support field: lala' - >>> Product({'name': 'Laptop PC', 'lala': 1500}) # warning: unknown field in dict - Traceback (most recent call last): - ... - KeyError: 'Product does not support field: lala' Extending Items =============== diff --git a/docs/topics/leaks.rst b/docs/topics/leaks.rst index 793636f59..87d9d262f 100644 --- a/docs/topics/leaks.rst +++ b/docs/topics/leaks.rst @@ -132,21 +132,21 @@ and check the code of the spider to discover the nasty line that is generating the leaks (passing response references inside requests). Sometimes extra information about live objects can be helpful. -Let's check the oldest response:: +Let's check the oldest response: - >>> from scrapy.utils.trackref import get_oldest - >>> r = get_oldest('HtmlResponse') - >>> r.url - 'http://www.somenastyspider.com/product.php?pid=123' +>>> from scrapy.utils.trackref import get_oldest +>>> r = get_oldest('HtmlResponse') +>>> r.url +'http://www.somenastyspider.com/product.php?pid=123' If you want to iterate over all objects, instead of getting the oldest one, you -can use the :func:`scrapy.utils.trackref.iter_all` function:: +can use the :func:`scrapy.utils.trackref.iter_all` function: - >>> from scrapy.utils.trackref import iter_all - >>> [r.url for r in iter_all('HtmlResponse')] - ['http://www.somenastyspider.com/product.php?pid=123', - 'http://www.somenastyspider.com/product.php?pid=584', - ... +>>> from scrapy.utils.trackref import iter_all +>>> [r.url for r in iter_all('HtmlResponse')] +['http://www.somenastyspider.com/product.php?pid=123', + 'http://www.somenastyspider.com/product.php?pid=584', +...] Too many spiders? ----------------- @@ -155,10 +155,10 @@ If your project has too many spiders executed in parallel, the output of :func:`prefs()` can be difficult to read. For this reason, that function has a ``ignore`` argument which can be used to ignore a particular class (and all its subclases). For -example, this won't show any live references to spiders:: +example, this won't show any live references to spiders: - >>> from scrapy.spiders import Spider - >>> prefs(ignore=Spider) +>>> from scrapy.spiders import Spider +>>> prefs(ignore=Spider) .. module:: scrapy.utils.trackref :synopsis: Track references of live objects @@ -214,41 +214,41 @@ If you use ``pip``, you can install Guppy with the following command:: The telnet console also comes with a built-in shortcut (``hpy``) for accessing Guppy heap objects. Here's an example to view all Python objects available in -the heap using Guppy:: +the heap using Guppy: - >>> x = hpy.heap() - >>> x.bytype - Partition of a set of 297033 objects. Total size = 52587824 bytes. - Index Count % Size % Cumulative % Type - 0 22307 8 16423880 31 16423880 31 dict - 1 122285 41 12441544 24 28865424 55 str - 2 68346 23 5966696 11 34832120 66 tuple - 3 227 0 5836528 11 40668648 77 unicode - 4 2461 1 2222272 4 42890920 82 type - 5 16870 6 2024400 4 44915320 85 function - 6 13949 5 1673880 3 46589200 89 types.CodeType - 7 13422 5 1653104 3 48242304 92 list - 8 3735 1 1173680 2 49415984 94 _sre.SRE_Pattern - 9 1209 0 456936 1 49872920 95 scrapy.http.headers.Headers - <1676 more rows. Type e.g. '_.more' to view.> +>>> x = hpy.heap() +>>> x.bytype +Partition of a set of 297033 objects. Total size = 52587824 bytes. + Index Count % Size % Cumulative % Type + 0 22307 8 16423880 31 16423880 31 dict + 1 122285 41 12441544 24 28865424 55 str + 2 68346 23 5966696 11 34832120 66 tuple + 3 227 0 5836528 11 40668648 77 unicode + 4 2461 1 2222272 4 42890920 82 type + 5 16870 6 2024400 4 44915320 85 function + 6 13949 5 1673880 3 46589200 89 types.CodeType + 7 13422 5 1653104 3 48242304 92 list + 8 3735 1 1173680 2 49415984 94 _sre.SRE_Pattern + 9 1209 0 456936 1 49872920 95 scrapy.http.headers.Headers +<1676 more rows. Type e.g. '_.more' to view.> You can see that most space is used by dicts. Then, if you want to see from -which attribute those dicts are referenced, you could do:: +which attribute those dicts are referenced, you could do: - >>> x.bytype[0].byvia - Partition of a set of 22307 objects. Total size = 16423880 bytes. - Index Count % Size % Cumulative % Referred Via: - 0 10982 49 9416336 57 9416336 57 '.__dict__' - 1 1820 8 2681504 16 12097840 74 '.__dict__', '.func_globals' - 2 3097 14 1122904 7 13220744 80 - 3 990 4 277200 2 13497944 82 "['cookies']" - 4 987 4 276360 2 13774304 84 "['cache']" - 5 985 4 275800 2 14050104 86 "['meta']" - 6 897 4 251160 2 14301264 87 '[2]' - 7 1 0 196888 1 14498152 88 "['moduleDict']", "['modules']" - 8 672 3 188160 1 14686312 89 "['cb_kwargs']" - 9 27 0 155016 1 14841328 90 '[1]' - <333 more rows. Type e.g. '_.more' to view.> +>>> x.bytype[0].byvia +Partition of a set of 22307 objects. Total size = 16423880 bytes. + Index Count % Size % Cumulative % Referred Via: + 0 10982 49 9416336 57 9416336 57 '.__dict__' + 1 1820 8 2681504 16 12097840 74 '.__dict__', '.func_globals' + 2 3097 14 1122904 7 13220744 80 + 3 990 4 277200 2 13497944 82 "['cookies']" + 4 987 4 276360 2 13774304 84 "['cache']" + 5 985 4 275800 2 14050104 86 "['meta']" + 6 897 4 251160 2 14301264 87 '[2]' + 7 1 0 196888 1 14498152 88 "['moduleDict']", "['modules']" + 8 672 3 188160 1 14686312 89 "['cb_kwargs']" + 9 27 0 155016 1 14841328 90 '[1]' +<333 more rows. Type e.g. '_.more' to view.> As you can see, the Guppy module is very powerful but also requires some deep knowledge about Python internals. For more info about Guppy, refer to the @@ -269,32 +269,32 @@ If you use ``pip``, you can install muppy with the following command:: pip install Pympler Here's an example to view all Python objects available in -the heap using muppy:: +the heap using muppy: - >>> from pympler import muppy - >>> all_objects = muppy.get_objects() - >>> len(all_objects) - 28667 - >>> from pympler import summary - >>> suml = summary.summarize(all_objects) - >>> summary.print_(suml) - types | # objects | total size - ==================================== | =========== | ============ - >> from pympler import muppy +>>> all_objects = muppy.get_objects() +>>> len(all_objects) +28667 +>>> from pympler import summary +>>> suml = summary.summarize(all_objects) +>>> summary.print_(suml) + types | # objects | total size +==================================== | =========== | ============ + >> from scrapy.loader import ItemLoader - >>> il = ItemLoader(item=Product()) - >>> il.add_value('name', [u'Welcome to my', u'website']) - >>> il.add_value('price', [u'€', u'1000']) - >>> il.load_item() - {'name': u'Welcome to my website', 'price': u'1000'} +>>> from scrapy.loader import ItemLoader +>>> il = ItemLoader(item=Product()) +>>> il.add_value('name', [u'Welcome to my', u'website']) +>>> il.add_value('price', [u'€', u'1000']) +>>> il.load_item() +{'name': u'Welcome to my website', 'price': u'1000'} The precedence order, for both input and output processors, is as follows: @@ -314,11 +312,11 @@ ItemLoader objects applied before processors :type re: str or compiled regex - Examples:: + Examples: - >>> from scrapy.loader.processors import TakeFirst - >>> loader.get_value(u'name: foo', TakeFirst(), unicode.upper, re='name: (.+)') - 'FOO` + >>> from scrapy.loader.processors import TakeFirst + >>> loader.get_value(u'name: foo', TakeFirst(), unicode.upper, re='name: (.+)') + 'FOO` .. method:: add_value(field_name, value, \*processors, \**kwargs) @@ -639,12 +637,12 @@ Here is a list of all built-in processors: values unchanged. It doesn't receive any ``__init__`` method arguments, nor does it accept Loader contexts. - Example:: + Example: - >>> from scrapy.loader.processors import Identity - >>> proc = Identity() - >>> proc(['one', 'two', 'three']) - ['one', 'two', 'three'] + >>> from scrapy.loader.processors import Identity + >>> proc = Identity() + >>> proc(['one', 'two', 'three']) + ['one', 'two', 'three'] .. class:: TakeFirst @@ -652,12 +650,12 @@ Here is a list of all built-in processors: so it's typically used as an output processor to single-valued fields. It doesn't receive any ``__init__`` method arguments, nor does it accept Loader contexts. - Example:: + Example: - >>> from scrapy.loader.processors import TakeFirst - >>> proc = TakeFirst() - >>> proc(['', 'one', 'two', 'three']) - 'one' + >>> from scrapy.loader.processors import TakeFirst + >>> proc = TakeFirst() + >>> proc(['', 'one', 'two', 'three']) + 'one' .. class:: Join(separator=u' ') @@ -667,15 +665,15 @@ Here is a list of all built-in processors: When using the default separator, this processor is equivalent to the function: ``u' '.join`` - Examples:: + Examples: - >>> from scrapy.loader.processors import Join - >>> proc = Join() - >>> proc(['one', 'two', 'three']) - 'one two three' - >>> proc = Join('
') - >>> proc(['one', 'two', 'three']) - 'one
two
three' + >>> from scrapy.loader.processors import Join + >>> proc = Join() + >>> proc(['one', 'two', 'three']) + 'one two three' + >>> proc = Join('
') + >>> proc(['one', 'two', 'three']) + 'one
two
three' .. class:: Compose(\*functions, \**default_loader_context) @@ -688,12 +686,12 @@ Here is a list of all built-in processors: By default, stop process on ``None`` value. This behaviour can be changed by passing keyword argument ``stop_on_none=False``. - Example:: + Example: - >>> from scrapy.loader.processors import Compose - >>> proc = Compose(lambda v: v[0], str.upper) - >>> proc(['hello', 'world']) - 'HELLO' + >>> from scrapy.loader.processors import Compose + >>> proc = Compose(lambda v: v[0], str.upper) + >>> proc(['hello', 'world']) + 'HELLO' Each function can optionally receive a ``loader_context`` parameter. For those which do, this processor will pass the currently active :ref:`Loader @@ -732,15 +730,15 @@ Here is a list of all built-in processors: :meth:`~scrapy.selector.Selector.extract` method of :ref:`selectors `, which returns a list of unicode strings. - The example below should clarify how it works:: + The example below should clarify how it works: - >>> def filter_world(x): - ... return None if x == 'world' else x - ... - >>> from scrapy.loader.processors import MapCompose - >>> proc = MapCompose(filter_world, str.upper) - >>> proc(['hello', 'world', 'this', 'is', 'scrapy']) - ['HELLO, 'THIS', 'IS', 'SCRAPY'] + >>> def filter_world(x): + ... return None if x == 'world' else x + ... + >>> from scrapy.loader.processors import MapCompose + >>> proc = MapCompose(filter_world, str.upper) + >>> proc(['hello', 'world', 'this', 'is', 'scrapy']) + ['HELLO, 'THIS', 'IS', 'SCRAPY'] As with the Compose processor, functions can receive Loader contexts, and ``__init__`` method keyword arguments are used as default context values. See @@ -752,21 +750,21 @@ Here is a list of all built-in processors: Requires jmespath (https://github.com/jmespath/jmespath.py) to run. This processor takes only one input at a time. - Example:: + Example: - >>> from scrapy.loader.processors import SelectJmes, Compose, MapCompose - >>> proc = SelectJmes("foo") #for direct use on lists and dictionaries - >>> proc({'foo': 'bar'}) - 'bar' - >>> proc({'foo': {'bar': 'baz'}}) - {'bar': 'baz'} + >>> from scrapy.loader.processors import SelectJmes, Compose, MapCompose + >>> proc = SelectJmes("foo") #for direct use on lists and dictionaries + >>> proc({'foo': 'bar'}) + 'bar' + >>> proc({'foo': {'bar': 'baz'}}) + {'bar': 'baz'} - Working with Json:: + Working with Json: - >>> import json - >>> proc_single_json_str = Compose(json.loads, SelectJmes("foo")) - >>> proc_single_json_str('{"foo": "bar"}') - 'bar' - >>> proc_json_list = Compose(json.loads, MapCompose(SelectJmes('foo'))) - >>> proc_json_list('[{"foo":"bar"}, {"baz":"tar"}]') - ['bar'] + >>> import json + >>> proc_single_json_str = Compose(json.loads, SelectJmes("foo")) + >>> proc_single_json_str('{"foo": "bar"}') + 'bar' + >>> proc_json_list = Compose(json.loads, MapCompose(SelectJmes('foo'))) + >>> proc_json_list('[{"foo":"bar"}, {"baz":"tar"}]') + ['bar'] diff --git a/docs/topics/selectors.rst b/docs/topics/selectors.rst index 282a585d4..8ec758b0e 100644 --- a/docs/topics/selectors.rst +++ b/docs/topics/selectors.rst @@ -51,18 +51,18 @@ Constructing selectors .. highlight:: python Response objects expose a :class:`~scrapy.selector.Selector` instance -on ``.selector`` attribute:: +on ``.selector`` attribute: - >>> response.selector.xpath('//span/text()').get() - 'good' +>>> response.selector.xpath('//span/text()').get() +'good' Querying responses using XPath and CSS is so common that responses include two -more shortcuts: ``response.xpath()`` and ``response.css()``:: +more shortcuts: ``response.xpath()`` and ``response.css()``: - >>> response.xpath('//span/text()').get() - 'good' - >>> response.css('span::text').get() - 'good' +>>> response.xpath('//span/text()').get() +'good' +>>> response.css('span::text').get() +'good' Scrapy selectors are instances of :class:`~scrapy.selector.Selector` class constructed by passing either :class:`~scrapy.http.TextResponse` object or @@ -74,21 +74,21 @@ shortcuts. By using ``response.selector`` or one of these shortcuts you can also ensure the response body is parsed only once. But if required, it is possible to use ``Selector`` directly. -Constructing from text:: +Constructing from text: - >>> from scrapy.selector import Selector - >>> body = 'good' - >>> Selector(text=body).xpath('//span/text()').get() - 'good' +>>> from scrapy.selector import Selector +>>> body = 'good' +>>> Selector(text=body).xpath('//span/text()').get() +'good' Constructing from response - :class:`~scrapy.http.HtmlResponse` is one of -:class:`~scrapy.http.TextResponse` subclasses:: +:class:`~scrapy.http.TextResponse` subclasses: - >>> from scrapy.selector import Selector - >>> from scrapy.http import HtmlResponse - >>> response = HtmlResponse(url='http://example.com', body=body) - >>> Selector(response=response).xpath('//span/text()').get() - 'good' +>>> from scrapy.selector import Selector +>>> from scrapy.http import HtmlResponse +>>> response = HtmlResponse(url='http://example.com', body=body) +>>> Selector(response=response).xpath('//span/text()').get() +'good' ``Selector`` automatically chooses the best parsing rules (XML vs HTML) based on input type. @@ -123,118 +123,118 @@ Since we're dealing with HTML, the selector will automatically use an HTML parse .. highlight:: python So, by looking at the :ref:`HTML code ` of that -page, let's construct an XPath for selecting the text inside the title tag:: +page, let's construct an XPath for selecting the text inside the title tag: - >>> response.xpath('//title/text()') - [] +>>> response.xpath('//title/text()') +[] To actually extract the textual data, you must call the selector ``.get()`` -or ``.getall()`` methods, as follows:: +or ``.getall()`` methods, as follows: - >>> response.xpath('//title/text()').getall() - ['Example website'] - >>> response.xpath('//title/text()').get() - 'Example website' +>>> response.xpath('//title/text()').getall() +['Example website'] +>>> response.xpath('//title/text()').get() +'Example website' ``.get()`` always returns a single result; if there are several matches, content of a first match is returned; if there are no matches, None is returned. ``.getall()`` returns a list with all results. Notice that CSS selectors can select text or attribute nodes using CSS3 -pseudo-elements:: +pseudo-elements: - >>> response.css('title::text').get() - 'Example website' +>>> response.css('title::text').get() +'Example website' As you can see, ``.xpath()`` and ``.css()`` methods return a :class:`~scrapy.selector.SelectorList` instance, which is a list of new -selectors. This API can be used for quickly selecting nested data:: +selectors. This API can be used for quickly selecting nested data: - >>> response.css('img').xpath('@src').getall() - ['image1_thumb.jpg', - 'image2_thumb.jpg', - 'image3_thumb.jpg', - 'image4_thumb.jpg', - 'image5_thumb.jpg'] +>>> response.css('img').xpath('@src').getall() +['image1_thumb.jpg', + 'image2_thumb.jpg', + 'image3_thumb.jpg', + 'image4_thumb.jpg', + 'image5_thumb.jpg'] If you want to extract only the first matched element, you can call the selector ``.get()`` (or its alias ``.extract_first()`` commonly used in -previous Scrapy versions):: +previous Scrapy versions): - >>> response.xpath('//div[@id="images"]/a/text()').get() - 'Name: My image 1 ' +>>> response.xpath('//div[@id="images"]/a/text()').get() +'Name: My image 1 ' -It returns ``None`` if no element was found:: +It returns ``None`` if no element was found: - >>> response.xpath('//div[@id="not-exists"]/text()').get() is None - True +>>> response.xpath('//div[@id="not-exists"]/text()').get() is None +True A default return value can be provided as an argument, to be used instead of ``None``: - >>> response.xpath('//div[@id="not-exists"]/text()').get(default='not-found') - 'not-found' +>>> response.xpath('//div[@id="not-exists"]/text()').get(default='not-found') +'not-found' Instead of using e.g. ``'@src'`` XPath it is possible to query for attributes -using ``.attrib`` property of a :class:`~scrapy.selector.Selector`:: +using ``.attrib`` property of a :class:`~scrapy.selector.Selector`: - >>> [img.attrib['src'] for img in response.css('img')] - ['image1_thumb.jpg', - 'image2_thumb.jpg', - 'image3_thumb.jpg', - 'image4_thumb.jpg', - 'image5_thumb.jpg'] +>>> [img.attrib['src'] for img in response.css('img')] +['image1_thumb.jpg', + 'image2_thumb.jpg', + 'image3_thumb.jpg', + 'image4_thumb.jpg', + 'image5_thumb.jpg'] As a shortcut, ``.attrib`` is also available on SelectorList directly; -it returns attributes for the first matching element:: +it returns attributes for the first matching element: - >>> response.css('img').attrib['src'] - 'image1_thumb.jpg' +>>> response.css('img').attrib['src'] +'image1_thumb.jpg' This is most useful when only a single result is expected, e.g. when selecting -by id, or selecting unique elements on a web page:: +by id, or selecting unique elements on a web page: - >>> response.css('base').attrib['href'] - 'http://example.com/' +>>> response.css('base').attrib['href'] +'http://example.com/' -Now we're going to get the base URL and some image links:: +Now we're going to get the base URL and some image links: - >>> response.xpath('//base/@href').get() - 'http://example.com/' +>>> response.xpath('//base/@href').get() +'http://example.com/' - >>> response.css('base::attr(href)').get() - 'http://example.com/' +>>> response.css('base::attr(href)').get() +'http://example.com/' - >>> response.css('base').attrib['href'] - 'http://example.com/' +>>> response.css('base').attrib['href'] +'http://example.com/' - >>> response.xpath('//a[contains(@href, "image")]/@href').getall() - ['image1.html', - 'image2.html', - 'image3.html', - 'image4.html', - 'image5.html'] +>>> response.xpath('//a[contains(@href, "image")]/@href').getall() +['image1.html', + 'image2.html', + 'image3.html', + 'image4.html', + 'image5.html'] - >>> response.css('a[href*=image]::attr(href)').getall() - ['image1.html', - 'image2.html', - 'image3.html', - 'image4.html', - 'image5.html'] +>>> response.css('a[href*=image]::attr(href)').getall() +['image1.html', + 'image2.html', + 'image3.html', + 'image4.html', + 'image5.html'] - >>> response.xpath('//a[contains(@href, "image")]/img/@src').getall() - ['image1_thumb.jpg', - 'image2_thumb.jpg', - 'image3_thumb.jpg', - 'image4_thumb.jpg', - 'image5_thumb.jpg'] +>>> response.xpath('//a[contains(@href, "image")]/img/@src').getall() +['image1_thumb.jpg', + 'image2_thumb.jpg', + 'image3_thumb.jpg', + 'image4_thumb.jpg', + 'image5_thumb.jpg'] - >>> response.css('a[href*=image] img::attr(src)').getall() - ['image1_thumb.jpg', - 'image2_thumb.jpg', - 'image3_thumb.jpg', - 'image4_thumb.jpg', - 'image5_thumb.jpg'] +>>> response.css('a[href*=image] img::attr(src)').getall() +['image1_thumb.jpg', + 'image2_thumb.jpg', + 'image3_thumb.jpg', + 'image4_thumb.jpg', + 'image5_thumb.jpg'] .. _topics-selectors-css-extensions: @@ -259,47 +259,47 @@ that Scrapy (parsel) implements a couple of **non-standard pseudo-elements**: Examples: -* ``title::text`` selects children text nodes of a descendant ```` element:: +* ``title::text`` selects children text nodes of a descendant ``<title>`` element: - >>> response.css('title::text').get() - 'Example website' +>>> response.css('title::text').get() +'Example website' -* ``*::text`` selects all descendant text nodes of the current selector context:: +* ``*::text`` selects all descendant text nodes of the current selector context: - >>> response.css('#images *::text').getall() - ['\n ', - 'Name: My image 1 ', - '\n ', - 'Name: My image 2 ', - '\n ', - 'Name: My image 3 ', - '\n ', - 'Name: My image 4 ', - '\n ', - 'Name: My image 5 ', - '\n '] +>>> response.css('#images *::text').getall() +['\n ', + 'Name: My image 1 ', + '\n ', + 'Name: My image 2 ', + '\n ', + 'Name: My image 3 ', + '\n ', + 'Name: My image 4 ', + '\n ', + 'Name: My image 5 ', + '\n '] * ``foo::text`` returns no results if ``foo`` element exists, but contains - no text (i.e. text is empty):: + no text (i.e. text is empty): - >>> response.css('img::text').getall() - [] +>>> response.css('img::text').getall() +[] This means ``.css('foo::text').get()`` could return None even if an element - exists. Use ``default=''`` if you always want a string:: + exists. Use ``default=''`` if you always want a string: - >>> response.css('img::text').get() - >>> response.css('img::text').get(default='') - '' +>>> response.css('img::text').get() +>>> response.css('img::text').get(default='') +'' -* ``a::attr(href)`` selects the *href* attribute value of descendant links:: +* ``a::attr(href)`` selects the *href* attribute value of descendant links: - >>> response.css('a::attr(href)').getall() - ['image1.html', - 'image2.html', - 'image3.html', - 'image4.html', - 'image5.html'] +>>> response.css('a::attr(href)').getall() +['image1.html', + 'image2.html', + 'image3.html', + 'image4.html', + 'image5.html'] .. note:: See also: :ref:`selecting-attributes`. @@ -318,25 +318,24 @@ Nesting selectors The selection methods (``.xpath()`` or ``.css()``) return a list of selectors of the same type, so you can call the selection methods for those selectors -too. Here's an example:: +too. Here's an example: - >>> links = response.xpath('//a[contains(@href, "image")]') - >>> links.getall() - ['<a href="image1.html">Name: My image 1 <br><img src="image1_thumb.jpg"></a>', - '<a href="image2.html">Name: My image 2 <br><img src="image2_thumb.jpg"></a>', - '<a href="image3.html">Name: My image 3 <br><img src="image3_thumb.jpg"></a>', - '<a href="image4.html">Name: My image 4 <br><img src="image4_thumb.jpg"></a>', - '<a href="image5.html">Name: My image 5 <br><img src="image5_thumb.jpg"></a>'] +>>> links = response.xpath('//a[contains(@href, "image")]') +>>> links.getall() +['<a href="image1.html">Name: My image 1 <br><img src="image1_thumb.jpg"></a>', + '<a href="image2.html">Name: My image 2 <br><img src="image2_thumb.jpg"></a>', + '<a href="image3.html">Name: My image 3 <br><img src="image3_thumb.jpg"></a>', + '<a href="image4.html">Name: My image 4 <br><img src="image4_thumb.jpg"></a>', + '<a href="image5.html">Name: My image 5 <br><img src="image5_thumb.jpg"></a>'] - >>> for index, link in enumerate(links): - ... args = (index, link.xpath('@href').get(), link.xpath('img/@src').get()) - ... print('Link number %d points to url %r and image %r' % args) - - Link number 0 points to url 'image1.html' and image 'image1_thumb.jpg' - Link number 1 points to url 'image2.html' and image 'image2_thumb.jpg' - Link number 2 points to url 'image3.html' and image 'image3_thumb.jpg' - Link number 3 points to url 'image4.html' and image 'image4_thumb.jpg' - Link number 4 points to url 'image5.html' and image 'image5_thumb.jpg' +>>> for index, link in enumerate(links): +... args = (index, link.xpath('@href').get(), link.xpath('img/@src').get()) +... print('Link number %d points to url %r and image %r' % args) +Link number 0 points to url 'image1.html' and image 'image1_thumb.jpg' +Link number 1 points to url 'image2.html' and image 'image2_thumb.jpg' +Link number 2 points to url 'image3.html' and image 'image3_thumb.jpg' +Link number 3 points to url 'image4.html' and image 'image4_thumb.jpg' +Link number 4 points to url 'image5.html' and image 'image5_thumb.jpg' .. _selecting-attributes: @@ -344,42 +343,42 @@ Selecting element attributes ---------------------------- There are several ways to get a value of an attribute. First, one can use -XPath syntax:: +XPath syntax: - >>> response.xpath("//a/@href").getall() - ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] +>>> response.xpath("//a/@href").getall() +['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] XPath syntax has a few advantages: it is a standard XPath feature, and ``@attributes`` can be used in other parts of an XPath expression - e.g. it is possible to filter by attribute value. Scrapy also provides an extension to CSS selectors (``::attr(...)``) -which allows to get attribute values:: +which allows to get attribute values: - >>> response.css('a::attr(href)').getall() - ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] +>>> response.css('a::attr(href)').getall() +['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] In addition to that, there is a ``.attrib`` property of Selector. You can use it if you prefer to lookup attributes in Python -code, without using XPaths or CSS extensions:: +code, without using XPaths or CSS extensions: - >>> [a.attrib['href'] for a in response.css('a')] - ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] +>>> [a.attrib['href'] for a in response.css('a')] +['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] This property is also available on SelectorList; it returns a dictionary with attributes of a first matching element. It is convenient to use when a selector is expected to give a single result (e.g. when selecting by element -ID, or when selecting an unique element on a page):: +ID, or when selecting an unique element on a page): - >>> response.css('base').attrib - {'href': 'http://example.com/'} - >>> response.css('base').attrib['href'] - 'http://example.com/' +>>> response.css('base').attrib +{'href': 'http://example.com/'} +>>> response.css('base').attrib['href'] +'http://example.com/' -``.attrib`` property of an empty SelectorList is empty:: +``.attrib`` property of an empty SelectorList is empty: - >>> response.css('foo').attrib - {} +>>> response.css('foo').attrib +{} Using selectors with regular expressions ---------------------------------------- @@ -390,21 +389,21 @@ data using regular expressions. However, unlike using ``.xpath()`` or can't construct nested ``.re()`` calls. Here's an example used to extract image names from the :ref:`HTML code -<topics-selectors-htmlcode>` above:: +<topics-selectors-htmlcode>` above: - >>> response.xpath('//a[contains(@href, "image")]/text()').re(r'Name:\s*(.*)') - ['My image 1', - 'My image 2', - 'My image 3', - 'My image 4', - 'My image 5'] +>>> response.xpath('//a[contains(@href, "image")]/text()').re(r'Name:\s*(.*)') +['My image 1', + 'My image 2', + 'My image 3', + 'My image 4', + 'My image 5'] There's an additional helper reciprocating ``.get()`` (and its alias ``.extract_first()``) for ``.re()``, named ``.re_first()``. -Use it to extract just the first matching string:: +Use it to extract just the first matching string: - >>> response.xpath('//a[contains(@href, "image")]/text()').re_first(r'Name:\s*(.*)') - 'My image 1' +>>> response.xpath('//a[contains(@href, "image")]/text()').re_first(r'Name:\s*(.*)') +'My image 1' .. _old-extraction-api: @@ -422,28 +421,28 @@ and readable code. The following examples show how these methods map to each other. -1. ``SelectorList.get()`` is the same as ``SelectorList.extract_first()``:: +1. ``SelectorList.get()`` is the same as ``SelectorList.extract_first()``: - >>> response.css('a::attr(href)').get() - 'image1.html' - >>> response.css('a::attr(href)').extract_first() - 'image1.html' + >>> response.css('a::attr(href)').get() + 'image1.html' + >>> response.css('a::attr(href)').extract_first() + 'image1.html' -2. ``SelectorList.getall()`` is the same as ``SelectorList.extract()``:: +2. ``SelectorList.getall()`` is the same as ``SelectorList.extract()``: - >>> response.css('a::attr(href)').getall() - ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] - >>> response.css('a::attr(href)').extract() - ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] + >>> response.css('a::attr(href)').getall() + ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] + >>> response.css('a::attr(href)').extract() + ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] -3. ``Selector.get()`` is the same as ``Selector.extract()``:: +3. ``Selector.get()`` is the same as ``Selector.extract()``: - >>> response.css('a::attr(href)')[0].get() - 'image1.html' - >>> response.css('a::attr(href)')[0].extract() - 'image1.html' + >>> response.css('a::attr(href)')[0].get() + 'image1.html' + >>> response.css('a::attr(href)')[0].extract() + 'image1.html' -4. For consistency, there is also ``Selector.getall()``, which returns a list:: +4. For consistency, there is also ``Selector.getall()``, which returns a list: >>> response.css('a::attr(href)')[0].getall() ['image1.html'] @@ -481,26 +480,26 @@ with ``/``, that XPath will be absolute to the document and not relative to the ``Selector`` you're calling it from. For example, suppose you want to extract all ``<p>`` elements inside ``<div>`` -elements. First, you would get all ``<div>`` elements:: +elements. First, you would get all ``<div>`` elements: - >>> divs = response.xpath('//div') +>>> divs = response.xpath('//div') At first, you may be tempted to use the following approach, which is wrong, as it actually extracts all ``<p>`` elements from the document, not only those -inside ``<div>`` elements:: +inside ``<div>`` elements: - >>> for p in divs.xpath('//p'): # this is wrong - gets all <p> from the whole document - ... print(p.get()) +>>> for p in divs.xpath('//p'): # this is wrong - gets all <p> from the whole document +... print(p.get()) -This is the proper way to do it (note the dot prefixing the ``.//p`` XPath):: +This is the proper way to do it (note the dot prefixing the ``.//p`` XPath): - >>> for p in divs.xpath('.//p'): # extracts all <p> inside - ... print(p.get()) +>>> for p in divs.xpath('.//p'): # extracts all <p> inside +... print(p.get()) -Another common case would be to extract all direct ``<p>`` children:: +Another common case would be to extract all direct ``<p>`` children: - >>> for p in divs.xpath('p'): - ... print(p.get()) +>>> for p in divs.xpath('p'): +... print(p.get()) For more details about relative XPaths see the `Location Paths`_ section in the XPath specification. @@ -521,12 +520,12 @@ for that you may end up with more elements that you want, if they have a differe class name that shares the string ``someclass``. As it turns out, Scrapy selectors allow you to chain selectors, so most of the time -you can just select by class using CSS and then switch to XPath when needed:: +you can just select by class using CSS and then switch to XPath when needed: - >>> from scrapy import Selector - >>> sel = Selector(text='<div class="hero shout"><time datetime="2014-07-23 19:00">Special date</time></div>') - >>> sel.css('.shout').xpath('./time/@datetime').getall() - ['2014-07-23 19:00'] +>>> from scrapy import Selector +>>> sel = Selector(text='<div class="hero shout"><time datetime="2014-07-23 19:00">Special date</time></div>') +>>> sel.css('.shout').xpath('./time/@datetime').getall() +['2014-07-23 19:00'] This is cleaner than using the verbose XPath trick shown above. Just remember to use the ``.`` in the XPath expressions that will follow. @@ -538,41 +537,41 @@ Beware of the difference between //node[1] and (//node)[1] ``(//node)[1]`` selects all the nodes in the document, and then gets only the first of them. -Example:: +Example: - >>> from scrapy import Selector - >>> sel = Selector(text=""" - ....: <ul class="list"> - ....: <li>1</li> - ....: <li>2</li> - ....: <li>3</li> - ....: </ul> - ....: <ul class="list"> - ....: <li>4</li> - ....: <li>5</li> - ....: <li>6</li> - ....: </ul>""") - >>> xp = lambda x: sel.xpath(x).getall() +>>> from scrapy import Selector +>>> sel = Selector(text=""" +....: <ul class="list"> +....: <li>1</li> +....: <li>2</li> +....: <li>3</li> +....: </ul> +....: <ul class="list"> +....: <li>4</li> +....: <li>5</li> +....: <li>6</li> +....: </ul>""") +>>> xp = lambda x: sel.xpath(x).getall() -This gets all first ``<li>`` elements under whatever it is its parent:: +This gets all first ``<li>`` elements under whatever it is its parent: - >>> xp("//li[1]") - ['<li>1</li>', '<li>4</li>'] +>>> xp("//li[1]") +['<li>1</li>', '<li>4</li>'] -And this gets the first ``<li>`` element in the whole document:: +And this gets the first ``<li>`` element in the whole document: - >>> xp("(//li)[1]") - ['<li>1</li>'] +>>> xp("(//li)[1]") +['<li>1</li>'] -This gets all first ``<li>`` elements under an ``<ul>`` parent:: +This gets all first ``<li>`` elements under an ``<ul>`` parent: - >>> xp("//ul/li[1]") - ['<li>1</li>', '<li>4</li>'] +>>> xp("//ul/li[1]") +['<li>1</li>', '<li>4</li>'] -And this gets the first ``<li>`` element under an ``<ul>`` parent in the whole document:: +And this gets the first ``<li>`` element under an ``<ul>`` parent in the whole document: - >>> xp("(//ul/li)[1]") - ['<li>1</li>'] +>>> xp("(//ul/li)[1]") +['<li>1</li>'] Using text nodes in a condition ------------------------------- @@ -584,34 +583,34 @@ This is because the expression ``.//text()`` yields a collection of text element And when a node-set is converted to a string, which happens when it is passed as argument to a string function like ``contains()`` or ``starts-with()``, it results in the text for the first element only. -Example:: +Example: - >>> from scrapy import Selector - >>> sel = Selector(text='<a href="#">Click here to go to the <strong>Next Page</strong></a>') +>>> from scrapy import Selector +>>> sel = Selector(text='<a href="#">Click here to go to the <strong>Next Page</strong></a>') -Converting a *node-set* to string:: +Converting a *node-set* to string: - >>> sel.xpath('//a//text()').getall() # take a peek at the node-set - ['Click here to go to the ', 'Next Page'] - >>> sel.xpath("string(//a[1]//text())").getall() # convert it to string - ['Click here to go to the '] +>>> sel.xpath('//a//text()').getall() # take a peek at the node-set +['Click here to go to the ', 'Next Page'] +>>> sel.xpath("string(//a[1]//text())").getall() # convert it to string +['Click here to go to the '] -A *node* converted to a string, however, puts together the text of itself plus of all its descendants:: +A *node* converted to a string, however, puts together the text of itself plus of all its descendants: - >>> sel.xpath("//a[1]").getall() # select the first node - ['<a href="#">Click here to go to the <strong>Next Page</strong></a>'] - >>> sel.xpath("string(//a[1])").getall() # convert it to string - ['Click here to go to the Next Page'] +>>> sel.xpath("//a[1]").getall() # select the first node +['<a href="#">Click here to go to the <strong>Next Page</strong></a>'] +>>> sel.xpath("string(//a[1])").getall() # convert it to string +['Click here to go to the Next Page'] -So, using the ``.//text()`` node-set won't select anything in this case:: +So, using the ``.//text()`` node-set won't select anything in this case: - >>> sel.xpath("//a[contains(.//text(), 'Next Page')]").getall() - [] +>>> sel.xpath("//a[contains(.//text(), 'Next Page')]").getall() +[] -But using the ``.`` to mean the node, works:: +But using the ``.`` to mean the node, works: - >>> sel.xpath("//a[contains(., 'Next Page')]").getall() - ['<a href="#">Click here to go to the <strong>Next Page</strong></a>'] +>>> sel.xpath("//a[contains(., 'Next Page')]").getall() +['<a href="#">Click here to go to the <strong>Next Page</strong></a>'] .. _`XPath string function`: https://www.w3.org/TR/xpath/#section-String-Functions @@ -627,17 +626,17 @@ some arguments in your queries with placeholders like ``?``, which are then substituted with values passed with the query. Here's an example to match an element based on its "id" attribute value, -without hard-coding it (that was shown previously):: +without hard-coding it (that was shown previously): - >>> # `$val` used in the expression, a `val` argument needs to be passed - >>> response.xpath('//div[@id=$val]/a/text()', val='images').get() - 'Name: My image 1 ' +>>> # `$val` used in the expression, a `val` argument needs to be passed +>>> response.xpath('//div[@id=$val]/a/text()', val='images').get() +'Name: My image 1 ' Here's another example, to find the "id" attribute of a ``<div>`` tag containing -five ``<a>`` children (here we pass the value ``5`` as an integer):: +five ``<a>`` children (here we pass the value ``5`` as an integer): - >>> response.xpath('//div[count(a)=$cnt]/@id', cnt=5).get() - 'images' +>>> response.xpath('//div[count(a)=$cnt]/@id', cnt=5).get() +'images' All variable references must have a binding value when calling ``.xpath()`` (otherwise you'll get a ``ValueError: XPath error:`` exception). @@ -687,19 +686,19 @@ You can see several namespace declarations including a default .. highlight:: python Once in the shell we can try selecting all ``<link>`` objects and see that it -doesn't work (because the Atom XML namespace is obfuscating those nodes):: +doesn't work (because the Atom XML namespace is obfuscating those nodes): - >>> response.xpath("//link") - [] +>>> response.xpath("//link") +[] But once we call the :meth:`Selector.remove_namespaces` method, all -nodes can be accessed directly by their names:: +nodes can be accessed directly by their names: - >>> response.selector.remove_namespaces() - >>> response.xpath("//link") - [<Selector xpath='//link' data='<link rel="alternate" type="text/html" h'>, - <Selector xpath='//link' data='<link rel="next" type="application/atom+'>, - ... +>>> response.selector.remove_namespaces() +>>> response.xpath("//link") +[<Selector xpath='//link' data='<link rel="alternate" type="text/html" h'>, + <Selector xpath='//link' data='<link rel="next" type="application/atom+'>, + ... If you wonder why the namespace removal procedure isn't always called by default instead of having to call it manually, this is because of two reasons, which, in order @@ -734,26 +733,25 @@ Regular expressions The ``test()`` function, for example, can prove quite useful when XPath's ``starts-with()`` or ``contains()`` are not sufficient. -Example selecting links in list item with a "class" attribute ending with a digit:: +Example selecting links in list item with a "class" attribute ending with a digit: - >>> from scrapy import Selector - >>> doc = u""" - ... <div> - ... <ul> - ... <li class="item-0"><a href="link1.html">first item</a></li> - ... <li class="item-1"><a href="link2.html">second item</a></li> - ... <li class="item-inactive"><a href="link3.html">third item</a></li> - ... <li class="item-1"><a href="link4.html">fourth item</a></li> - ... <li class="item-0"><a href="link5.html">fifth item</a></li> - ... </ul> - ... </div> - ... """ - >>> sel = Selector(text=doc, type="html") - >>> sel.xpath('//li//@href').getall() - ['link1.html', 'link2.html', 'link3.html', 'link4.html', 'link5.html'] - >>> sel.xpath('//li[re:test(@class, "item-\d$")]//@href').getall() - ['link1.html', 'link2.html', 'link4.html', 'link5.html'] - >>> +>>> from scrapy import Selector +>>> doc = u""" +... <div> +... <ul> +... <li class="item-0"><a href="link1.html">first item</a></li> +... <li class="item-1"><a href="link2.html">second item</a></li> +... <li class="item-inactive"><a href="link3.html">third item</a></li> +... <li class="item-1"><a href="link4.html">fourth item</a></li> +... <li class="item-0"><a href="link5.html">fifth item</a></li> +... </ul> +... </div> +... """ +>>> sel = Selector(text=doc, type="html") +>>> sel.xpath('//li//@href').getall() +['link1.html', 'link2.html', 'link3.html', 'link4.html', 'link5.html'] +>>> sel.xpath('//li[re:test(@class, "item-\d$")]//@href').getall() +['link1.html', 'link2.html', 'link4.html', 'link5.html'] .. warning:: C library ``libxslt`` doesn't natively support EXSLT regular expressions so `lxml`_'s implementation uses hooks to Python's ``re`` module. @@ -849,7 +847,6 @@ with groups of itemscopes and corresponding itemprops:: current scope: ['http://schema.org/Rating'] properties: ['worstRating', 'ratingValue', 'bestRating'] - >>> Here we first iterate over ``itemscope`` elements, and for each one, we look for all ``itemprops`` elements and exclude those that are themselves @@ -877,15 +874,15 @@ For the following HTML:: .. highlight:: python -You can use it like this:: +You can use it like this: - >>> response.xpath('//p[has-class("foo")]') - [<Selector xpath='//p[has-class("foo")]' data='<p class="foo bar-baz">First</p>'>, - <Selector xpath='//p[has-class("foo")]' data='<p class="foo">Second</p>'>] - >>> response.xpath('//p[has-class("foo", "bar-baz")]') - [<Selector xpath='//p[has-class("foo", "bar-baz")]' data='<p class="foo bar-baz">First</p>'>] - >>> response.xpath('//p[has-class("foo", "bar")]') - [] +>>> response.xpath('//p[has-class("foo")]') +[<Selector xpath='//p[has-class("foo")]' data='<p class="foo bar-baz">First</p>'>, + <Selector xpath='//p[has-class("foo")]' data='<p class="foo">Second</p>'>] +>>> response.xpath('//p[has-class("foo", "bar-baz")]') +[<Selector xpath='//p[has-class("foo", "bar-baz")]' data='<p class="foo bar-baz">First</p>'>] +>>> response.xpath('//p[has-class("foo", "bar")]') +[] So XPath ``//p[has-class("foo", "bar-baz")]`` is roughly equivalent to CSS ``p.foo.bar-baz``. Please note, that it is slower in most of the cases, diff --git a/docs/topics/shell.rst b/docs/topics/shell.rst index 68a0b19b5..c1fdfd221 100644 --- a/docs/topics/shell.rst +++ b/docs/topics/shell.rst @@ -177,47 +177,46 @@ all start with the ``[s]`` prefix):: >>> -After that, we can start playing with the objects:: +After that, we can start playing with the objects: - >>> response.xpath('//title/text()').get() - 'Scrapy | A Fast and Powerful Scraping and Web Crawling Framework' +>>> response.xpath('//title/text()').get() +'Scrapy | A Fast and Powerful Scraping and Web Crawling Framework' - >>> fetch("https://reddit.com") +>>> fetch("https://reddit.com") - >>> response.xpath('//title/text()').get() - 'reddit: the front page of the internet' +>>> response.xpath('//title/text()').get() +'reddit: the front page of the internet' - >>> request = request.replace(method="POST") +>>> request = request.replace(method="POST") - >>> fetch(request) +>>> fetch(request) - >>> response.status - 404 +>>> response.status +404 - >>> from pprint import pprint +>>> from pprint import pprint - >>> pprint(response.headers) - {'Accept-Ranges': ['bytes'], - 'Cache-Control': ['max-age=0, must-revalidate'], - 'Content-Type': ['text/html; charset=UTF-8'], - 'Date': ['Thu, 08 Dec 2016 16:21:19 GMT'], - 'Server': ['snooserv'], - 'Set-Cookie': ['loid=KqNLou0V9SKMX4qb4n; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', - 'loidcreated=2016-12-08T16%3A21%3A19.445Z; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', - 'loid=vi0ZVe4NkxNWdlH7r7; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', - 'loidcreated=2016-12-08T16%3A21%3A19.459Z; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure'], - 'Vary': ['accept-encoding'], - 'Via': ['1.1 varnish'], - 'X-Cache': ['MISS'], - 'X-Cache-Hits': ['0'], - 'X-Content-Type-Options': ['nosniff'], - 'X-Frame-Options': ['SAMEORIGIN'], - 'X-Moose': ['majestic'], - 'X-Served-By': ['cache-cdg8730-CDG'], - 'X-Timer': ['S1481214079.394283,VS0,VE159'], - 'X-Ua-Compatible': ['IE=edge'], - 'X-Xss-Protection': ['1; mode=block']} - >>> +>>> pprint(response.headers) +{'Accept-Ranges': ['bytes'], + 'Cache-Control': ['max-age=0, must-revalidate'], + 'Content-Type': ['text/html; charset=UTF-8'], + 'Date': ['Thu, 08 Dec 2016 16:21:19 GMT'], + 'Server': ['snooserv'], + 'Set-Cookie': ['loid=KqNLou0V9SKMX4qb4n; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', + 'loidcreated=2016-12-08T16%3A21%3A19.445Z; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', + 'loid=vi0ZVe4NkxNWdlH7r7; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', + 'loidcreated=2016-12-08T16%3A21%3A19.459Z; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure'], + 'Vary': ['accept-encoding'], + 'Via': ['1.1 varnish'], + 'X-Cache': ['MISS'], + 'X-Cache-Hits': ['0'], + 'X-Content-Type-Options': ['nosniff'], + 'X-Frame-Options': ['SAMEORIGIN'], + 'X-Moose': ['majestic'], + 'X-Served-By': ['cache-cdg8730-CDG'], + 'X-Timer': ['S1481214079.394283,VS0,VE159'], + 'X-Ua-Compatible': ['IE=edge'], + 'X-Xss-Protection': ['1; mode=block']} .. _topics-shell-inspect-response: @@ -263,16 +262,16 @@ When you run the spider, you will get something similar to this:: >>> response.url 'http://example.org' -Then, you can check if the extraction code is working:: +Then, you can check if the extraction code is working: - >>> response.xpath('//h1[@class="fn"]') - [] +>>> response.xpath('//h1[@class="fn"]') +[] Nope, it doesn't. So you can open the response in your web browser and see if -it's the response you were expecting:: +it's the response you were expecting: - >>> view(response) - True +>>> view(response) +True Finally you hit Ctrl-D (or Ctrl-Z in Windows) to exit the shell and resume the crawling:: diff --git a/docs/topics/stats.rst b/docs/topics/stats.rst index 38648ec55..3dd829ebe 100644 --- a/docs/topics/stats.rst +++ b/docs/topics/stats.rst @@ -57,15 +57,15 @@ Set stat value only if lower than previous:: stats.min_value('min_free_memory_percent', value) -Get stat value:: +Get stat value: - >>> stats.get_value('custom_count') - 1 +>>> stats.get_value('custom_count') +1 -Get all stats:: +Get all stats: - >>> stats.get_stats() - {'custom_count': 1, 'start_time': datetime.datetime(2009, 7, 14, 21, 47, 28, 977139)} +>>> stats.get_stats() +{'custom_count': 1, 'start_time': datetime.datetime(2009, 7, 14, 21, 47, 28, 977139)} Available Stats Collectors ==========================