mirror of https://github.com/scrapy/scrapy.git
change tutorial to follow changes on dmoz site
This commit is contained in:
parent
98f3f87530
commit
bcb31988f2
|
|
@ -268,19 +268,19 @@ The shell also instantiates two selectors, one for HTML (in the ``hxs``
|
|||
variable) and one for XML (in the ``xxs`` variable) with this response. So let's
|
||||
try them::
|
||||
|
||||
In [1]: hxs.select('/html/head/title')
|
||||
Out[1]: [<HtmlXPathSelector (title) xpath=/html/head/title>]
|
||||
In [1]: hxs.select('//title')
|
||||
Out[1]: [<HtmlXPathSelector (title) xpath=//title>]
|
||||
|
||||
In [2]: hxs.select('/html/head/title').extract()
|
||||
In [2]: hxs.select('//title').extract()
|
||||
Out[2]: [u'<title>Open Directory - Computers: Programming: Languages: Python: Books</title>']
|
||||
|
||||
In [3]: hxs.select('/html/head/title/text()')
|
||||
Out[3]: [<HtmlXPathSelector (text) xpath=/html/head/title/text()>]
|
||||
In [3]: hxs.select('//title/text()')
|
||||
Out[3]: [<HtmlXPathSelector (text) xpath=//title/text()>]
|
||||
|
||||
In [4]: hxs.select('/html/head/title/text()').extract()
|
||||
In [4]: hxs.select('//title/text()').extract()
|
||||
Out[4]: [u'Open Directory - Computers: Programming: Languages: Python: Books']
|
||||
|
||||
In [5]: hxs.select('/html/head/title/text()').re('(\w+):')
|
||||
In [5]: hxs.select('//title/text()').re('(\w+):')
|
||||
Out[5]: [u'Computers', u'Programming', u'Languages', u'Python']
|
||||
|
||||
Extracting the data
|
||||
|
|
|
|||
Loading…
Reference in New Issue