diff --git a/docs/intro/overview.rst b/docs/intro/overview.rst index a199b1baf..d10c20138 100644 --- a/docs/intro/overview.rst +++ b/docs/intro/overview.rst @@ -50,7 +50,7 @@ that to construct the regular expression for the links to follow: ``/tor/\d+``. For extracting data we'll use `XPath`_ to select the part of the document where the data is to be extracted. Let's take one of those torrent pages: - http://www.mininova.org/tor/2004522 + http://www.mininova.org/tor/2657665 .. _XPath: http://www.w3.org/TR/xpath @@ -62,7 +62,7 @@ want to extract which is: torrent name, description and size. By looking at the page HTML source we can see that the file name is contained inside a ``
`` tag inside the ``
Category: - Movies > Action + Movies > Documentary
Total size: - 801.44 megabyte
+ 699.79 megabyte + .. highlight:: none An XPath expression to select the description could be:: - //div[@id='info-left']/p[2]/text()[2] + //div[@id='specifications']/p[2]/text()[2] .. highlight:: python