diff --git a/docs/topics/link-extractors.rst b/docs/topics/link-extractors.rst index 79af6db35..1a3a2a8b7 100644 --- a/docs/topics/link-extractors.rst +++ b/docs/topics/link-extractors.rst @@ -16,7 +16,7 @@ The only public method that every LinkExtractor has is ``extract_links``, which receives a :class:`~scrapy.http.Response` object and returns a list of :class:`scrapy.link.Link` objects. Link Extractors are meant to be instantiated once and their ``extract_links`` method called several times with different responses, to -extract links to follow. +extract links to follow. Link extractors are used in the :class:`~scrapy.contrib.spiders.CrawlSpider` class (available in Scrapy), through a set of rules, but you can also use it in @@ -36,16 +36,82 @@ Built-in link extractors reference All available link extractors classes bundled with Scrapy are provided in the :mod:`scrapy.contrib.linkextractors` module. +.. module:: scrapy.contrib.linkextractors.lxmlhtml + :synopsis: lxml's HTMLParser-based link extractors + .. module:: scrapy.contrib.linkextractors.sgml :synopsis: SGMLParser-based link extractors +LxmlLinkExtractor +----------------- + +.. class:: LxmlLinkExtractor(allow=(), deny=(), allow_domains=(), deny_domains=(), deny_extensions=None, restrict_xpaths=(), tags=('a', 'area'), attrs=('href',), canonicalize=True, unique=True, process_value=None) + + LxmlLinkExtractor provides the same API as :class:`SgmlLinkExtractor`, + and therefore has the same handy filtering options as constructor parameters, + but the implementation underneath uses lxml's more robust ``HTMLParser`` + to parse HTML documents: + + :param allow: a single regular expression (or list of regular expressions) + that the (absolute) urls must match in order to be extracted. If not + given (or empty), it will match all links. + :type allow: a regular expression (or list of) + + :param deny: a single regular expression (or list of regular expressions) + that the (absolute) urls must match in order to be excluded (ie. not + extracted). It has precedence over the ``allow`` parameter. If not + given (or empty) it won't exclude any links. + :type deny: a regular expression (or list of) + + :param allow_domains: a single value or a list of string containing + domains which will be considered for extracting the links + :type allow_domains: str or list + + :param deny_domains: a single value or a list of strings containing + domains which won't be considered for extracting the links + :type deny_domains: str or list + + :param deny_extensions: a single value or list of strings containing + extensions that should be ignored when extracting links. + If not given, it will default to the + ``IGNORED_EXTENSIONS`` list defined in the `scrapy.linkextractor`_ + module. + :type deny_extensions: list + + :param restrict_xpaths: is a XPath (or list of XPath's) which defines + regions inside the response where links should be extracted from. + If given, only the text selected by those XPath will be scanned for + links. See examples below. + :type restrict_xpaths: str or list + + :param tags: a tag or a list of tags to consider when extracting links. + Defaults to ``('a', 'area')``. + :type tags: str or list + + :param attrs: an attribute or list of attributes which should be considered when looking + for links to extract (only for those tags specified in the ``tags`` + parameter). Defaults to ``('href',)`` + :type attrs: list + + :param canonicalize: canonicalize each extracted url (using + scrapy.utils.url.canonicalize_url). Defaults to ``True``. + :type canonicalize: boolean + + :param unique: whether duplicate filtering should be applied to extracted + links. + :type unique: boolean + + :param process_value: see ``process_value`` argument of + :class:`BaseSgmlLinkExtractor` class constructor + :type process_value: callable + SgmlLinkExtractor ----------------- .. class:: SgmlLinkExtractor(allow=(), deny=(), allow_domains=(), deny_domains=(), deny_extensions=None, restrict_xpaths=(), tags=('a', 'area'), attrs=('href'), canonicalize=True, unique=True, process_value=None) - The SgmlLinkExtractor extends the base :class:`BaseSgmlLinkExtractor` by - providing additional filters that you can specify to extract links, + The SgmlLinkExtractor is built upon the base :class:`BaseSgmlLinkExtractor` + and provides additional filters that you can specify to extract links, including regular expressions patterns that the links must match to be extracted. All those filters are configured through these constructor parameters: @@ -70,14 +136,14 @@ SgmlLinkExtractor :type deny_domains: str or list :param deny_extensions: a single value or list of strings containing - extensions that should be ignored when extracting links. + extensions that should be ignored when extracting links. If not given, it will default to the ``IGNORED_EXTENSIONS`` list defined in the `scrapy.linkextractor`_ module. :type deny_extensions: list :param restrict_xpaths: is a XPath (or list of XPath's) which defines - regions inside the response where links should be extracted from. + regions inside the response where links should be extracted from. If given, only the text selected by those XPath will be scanned for links. See examples below. :type restrict_xpaths: str or list @@ -110,7 +176,7 @@ BaseSgmlLinkExtractor The purpose of this Link Extractor is only to serve as a base class for the :class:`SgmlLinkExtractor`. You should use that one instead. - + The constructor arguments are: :param tag: either a string (with the name of a tag) or a function that @@ -140,15 +206,15 @@ BaseSgmlLinkExtractor For example, to extract links from this code:: Link text - + .. highlight:: python You can use the following function in ``process_value``:: - + def process_value(value): m = re.search("javascript:goToPage\('(.*?)'", value) if m: - return m.group(1) + return m.group(1) :type process_value: callable