mirror of https://github.com/scrapy/scrapy.git
some improvements to XMLFeedSpider including namespaces patch submitted by talboito, refs #81
--HG-- extra : convert_revision : svn%3Ab85faa78-f9eb-468e-a121-7cced6da292c%401096
This commit is contained in:
parent
560dd3fd8e
commit
f4e7ee5dac
|
|
@ -191,63 +191,84 @@ XMLFeedSpider
|
|||
|
||||
.. class:: XMLFeedSpider
|
||||
|
||||
XMLFeedSpider is designed for parsing XML feeds by iterating through them by a
|
||||
certain node name. The iterator can be chosen from: ``iternodes``, ``xml``,
|
||||
and ``html``. It's recommended to use the ``iternodes`` iterator for
|
||||
performance reasons, since the ``xml`` and ``html`` iterators generate the
|
||||
whole DOM at once in order to parse it. However, using ``html`` as the
|
||||
iterator may be useful when parsing XML with bad markup.
|
||||
XMLFeedSpider is designed for parsing XML feeds by iterating through them by a
|
||||
certain node name. The iterator can be chosen from: ``iternodes``, ``xml``,
|
||||
and ``html``. It's recommended to use the ``iternodes`` iterator for
|
||||
performance reasons, since the ``xml`` and ``html`` iterators generate the
|
||||
whole DOM at once in order to parse it. However, using ``html`` as the
|
||||
iterator may be useful when parsing XML with bad markup.
|
||||
|
||||
For setting the iterator and the tag name, you must define the following class
|
||||
attributes:
|
||||
For setting the iterator and the tag name, you must define the following class
|
||||
attributes:
|
||||
|
||||
.. attribute:: XMLFeedSpider.iterator
|
||||
.. attribute:: iterator
|
||||
|
||||
A string which defines the iterator to use. It can be either:
|
||||
A string which defines the iterator to use. It can be either:
|
||||
|
||||
- ``'iternodes'`` - a fast iterator based on regular expressions
|
||||
- ``'iternodes'`` - a fast iterator based on regular expressions
|
||||
|
||||
- ``'html'`` - an iterator which uses HtmlXPathSelector. Keep in mind
|
||||
this uses DOM parsing and must load all DOM in memory which could be a
|
||||
problem for big feeds
|
||||
- ``'html'`` - an iterator which uses HtmlXPathSelector. Keep in mind
|
||||
this uses DOM parsing and must load all DOM in memory which could be a
|
||||
problem for big feeds
|
||||
|
||||
- ``'xml'`` - an iterator which uses XmlXPathSelector. Keep in mind
|
||||
this uses DOM parsing and must load all DOM in memory which could be a
|
||||
problem for big feeds
|
||||
- ``'xml'`` - an iterator which uses XmlXPathSelector. Keep in mind
|
||||
this uses DOM parsing and must load all DOM in memory which could be a
|
||||
problem for big feeds
|
||||
|
||||
It defaults to: ``'iternodes'``.
|
||||
It defaults to: ``'iternodes'``.
|
||||
|
||||
.. attribute:: XMLFeedSpider.itertag
|
||||
.. attribute:: itertag
|
||||
|
||||
A stirng with the name of the node (or element) to iterate in.
|
||||
A string with the name of the node (or element) to iterate in. Example::
|
||||
|
||||
Apart from these new attributes, this spider has the following overrideable
|
||||
methods too:
|
||||
itertag = 'product'
|
||||
|
||||
.. method:: XMLFeedSpider.adapt_response(response)
|
||||
.. attribute:: namespaces
|
||||
|
||||
A method that receives the response as soon as it arrives from the spider
|
||||
middleware and before start parsing it. It can be used used for modifying
|
||||
the response body before parsing it. This method receives a response and
|
||||
returns response (it could be the same or another one).
|
||||
A list of ``(prefix, uri)`` tuples which define the namespaces
|
||||
available in that document that will be processed with this spider. The
|
||||
``prefix`` and ``uri`` will be used to automatically register
|
||||
namespaces using the
|
||||
:meth:`~scrapy.xpath.XPathSelector.register_namespace` method.
|
||||
|
||||
.. method:: XMLFeedSpider.parse_item(response, selector)
|
||||
|
||||
This method is called for the nodes matching the provided tag name
|
||||
(``itertag``). Receives the response and an XPathSelector for each node.
|
||||
Overriding this method is mandatory. Otherwise, you spider won't work.
|
||||
This method must return either a ScrapedItem, a Request, or a list
|
||||
containing any of them.
|
||||
You can then specify nodes with namespaces in the :attr:`itertag`
|
||||
attribute.
|
||||
|
||||
.. warning:: This method will soon change its name to ``parse_node``
|
||||
Example::
|
||||
|
||||
class YourSpider(XMLFeedSpider):
|
||||
|
||||
.. method:: XMLFeedSpider.process_results(response, results)
|
||||
|
||||
This method is called for each result (item or request) returned by the
|
||||
spider, and it's intended to perform any last time processing required
|
||||
before returning the results to the framework core, for example setting the
|
||||
item IDs. It receives a list of results and the response which originated
|
||||
that results. It must return a list of results (Items or Requests)."""
|
||||
namespaces = [('n', 'http://www.sitemaps.org/schemas/sitemap/0.9')]
|
||||
itertag = 'n:url'
|
||||
# ...
|
||||
|
||||
Apart from these new attributes, this spider has the following overrideable
|
||||
methods too:
|
||||
|
||||
.. method:: adapt_response(response)
|
||||
|
||||
A method that receives the response as soon as it arrives from the spider
|
||||
middleware and before start parsing it. It can be used used for modifying
|
||||
the response body before parsing it. This method receives a response and
|
||||
returns response (it could be the same or another one).
|
||||
|
||||
.. method:: parse_item(response, selector)
|
||||
|
||||
This method is called for the nodes matching the provided tag name
|
||||
(``itertag``). Receives the response and an XPathSelector for each node.
|
||||
Overriding this method is mandatory. Otherwise, you spider won't work.
|
||||
This method must return either a ScrapedItem, a Request, or a list
|
||||
containing any of them.
|
||||
|
||||
.. warning:: This method will soon change its name to ``parse_node``
|
||||
|
||||
.. method:: process_results(response, results)
|
||||
|
||||
This method is called for each result (item or request) returned by the
|
||||
spider, and it's intended to perform any last time processing required
|
||||
before returning the results to the framework core, for example setting the
|
||||
item IDs. It receives a list of results and the response which originated
|
||||
that results. It must return a list of results (Items or Requests)."""
|
||||
|
||||
|
||||
XMLFeedSpider example
|
||||
|
|
|
|||
|
|
@ -24,6 +24,7 @@ class XMLFeedSpider(BaseSpider):
|
|||
|
||||
iterator = 'iternodes'
|
||||
itertag = 'item'
|
||||
namespaces = ()
|
||||
|
||||
def process_results(self, response, results):
|
||||
"""This overridable method is called for each result (item or request)
|
||||
|
|
@ -42,6 +43,10 @@ class XMLFeedSpider(BaseSpider):
|
|||
"""
|
||||
return response
|
||||
|
||||
def parse_item(self, response, selector):
|
||||
"""This method must be overriden with your custom spider functionality"""
|
||||
raise NotImplemented
|
||||
|
||||
def parse_nodes(self, response, nodes):
|
||||
"""This method is called for the nodes matching the provided tag name
|
||||
(itertag). Receives the response and an XPathSelector for each node.
|
||||
|
|
@ -67,14 +72,22 @@ class XMLFeedSpider(BaseSpider):
|
|||
if self.iterator == 'iternodes':
|
||||
nodes = xmliter(response, self.itertag)
|
||||
elif self.iterator == 'xml':
|
||||
nodes = XmlXPathSelector(response).x('//%s' % self.itertag)
|
||||
selector = XmlXPathSelector(response)
|
||||
self._register_namespaces(selector)
|
||||
nodes = selector.x('//%s' % self.itertag)
|
||||
elif self.iterator == 'html':
|
||||
nodes = HtmlXPathSelector(response).x('//%s' % self.itertag)
|
||||
selector = HtmlXPathSelector(response)
|
||||
self._register_namespaces(selector)
|
||||
nodes = selector.x('//%s' % self.itertag)
|
||||
else:
|
||||
raise NotSupported('Unsupported node iterator')
|
||||
|
||||
return self.parse_nodes(response, nodes)
|
||||
|
||||
def _register_namespaces(self, selector):
|
||||
for (prefix, uri) in self.namespaces:
|
||||
selector.register_namespace(prefix, uri)
|
||||
|
||||
class CSVFeedSpider(BaseSpider):
|
||||
"""Spider for parsing CSV feeds.
|
||||
It receives a CSV file in a response; iterates through each of its rows,
|
||||
|
|
@ -95,6 +108,10 @@ class CSVFeedSpider(BaseSpider):
|
|||
"""This method has the same purpose as the one in XMLFeedSpider"""
|
||||
return response
|
||||
|
||||
def parse_row(self, row):
|
||||
"""This method must be overriden with your custom spider functionality"""
|
||||
raise NotImplemented
|
||||
|
||||
def parse_rows(self, response):
|
||||
"""Receives a response and a dict (representing each row) with a key for
|
||||
each provided (or detected) header of the CSV file. This spider also
|
||||
|
|
|
|||
Loading…
Reference in New Issue