mirror of https://github.com/scrapy/scrapy.git
Added documentation for Items and Loaders, removed obsolete Item Adaptors documentation
--HG-- rename : docs/experimental/topics/newitem/index.rst => docs/experimental/newitem.rst
This commit is contained in:
parent
efa08318be
commit
d95e99f585
|
|
@ -3,26 +3,21 @@
|
|||
Experimental features
|
||||
=====================
|
||||
|
||||
This section documents experimental features that may become stable in the
|
||||
future, but whose API is not yet stable. Use them with caution, and subscribe
|
||||
to the `mailing lists <http://scrapy.org/community/>`_ to be notified of any
|
||||
changes.
|
||||
This section documents experimental Scrapy features that may become stable in
|
||||
future releases, but whose API is not yet stable. Use them with caution, and
|
||||
subscribe to the `mailing lists <http://scrapy.org/community/>`_ to get
|
||||
notified of any changes.
|
||||
|
||||
This section may also contain documentation which is outdated or incomplete (as
|
||||
it's not revised so frequently), or documentation which overlaps with existing
|
||||
(stable) documentation and needs to be manually merged, use at your own risk.
|
||||
Since it's not revised so frequently, this section may contain documentation
|
||||
which is outdated, incomplete or overlapping with stable documentation (until
|
||||
it's properly merged) . Use at your own risk.
|
||||
|
||||
.. warning::
|
||||
|
||||
This documentation is a work in progress, use at your own risk.
|
||||
This documentation is a work in progress. Use at your own risk.
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 1
|
||||
|
||||
topics/index
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 2
|
||||
|
||||
ref/index
|
||||
|
||||
newitem
|
||||
newitem-loader
|
||||
|
|
|
|||
|
|
@ -0,0 +1,357 @@
|
|||
.. _topics-loader:
|
||||
|
||||
============
|
||||
Item Loaders
|
||||
============
|
||||
|
||||
.. module:: scrapy.newitem.loader
|
||||
:synopsis: Item Loader class
|
||||
|
||||
Item Loaders (or Loaders, for short) provide a convenient mechanism for
|
||||
populating scraped :ref:`Items <topics-newitem>`. Even though Items can be
|
||||
populated using their own dictionary-like API, the Loaders provide a much more
|
||||
convenient API for populating them from a scraping process, by automating some
|
||||
common tasks like parsing the raw extracted data before assigning it.
|
||||
|
||||
In other words, :ref:`Items <topics-newitem>` provide the *container* of
|
||||
scraped data, while Loaders provide the mechanism for *populating* that
|
||||
container.
|
||||
|
||||
Loaders are designed to provide a flexible, efficient and easy mechanism for
|
||||
extending and overriding different field parsing rules (either by spider, or by
|
||||
source format) without becoming a nightmare to maintain
|
||||
|
||||
Using Loaders to populate items
|
||||
===============================
|
||||
|
||||
To use a Loader, you must first instantiate it. You can either instantiate it
|
||||
with an Item object or without one, in which case an Item is automatically
|
||||
instantiated in the Loader constructor using the Item class specified in the
|
||||
:attr:`ItemLoader.default_item_class` attribute.
|
||||
|
||||
Then, you start adding values to the Loader, typically collecting them using
|
||||
:ref:`Selectors <topics-selectors>`. You can add more than one value to the
|
||||
same item field, the Loader will know how to "join" those values later using a
|
||||
Reducer.
|
||||
|
||||
Here's a typical Loader usage in a :ref:`Spider <topics-spiders>` the
|
||||
:ref:`Product item defined in the Items chapter <topics-newitem-declaring>`.::
|
||||
|
||||
from scrapy.item.loader import Loader
|
||||
from scrapy.xpath import HtmlXPathSelector
|
||||
from myproject.items import Product
|
||||
|
||||
def parse(self, response):
|
||||
x = HtmlXPathSelector(response)
|
||||
l = ItemLoader(item=Product())
|
||||
l.add_value('name', x.x('//div[@class="product_name"]').extract())
|
||||
l.add_value('name', x.x('//div[@class="product_title"]').extract())
|
||||
l.add_value('price', x.x('//p[@id="price"]').extract())
|
||||
l.add_value('stock', x.x('//p[@id="stock"]').extract())
|
||||
l.add_value('last_updated', 'today')
|
||||
return l.get_item()
|
||||
|
||||
By looking at that code we can see the ``name`` field is being extracted from
|
||||
two different XPath locations in the page:
|
||||
|
||||
* ``//div[@class="product_name"]``
|
||||
* ``//div[@class="product_title"]``
|
||||
|
||||
So both XPaths are used for extracting data from the page, and the data
|
||||
returned by them is collected to be assigned to the ``name`` attribute.
|
||||
|
||||
Afterwards, similar calls are used for ``price`` and ``stock`` fields, and
|
||||
finally the ``last_update`` field is populated directly with a literal value
|
||||
(``today``).
|
||||
|
||||
Finally, when all data is collected, the :meth:`ItemLoader.get_item` method is
|
||||
called which actually populates and returns the item populated with the data
|
||||
previously extracted with the ``add_value`` calls.
|
||||
|
||||
Expanders and Reducers
|
||||
======================
|
||||
|
||||
A Loader is composed of one expander and reducer for each item field. The
|
||||
Expander processes the extracted data as soon as it's received through the
|
||||
:meth:`ItemLoader.add_value` method, and the result of the expander is collected
|
||||
and kept inside the Loader. After collecting all data, the
|
||||
:meth:`ItemLoader.get_item` method is called to actually populate and get the Item.
|
||||
That's when the Reducers are called with the data previously collected (using
|
||||
the Expanders) and the output of the Reducers are the actual values that get
|
||||
assigned to the item.
|
||||
|
||||
Let's see an example to illustrate how Expanders and Reducers are called, for a
|
||||
particular field (the same applies for any other field)::
|
||||
|
||||
l = ItemLoader(Product())
|
||||
l.add_value('name', x.x(xpath1).extract()) # (1)
|
||||
l.add_value('name', x.x(xpath2).extract()) # (2)
|
||||
return l.get_item() # (3)
|
||||
|
||||
So what happens is:
|
||||
|
||||
1. Data from ``xpath1`` is extracted, and passed through the Expander of the
|
||||
``name`` field. The output of the expander is collected and kept in the
|
||||
loader (but not yet assigned to the item).
|
||||
|
||||
2. Data from ``xpath2`` is extracted, and passed through the same Expander used
|
||||
in (1). The output of the expander is appended to the data collected in (1)
|
||||
(if any).
|
||||
|
||||
3. The data collected in (1) and (2) is passed through the Reducer of the
|
||||
``name`` field. The output of the Reducer is the value assigned to the
|
||||
``name`` field in the item.
|
||||
|
||||
Scrapy comes with one major expander built-in, the :ref:`Tree Expander
|
||||
<topics-loader-tree-expander>`, and :ref:`a couple of commonly used reducers
|
||||
<topics-loader-reducers>`.
|
||||
|
||||
Declaring Loaders
|
||||
=================
|
||||
|
||||
Loaders are declared like Items, by using a class definition syntax. Here is an
|
||||
example::
|
||||
|
||||
from scrapy.newitem.loader import ItemLoader
|
||||
from scrapy.newitem.loader.expanders import TreeExpander
|
||||
from scrapy.newitem.loader.reducers import Join, TakeFirst
|
||||
|
||||
class ProductLoader(ItemLoader):
|
||||
|
||||
default_expander = TakeFirst()
|
||||
|
||||
name_exp = TreeExpander(unicode.title)
|
||||
name_red = Join()
|
||||
|
||||
price_exp = TreeExpander(unicode.strip)
|
||||
price_red = TakeFirst()
|
||||
|
||||
...
|
||||
|
||||
As you can see, expanders are declared using the ``_exp`` suffix while reducers
|
||||
are declared using the ``_red`` suffix. And you can also declare a default
|
||||
expander using the :attr:`ItemLoader.default_expander` attribute.
|
||||
|
||||
.. _topics-loader-args:
|
||||
|
||||
Item Loader arguments
|
||||
=====================
|
||||
|
||||
The Loader arguments is a dict of arbitrary key/values which can be passed when
|
||||
declaring, instantiating or using Loaders. They are used modify the behaviour
|
||||
of the expanders.
|
||||
|
||||
For example, suppose you have a function ``parse_length`` which receives a text
|
||||
value and extracts a length from it::
|
||||
|
||||
def parse_length(text, loader_args):
|
||||
unit = loader_args('unit', 'm')
|
||||
# ... length parsing code goes here ...
|
||||
return parsed_length
|
||||
|
||||
Since it receives a ``loader_args`` the Expander will pass the currently active
|
||||
loader arguments when calling it.
|
||||
|
||||
There are seveal ways to pass loader arguments:
|
||||
|
||||
1. Passing arguments on Loader declaration::
|
||||
|
||||
class ProductLoader(ItemLoader):
|
||||
length_exp = TreeExpander(parse_length, unit='cm')
|
||||
|
||||
2. Passing arguments on Loader instantiation::
|
||||
|
||||
l = ItemLoader(product, unit='cm')
|
||||
|
||||
3. Passing arguments on Loader usage::
|
||||
|
||||
l.add_value('length', x.x('//div').extract(), unit='cm')
|
||||
|
||||
ItemLoader objects
|
||||
==================
|
||||
|
||||
.. class:: ItemLoader([item], \**loader_args)
|
||||
|
||||
Return a new Item Loader for populating the given Item. If no item is
|
||||
given, one is instantiated using the class in :attr:`default_item_class`.
|
||||
|
||||
.. method:: add_value(field_name, value, \**new_loader_args)
|
||||
|
||||
Add the given ``value`` for the given field.
|
||||
|
||||
The value is passed through the field expander and its output appened
|
||||
to the data collected for that field. If the field already contains
|
||||
collected data, the new data is added.
|
||||
|
||||
If any keyword arguments are passed, they're used as :ref:`Loader
|
||||
arguments <topics-loader-args>` when calling the expanders.
|
||||
|
||||
.. method:: replace_value(field_name, value, \**new_loader_args)
|
||||
|
||||
Similar to :meth:`add_value` but replaces collected data instead of
|
||||
adding it.
|
||||
|
||||
|
||||
.. method:: get_item()
|
||||
|
||||
Populate the item with the data collected so far, and return it.
|
||||
|
||||
.. method:: get_expanded_value(field_name)
|
||||
|
||||
Return the expanded data for the given field. In other words, return
|
||||
the dat collected so far for the given field, without reducing it.
|
||||
|
||||
.. method:: get_reduced_value(field_name)
|
||||
|
||||
Return the reduced value for the given field, without modifying the
|
||||
item.
|
||||
|
||||
.. method:: get_expander(field_name)
|
||||
|
||||
Return the expander for the given field.
|
||||
|
||||
.. method:: get_reducer(field_name)
|
||||
|
||||
Return the reducer for the given field.
|
||||
|
||||
.. attribute:: default_item_class
|
||||
|
||||
An Item class (or factory), used to instantiate items when not given in
|
||||
the constructor.
|
||||
|
||||
.. attribute:: default_expander
|
||||
|
||||
The default expander to use for those fields which don't define a
|
||||
specific expander
|
||||
|
||||
.. attribute:: default_reducer
|
||||
|
||||
The default reducer to use for those fields which don't define a
|
||||
specific expander
|
||||
|
||||
Reusing and extending Loaders
|
||||
=============================
|
||||
|
||||
As your project grows bigger and acquires more and more spiders, maintenance
|
||||
becomes a fundamental problem, specially when you have to deal with many
|
||||
different parsing rules per spider, a lot of exceptions, but also want to reuse
|
||||
the common cases.
|
||||
|
||||
Loaders are designed to ease the maintenance of parsing rules, without loosing
|
||||
flexibility and, at the same time, providing a convenient mechanism for
|
||||
extending and overriding them. For this reason Loaders support traditional
|
||||
class inheritance for for dealing with differences of specific spiders (or
|
||||
group of spiders).
|
||||
|
||||
Suppose, for example, that some particular site encloses their product names
|
||||
between three dashes (ie. ``---Plasma TV---``) and you don't want to end up
|
||||
scraping those dashes in the final product names.
|
||||
|
||||
Here's how you can remove those dashes by reusing and extending the default
|
||||
Product Loader::
|
||||
|
||||
strip_dashes = lambda x: x.strip('-')
|
||||
|
||||
class SiteSpecificLoader(ProductLoader):
|
||||
name_exp = TreeExpander(ProductLoader.name_exp, strip_dashes)
|
||||
|
||||
Another case where extending Loaders can be very helpful is when you have
|
||||
multiple source formats, for example XML and HTML. In the XML version you may
|
||||
want to remove ``CDATA`` occurrences. Here's an example of how to do it::
|
||||
|
||||
from myproject.utils.xml import remove_cdata
|
||||
|
||||
class XmlLoader(ProductLoader):
|
||||
name_exp = TreeExpander(remove_cdata, ProductLoader.name_exp)
|
||||
|
||||
There are many other possible ways to extend, inherit and override your
|
||||
Loaders, and different Loader hierarchies may fit better for different
|
||||
projects. Scrapy only provides the mechanism, it doesn't impose any specific
|
||||
organization of your Loaders collection - that's up to you and your project
|
||||
needs.
|
||||
|
||||
Available Expanders
|
||||
===================
|
||||
|
||||
.. _topics-loader-tree-expander:
|
||||
|
||||
Tree Expander
|
||||
-------------
|
||||
|
||||
The Tree Expander is the recommended Expander to use and the only really useful
|
||||
one, as the other is just an identity expander.
|
||||
|
||||
.. module:: scrapy.newitem.loader.expanders
|
||||
:synopsis: Expander classes to use with Item Loaders
|
||||
|
||||
.. class:: TreeExpander(\*functions, \**default_loader_arguments)
|
||||
|
||||
An expander which applies the given functions consecutively, in order, to
|
||||
each value returned by the previous function.
|
||||
|
||||
The algorithm consists in an ordered list of functions, each of which
|
||||
receives one value and can return zero, one or more values (as a list or
|
||||
iterable). If a function returns more than one value, the next function in
|
||||
the list will be called with each of those values, potentially returning
|
||||
more values and thus expanding the execution into different branches, which
|
||||
is why this expander is called Tree Expander.
|
||||
|
||||
Each expander function can optionally receive a ``loader_args`` argument,
|
||||
which will contain the currently active loader arguments.
|
||||
|
||||
The keyword arguments passed in the consturctor are used as the default
|
||||
loader arguments passed to on each expander call. This arguments can be
|
||||
overriden with specific loader arguments passed on each expander call.
|
||||
|
||||
IdentityExpander
|
||||
----------------
|
||||
|
||||
.. class:: IdentityExpander
|
||||
|
||||
An expander which returns the original values unchanged. It doesn't support
|
||||
any constructor arguments.
|
||||
|
||||
.. _topics-loader-reducers:
|
||||
|
||||
Available Reducers
|
||||
==================
|
||||
|
||||
.. module:: scrapy.newitem.loader.reducers
|
||||
:synopsis: Reducer classes to use with Item Loaders
|
||||
|
||||
Reducers are callable objects which are called with a list of values (to be
|
||||
reduced) as their first and only argument. Scrapy provides some simple,
|
||||
commonly used reducers, which are described below. But you can use any function
|
||||
or callable as reducer.
|
||||
|
||||
.. class:: TakeFirst
|
||||
|
||||
Return the first non-null value from the values to reduce, so it's used for
|
||||
single-valued fields. It doesn't receive any constructor arguments.
|
||||
|
||||
Example::
|
||||
|
||||
name_red = TakeFirst()
|
||||
|
||||
.. class:: Identity
|
||||
|
||||
Return the values to reduce unchanged, so it's used for multi-valued
|
||||
fields. It doesn't receive any constructor arguments.
|
||||
|
||||
Example::
|
||||
|
||||
features_red = Identity()
|
||||
|
||||
.. class:: Join(separator=u' ')
|
||||
|
||||
Return a the values to reduce joined with the separator given in the
|
||||
constructor, which defaults to ``u' '``.
|
||||
|
||||
When using the default separator, this reducer is equivalent to the
|
||||
function: ``u' '.join``
|
||||
|
||||
Examples::
|
||||
|
||||
name_red = Join()
|
||||
name_red = Join('<br>')
|
||||
|
||||
|
|
@ -0,0 +1,239 @@
|
|||
.. _topics-newitem:
|
||||
|
||||
=====
|
||||
Items
|
||||
=====
|
||||
|
||||
.. module:: scrapy.newitem
|
||||
:synopsis: Item and Field classes
|
||||
|
||||
The main goal in scraping is to extract structured data from unstructured
|
||||
sources, typically, web pages. Scrapy provides the :class:`Item` class for this
|
||||
purpose.
|
||||
|
||||
:class:`Item` objects are simple containers used to collect the scraped data.
|
||||
They provide a `dictionary-like`_ API with a convenient syntax for declaring
|
||||
their available fields.
|
||||
|
||||
.. _dictionary-like: http://docs.python.org/library/stdtypes.html#dict
|
||||
|
||||
.. _topics-newitem-declaring:
|
||||
|
||||
Declaring Items
|
||||
===============
|
||||
|
||||
Items are declared using a simple class definition syntax and :class:`Field`
|
||||
objects. Here is an example::
|
||||
|
||||
from scrapy.newitem import Item, Field
|
||||
|
||||
class Product(Item):
|
||||
name = Field()
|
||||
price = Field()
|
||||
stock = Field(default=0)
|
||||
last_updated = Field()
|
||||
|
||||
.. note:: Those familiar with `Django`_ will notice that Scrapy Items are
|
||||
declared similar to `Django Models`_, except that Scrapy Items are much
|
||||
simpler as there is no concept of different field types.
|
||||
|
||||
.. _Django: http://www.djangoproject.com/
|
||||
.. _Django Models: http://docs.djangoproject.com/en/dev/topics/db/models/
|
||||
|
||||
:class:`Field` objects are used to specify metadata for each field. For
|
||||
example, the default value for the ``stock`` field illustrated in the example
|
||||
above.
|
||||
|
||||
You can specify any kind of metadata for each field. There is no restriction on
|
||||
the values accepted by :class:`Field` objects. For this same
|
||||
reason, there isn't a reference list of all available metadata keys. Each key
|
||||
defined in :class:`Field` objects could be used by a different components, and
|
||||
only those components know about it. You can also define and use any other
|
||||
:class:`Field` key in your project too, for your own needs. The main goal of
|
||||
:class:`Field` objects is to provide a way to define all field metadata in one
|
||||
place. Typically, those components whose behaviour depends on each field, use
|
||||
certain field keys to configure that behaviour. You must refer to their
|
||||
documentation to see which metadata keys are used by each component.
|
||||
|
||||
It's important to note that the :class:`Field` objects used to declare the item
|
||||
do not stay assigned as class attributes. Instead, they can be accesed through
|
||||
the :attr:`Item.fields` attribute.
|
||||
|
||||
And that's all you need to know about declaring items.
|
||||
|
||||
Working with Items
|
||||
==================
|
||||
|
||||
Here are some examples of common tasks performed with items, using the
|
||||
``Product`` item :ref:`declared above <topics-newitem-declaring>`. You will
|
||||
notice the API is very similar to the `dict API`_.
|
||||
|
||||
Creating items
|
||||
--------------
|
||||
|
||||
::
|
||||
|
||||
>>> product = Product(name='Desktop PC', price=1000)
|
||||
>>> print product
|
||||
Product(name='Desktop PC', price=1000)
|
||||
|
||||
Getting field values
|
||||
--------------------
|
||||
|
||||
::
|
||||
|
||||
>>> product['name']
|
||||
Desktop PC
|
||||
>>> product.get('name')
|
||||
Desktop PC
|
||||
|
||||
>>> product['price']
|
||||
1000
|
||||
|
||||
>>> product['stock'] # getting field with default value
|
||||
0
|
||||
|
||||
>>> product['last_updated'] # getting field with no default value
|
||||
Traceback (most recent call last):
|
||||
...
|
||||
KeyError: 'last_updated'
|
||||
|
||||
>>> product.get('last_updated', 'not set')
|
||||
not set
|
||||
|
||||
>>> product['lala'] # getting unknown field
|
||||
Traceback (most recent call last):
|
||||
...
|
||||
KeyError: 'lala'
|
||||
|
||||
>>> product.get('lala', 'unknown field')
|
||||
'unknown field'
|
||||
|
||||
>>> 'name' in product # is name field populated?
|
||||
True
|
||||
|
||||
>>> 'last_updated' in product # is last_updated populated?
|
||||
False
|
||||
|
||||
>>> 'last_updated' in product.fields # is last_updated a declared field?
|
||||
True
|
||||
|
||||
>>> 'lala' in product.fields # is lala a declared field?
|
||||
False
|
||||
|
||||
Setting field values
|
||||
--------------------
|
||||
|
||||
::
|
||||
|
||||
>>> product['last_updated'] = 'today'
|
||||
>>> product['last_updated']
|
||||
today
|
||||
|
||||
>>> product['lala'] = 'test' # setting unknown field
|
||||
Traceback (most recent call last):
|
||||
...
|
||||
KeyError: 'Product does not support field: lala'
|
||||
|
||||
Accesing all populated values
|
||||
-----------------------------
|
||||
|
||||
To access all populated values just use the typical `dict API`_::
|
||||
|
||||
>>> product.keys()
|
||||
['price', 'name']
|
||||
|
||||
>>> product.items()
|
||||
[('price', 1000), ('name', 'Desktop PC')]
|
||||
|
||||
Other common tasks
|
||||
------------------
|
||||
|
||||
Copying items::
|
||||
|
||||
>>> product2 = Product(product)
|
||||
>>> print product2
|
||||
Product(name='Desktop PC', price=1000)
|
||||
|
||||
Creating dicts from items::
|
||||
|
||||
>>> dict(product) # create a dict from all populated values
|
||||
{'price': 1000, 'name': 'Desktop PC'}
|
||||
|
||||
Creating items from dicts::
|
||||
|
||||
>>> Product({'name': 'Laptop PC', 'price': 1500})
|
||||
Product(price=1500, name='Laptop PC')
|
||||
|
||||
>>> Product({'name': 'Laptop PC', 'lala': 1500}) # warning: unknown field in dict
|
||||
Traceback (most recent call last):
|
||||
...
|
||||
KeyError: 'Product does not support field: lala'
|
||||
|
||||
Default values
|
||||
==============
|
||||
|
||||
The only field metadata key supported by Items themselves is ``default``, which
|
||||
specifies the default value to return when trying to access a field which
|
||||
wasn't populated before.
|
||||
|
||||
So, for the ``Product`` item declared above::
|
||||
|
||||
>>> product = Product()
|
||||
|
||||
>>> product['stock'] # field with default value
|
||||
0
|
||||
|
||||
>>> product['name'] # field with no default value
|
||||
Traceback (most recent call last):
|
||||
...
|
||||
KeyError: 'name'
|
||||
|
||||
>>> product.get('name') is None
|
||||
True
|
||||
|
||||
Extending Items
|
||||
===============
|
||||
|
||||
You can extend Items (to add more fields or to change some metadata for some
|
||||
fields) by declaring a subclass of your original Item.
|
||||
|
||||
For example::
|
||||
|
||||
class DiscountedProduct(Product):
|
||||
discount_percent = Field(default=0)
|
||||
discount_expiration_date = Field()
|
||||
|
||||
Item objects
|
||||
============
|
||||
|
||||
.. class:: Item([arg])
|
||||
|
||||
Return a new Item optionally initialized from the given argument.
|
||||
|
||||
Items replicate the standard `dict API`_, including its constructor. The
|
||||
only additional attribute provided by Items is:
|
||||
|
||||
.. attribute:: fields
|
||||
|
||||
A dictionary containing *all declared fields* for this Item, not only
|
||||
those populated. The keys are the field names and the values are the
|
||||
:class:`Field` objects used in the :ref:`Item declaration
|
||||
<topics-newitem-declaring>`.
|
||||
|
||||
.. _dict API: http://docs.python.org/library/stdtypes.html#dict
|
||||
|
||||
Field objects
|
||||
=============
|
||||
|
||||
.. class:: Field([arg])
|
||||
|
||||
The :class:`Field` class is just an alias to the built-in `dict`_ class and
|
||||
doesn't provide any extra functionality or attributes. In other words,
|
||||
:class:`Field` objects are plain-old Python dicts. A separate class is used
|
||||
to support the :ref:`item declaration syntax <topics-newitem-declaring>`
|
||||
based on class attributes.
|
||||
|
||||
.. _dict: http://docs.python.org/library/stdtypes.html#dict
|
||||
|
||||
|
||||
|
|
@ -1,3 +0,0 @@
|
|||
.. toctree::
|
||||
|
||||
newitem/index
|
||||
|
|
@ -1,200 +0,0 @@
|
|||
.. _ref-newitem-fields:
|
||||
|
||||
===========
|
||||
Item Fields
|
||||
===========
|
||||
|
||||
.. module:: scrapy.newitem.fields
|
||||
|
||||
|
||||
Field options
|
||||
=============
|
||||
|
||||
Every ``Field`` class constructor accepts these arguments.
|
||||
|
||||
|
||||
``default``
|
||||
-----------
|
||||
|
||||
The default value for the field. See :ref:`topics-newitem-index-defaults`.
|
||||
|
||||
|
||||
Field types
|
||||
===========
|
||||
|
||||
These are the available built-in ``Field`` types. See
|
||||
:ref:`ref-newitem-fields-custom-fields` for info on creating your own field types.
|
||||
|
||||
|
||||
``BooleanField``
|
||||
----------------
|
||||
|
||||
.. class:: BooleanField
|
||||
|
||||
A boolean (true/false) field.
|
||||
|
||||
|
||||
``DateField``
|
||||
-------------
|
||||
|
||||
.. class:: DateField
|
||||
|
||||
A date, represented in Python by a `datetime.date`_ instance.
|
||||
|
||||
.. _datetime.date: http://docs.python.org/library/datetime.html#datetime.date
|
||||
|
||||
|
||||
``DateTimeField``
|
||||
-----------------
|
||||
|
||||
.. class:: DateTimeField
|
||||
|
||||
A date with time, represented in Python by a `datetime.datetime`_ instance.
|
||||
|
||||
.. _datetime.datetime: http://docs.python.org/library/datetime.html#datetime.datetime
|
||||
|
||||
|
||||
``DecimalField``
|
||||
----------------
|
||||
|
||||
.. class:: DecimalField
|
||||
|
||||
A fixed-precision decimal number, represented in Python by a `Decimal`_
|
||||
instance.
|
||||
|
||||
.. _Decimal: http://docs.python.org/library/decimal.html#decimal.Decimal
|
||||
|
||||
|
||||
``FloatField``
|
||||
--------------
|
||||
|
||||
.. class:: FloatField
|
||||
|
||||
A floating-point number represented in Python by a ``float`` instance.
|
||||
|
||||
|
||||
``IntegerField``
|
||||
----------------
|
||||
|
||||
.. class:: IntegerField
|
||||
|
||||
An integer.
|
||||
|
||||
|
||||
``ListField``
|
||||
-------------
|
||||
|
||||
.. class:: ListField(field)
|
||||
|
||||
A special field that works like a list of fields of another provided field kind.
|
||||
|
||||
:param field: The field which the elements of this list must conform to.
|
||||
:type field: a :class:`~scrapy.newitem.fields.BaseField` object
|
||||
|
||||
Usage example::
|
||||
|
||||
class ExampleItem(Item)
|
||||
names = fields.ListField(fields.TextField())
|
||||
|
||||
item = ExampleItem()
|
||||
item['names'] = [u'John', u'Jeena']
|
||||
|
||||
|
||||
``TextField``
|
||||
-------------
|
||||
|
||||
.. class:: TextField
|
||||
|
||||
A unicode text.
|
||||
|
||||
This class overrides the following methods from :class:`BaseField`:
|
||||
|
||||
.. method:: from_unicode_list(unicode_list)
|
||||
|
||||
Return a unicode string composed by joining the elements of
|
||||
``unicode_list`` with spaces.
|
||||
|
||||
For more info about this method see :class:`BaseField.from_unicode_list`.
|
||||
|
||||
|
||||
``TimeField``
|
||||
-------------
|
||||
|
||||
.. class:: TimeField
|
||||
|
||||
A time, represented in Python by a `datetime.time`_ instance.
|
||||
|
||||
.. _datetime.time: http://docs.python.org/library/datetime.html#datetime.time
|
||||
|
||||
.. _ref-newitem-fields-custom-fields:
|
||||
|
||||
|
||||
Creating custom fields
|
||||
======================
|
||||
|
||||
All field classes are subclasses of the :class:`BaseField` class (see below)
|
||||
which you can also subclass to create your own custom fields.
|
||||
|
||||
You can also subclass a more specific field class, say :class:`DecimalField`,
|
||||
to implement a ``PriceField``, for example.
|
||||
|
||||
|
||||
BaseField class
|
||||
---------------
|
||||
|
||||
.. class:: BaseField(default=None)
|
||||
|
||||
The base class for all fields. It only provides code for handling default
|
||||
values, not any particular type. It cannot be used directly either, as its
|
||||
:meth:`BaseField.to_python` method is not implemented.
|
||||
|
||||
The ``default`` argument (if given) must be of the type expected by this
|
||||
field, or any type that is accepted by the :meth:`BaseField.to_python`
|
||||
method of this field.
|
||||
|
||||
For example::
|
||||
|
||||
class NewsItem(Item):
|
||||
content = fields.TextField() # correct, no default value
|
||||
author = fields.TextField(default=u'Myself") # correct, with default value
|
||||
published = fields.DateField(default=23) # wrong default type (will raise TypeError)
|
||||
|
||||
.. method:: to_python(value)
|
||||
|
||||
Convert the input value to the type expected by this field and return
|
||||
it.
|
||||
|
||||
For example, :class:`IntegerField` would convert ``'1'`` to ``1``, while
|
||||
:class:`DecimalField` would convert ``'1'`` to ``Decimal('1')`` and so
|
||||
on.
|
||||
|
||||
This method is not implemented in the :class:`BaseField` class, so it
|
||||
must always be implemented in all its subclasses, in order to be usable.
|
||||
|
||||
This method should raise ``TypeError`` if the input type is not
|
||||
supported, and ``ValueError`` if the input type is support but its value
|
||||
is not appropriate (for example, an integer outside a given range).
|
||||
|
||||
This method must always return object of the expected field type.
|
||||
|
||||
.. method:: from_unicode_list(unicode_list)
|
||||
|
||||
Take the input list of unicode strings and convert it to a proper value
|
||||
with the type expected by this field. If no proper value if found,
|
||||
``None`` is returned instead.
|
||||
|
||||
The default behaviour is to return the value of the first item of the
|
||||
list, passed through the :meth:`to_python` method, or ``None`` if the
|
||||
list is empty::
|
||||
|
||||
return self.to_python(unicode_list[0]) if unicode_list else None
|
||||
|
||||
This default behaviour is provided because it's the more common one, but
|
||||
it's typical for :class:`BaseField` subclasses to override this method,
|
||||
such as the :meth:`TextField.from_unicode_list` method.
|
||||
|
||||
.. method:: get_default()
|
||||
|
||||
Return the default value for this field, or ``None`` if the field
|
||||
doesn't specify any.
|
||||
|
||||
|
|
@ -1,12 +0,0 @@
|
|||
.. _ref-newitems-index:
|
||||
|
||||
Items Reference
|
||||
===============
|
||||
|
||||
This is the Items reference, for introductory material see :ref:`topics-newitem-index`
|
||||
|
||||
.. toctree::
|
||||
:maxdepth: 1
|
||||
|
||||
fields
|
||||
|
||||
|
|
@ -1,4 +0,0 @@
|
|||
.. toctree::
|
||||
|
||||
newitem/index
|
||||
newitem/adaptors
|
||||
|
|
@ -1,139 +0,0 @@
|
|||
.. _topics-newitem-adaptors:
|
||||
|
||||
=============
|
||||
Item Adaptors
|
||||
=============
|
||||
|
||||
.. class:: scrapy.contrib_exp.newitem.adaptors.ItemAdaptor
|
||||
|
||||
As you probably want to scrape the same kind of Items from many sources
|
||||
(different websites, RSS feeds, etc.), Scrapy implements ItemAdaptors, they
|
||||
allow you to adapt chunks of HTML or XML (selected using Selectors) to the
|
||||
expected format of your Item fields.
|
||||
|
||||
An ItemAdaptor acts like a wrapper of an Item, you define an ItemAdaptor class,
|
||||
set the Item class to wrap and assign a set of functions (adaptor functions) to be called when you assign a value to a field.
|
||||
|
||||
Here's an example of an ItemAdaptor for our previously created Item::
|
||||
|
||||
class NewsAdaptor(ItemAdaptor):
|
||||
item_class = NewsItem
|
||||
|
||||
url = adaptor(extract, remove_tags(), unquote(), strip)
|
||||
headline = adaptor(extract, remove_tags(), unquote(), strip)
|
||||
summary = adaptor(extract, remove_tags(), unquote(), strip)
|
||||
content = adaptor(extract, remove_tags(), unquote(), strip)
|
||||
|
||||
How do we use it? Let's see it in action in a Spider::
|
||||
|
||||
def parse_newspage(self, response):
|
||||
xhs = HtmlXPathSelector(response)
|
||||
i = NewsAdaptor(response)
|
||||
|
||||
i.url = response.url
|
||||
i.headline = xhs.x('//h1[@class="headline"]')
|
||||
i.summary = xhs.x('//div[@class="summary"]')
|
||||
i.content = xhs.x('//div[@id="body"]')
|
||||
# published attribute is intentionally omitted, see below for site-specific adaptors
|
||||
return [i]
|
||||
|
||||
What happens underneath?
|
||||
|
||||
When we assign a value to a ItemAdaptor field it passes for the chain of
|
||||
functions defined previously in it's class, in this case, the value gets
|
||||
extracted (note that we assign directly the value obtained from the Selector),
|
||||
then tags will be removed, then the result will be unquoted, stripped and
|
||||
finally assigned to the Item Field.
|
||||
|
||||
This final assignment is done in an internal instance of the Item on the
|
||||
ItemAdaptor, that's why we can return an ItemAdaptor instead of an Item and
|
||||
Scrapy will know how to extract the item from it.
|
||||
|
||||
An Item can have as many ItemAdaptors as you want. It generally depends on how
|
||||
many sources and formats are you scraping from.
|
||||
|
||||
ItemAdaptor inheritance
|
||||
=======================
|
||||
|
||||
As we said before you generally want an ItemAdaptor for each different source of
|
||||
data and maybe some for specific sites, inheritance make this really easy, let's
|
||||
see an example of adapting HTML and XML::
|
||||
|
||||
class NewsAdaptor(ItemAdaptor):
|
||||
item_class = NewsItem
|
||||
|
||||
|
||||
class HtmlNewsAdaptor(NewsAdaptor):
|
||||
url = adaptor(extract, remove_tags(), unquote(), strip)
|
||||
headline = adaptor(extract, remove_tags(), unquote(), strip)
|
||||
summary = adaptor(extract, remove_tags(), unquote(), strip)
|
||||
content = adaptor(extract, remove_tags(), unquote(), strip)
|
||||
published = adaptor(extract, remove_tags(), unquote(), strip)
|
||||
|
||||
|
||||
class XmlNewsAdaptor(HtmlNewsAdaptor):
|
||||
url = adaptor(extract, remove_root, strip)
|
||||
headline = adaptor(extract, remove_root, strip)
|
||||
summary = adaptor(extract, remove_root, strip)
|
||||
content = adaptor(extract, remove_root, strip)
|
||||
published = adaptor(extract, remove_root, strip)
|
||||
|
||||
|
||||
Site specific ItemAdaptors
|
||||
==========================
|
||||
|
||||
For the moment we have covered adapting information from different sources, but
|
||||
other common case is adapting information for specific sites, think for example
|
||||
in our published field, it keeps the publication date of the news article.
|
||||
|
||||
As sites offer this information in different formats, we will have to make
|
||||
custom adaptors for it, let's see an example using our Item published field::
|
||||
|
||||
class SpecificSiteNewsAdaptor(HtmlNewsAdaptor):
|
||||
published = adaptor(HtmlNewsAdaptor.published, to_date('%d.%m.%Y'))
|
||||
|
||||
|
||||
The ``to_date`` adaptor function converts a string with the format specified in
|
||||
its parameter to one in 'YYYY-mm-dd' format (the one that DateField expects).
|
||||
|
||||
And in this example we're appending it to the of the chain of adaptor functions
|
||||
of published.
|
||||
|
||||
Note that ``SpecificSiteNewsAdaptor`` will inherit the field adaptations from
|
||||
``HtmlNewsAdaptor``.
|
||||
|
||||
Let's see it in action::
|
||||
|
||||
def parse_newspage(self, response):
|
||||
xhs = HtmlXPathSelector(response)
|
||||
i = SpecificSiteNewsAdaptor(response)
|
||||
|
||||
i.url = response.url
|
||||
i.headline = xhs.x('//h1[@class="headline"]')
|
||||
i.summary = xhs.x('//div[@class="summary"]')
|
||||
i.content = xhs.x('//div[@id="body"]')
|
||||
i.published = xhs.x('//h1[@class="date"]').re('\d{2}\.\d{2}\.\d{4}')
|
||||
return [i]
|
||||
|
||||
ItemAdaptor default_adaptor
|
||||
===========================
|
||||
|
||||
If you look closely at the code for our ItemAdaptors you can see that we're
|
||||
using the same set of adaptation functions in every field.
|
||||
|
||||
It is common for ItemAdaptors to have a basic set of adaptor functions that
|
||||
will be applied to almost every Field in the Item. To avoid repeating the same
|
||||
code, ItemAdaptor implements the ``default_adaptor`` shortcut.
|
||||
|
||||
``default_adaptor`` (if set) will be called when assigning a value for an Item
|
||||
Field that has no adaptor, so the process for determining what value gets
|
||||
assigned to an item when you assign a value to an ItemAdaptor field is as
|
||||
follows:
|
||||
|
||||
1. If there's an adaptor function for this field its called before assigning
|
||||
the value to the item.
|
||||
2. If no adaptor function if set and default_adaptor is, the value passes for
|
||||
``default_adaptor`` before being assigned.
|
||||
3. If no adaptor is defined for that field and no ``default_adaptor`` is set,
|
||||
the value is assigned directly.
|
||||
|
||||
|
|
@ -1,103 +0,0 @@
|
|||
.. _topics-newitem-index:
|
||||
|
||||
.. _topics-newitem-scrapeditem:
|
||||
|
||||
=====
|
||||
Items
|
||||
=====
|
||||
|
||||
.. currentmodule:: scrapy.item
|
||||
|
||||
The goal of the scraping process is to obtain scraped items from scraped pages.
|
||||
|
||||
Basic scraped items
|
||||
===================
|
||||
|
||||
In Scrapy the items are represented by a :class:`ScrapedItem` (almost an empty
|
||||
class) or any subclass of it.
|
||||
|
||||
To use :class:`ScrapedItem` you simply instantiate it and use instance
|
||||
attributes to store the information.
|
||||
|
||||
>>> from scrapy.item import ScrapedItem
|
||||
>>> item = ScrapedItem()
|
||||
>>> item.headline = 'Headline'
|
||||
>>> item.content = 'Content'
|
||||
>>> item.published = '2009-07-08'
|
||||
>>> item
|
||||
ScrapedItem({'headline': 'Headline', 'content': 'Content', 'published': '2009-07-08'})
|
||||
|
||||
Or you can use your own class to represent items, just be sure it inherits from
|
||||
:class:`ScrapedItem`.
|
||||
|
||||
.. _topics-newitem-index-item:
|
||||
|
||||
More advanced items
|
||||
===================
|
||||
|
||||
.. currentmodule:: scrapy.newitem
|
||||
|
||||
Scrapy provides :class:`Item` (a subclass of :class:`~scrapy.item.ScrapedItem`)
|
||||
that works like a form with fields to store the item's data.
|
||||
|
||||
To use this items you first define the item's fields as class attributes::
|
||||
|
||||
from scrapy.newitem import Item
|
||||
from scrapy.newitem import fields
|
||||
|
||||
class NewsItem(Item):
|
||||
headline = fields.TextField()
|
||||
content = fields.TextField()
|
||||
published = fields.DateField()
|
||||
|
||||
And then you instantiate the item and assign values to its fields, which will be
|
||||
converted to the expected Python types depending of their class::
|
||||
|
||||
>>> item = NewsItem()
|
||||
>>> item['headline'] = u'Headline'
|
||||
>>> item['content'] = u'Content'
|
||||
>>> item['published'] = '2009-07-08'
|
||||
>>> item
|
||||
NewsItem(headline=u'Headline', content=u'Content', published=datetime.date(2009, 7, 8))
|
||||
|
||||
Using this may seen complicated at first, but gives you much power over scraped
|
||||
data, like :ref:`topics-newitem-index-defaults`,
|
||||
:ref:`topics-newitem-adaptors`, etc.
|
||||
|
||||
.. _topics-newitem-index-defaults:
|
||||
|
||||
Default values for fields
|
||||
-------------------------
|
||||
|
||||
Each field accepts a ``default`` argument, that sets the default value of the
|
||||
field.
|
||||
|
||||
Fields which contain a default value will always return that value when not
|
||||
set, while fields which don't contain a default value will raise ``KeyError``,
|
||||
you can use :meth:`Item.get` method to avoid this.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
from scrapy.newitem import Item, fields
|
||||
|
||||
class NewsItem(Item):
|
||||
headline = fields.TextField()
|
||||
content = fields.TextField()
|
||||
published = fields.DateField()
|
||||
author = fields.TextField(default=u'Myself')
|
||||
views = fields.IntegerField(default=0)
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
>>> item = NewsItem()
|
||||
>>> item['author']
|
||||
u'Myself'
|
||||
>>> item['views']
|
||||
0
|
||||
>>> item['headline']
|
||||
Traceback (most recent call last):
|
||||
...
|
||||
KeyError: 'headline
|
||||
>>> item.get('headline') is None
|
||||
True
|
||||
|
||||
|
|
@ -51,7 +51,7 @@ to steal from us!
|
|||
Does Scrapy work with HTTP proxies?
|
||||
-----------------------------------
|
||||
|
||||
No. support for HTTP proxies is not currently implemented in Scrapy, but it
|
||||
No. Support for HTTP proxies is not currently implemented in Scrapy, but it
|
||||
will be in the future. For more information about this, follow `this ticket
|
||||
<http://dev.scrapy.org/ticket/71>`_. Setting the ``http_proxy`` environment
|
||||
variable won't work because Twisted (the library used by Scrapy to download
|
||||
|
|
|
|||
|
|
@ -57,7 +57,7 @@ RobustScrapedItems
|
|||
.. warning::
|
||||
|
||||
RobustScapedItems are deprecated and will be replaced by the :ref:`New item
|
||||
API <topics-newitem-index-item>` (still in development).
|
||||
API <topics-newitem>` (still in development).
|
||||
|
||||
.. module:: scrapy.contrib.item
|
||||
:synopsis: Objects for storing scraped data
|
||||
|
|
|
|||
Loading…
Reference in New Issue