diff --git a/docs/index.rst b/docs/index.rst index 507b9bea9..de3e015d5 100644 --- a/docs/index.rst +++ b/docs/index.rst @@ -56,9 +56,9 @@ Basic concepts :hidden: topics/commands - topics/items topics/spiders topics/selectors + topics/items topics/loaders topics/shell topics/item-pipeline @@ -72,9 +72,6 @@ Basic concepts :doc:`topics/commands` Learn about the command-line tool used to manage your Scrapy project. -:doc:`topics/items` - Define the data you want to scrape. - :doc:`topics/spiders` Write the rules to crawl your websites. @@ -84,6 +81,9 @@ Basic concepts :doc:`topics/shell` Test your extraction code in an interactive environment. +:doc:`topics/items` + Define the data you want to scrape. + :doc:`topics/loaders` Populate your items with the extracted data. diff --git a/docs/topics/architecture.rst b/docs/topics/architecture.rst index 700e6d92d..dad171a07 100644 --- a/docs/topics/architecture.rst +++ b/docs/topics/architecture.rst @@ -102,10 +102,10 @@ this: 6. The Engine receives the Response from the Downloader and sends it to the Spider for processing, passing through the Spider Middleware (input direction). -7. The Spider processes the Response and returns scraped Items and new Requests +7. The Spider processes the Response and returns scraped items and new Requests (to follow) to the Engine. -8. The Engine sends scraped Items (returned by the Spider) to the Item Pipeline +8. The Engine sends scraped items (returned by the Spider) to the Item Pipeline and Requests (returned by spider) to the Scheduler 9. The process repeats (from step 2) until there are no more requests from the diff --git a/docs/topics/exporters.rst b/docs/topics/exporters.rst index f7feed4af..43846852b 100644 --- a/docs/topics/exporters.rst +++ b/docs/topics/exporters.rst @@ -7,7 +7,7 @@ Item Exporters .. module:: scrapy.contrib.exporter :synopsis: Item Exporters -Once you have scraped your Items, you often want to persist or export those +Once you have scraped your items, you often want to persist or export those items, to use the data in some other application. That is, after all, the whole purpose of the scraping process. @@ -90,9 +90,9 @@ described next. 1. Declaring a serializer in the field -------------------------------------- -You can declare a serializer in the :ref:`field metadata -`. The serializer must be a callable which receives a -value and returns its serialized form. +If you use :class:`~.Item` you can declare a serializer in the +:ref:`field metadata `. The serializer must be +a callable which receives a value and returns its serialized form. Example:: @@ -167,8 +167,9 @@ BaseItemExporter value unchanged except for ``unicode`` values which are encoded to ``str`` using the encoding declared in the :attr:`encoding` attribute. - :param field: the field being serialized - :type field: :class:`~scrapy.item.Field` object + :param field: the field being serialized. If a raw dict is being + exported (not :class:`~.Item`) *field* value is an empty dict. + :type field: :class:`~scrapy.item.Field` object or an empty dict :param name: the name of the field being serialized :type name: str @@ -197,12 +198,17 @@ BaseItemExporter Some exporters (like :class:`CsvItemExporter`) respect the order of the fields defined in this attribute. + Some exporters may require fields_to_export list in order to export the + data properly when spiders return dicts (not :class:`~Item` instances). + .. attribute:: export_empty_fields Whether to include empty/unpopulated item fields in the exported data. Defaults to ``False``. Some exporters (like :class:`CsvItemExporter`) ignore this attribute and always export all empty fields. + This option is ignored for dict items. + .. attribute:: encoding The encoding that will be used to encode unicode values. This only diff --git a/docs/topics/images.rst b/docs/topics/images.rst index 4b07300eb..890c7fd4a 100644 --- a/docs/topics/images.rst +++ b/docs/topics/images.rst @@ -63,9 +63,14 @@ this: Usage example ============= -In order to use the image pipeline you just need to :ref:`enable it -` and define an item with the ``image_urls`` and -``images`` fields:: +In order to use the image pipeline first +:ref:`enable it `. + +Then, if a spider returns a dict with 'image_urls' key, +the pipeline will put the results under 'images' key. + +If you prefer to use :class:`~.Item` then define a custom +item with the ``image_urls`` and ``images`` fields:: import scrapy @@ -74,7 +79,7 @@ In order to use the image pipeline you just need to :ref:`enable it # ... other item fields ... image_urls = scrapy.Field() images = scrapy.Field() - + If you need something more complex and want to override the custom images pipeline behaviour, see :ref:`topics-images-override`. diff --git a/docs/topics/item-pipeline.rst b/docs/topics/item-pipeline.rst index 7b66753b8..973c77516 100644 --- a/docs/topics/item-pipeline.rst +++ b/docs/topics/item-pipeline.rst @@ -8,8 +8,8 @@ After an item has been scraped by a spider, it is sent to the Item Pipeline which process it through several components that are executed sequentially. Each item pipeline component (sometimes referred as just "Item Pipeline") is a -Python class that implements a simple method. They receive an Item and perform -an action over it, also deciding if the Item should continue through the +Python class that implements a simple method. They receive an item and perform +an action over it, also deciding if the item should continue through the pipeline or be dropped and no longer processed. Typical use for item pipelines are: @@ -28,12 +28,12 @@ Each item pipeline component is a Python class that must implement the following .. method:: process_item(self, item, spider) This method is called for every item pipeline component and must either return - a :class:`~scrapy.item.Item` (or any descendant class) object or raise a - :exc:`~scrapy.exceptions.DropItem` exception. Dropped items are no longer + a dict with data, :class:`~scrapy.item.Item` (or any descendant class) object + or raise a :exc:`~scrapy.exceptions.DropItem` exception. Dropped items are no longer processed by further pipeline components. :param item: the item scraped - :type item: :class:`~scrapy.item.Item` object + :type item: :class:`~scrapy.item.Item` object or a dict :param spider: the spider which scraped the item :type spider: :class:`~scrapy.spider.Spider` object @@ -135,6 +135,8 @@ method and how to clean up the resources properly. import pymongo class MongoPipeline(object): + + collection_name = 'scrapy_items' def __init__(self, mongo_uri, mongo_db): self.mongo_uri = mongo_uri @@ -155,8 +157,7 @@ method and how to clean up the resources properly. self.client.close() def process_item(self, item, spider): - collection_name = item.__class__.__name__ - self.db[collection_name].insert(dict(item)) + self.db[self.collection_name].insert(dict(item)) return item .. _MongoDB: http://www.mongodb.org/ diff --git a/docs/topics/items.rst b/docs/topics/items.rst index 17f10a88c..ac3eb6699 100644 --- a/docs/topics/items.rst +++ b/docs/topics/items.rst @@ -8,12 +8,21 @@ Items :synopsis: Item and Field classes The main goal in scraping is to extract structured data from unstructured -sources, typically, web pages. Scrapy provides the :class:`Item` class for this -purpose. +sources, typically, web pages. Scrapy spiders can return the extracted data +as Python dicts. While convenient and familiar, Python dicts lack structure: +it is easy to make a typo in a field name or return inconsistent data, +especially in a larger project with many spiders. +To define common output data format Scrapy provides the :class:`Item` class. :class:`Item` objects are simple containers used to collect the scraped data. They provide a `dictionary-like`_ API with a convenient syntax for declaring -their available fields. +their available fields. + +Various Scrapy components use extra information provided by Items: +exporters look at declared fields to figure out columns to export, +serialization can be customized using Item fields metadata, :mod:`trackref` +tracks Item instances to help finding memory leaks +(see :ref:`topics-leaks-trackrefs`_), etc. .. _dictionary-like: https://docs.python.org/2/library/stdtypes.html#dict @@ -64,8 +73,6 @@ It's important to note that the :class:`Field` objects used to declare the item do not stay assigned as class attributes. Instead, they can be accessed through the :attr:`Item.fields` attribute. -And that's all you need to know about declaring items. - Working with Items ================== diff --git a/docs/topics/practices.rst b/docs/topics/practices.rst index 13dde52a3..3ec7bc29b 100644 --- a/docs/topics/practices.rst +++ b/docs/topics/practices.rst @@ -183,22 +183,3 @@ If you are still unable to prevent your bot getting banned, consider contacting .. _testspiders: https://github.com/scrapinghub/testspiders .. _Twisted Reactor Overview: http://twistedmatrix.com/documents/current/core/howto/reactor-basics.html .. _Crawlera: http://crawlera.com - -.. _dynamic-item-classes: - -Dynamic Creation of Item Classes -================================ - -For applications in which the structure of item class is to be determined by -user input, or other changing conditions, you can dynamically create item -classes instead of manually coding them. - -:: - - - from scrapy.item import DictItem, Field - - def create_item_class(class_name, field_list): - fields = {field_name: Field() for field_name in field_list} - - return type(class_name, (DictItem,), {'fields': fields}) diff --git a/docs/topics/signals.rst b/docs/topics/signals.rst index 405b131ed..85cf43c76 100644 --- a/docs/topics/signals.rst +++ b/docs/topics/signals.rst @@ -71,7 +71,7 @@ item_scraped This signal supports returning deferreds from their handlers. :param item: the item scraped - :type item: :class:`~scrapy.item.Item` object + :type item: dict or :class:`~scrapy.item.Item` object :param spider: the spider which scraped the item :type spider: :class:`~scrapy.spider.Spider` object @@ -91,7 +91,7 @@ item_dropped This signal supports returning deferreds from their handlers. :param item: the item dropped from the :ref:`topics-item-pipeline` - :type item: :class:`~scrapy.item.Item` object + :type item: dict or :class:`~scrapy.item.Item` object :param spider: the spider which scraped the item :type spider: :class:`~scrapy.spider.Spider` object diff --git a/docs/topics/spider-middleware.rst b/docs/topics/spider-middleware.rst index 6f14567fc..abeae2bce 100644 --- a/docs/topics/spider-middleware.rst +++ b/docs/topics/spider-middleware.rst @@ -90,15 +90,16 @@ following methods: it has processed the response. :meth:`process_spider_output` must return an iterable of - :class:`~scrapy.http.Request` or :class:`~scrapy.item.Item` objects. + :class:`~scrapy.http.Request`, dict or :class:`~scrapy.item.Item` + objects. :param response: the response which generated this output from the spider :type response: :class:`~scrapy.http.Response` object :param result: the result returned by the spider - :type result: an iterable of :class:`~scrapy.http.Request` or - :class:`~scrapy.item.Item` objects + :type result: an iterable of :class:`~scrapy.http.Request`, dict + or :class:`~scrapy.item.Item` objects :param spider: the spider whose result is being processed :type spider: :class:`~scrapy.spider.Spider` object @@ -110,7 +111,7 @@ following methods: method (from other spider middleware) raises an exception. :meth:`process_spider_exception` should return either ``None`` or an - iterable of :class:`~scrapy.http.Response` or + iterable of :class:`~scrapy.http.Response`, dict or :class:`~scrapy.item.Item` objects. If it returns ``None``, Scrapy will continue processing this exception, diff --git a/docs/topics/spiders.rst b/docs/topics/spiders.rst index a7e7d2746..e395f36d5 100644 --- a/docs/topics/spiders.rst +++ b/docs/topics/spiders.rst @@ -24,8 +24,9 @@ For spiders, the scraping cycle goes through something like this: Requests. 2. In the callback function, you parse the response (web page) and return either - :class:`~scrapy.item.Item` objects, :class:`~scrapy.http.Request` objects, - or an iterable of both. Those Requests will also contain a callback (maybe + dicts with extracted data, :class:`~scrapy.item.Item` objects, + :class:`~scrapy.http.Request` objects, or an iterable of these objects. + Those Requests will also contain a callback (maybe the same) and will then be downloaded by Scrapy and then their response handled by the specified callback. @@ -41,70 +42,22 @@ Even though this cycle applies (more or less) to any kind of spider, there are different kinds of default spiders bundled into Scrapy for different purposes. We will talk about those types here. -.. _spiderargs: - -Spider arguments -================ - -Spiders can receive arguments that modify their behaviour. Some common uses for -spider arguments are to define the start URLs or to restrict the crawl to -certain sections of the site, but they can be used to configure any -functionality of the spider. - -Spider arguments are passed through the :command:`crawl` command using the -``-a`` option. For example:: - - scrapy crawl myspider -a category=electronics - -Spiders receive arguments in their constructors:: - - import scrapy - - class MySpider(scrapy.Spider): - name = 'myspider' - - def __init__(self, category=None, *args, **kwargs): - super(MySpider, self).__init__(*args, **kwargs) - self.start_urls = ['http://www.example.com/categories/%s' % category] - # ... - -Spider arguments can also be passed through the Scrapyd ``schedule.json`` API. -See `Scrapyd documentation`_. - -.. _topics-spiders-ref: - -Built-in spiders reference -========================== - -Scrapy comes with some useful generic spiders that you can use, to subclass -your spiders from. Their aim is to provide convenient functionality for a few -common scraping cases, like following all links on a site based on certain -rules, crawling from `Sitemaps`_, or parsing a XML/CSV feed. - -For the examples used in the following spiders, we'll assume you have a project -with a ``TestItem`` declared in a ``myproject.items`` module:: - - import scrapy - - class TestItem(scrapy.Item): - id = scrapy.Field() - name = scrapy.Field() - description = scrapy.Field() - - .. module:: scrapy.spider :synopsis: Spiders base class, spider manager and spider middleware -Spider ------- +.. _topics-spiders-ref: + +scrapy.Spider +============= .. class:: Spider() This is the simplest spider, and the one from which every other spider must inherit from (either the ones that come bundled with Scrapy, or the ones that you write yourself). It doesn't provide any special functionality. It just - requests the given ``start_urls``/``start_requests``, and calls the spider's - method ``parse`` for each of the resulting responses. + provides a default :meth:`start_requests` implementation which sends requests from + the :attr:`start_urls` spider attribute and calls the spider's method ``parse`` + for each of the resulting responses. .. attribute:: name @@ -198,15 +151,18 @@ Spider the method to override. For example, if you need to start by logging in using a POST request, you could do:: - def start_requests(self): - return [scrapy.FormRequest("http://www.example.com/login", - formdata={'user': 'john', 'pass': 'secret'}, - callback=self.logged_in)] + class MySpider(scrapy.Spider): + name = 'myspider' + + def start_requests(self): + return [scrapy.FormRequest("http://www.example.com/login", + formdata={'user': 'john', 'pass': 'secret'}, + callback=self.logged_in)] - def logged_in(self, response): - # here you would extract links to follow and return Requests for - # each of them, with another callback - pass + def logged_in(self, response): + # here you would extract links to follow and return Requests for + # each of them, with another callback + pass .. method:: make_requests_from_url(url) @@ -231,7 +187,7 @@ Spider This method, as well as any other Request callback, must return an iterable of :class:`~scrapy.http.Request` and/or - :class:`~scrapy.item.Item` objects. + dicts or :class:`~scrapy.item.Item` objects. :param response: the response to parse :type response: :class:~scrapy.http.Response` @@ -247,10 +203,6 @@ Spider Called when the spider closes. This method provides a shortcut to signals.connect() for the :signal:`spider_closed` signal. - -Spider example -~~~~~~~~~~~~~~ - Let's see an example:: import scrapy @@ -268,10 +220,9 @@ Let's see an example:: def parse(self, response): self.log('A response from %s just arrived!' % response.url) -Another example returning multiple Requests and Items from a single callback:: +Return multiple Requests and items from a single callback:: import scrapy - from myproject.items import MyItem class MySpider(scrapy.Spider): name = 'example.com' @@ -282,12 +233,85 @@ Another example returning multiple Requests and Items from a single callback:: 'http://www.example.com/3.html', ] + def parse(self, response): + for h3 in response.xpath('//h3').extract(): + yield {"title": h3} + + for url in response.xpath('//a/@href').extract(): + yield scrapy.Request(url, callback=self.parse) + +Instead of :attr:`~.start_urls` you can use :meth:`~.start_requests` directly; +to give data more structure you can use :ref:`topics-items`:: + + import scrapy + from myproject.items import MyItem + + class MySpider(scrapy.Spider): + name = 'example.com' + allowed_domains = ['example.com'] + + def start_requests(self): + yield scrapy.Request('http://www.example.com/1.html', self.parse) + yield scrapy.Request('http://www.example.com/2.html', self.parse) + yield scrapy.Request('http://www.example.com/3.html', self.parse) + def parse(self, response): for h3 in response.xpath('//h3').extract(): yield MyItem(title=h3) for url in response.xpath('//a/@href').extract(): yield scrapy.Request(url, callback=self.parse) + +.. _spiderargs: + +Spider arguments +================ + +Spiders can receive arguments that modify their behaviour. Some common uses for +spider arguments are to define the start URLs or to restrict the crawl to +certain sections of the site, but they can be used to configure any +functionality of the spider. + +Spider arguments are passed through the :command:`crawl` command using the +``-a`` option. For example:: + + scrapy crawl myspider -a category=electronics + +Spiders receive arguments in their constructors:: + + import scrapy + + class MySpider(scrapy.Spider): + name = 'myspider' + + def __init__(self, category=None, *args, **kwargs): + super(MySpider, self).__init__(*args, **kwargs) + self.start_urls = ['http://www.example.com/categories/%s' % category] + # ... + +Spider arguments can also be passed through the Scrapyd ``schedule.json`` API. +See `Scrapyd documentation`_. + +.. _builtin-spiders: + +Generic Spiders +=============== + +Scrapy comes with some useful generic spiders that you can use, to subclass +your spiders from. Their aim is to provide convenient functionality for a few +common scraping cases, like following all links on a site based on certain +rules, crawling from `Sitemaps`_, or parsing a XML/CSV feed. + +For the examples used in the following spiders, we'll assume you have a project +with a ``TestItem`` declared in a ``myproject.items`` module:: + + import scrapy + + class TestItem(scrapy.Item): + id = scrapy.Field() + name = scrapy.Field() + description = scrapy.Field() + .. module:: scrapy.contrib.spiders :synopsis: Collection of generic spiders diff --git a/scrapy/commands/parse.py b/scrapy/commands/parse.py index 01c7fff0a..b8cc140d4 100644 --- a/scrapy/commands/parse.py +++ b/scrapy/commands/parse.py @@ -107,7 +107,7 @@ class Command(ScrapyCommand): items, requests = [], [] for x in iterate_spider_output(cb(response)): - if isinstance(x, BaseItem): + if isinstance(x, (BaseItem, dict)): items.append(x) elif isinstance(x, Request): requests.append(x) diff --git a/scrapy/contracts/default.py b/scrapy/contracts/default.py index 1d8367f82..20582503d 100644 --- a/scrapy/contracts/default.py +++ b/scrapy/contracts/default.py @@ -35,8 +35,8 @@ class ReturnsContract(Contract): objects = { 'request': Request, 'requests': Request, - 'item': BaseItem, - 'items': BaseItem, + 'item': (BaseItem, dict), + 'items': (BaseItem, dict), } def __init__(self, *args, **kwargs): @@ -83,7 +83,7 @@ class ScrapesContract(Contract): def post_process(self, output): for x in output: - if isinstance(x, BaseItem): + if isinstance(x, (BaseItem, dict)): for arg in self.args: if not arg in x: raise ContractFail("'%s' field is missing" % arg) diff --git a/scrapy/contrib/exporter/__init__.py b/scrapy/contrib/exporter/__init__.py index cc88f8792..7e1d01a0a 100644 --- a/scrapy/contrib/exporter/__init__.py +++ b/scrapy/contrib/exporter/__init__.py @@ -9,6 +9,7 @@ import marshal import six from six.moves import cPickle as pickle from xml.sax.saxutils import XMLGenerator + from scrapy.utils.serialize import ScrapyJSONEncoder from scrapy.item import BaseItem @@ -50,13 +51,13 @@ class BaseItemExporter(object): return value.encode(self.encoding) if isinstance(value, unicode) else value def _get_serialized_fields(self, item, default_value=None, include_empty=None): - """Return the fields to export as an iterable of tuples (name, - serialized_value) + """Return the fields to export as an iterable of tuples + (name, serialized_value) """ if include_empty is None: include_empty = self.export_empty_fields if self.fields_to_export is None: - if include_empty: + if include_empty and not isinstance(item, dict): field_iter = six.iterkeys(item.fields) else: field_iter = six.iterkeys(item) @@ -64,12 +65,11 @@ class BaseItemExporter(object): if include_empty: field_iter = self.fields_to_export else: - nonempty_fields = set(item.keys()) - field_iter = (x for x in self.fields_to_export if x in - nonempty_fields) + field_iter = (x for x in self.fields_to_export if x in item) + for field_name in field_iter: if field_name in item: - field = item.fields[field_name] + field = {} if isinstance(item, dict) else item.fields[field_name] value = self.serialize_field(field, field_name, item[field_name]) else: value = default_value @@ -191,7 +191,12 @@ class CsvItemExporter(BaseItemExporter): def _write_headers_and_set_fields_to_export(self, item): if self.include_headers_line: if not self.fields_to_export: - self.fields_to_export = item.fields.keys() + if isinstance(item, dict): + # for dicts try using fields of the first item + self.fields_to_export = list(item.keys()) + else: + # use fields declared in Item + self.fields_to_export = list(item.fields.keys()) self.csv_writer.writerow(self.fields_to_export) diff --git a/scrapy/contrib/pipeline/files.py b/scrapy/contrib/pipeline/files.py index db8cf8b76..9e803aca0 100644 --- a/scrapy/contrib/pipeline/files.py +++ b/scrapy/contrib/pipeline/files.py @@ -267,7 +267,7 @@ class FilesPipeline(MediaPipeline): return checksum def item_completed(self, results, item, info): - if self.FILES_RESULT_FIELD in item.fields: + if isinstance(item, dict) or self.FILES_RESULT_FIELD in item.fields: item[self.FILES_RESULT_FIELD] = [x for ok, x in results if ok] return item diff --git a/scrapy/contrib/pipeline/images.py b/scrapy/contrib/pipeline/images.py index 9c1a54455..b12995f09 100644 --- a/scrapy/contrib/pipeline/images.py +++ b/scrapy/contrib/pipeline/images.py @@ -109,7 +109,7 @@ class ImagesPipeline(FilesPipeline): return [Request(x) for x in item.get(self.IMAGES_URLS_FIELD, [])] def item_completed(self, results, item, info): - if self.IMAGES_RESULT_FIELD in item.fields: + if isinstance(item, dict) or self.IMAGES_RESULT_FIELD in item.fields: item[self.IMAGES_RESULT_FIELD] = [x for ok, x in results if ok] return item diff --git a/scrapy/core/scraper.py b/scrapy/core/scraper.py index 3409a0e7c..b301aa962 100644 --- a/scrapy/core/scraper.py +++ b/scrapy/core/scraper.py @@ -174,7 +174,7 @@ class Scraper(object): """ if isinstance(output, Request): self.crawler.engine.crawl(request=output, spider=spider) - elif isinstance(output, BaseItem): + elif isinstance(output, (BaseItem, dict)): self.slot.itemproc_size += 1 dfd = self.itemproc.process_item(output, spider) dfd.addBoth(self._itemproc_finished, output, response, spider) @@ -183,7 +183,7 @@ class Scraper(object): pass else: typename = type(output).__name__ - log.msg(format='Spider must return Request, BaseItem or None, ' + log.msg(format='Spider must return Request, BaseItem, dict or None, ' 'got %(typename)r in %(request)s', level=log.ERROR, spider=spider, request=request, typename=typename) diff --git a/tests/spiders.py b/tests/spiders.py index 83d767f5c..86ace9d6e 100644 --- a/tests/spiders.py +++ b/tests/spiders.py @@ -85,6 +85,7 @@ class ItemSpider(FollowAllSpider): for request in super(ItemSpider, self).parse(response): yield request yield Item() + yield {} class DefaultError(Exception): diff --git a/tests/test_commands.py b/tests/test_commands.py index 70b4e74dc..eb3556b62 100644 --- a/tests/test_commands.py +++ b/tests/test_commands.py @@ -127,6 +127,7 @@ class MiscCommandsTest(CommandTest): def test_list(self): self.assertEqual(0, self.call('list')) + class RunSpiderCommandTest(CommandTest): def test_runspider(self): @@ -135,10 +136,10 @@ class RunSpiderCommandTest(CommandTest): fname = abspath(join(tmpdir, 'myspider.py')) with open(fname, 'w') as f: f.write(""" +import scrapy from scrapy import log -from scrapy.spider import Spider -class MySpider(Spider): +class MySpider(scrapy.Spider): name = 'myspider' def start_requests(self): @@ -192,16 +193,15 @@ class ParseCommandTest(ProcessTest, SiteTest, CommandTest): with open(fname, 'w') as f: f.write(""" from scrapy import log -from scrapy.spider import Spider -from scrapy.item import Item +import scrapy -class MySpider(Spider): +class MySpider(scrapy.Spider): name = '{0}' def parse(self, response): if getattr(self, 'test_arg', None): self.log('It Works!') - return [Item()] + return [scrapy.Item(), dict(foo='bar')] """.format(self.spider_name)) fname = abspath(join(self.proj_mod_path, 'pipelines.py')) @@ -239,6 +239,14 @@ ITEM_PIPELINES = {'%s.pipelines.MyPipeline': 1} self.url('/html')]) self.assert_("[scrapy] INFO: It Works!" in stderr, stderr) + @defer.inlineCallbacks + def test_parse_items(self): + status, out, stderr = yield self.execute( + ['--spider', self.spider_name, '-c', 'parse', self.url('/html')] + ) + self.assertIn("""[{}, {'foo': 'bar'}]""", out) + + class BenchCommandTest(CommandTest): diff --git a/tests/test_contracts.py b/tests/test_contracts.py index a651576a5..d7732f55d 100644 --- a/tests/test_contracts.py +++ b/tests/test_contracts.py @@ -39,6 +39,13 @@ class TestSpider(Spider): """ return TestItem(url=response.url) + def returns_dict_item(self, response): + """ method which returns item + @url http://scrapy.org + @returns items 1 1 + """ + return {"url": response.url} + def returns_fail(self, response): """ method which returns item @url http://scrapy.org @@ -46,6 +53,13 @@ class TestSpider(Spider): """ return TestItem(url=response.url) + def returns_dict_fail(self, response): + """ method which returns item + @url http://scrapy.org + @returns items 0 0 + """ + return {'url': response.url} + def scrapes_item_ok(self, response): """ returns item with name and url @url http://scrapy.org @@ -54,6 +68,14 @@ class TestSpider(Spider): """ return TestItem(name='test', url=response.url) + def scrapes_dict_item_ok(self, response): + """ returns item with name and url + @url http://scrapy.org + @returns items 1 1 + @scrapes name url + """ + return {'name': 'test', 'url': response.url} + def scrapes_item_fail(self, response): """ returns item with no name @url http://scrapy.org @@ -62,6 +84,14 @@ class TestSpider(Spider): """ return TestItem(url=response.url) + def scrapes_dict_item_fail(self, response): + """ returns item with no name + @url http://scrapy.org + @returns items 1 1 + @scrapes name url + """ + return {'url': response.url} + def parse_no_url(self, response): """ method with no url @returns items 1 1 @@ -110,6 +140,11 @@ class ContractsManagerTest(unittest.TestCase): request.callback(response) self.should_succeed() + # returns_dict_item + request = self.conman.from_method(spider.returns_dict_item, self.results) + request.callback(response) + self.should_succeed() + # returns_request request = self.conman.from_method(spider.returns_request, self.results) request.callback(response) @@ -120,6 +155,11 @@ class ContractsManagerTest(unittest.TestCase): request.callback(response) self.should_fail() + # returns_dict_fail + request = self.conman.from_method(spider.returns_dict_fail, self.results) + request.callback(response) + self.should_fail() + def test_scrapes(self): spider = TestSpider() response = ResponseMock() @@ -129,8 +169,19 @@ class ContractsManagerTest(unittest.TestCase): request.callback(response) self.should_succeed() + # scrapes_dict_item_ok + request = self.conman.from_method(spider.scrapes_dict_item_ok, self.results) + request.callback(response) + self.should_succeed() + # scrapes_item_fail request = self.conman.from_method(spider.scrapes_item_fail, self.results) request.callback(response) self.should_fail() + + # scrapes_dict_item_fail + request = self.conman.from_method(spider.scrapes_dict_item_fail, + self.results) + request.callback(response) + self.should_fail() diff --git a/tests/test_contrib_exporter.py b/tests/test_contrib_exporter.py index 9092007e5..746aeb65b 100644 --- a/tests/test_contrib_exporter.py +++ b/tests/test_contrib_exporter.py @@ -1,14 +1,19 @@ -import unittest, json +from __future__ import absolute_import +import re +import json +import unittest from io import BytesIO from six.moves import cPickle as pickle + import lxml.etree -import re from scrapy.item import Item, Field from scrapy.utils.python import str_to_unicode -from scrapy.contrib.exporter import BaseItemExporter, PprintItemExporter, \ - PickleItemExporter, CsvItemExporter, XmlItemExporter, JsonLinesItemExporter, \ - JsonItemExporter, PythonItemExporter +from scrapy.contrib.exporter import ( + BaseItemExporter, PprintItemExporter, PickleItemExporter, CsvItemExporter, + XmlItemExporter, JsonLinesItemExporter, JsonItemExporter, PythonItemExporter +) + class TestItem(Item): name = Field() @@ -33,21 +38,28 @@ class BaseItemExporterTest(unittest.TestCase): exported_dict[k] = str_to_unicode(v) self.assertEqual(self.i, exported_dict) - def test_export_item(self): + def assertItemExportWorks(self, item): self.ie.start_exporting() try: - self.ie.export_item(self.i) + self.ie.export_item(item) except NotImplementedError: if self.ie.__class__ is not BaseItemExporter: raise self.ie.finish_exporting() self._check_output() + def test_export_item(self): + self.assertItemExportWorks(self.i) + + def test_export_dict_item(self): + self.assertItemExportWorks(dict(self.i)) + def test_serialize_field(self): - self.assertEqual(self.ie.serialize_field( \ - self.i.fields['name'], 'name', self.i['name']), 'John\xc2\xa3') - self.assertEqual( \ - self.ie.serialize_field(self.i.fields['age'], 'age', self.i['age']), '22') + res = self.ie.serialize_field(self.i.fields['name'], 'name', self.i['name']) + self.assertEqual(res, 'John\xc2\xa3') + + res = self.ie.serialize_field(self.i.fields['age'], 'age', self.i['age']) + self.assertEqual(res, '22') def test_fields_to_export(self): ie = self._get_exporter(fields_to_export=['name']) @@ -72,13 +84,14 @@ class BaseItemExporterTest(unittest.TestCase): self.assertEqual(ie.serialize_field(i.fields['name'], 'name', i['name']), 'John\xc2\xa3') self.assertEqual(ie.serialize_field(i.fields['age'], 'age', i['age']), '24') + class PythonItemExporterTest(BaseItemExporterTest): def _get_exporter(self, **kwargs): return PythonItemExporter(**kwargs) def test_nested_item(self): i1 = TestItem(name=u'Joseph', age='22') - i2 = TestItem(name=u'Maria', age=i1) + i2 = dict(name=u'Maria', age=i1) i3 = TestItem(name=u'Jesus', age=i2) ie = self._get_exporter() exported = ie.export_item(i3) @@ -107,6 +120,7 @@ class PythonItemExporterTest(BaseItemExporterTest): self.assertEqual(type(exported['age'][0]), dict) self.assertEqual(type(exported['age'][0]['age'][0]), dict) + class PprintItemExporterTest(BaseItemExporterTest): def _get_exporter(self, **kwargs): @@ -115,6 +129,7 @@ class PprintItemExporterTest(BaseItemExporterTest): def _check_output(self): self._assert_expected_item(eval(self.output.getvalue())) + class PickleItemExporterTest(BaseItemExporterTest): def _get_exporter(self, **kwargs): @@ -150,48 +165,65 @@ class CsvItemExporterTest(BaseItemExporterTest): def _check_output(self): self.assertCsvEqual(self.output.getvalue(), 'age,name\r\n22,John\xc2\xa3\r\n') - def test_header(self): - output = BytesIO() - ie = CsvItemExporter(output, fields_to_export=self.i.fields.keys()) + def assertExportResult(self, item, expected, **kwargs): + fp = BytesIO() + ie = CsvItemExporter(fp, **kwargs) ie.start_exporting() - ie.export_item(self.i) + ie.export_item(item) ie.finish_exporting() - self.assertCsvEqual(output.getvalue(), 'age,name\r\n22,John\xc2\xa3\r\n') + self.assertCsvEqual(fp.getvalue(), expected) - output = BytesIO() - ie = CsvItemExporter(output, fields_to_export=['age']) - ie.start_exporting() - ie.export_item(self.i) - ie.finish_exporting() - self.assertCsvEqual(output.getvalue(), 'age\r\n22\r\n') + def test_header_export_all(self): + self.assertExportResult( + item=self.i, + fields_to_export=self.i.fields.keys(), + expected='age,name\r\n22,John\xc2\xa3\r\n', + ) - output = BytesIO() - ie = CsvItemExporter(output) - ie.start_exporting() - ie.export_item(self.i) - ie.export_item(self.i) - ie.finish_exporting() - self.assertCsvEqual(output.getvalue(), 'age,name\r\n22,John\xc2\xa3\r\n22,John\xc2\xa3\r\n') + def test_header_export_all_dict(self): + self.assertExportResult( + item=dict(self.i), + expected='age,name\r\n22,John\xc2\xa3\r\n', + ) - output = BytesIO() - ie = CsvItemExporter(output, include_headers_line=False) - ie.start_exporting() - ie.export_item(self.i) - ie.finish_exporting() - self.assertCsvEqual(output.getvalue(), '22,John\xc2\xa3\r\n') + def test_header_export_single_field(self): + for item in [self.i, dict(self.i)]: + self.assertExportResult( + item=item, + fields_to_export=['age'], + expected='age\r\n22\r\n', + ) + + def test_header_export_two_items(self): + for item in [self.i, dict(self.i)]: + output = BytesIO() + ie = CsvItemExporter(output) + ie.start_exporting() + ie.export_item(item) + ie.export_item(item) + ie.finish_exporting() + self.assertCsvEqual(output.getvalue(), 'age,name\r\n22,John\xc2\xa3\r\n22,John\xc2\xa3\r\n') + + def test_header_no_header_line(self): + for item in [self.i, dict(self.i)]: + self.assertExportResult( + item=item, + include_headers_line=False, + expected='22,John\xc2\xa3\r\n', + ) def test_join_multivalue(self): class TestItem2(Item): name = Field() friends = Field() - i = TestItem2(name='John', friends=['Mary', 'Paul']) - output = BytesIO() - ie = CsvItemExporter(output, include_headers_line=False) - ie.start_exporting() - ie.export_item(i) - ie.finish_exporting() - self.assertCsvEqual(output.getvalue(), '"Mary,Paul",John\r\n') + for cls in TestItem2, dict: + self.assertExportResult( + item=cls(name='John', friends=['Mary', 'Paul']), + include_headers_line=False, + expected='"Mary,Paul",John\r\n', + ) + class XmlItemExporterTest(BaseItemExporterTest): @@ -211,60 +243,62 @@ class XmlItemExporterTest(BaseItemExporterTest): return xmltuple(doc) return self.assertEqual(xmlsplit(first), xmlsplit(second), msg) + def assertExportResult(self, item, expected_value): + fp = BytesIO() + ie = XmlItemExporter(fp) + ie.start_exporting() + ie.export_item(item) + ie.finish_exporting() + self.assertXmlEquivalent(fp.getvalue(), expected_value) + def _check_output(self): expected_value = '\n22John\xc2\xa3' self.assertXmlEquivalent(self.output.getvalue(), expected_value) def test_multivalued_fields(self): - output = BytesIO() - item = TestItem(name=[u'John\xa3', u'Doe']) - ie = XmlItemExporter(output) - ie.start_exporting() - ie.export_item(item) - ie.finish_exporting() - expected_value = '\nJohn\xc2\xa3Doe' - self.assertXmlEquivalent(output.getvalue(), expected_value) + self.assertExportResult( + TestItem(name=[u'John\xa3', u'Doe']), + '\nJohn\xc2\xa3Doe' + ) def test_nested_item(self): - output = BytesIO() i1 = TestItem(name=u'foo\xa3hoo', age='22') - i2 = TestItem(name=u'bar', age=i1) + i2 = dict(name=u'bar', age=i1) i3 = TestItem(name=u'buz', age=i2) - ie = XmlItemExporter(output) - ie.start_exporting() - ie.export_item(i3) - ie.finish_exporting() - expected_value = '\n'\ - ''\ - ''\ - ''\ - '22'\ - 'foo\xc2\xa3hoo'\ - ''\ - 'bar'\ - ''\ - 'buz'\ - '' - self.assertXmlEquivalent(output.getvalue(), expected_value) + + self.assertExportResult(i3, + '\n' + '' + '' + '' + '' + '22' + 'foo\xc2\xa3hoo' + '' + 'bar' + '' + 'buz' + '' + '' + ) def test_nested_list_item(self): - output = BytesIO() i1 = TestItem(name=u'foo') - i2 = TestItem(name=u'bar') + i2 = dict(name=u'bar', v2={"egg": ["spam"]}) i3 = TestItem(name=u'buz', age=[i1, i2]) - ie = XmlItemExporter(output) - ie.start_exporting() - ie.export_item(i3) - ie.finish_exporting() - expected_value = '\n'\ - ''\ - ''\ - 'foo'\ - 'bar'\ - ''\ - 'buz'\ - '' - self.assertXmlEquivalent(output.getvalue(), expected_value) + + self.assertExportResult(i3, + '\n' + '' + '' + '' + 'foo' + 'barspam' + '' + 'buz' + '' + '' + ) class JsonLinesItemExporterTest(BaseItemExporterTest): @@ -280,7 +314,7 @@ class JsonLinesItemExporterTest(BaseItemExporterTest): def test_nested_item(self): i1 = TestItem(name=u'Joseph', age='22') - i2 = TestItem(name=u'Maria', age=i1) + i2 = dict(name=u'Maria', age=i1) i3 = TestItem(name=u'Jesus', age=i2) self.ie.start_exporting() self.ie.export_item(i3) @@ -306,13 +340,19 @@ class JsonItemExporterTest(JsonLinesItemExporterTest): exported = json.loads(self.output.getvalue().strip()) self.assertEqual(exported, [dict(self.i)]) - def test_two_items(self): + def assertTwoItemsExported(self, item): self.ie.start_exporting() - self.ie.export_item(self.i) - self.ie.export_item(self.i) + self.ie.export_item(item) + self.ie.export_item(item) self.ie.finish_exporting() exported = json.loads(self.output.getvalue()) - self.assertEqual(exported, [dict(self.i), dict(self.i)]) + self.assertEqual(exported, [dict(item), dict(item)]) + + def test_two_items(self): + self.assertTwoItemsExported(self.i) + + def test_two_dict_items(self): + self.assertTwoItemsExported(dict(self.i)) def test_nested_item(self): i1 = TestItem(name=u'Joseph\xa3', age='22') @@ -325,6 +365,18 @@ class JsonItemExporterTest(JsonLinesItemExporterTest): expected = {'name': u'Jesus', 'age': {'name': 'Maria', 'age': dict(i1)}} self.assertEqual(exported, [expected]) + def test_nested_dict_item(self): + i1 = dict(name=u'Joseph\xa3', age='22') + i2 = TestItem(name=u'Maria', age=i1) + i3 = dict(name=u'Jesus', age=i2) + self.ie.start_exporting() + self.ie.export_item(i3) + self.ie.finish_exporting() + exported = json.loads(self.output.getvalue()) + expected = {'name': u'Jesus', 'age': {'name': 'Maria', 'age': i1}} + self.assertEqual(exported, [expected]) + + class CustomItemExporterTest(unittest.TestCase): def test_exporter_custom_serializer(self): @@ -333,16 +385,17 @@ class CustomItemExporterTest(unittest.TestCase): if name == 'age': return str(int(value) + 1) else: - return super(CustomItemExporter, self).serialize_field(field, \ - name, value) + return super(CustomItemExporter, self).serialize_field(field, name, value) i = TestItem(name=u'John', age='22') ie = CustomItemExporter() - self.assertEqual( \ - ie.serialize_field(i.fields['name'], 'name', i['name']), 'John') - self.assertEqual( - ie.serialize_field(i.fields['age'], 'age', i['age']), '23') + self.assertEqual(ie.serialize_field(i.fields['name'], 'name', i['name']), 'John') + self.assertEqual(ie.serialize_field(i.fields['age'], 'age', i['age']), '23') + + i2 = {'name': u'John', 'age': '22'} + self.assertEqual(ie.serialize_field({}, 'name', i2['name']), 'John') + self.assertEqual(ie.serialize_field({}, 'age', i2['age']), '23') if __name__ == '__main__': diff --git a/tests/test_engine.py b/tests/test_engine.py index 52c8e5752..04fae02c0 100644 --- a/tests/test_engine.py +++ b/tests/test_engine.py @@ -28,11 +28,13 @@ from scrapy.contrib.linkextractors import LinkExtractor from scrapy.http import Request from scrapy.utils.signal import disconnect_all + class TestItem(Item): name = Field() url = Field() price = Field() + class TestSpider(Spider): name = "scrapytest.org" allowed_domains = ["scrapytest.org", "localhost"] @@ -41,6 +43,8 @@ class TestSpider(Spider): name_re = re.compile("

(.*?)

", re.M) price_re = re.compile(">Price: \$(.*?)<", re.M) + item_cls = TestItem + def parse(self, response): xlink = LinkExtractor() itemre = re.compile(self.itemurl_re) @@ -49,7 +53,7 @@ class TestSpider(Spider): yield Request(url=link.url, callback=self.parse_item) def parse_item(self, response): - item = TestItem() + item = self.item_cls() m = self.name_re.search(response.body) if m: item['name'] = m.group(1) @@ -65,6 +69,10 @@ class TestDupeFilterSpider(TestSpider): return Request(url) # dont_filter=False +class DictItemsSpider(TestSpider): + item_cls = dict + + def start_test_site(debug=False): root_dir = os.path.join(tests_datadir, "test_site") r = static.File(root_dir) @@ -81,15 +89,14 @@ def start_test_site(debug=False): class CrawlerRun(object): """A class to run the crawler and keep track of events occurred""" - def __init__(self, with_dupefilter=False): + def __init__(self, spider_class): self.spider = None self.respplug = [] self.reqplug = [] self.reqdropped = [] self.itemresp = [] self.signals_catched = {} - self.spider_class = TestSpider if not with_dupefilter else \ - TestDupeFilterSpider + self.spider_class = spider_class def run(self): self.port = start_test_site() @@ -152,14 +159,17 @@ class EngineTest(unittest.TestCase): @defer.inlineCallbacks def test_crawler(self): - self.run = CrawlerRun() - yield self.run.run() - self._assert_visited_urls() - self._assert_scheduled_requests(urls_to_visit=8) - self._assert_downloaded_responses() - self._assert_scraped_items() - self._assert_signals_catched() - self.run = CrawlerRun(with_dupefilter=True) + + for spider in TestSpider, DictItemsSpider: + self.run = CrawlerRun(spider) + yield self.run.run() + self._assert_visited_urls() + self._assert_scheduled_requests(urls_to_visit=8) + self._assert_downloaded_responses() + self._assert_scraped_items() + self._assert_signals_catched() + + self.run = CrawlerRun(TestDupeFilterSpider) yield self.run.run() self._assert_scheduled_requests(urls_to_visit=7) self._assert_dropped_requests() diff --git a/tests/test_pipeline_files.py b/tests/test_pipeline_files.py index 0a1737c44..84fe4927d 100644 --- a/tests/test_pipeline_files.py +++ b/tests/test_pipeline_files.py @@ -142,35 +142,40 @@ class DeprecatedFilesPipelineTestCase(unittest.TestCase): class FilesPipelineTestCaseFields(unittest.TestCase): def test_item_fields_default(self): - from scrapy.contrib.pipeline.files import FilesPipeline class TestItem(Item): name = Field() file_urls = Field() files = Field() - url = 'http://www.example.com/files/1.txt' - item = TestItem({'name': 'item1', 'file_urls': [url]}) - pipeline = FilesPipeline.from_settings(Settings({'FILES_STORE': 's3://example/files/'})) - requests = list(pipeline.get_media_requests(item, None)) - self.assertEqual(requests[0].url, url) - results = [(True, {'url': url})] - pipeline.item_completed(results, item, None) - self.assertEqual(item['files'], [results[0][1]]) + + for cls in TestItem, dict: + url = 'http://www.example.com/files/1.txt' + item = cls({'name': 'item1', 'file_urls': [url]}) + pipeline = FilesPipeline.from_settings(Settings({'FILES_STORE': 's3://example/files/'})) + requests = list(pipeline.get_media_requests(item, None)) + self.assertEqual(requests[0].url, url) + results = [(True, {'url': url})] + pipeline.item_completed(results, item, None) + self.assertEqual(item['files'], [results[0][1]]) def test_item_fields_override_settings(self): - from scrapy.contrib.pipeline.files import FilesPipeline class TestItem(Item): name = Field() files = Field() stored_file = Field() - url = 'http://www.example.com/files/1.txt' - item = TestItem({'name': 'item1', 'files': [url]}) - pipeline = FilesPipeline.from_settings(Settings({'FILES_STORE': 's3://example/files/', - 'FILES_URLS_FIELD': 'files', 'FILES_RESULT_FIELD': 'stored_file'})) - requests = list(pipeline.get_media_requests(item, None)) - self.assertEqual(requests[0].url, url) - results = [(True, {'url': url})] - pipeline.item_completed(results, item, None) - self.assertEqual(item['stored_file'], [results[0][1]]) + + for cls in TestItem, dict: + url = 'http://www.example.com/files/1.txt' + item = cls({'name': 'item1', 'files': [url]}) + pipeline = FilesPipeline.from_settings(Settings({ + 'FILES_STORE': 's3://example/files/', + 'FILES_URLS_FIELD': 'files', + 'FILES_RESULT_FIELD': 'stored_file' + })) + requests = list(pipeline.get_media_requests(item, None)) + self.assertEqual(requests[0].url, url) + results = [(True, {'url': url})] + pipeline.item_completed(results, item, None) + self.assertEqual(item['stored_file'], [results[0][1]]) class ItemWithFiles(Item): diff --git a/tests/test_pipeline_images.py b/tests/test_pipeline_images.py index a3b1059ef..f5750b4fc 100644 --- a/tests/test_pipeline_images.py +++ b/tests/test_pipeline_images.py @@ -168,35 +168,40 @@ class DeprecatedImagesPipelineTestCase(unittest.TestCase): class ImagesPipelineTestCaseFields(unittest.TestCase): def test_item_fields_default(self): - from scrapy.contrib.pipeline.images import ImagesPipeline class TestItem(Item): name = Field() image_urls = Field() images = Field() - url = 'http://www.example.com/images/1.jpg' - item = TestItem({'name': 'item1', 'image_urls': [url]}) - pipeline = ImagesPipeline.from_settings(Settings({'IMAGES_STORE': 's3://example/images/'})) - requests = list(pipeline.get_media_requests(item, None)) - self.assertEqual(requests[0].url, url) - results = [(True, {'url': url})] - pipeline.item_completed(results, item, None) - self.assertEqual(item['images'], [results[0][1]]) + + for cls in TestItem, dict: + url = 'http://www.example.com/images/1.jpg' + item = cls({'name': 'item1', 'image_urls': [url]}) + pipeline = ImagesPipeline.from_settings(Settings({'IMAGES_STORE': 's3://example/images/'})) + requests = list(pipeline.get_media_requests(item, None)) + self.assertEqual(requests[0].url, url) + results = [(True, {'url': url})] + pipeline.item_completed(results, item, None) + self.assertEqual(item['images'], [results[0][1]]) def test_item_fields_override_settings(self): - from scrapy.contrib.pipeline.images import ImagesPipeline class TestItem(Item): name = Field() image = Field() stored_image = Field() - url = 'http://www.example.com/images/1.jpg' - item = TestItem({'name': 'item1', 'image': [url]}) - pipeline = ImagesPipeline.from_settings(Settings({'IMAGES_STORE': 's3://example/images/', - 'IMAGES_URLS_FIELD': 'image', 'IMAGES_RESULT_FIELD': 'stored_image'})) - requests = list(pipeline.get_media_requests(item, None)) - self.assertEqual(requests[0].url, url) - results = [(True, {'url': url})] - pipeline.item_completed(results, item, None) - self.assertEqual(item['stored_image'], [results[0][1]]) + + for cls in TestItem, dict: + url = 'http://www.example.com/images/1.jpg' + item = cls({'name': 'item1', 'image': [url]}) + pipeline = ImagesPipeline.from_settings(Settings({ + 'IMAGES_STORE': 's3://example/images/', + 'IMAGES_URLS_FIELD': 'image', + 'IMAGES_RESULT_FIELD': 'stored_image' + })) + requests = list(pipeline.get_media_requests(item, None)) + self.assertEqual(requests[0].url, url) + results = [(True, {'url': url})] + pipeline.item_completed(results, item, None) + self.assertEqual(item['stored_image'], [results[0][1]]) def _create_image(format, *a, **kw):