mirror of https://github.com/scrapy/scrapy.git
Implements a DuplicatesFilterMiddleware as spidermiddleware, a wraper using a minimal defined API of a filtering class configurable by settings. Enabling this middleware doesn't gives us same functionality compared to scheduler duplicate filter builtin, but it filter the most important source for duplicate requests, the spiders. What requests aren't filtered by new middleware? The ones originated from any part of scrapy outside of spiders, like S3 images requests or any other request manually schedule using ``scrapyengine.schedule()`` method. Previously, we usually added dont_filter=True to requests created outside of spiders to avoid collisions downloading same pages than spider. Now, this is not required anymore because new middleware filters just the spider generated requests. There is a caveat, as usual downloadmiddlewares can returns a Request object at any point of the chain, and that request is scheduled and downloaded as usual too. One of the downloadmiddlewares using this feature is RedirectMiddleware that counts on scheduler filtering builtin to avoid redirection loops. I think we can implement a request time to live decreasing counter and add it to request's ``meta`` attribute with a default value if not present, and decrement each time the request is redirected. --HG-- extra : convert_revision : svn%3Ab85faa78-f9eb-468e-a121-7cced6da292c%40846 |
||
|---|---|---|
| scrapy | ||
| sites | ||
| .hgignore | ||