mirror of https://github.com/scrapy/scrapy.git
crawls. * requests are serialized (using marshal by default) and stored on disk, using one queue per priority * request priorities must be integers now * breadh-first and depth-first crawling orders can now be configured through a new DEPTH_PRIORITY setting (see doc). backwards compatilibty with SCHEDULER_ORDER was kept. * requests that can't be serialized (for example, non serializable callbacks) are always kept in memory queues * adapted crawl spider to work with persitent scheduler |
||
|---|---|---|
| .. | ||
| downloader | ||
| __init__.py | ||
| engine.py | ||
| scheduler.py | ||
| scraper.py | ||
| spidermw.py | ||