Scrapy, a fast high-level web crawling & scraping framework for Python.
Go to file
Vostretsov Nikita 67ee0b097f
Remove specifics of downstream request queues from scheduler (#3884)
* move serialization/deserialization logic to downstream queues

* make memory queues conform to common interface

* make ScrapyPriorityQueue conform common interface

* ScrapyPriorityQueue works with disk

* make key as string

* return list instead of dict as earlier

* downloader aware pq works with new interface

* we don`t need these methods anymore

* create directories for files

* remove dummy priority

* remove priority as parameter, let every queue decide for itself

* rename obj to request

* DownloaderAwarePriorityQueue is too thin wrapper around _SlotPriorityQueues, just remove second one

* remove priority as parameter, let every queue decide for itself

* rename argument

* more granular class separation

* python2 compatible

* one more argument for common interface

* more simple downstream queue interface

* single place for easier customization

* rename function

* shorter

* shorter

* use named arguments

* fix typo

* add docstring

* Update scrapy/pqueues.py

Co-Authored-By: Mikhail Korobov <kmike84@gmail.com>

* Update scrapy/pqueues.py

Co-Authored-By: Mikhail Korobov <kmike84@gmail.com>

* 4 spaces indentation

* we ok with existing directories

* remove unused import

* rename method

* remove unused imports

* it has no sense now

* relining

* note about queues

* add value

* Revert "it has no sense now"

This reverts commit b61604275b.

* pep8 E261

* pep8 E303

* pep8 E501

* pep8 E123

* pep8 E123

* use create instance

* remove excessive import

Co-authored-by: Mikhail Korobov <kmike84@gmail.com>
2020-02-22 17:02:57 +05:00
.github/ISSUE_TEMPLATE Fix case of GitHub. 2019-10-05 10:09:14 +10:00
artwork Fixed artwork/README formatting 2020-01-15 08:54:25 +04:00
debian Remove deprecated xlib module 2019-09-13 14:32:05 -03:00
docs Merge pull request #4331 from Gallaecio/response-cb-kwargs 2020-02-19 22:40:14 +05:00
extras add zsh -h autocomplete option 2020-01-27 18:24:57 +03:00
scrapy Remove specifics of downstream request queues from scheduler (#3884) 2020-02-22 17:02:57 +05:00
sep Fix a spelling error: ie. → i.e. (#4338) 2020-02-18 17:58:31 +01:00
tests Remove specifics of downstream request queues from scheduler (#3884) 2020-02-22 17:02:57 +05:00
.bandit.yml Mark bandit’s 402 check as addressed by #4180 (#4181) 2019-12-05 14:48:31 +01:00
.bumpversion.cfg Bump version: 1.7.0 → 1.8.0 2019-10-29 12:57:02 +01:00
.coveragerc Remove deprecated xlib module 2019-09-13 14:32:05 -03:00
.gitignore remove duplicated entry in gitignore 2019-03-22 18:56:25 +01:00
.readthedocs.yml Use Python 3.7 to build the documentation 2019-12-19 12:06:15 +01:00
.travis.yml Remove the py35-asyncio env for 3.5 from Travis. 2020-01-03 21:38:05 +05:00
AUTHORS added Nicolas Ramirez to AUTHORS 2013-03-14 12:44:39 -03:00
CODE_OF_CONDUCT.md Make punctuation consistent 2019-10-16 22:44:42 +02:00
CONTRIBUTING.md Be consistent with domain used for links to documentation website 2019-01-31 01:28:53 -03:00
INSTALL Be consistent with domain used for links to documentation website 2019-01-31 01:28:53 -03:00
LICENSE added oxford commas to LICENSE 2018-06-01 21:48:43 -03:00
MANIFEST.in Include additional files in sdists 2018-11-16 13:38:19 -05:00
Makefile.buildbot Generated version as pep440 and dpkg compatible 2015-06-16 00:16:09 +00:00
NEWS added NEWS file pointing to docs/news.rst 2012-04-28 23:32:51 -03:00
README.rst Merge pull request #4059 from josealberto4444/master 2019-11-15 15:39:27 +05:00
appveyor.yml Set the cloned directory as PYTHONPATH in appveyor.yml 2019-06-11 15:56:27 +02:00
codecov.yml codecov config: disable project check, tweak PR comments 2017-05-19 00:01:27 +05:00
conftest.py Support yield in async def callbacks. 2020-02-07 21:32:45 +05:00
pytest.ini fix E22X flake8 2020-02-21 08:39:14 +01:00
setup.cfg Build universal wheels 2016-03-01 11:00:20 +01:00
setup.py Remove six from requirements and setup files 2019-11-03 12:30:34 -03:00
tox.ini Remove elusive six occurrence from tox.ini 2020-02-05 13:27:54 -03:00

README.rst

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>

Scrapy

Overview

Scrapy is a fast high-level web crawling and web scraping framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

Check the Scrapy homepage at https://scrapy.org for more information, including a list of features.

Requirements

  • Python 3.5+
  • Works on Linux, Windows, Mac OSX, BSD

Install

The quick way:

pip install scrapy

See the install section in the documentation at https://docs.scrapy.org/en/latest/intro/install.html for more details.

Documentation

Documentation is available online at https://docs.scrapy.org/ and in the docs directory.

Releases

You can check https://docs.scrapy.org/en/latest/news.html for the release notes.

Community (blog, twitter, mail list, IRC)

See https://scrapy.org/community/ for details.

Contributing

See https://docs.scrapy.org/en/master/contributing.html for details.

Code of Conduct

Please note that this project is released with a Contributor Code of Conduct (see https://github.com/scrapy/scrapy/blob/master/CODE_OF_CONDUCT.md).

By participating in this project you agree to abide by its terms. Please report unacceptable behavior to opensource@scrapinghub.com.

Companies using Scrapy

See https://scrapy.org/companies/ for a list.

Commercial Support

See https://scrapy.org/support/ for details.

</html>