Commit Graph

7772 Commits

Author SHA1 Message Date
Mikhail Korobov 5f407cf657
Merge pull request #3961 from OmarFarrag/ftp_files#3928
Add FTPFileStore to FilesPipeline
2020-01-25 04:58:23 +05:00
Mikhail Korobov 7a62bd310c
Merge pull request #4126 from elacuesta/from_crawler_downloader_handlers
Download handlers: from_crawler factory method, take crawler in __init__
2020-01-25 04:52:56 +05:00
OmarFarrag 9e6d5573f1 Fix Flake8 errors 2020-01-24 15:58:52 +02:00
OmarFarrag 40e0a11aa8 Fix Flake8 errors 2020-01-24 15:51:48 +02:00
OmarFarrag f5d9eb15f8 use `__future__` imports at the begining of the file 2020-01-24 15:06:40 +02:00
OmarFarrag fc98aa6b67
Merge branch 'master' into ftp_files#3928 2020-01-24 14:52:40 +02:00
OmarFarrag c544c0d2b8 Use context management with `FTP` 2020-01-24 14:36:16 +02:00
Mikhail Korobov bd54f22fef
Merge pull request #4282 from petervandenabeele/patch-1
fix logical documentation error with PER_DOMAIN or PER_DOMAIN
2020-01-24 02:01:22 +05:00
Mikhail Korobov 6a98d660e5
Merge pull request #3551 from jpbalarini/change_scraper_slot
[MRG+1] Add ability to change max_active_size by setting
2020-01-23 23:40:05 +05:00
Mikhail Korobov f80c7776ae
Merge pull request #4008 from elacuesta/docs_request_errback
Request: remove restriction about errback without callback
2020-01-23 23:12:44 +05:00
Peter Vandenabeele 7d5cebcf77
fix logical documentation error with PER_DOMAIN or PER_DOMAIN 2020-01-23 09:08:21 +01:00
Mikhail Korobov c0a7dfbc01
Merge pull request #4057 from elacuesta/response_follow_all
Response.follow_all
2020-01-23 02:15:24 +05:00
Eugenio Lacuesta c75cf15b7a
Update CSS selectors in tutorial 2020-01-22 10:38:59 -03:00
OmarFarrag 06ab668ec7 Use kwargs-only parameters in `ftp_store_file` 2020-01-22 03:48:07 +02:00
OmarFarrag 8ea8f14827
Update scrapy/utils/ftp.py
Co-Authored-By: Mikhail Korobov <kmike84@gmail.com>
2020-01-20 18:19:36 +02:00
JP Balarini 0f2d871d88 Use PEP 515 style for SCRAPER_SLOT_MAX_ACTIVE_SIZE documentation 2020-01-20 11:28:28 -03:00
Juan Pablo Balarini eaa8ed02d0 Add ability to change max_active_size by settings 2020-01-20 11:27:58 -03:00
Mikhail Korobov 50310fc0f9
Merge pull request #4270 from wRAR/asyncio-pipelines
async def support in pipelines
2020-01-16 03:28:09 +05:00
Andrey Rakhmatullin 7d85984880 Use get_from_asyncio_queue in the pipeline test. 2020-01-09 14:49:16 +05:00
Andrey Rakhmatullin 9d8c54c0f2 Fix/ignore flake8 problems. 2020-01-09 14:49:02 +05:00
Andrey Rakhmatullin bdef948aae Mark the asyncio pipelines test as only_asyncio. 2020-01-09 14:19:02 +05:00
Andrey Rakhmatullin bfdd552a32 Add a test for pipelines using asyncio. 2020-01-09 14:19:02 +05:00
Andrey Rakhmatullin 1f9cef787d Add async def support to pipelines. 2020-01-09 14:19:02 +05:00
Andrey Rakhmatullin 8117566974 Add utils.defer.deferred_f_from_coro_f. 2020-01-09 14:19:02 +05:00
Eugenio Lacuesta 2e405d2d5c
Merge branch 'master' into response_follow_all 2020-01-05 00:33:19 -03:00
Mikhail Korobov ce618fb6f2
Merge pull request #4259 from scrapy/asyncio-mw
Asyncio support in downloader middlewares
2020-01-03 22:28:41 +05:00
Andrey Rakhmatullin b2dd379bc2 Remove the py35-asyncio env for 3.5 from Travis. 2020-01-03 21:38:05 +05:00
Andrey Rakhmatullin 2b9254c2bd Add a test function that uses asyncio.Queue(). 2019-12-31 17:54:41 +05:00
Andrey Rakhmatullin e3b8ba6188 Run py35-asyncio also on 3.5.2 to test Xenial. 2019-12-31 17:54:01 +05:00
Andrey Rakhmatullin 16787f5bf4 Merge middleware tests back as we don't need to set the setting anymore. 2019-12-30 12:02:19 +05:00
Andrey Rakhmatullin 50aa6ef22c Add deferred_from_coro. 2019-12-30 11:46:45 +05:00
Andrey Rakhmatullin 5cf1ac0005 Move the asyncio downloader mw test to a separate class. 2019-12-30 11:46:45 +05:00
Andrey Rakhmatullin 3603644552 Add a non-asyncio async def middleware test. 2019-12-30 11:46:45 +05:00
Andrey Rakhmatullin 21f50c795a Add async def support to downloader middlewares. 2019-12-30 11:46:45 +05:00
1um0s 14d4428e70 Rephrasing documentation for image and file pipelines (#4252)
* scrapy#4034 Clarify documentation for image and file pipelines

* scrapy#4034 Clarify documentation for file pipeline

* scrapy#4034 Simplify documentation for pipeline

* scrapy#4034 Simplify documentation for pipeline

* scrapy#4034 Clarify documentation for image and file pipelines

* scrapy#4034 Clarify documentation for file pipeline

* scrapy#4034 Simplify documentation for pipeline

* scrapy#4034 Simplify documentation for pipeline

* scrapy#4034 Revert image, file pipeline docs. Enhance custom media pipeline docs.

* scrapy#4034 rebase master

* scrapy#4034 Clarify documentation for image and file pipelines

* scrapy#4034 Clarify documentation for file pipeline

* scrapy#4034 Simplify documentation for pipeline

* scrapy#4034 Simplify documentation for pipeline

* scrapy#4034 Clarify documentation for image and file pipelines

* scrapy#4034 Clarify documentation for file pipeline

* scrapy#4034 Simplify documentation for pipeline

* scrapy#4034 Simplify documentation for pipeline

* scrapy#4034 Revert image, file pipeline docs. Enhance custom media pipeline docs.

* scrapy#4034 rebase master

* Rebase master

* Add class to media pipeline docs

Co-Authored-By: elacuesta <elacuesta@users.noreply.github.com>

Co-authored-by: elacuesta <elacuesta@users.noreply.github.com>
2019-12-30 00:56:22 +05:00
Mikhail Korobov f0ae673452
Merge pull request #4258 from atul-g/patch-1
Edited the link provided to homepage of lxml's website
2019-12-30 00:55:15 +05:00
Mikhail Korobov bb991cd303
Merge pull request #4010 from scrapy/asyncio-base
Base support for asyncio
2019-12-30 00:51:28 +05:00
Atul Gopinathan 82861c73c8
Edited the link of the homepage of lxml website
The link "https://lxml.de" is redirecting to a completely different and unintended website. I changed the link to the index page of lxml's official website. I thought of changing it to the PyPi page of lxml, but even they are providing the same "https://lxml.de" link which doesn't seem to be working now.
2019-12-27 22:57:58 +05:30
Andrey Rakhmatullin dc1ee09481 Rename ASYNCIO_ENABLED to ASYNCIO_REACTOR, change the logic accordingly. 2019-12-27 21:56:28 +05:00
Andrey Rakhmatullin f75ccc997a FIx a typo in the only_asyncio fixture. 2019-12-27 19:48:54 +05:00
Andrey Rakhmatullin 30ebd05a5f Simplify the tox asyncio entries. 2019-12-27 00:05:14 +05:00
Andrey Rakhmatullin 37ac47ff80 Fix a deprecation warning. 2019-12-26 20:46:54 +05:00
Andrey Rakhmatullin 87ece066ca Remove conditional asyncio imports. 2019-12-26 20:41:06 +05:00
Eugenio Lacuesta ab54e0d33e
Keyword-only args for S3DownloadHandler 2019-12-23 20:37:18 -03:00
Eugenio Lacuesta 982a66f9fb
[test] Download handler: avoid passing settings if not necessary 2019-12-23 20:28:17 -03:00
Eugenio Lacuesta 9a75b46fb8
Explicit argument names 2019-12-23 20:26:58 -03:00
Eugenio Lacuesta 2fb160e3ba
Use settings instead of crawler 2019-12-23 20:24:16 -03:00
Mikhail Korobov 4d594b8c2b
Merge pull request #4193 from Gallaecio/lgtm
Use super().__init__ in BaseItemExporter subclasses
2019-12-23 23:52:43 +05:00
Eugenio Lacuesta e2e15d6651
Downloader handlers: sort imports 2019-12-23 10:48:19 -03:00
Eugenio Lacuesta a6ec89251e
Downloader handlers: crawler=None in __init__ 2019-12-23 10:47:08 -03:00