Commit Graph

4159 Commits

Author SHA1 Message Date
Eugenio Lacuesta 60c2ef86f0 Revert "Default values for OffsiteMiddleware"
This reverts commit ba29435138.
2018-07-15 16:47:55 -03:00
Eugenio Lacuesta c5fa0ae6bc Untested experiment 2018-07-14 20:15:53 -03:00
Eugenio Lacuesta 985ab636cf Store output methods on the 'methods' dict 2018-07-01 19:11:59 -03:00
Eugenio Lacuesta ba29435138 Default values for OffsiteMiddleware
For some reason test_crawl.py seems to be skipping the spider_opened
method, which initializes the host_regex instance variable
2018-06-25 15:02:10 -03:00
Eugenio Lacuesta 4740dca8f2 Deferred-like process_output/process_exception chain 2018-06-25 11:16:19 -03:00
Eugenio Lacuesta ab48837f09 Merge branch 'master' 2018-06-24 20:31:37 -03:00
Vostretsov Nikita 72d0899bce Return non-zero exit code from scrapy commands in case of spider bootstrap errors
* method to detect spider creation in crawler

* correct method name

* method to know if crawlers has spiders

* we do not need to issue requests

* set exit code accordingly to spiders in crawlers

* more portable way to check ofr exceptions

* more clear way

* test cases for several spiders per crawler

* grammatically correct name for method

* method is private

* grammatically correct name for method

* method is private

* remove unused import

* correct order of imports

* changes mechanism of obtaining spider status from method to object member

* rename tests
2018-06-14 19:58:48 +05:00
Daniel Graña c6030ce8c6
Merge pull request #3231 from starrify/updating-argument-of-cookiejar-clear
[MRG+2] Added: Allowing optional arguments for `scrapy.http.cookies.CookieJar.clear`
2018-06-01 21:52:32 -03:00
Fredrik Bergenlid 6a2d2c3b77 Improve gunzip performance for big files 2018-06-01 21:38:07 +02:00
Pengyu Chen e75f721c04
Added: Allowing optional arguments for `scrapy.http.cookies.CookieJar.clear` 2018-04-23 22:08:28 +08:00
rhoboro 560ee623fd set defalut value "" to FILES_STORE_GCS_ACL 2018-04-13 19:00:27 +09:00
rhoboro 464973489e Using bucket's default object ACL 2018-04-13 12:06:39 +09:00
rhoboro 5254ac393b added test for gcs policy 2018-04-03 18:06:34 +09:00
rhoboro 8e8994c6b5 add acl support for gcs 2018-04-02 15:36:47 +09:00
Daniel Graña 6c3970e672
Merge pull request #3153 from virmht/new_bug
[MRG+1] Fixed bug FormRequest.from_response() clickdata ignores input[type=image]
2018-03-21 16:32:12 -03:00
Viral Mehta dd064413a4 corrected syntax error in XPath 2018-03-19 19:28:41 +05:30
Viral Mehta a5acc9373f Resolving Comments 2018-03-19 18:19:39 +05:30
Viral Mehta e25e2afe17 Removed unnecessary print statements 2018-03-17 18:20:14 +05:30
Viral Mehta ff5f717f7a Fixed formatting issues 2018-03-17 18:17:48 +05:30
Daniel Graña 6cc6bbb5fc
Merge pull request #3166 from lucywang000/catch-tls-certificate-error
catch CertificateError in tls verification
2018-03-14 11:47:58 -03:00
Lucy Wang 2c58da19a6 update docstring of ScrapyClientTLSOptions 2018-03-14 09:27:59 +08:00
Lucy Wang 1a2f0193a3 fix tests on jessie 2018-03-13 19:14:52 +08:00
siulkilulki 6a7cdf9a6c [MRG+1] Add 'flv' to ignored video extensions. (#3165) 2018-03-13 10:35:27 +03:00
Lucy Wang 13a74d77e2 catch CertificateError in tls verification 2018-03-12 22:25:19 +08:00
Viral Mehta d5b7ebcfdc Fixed bug FormRequest.from_response() clickdata ignores input[type=image] 2018-03-03 18:17:49 +05:30
NewUserHa acd2b8d43b [MRG+1] Fix part of issue #3128 - None should not be a valid type for 'url' in Response.follow (#3131)
* fix one issue of issue#3128

because @kmike posted: 'If url is '', Scrapy should follow the same page, this is an intended behavior.'

*  fix one issue of issue#3128

because @kmike posted: 'If url is '', Scrapy should follow the same page, this is an intended behavior.'
2018-02-22 03:37:26 +05:00
Daniel Graña c56f7b3c8d
Merge pull request #3113 from WenbinZhang/master
[MRG+1] Update robotstxt.py
2018-02-08 18:22:56 -03:00
Daniel Graña 68e45d32e0
Merge pull request #3115 from scrapy/telnet-log-level
[MRG+1] use INFO log level to show telnet host/port
2018-02-08 18:22:36 -03:00
Konstantin Lopuhin 936dbc7bf6
Merge branch 'master' into master 2018-02-08 23:43:29 +03:00
Eugenio Lacuesta 6edd4114c4 Clarify comment about Pyhton versions 2018-02-08 15:47:20 -03:00
Eugenio Lacuesta a56540877c Do not serialize unpickable objects (py3) 2018-02-08 15:03:57 -03:00
Konstantin Lopuhin eff469245a
Merge pull request #3100 from scrapy/robots-stats
[MRG+1] more stats for RobotsTxtMiddleware
2018-02-08 12:01:36 +03:00
Mikhail Korobov 0c374c00fb use INFO log level to show telnet host/port 2018-02-08 05:09:02 +05:00
Wenbin Zhang 4d5e5378bd
Update robotstxt.py
Add message to IgnoreRequest exception so that it can be detectedin the errbak method of a spider
2018-02-07 10:59:32 -05:00
Mikhail Korobov 6f264ab190 more stats for RobotsTxtMiddleware 2018-01-30 05:47:28 +05:00
Jesús Losada c1916626c1 Fix OS signal names 2018-01-27 21:24:15 +00:00
Yash Sharma 1d1581266c Changed some documentations (#3089)
DOC typo fix in defer_fail docstring
2018-01-26 01:12:17 +05:00
Mikhail Korobov 7c9e32213d
Merge pull request #3059 from jesuslosada/fix-typo
Fix typo in comment
2018-01-11 01:34:11 +05:00
Jesús Losada 61c0b14782 Fix typo in comment 2018-01-01 16:03:55 +00:00
Daniel Graña 2dee191374
Merge branch 'master' into extending-s3-files-store 2017-12-31 16:44:38 -03:00
Mikhail Korobov aa83e159c9 Bump version: 1.4.0 → 1.5.0 2017-12-30 02:09:52 +05:00
Daniel Graña 57d04aa960
Merge pull request #2767 from redapple/http-proxy-endpoint-key
[MRG+1] Use HTTP pool and proper endpoint key for ProxyAgent
2017-12-26 14:12:47 -03:00
Konstantin Lopuhin a21b800419
Merge pull request #3011 from Jane222/master
[MRG+1] Issues a warning when user puts a URL into allowed_domains (#2250)
2017-12-08 15:45:16 +03:00
Jana Cavojska 22c68baf99 url_pattern is now being compiled before entering the loop 2017-12-07 18:38:29 +01:00
Daniel Graña 3cf0332ec3
Merge pull request #2957 from ScrapingLab/add_meta_json_to_parse_command
[MRG+1] Scrapy Command: add --meta/-m to the "parse" command to pass additional meta data into the request
2017-11-29 16:26:48 -03:00
Jana Cavojska 454d5e5733 checking for subclass of URLWarning instead of checking error message text when URL in allowed_domains 2017-11-26 20:07:04 +01:00
Jana Cavojska 8ec3b476b0 triggering a warning when user puts URL in allowed_domains now covered by test 2017-11-26 16:36:15 +01:00
Jana Cavojska 91ff194d1e looping over allowed_domains directly instead of via index 2017-11-20 21:23:31 +01:00
KosayJabre 5441cc18e4
Separated import statements
Just separated the import statements. Tiny change - testing GitHub!
2017-11-19 18:09:38 -04:00
Jana Cavojska 62a6261028 Issues a warning when user puts a URL into allowed_domains (#2250) 2017-11-18 20:03:59 +01:00