Eugenio Lacuesta
60c2ef86f0
Revert "Default values for OffsiteMiddleware"
...
This reverts commit ba29435138 .
2018-07-15 16:47:55 -03:00
Eugenio Lacuesta
c5fa0ae6bc
Untested experiment
2018-07-14 20:15:53 -03:00
Eugenio Lacuesta
985ab636cf
Store output methods on the 'methods' dict
2018-07-01 19:11:59 -03:00
Eugenio Lacuesta
ba29435138
Default values for OffsiteMiddleware
...
For some reason test_crawl.py seems to be skipping the spider_opened
method, which initializes the host_regex instance variable
2018-06-25 15:02:10 -03:00
Eugenio Lacuesta
4740dca8f2
Deferred-like process_output/process_exception chain
2018-06-25 11:16:19 -03:00
Eugenio Lacuesta
ab48837f09
Merge branch 'master'
2018-06-24 20:31:37 -03:00
Vostretsov Nikita
72d0899bce
Return non-zero exit code from scrapy commands in case of spider bootstrap errors
...
* method to detect spider creation in crawler
* correct method name
* method to know if crawlers has spiders
* we do not need to issue requests
* set exit code accordingly to spiders in crawlers
* more portable way to check ofr exceptions
* more clear way
* test cases for several spiders per crawler
* grammatically correct name for method
* method is private
* grammatically correct name for method
* method is private
* remove unused import
* correct order of imports
* changes mechanism of obtaining spider status from method to object member
* rename tests
2018-06-14 19:58:48 +05:00
Daniel Graña
c6030ce8c6
Merge pull request #3231 from starrify/updating-argument-of-cookiejar-clear
...
[MRG+2] Added: Allowing optional arguments for `scrapy.http.cookies.CookieJar.clear`
2018-06-01 21:52:32 -03:00
Fredrik Bergenlid
6a2d2c3b77
Improve gunzip performance for big files
2018-06-01 21:38:07 +02:00
Pengyu Chen
e75f721c04
Added: Allowing optional arguments for `scrapy.http.cookies.CookieJar.clear`
2018-04-23 22:08:28 +08:00
rhoboro
560ee623fd
set defalut value "" to FILES_STORE_GCS_ACL
2018-04-13 19:00:27 +09:00
rhoboro
464973489e
Using bucket's default object ACL
2018-04-13 12:06:39 +09:00
rhoboro
5254ac393b
added test for gcs policy
2018-04-03 18:06:34 +09:00
rhoboro
8e8994c6b5
add acl support for gcs
2018-04-02 15:36:47 +09:00
Daniel Graña
6c3970e672
Merge pull request #3153 from virmht/new_bug
...
[MRG+1] Fixed bug FormRequest.from_response() clickdata ignores input[type=image]
2018-03-21 16:32:12 -03:00
Viral Mehta
dd064413a4
corrected syntax error in XPath
2018-03-19 19:28:41 +05:30
Viral Mehta
a5acc9373f
Resolving Comments
2018-03-19 18:19:39 +05:30
Viral Mehta
e25e2afe17
Removed unnecessary print statements
2018-03-17 18:20:14 +05:30
Viral Mehta
ff5f717f7a
Fixed formatting issues
2018-03-17 18:17:48 +05:30
Daniel Graña
6cc6bbb5fc
Merge pull request #3166 from lucywang000/catch-tls-certificate-error
...
catch CertificateError in tls verification
2018-03-14 11:47:58 -03:00
Lucy Wang
2c58da19a6
update docstring of ScrapyClientTLSOptions
2018-03-14 09:27:59 +08:00
Lucy Wang
1a2f0193a3
fix tests on jessie
2018-03-13 19:14:52 +08:00
siulkilulki
6a7cdf9a6c
[MRG+1] Add 'flv' to ignored video extensions. ( #3165 )
2018-03-13 10:35:27 +03:00
Lucy Wang
13a74d77e2
catch CertificateError in tls verification
2018-03-12 22:25:19 +08:00
Viral Mehta
d5b7ebcfdc
Fixed bug FormRequest.from_response() clickdata ignores input[type=image]
2018-03-03 18:17:49 +05:30
NewUserHa
acd2b8d43b
[MRG+1] Fix part of issue #3128 - None should not be a valid type for 'url' in Response.follow ( #3131 )
...
* fix one issue of issue#3128
because @kmike posted: 'If url is '', Scrapy should follow the same page, this is an intended behavior.'
* fix one issue of issue#3128
because @kmike posted: 'If url is '', Scrapy should follow the same page, this is an intended behavior.'
2018-02-22 03:37:26 +05:00
Daniel Graña
c56f7b3c8d
Merge pull request #3113 from WenbinZhang/master
...
[MRG+1] Update robotstxt.py
2018-02-08 18:22:56 -03:00
Daniel Graña
68e45d32e0
Merge pull request #3115 from scrapy/telnet-log-level
...
[MRG+1] use INFO log level to show telnet host/port
2018-02-08 18:22:36 -03:00
Konstantin Lopuhin
936dbc7bf6
Merge branch 'master' into master
2018-02-08 23:43:29 +03:00
Eugenio Lacuesta
6edd4114c4
Clarify comment about Pyhton versions
2018-02-08 15:47:20 -03:00
Eugenio Lacuesta
a56540877c
Do not serialize unpickable objects (py3)
2018-02-08 15:03:57 -03:00
Konstantin Lopuhin
eff469245a
Merge pull request #3100 from scrapy/robots-stats
...
[MRG+1] more stats for RobotsTxtMiddleware
2018-02-08 12:01:36 +03:00
Mikhail Korobov
0c374c00fb
use INFO log level to show telnet host/port
2018-02-08 05:09:02 +05:00
Wenbin Zhang
4d5e5378bd
Update robotstxt.py
...
Add message to IgnoreRequest exception so that it can be detectedin the errbak method of a spider
2018-02-07 10:59:32 -05:00
Mikhail Korobov
6f264ab190
more stats for RobotsTxtMiddleware
2018-01-30 05:47:28 +05:00
Jesús Losada
c1916626c1
Fix OS signal names
2018-01-27 21:24:15 +00:00
Yash Sharma
1d1581266c
Changed some documentations ( #3089 )
...
DOC typo fix in defer_fail docstring
2018-01-26 01:12:17 +05:00
Mikhail Korobov
7c9e32213d
Merge pull request #3059 from jesuslosada/fix-typo
...
Fix typo in comment
2018-01-11 01:34:11 +05:00
Jesús Losada
61c0b14782
Fix typo in comment
2018-01-01 16:03:55 +00:00
Daniel Graña
2dee191374
Merge branch 'master' into extending-s3-files-store
2017-12-31 16:44:38 -03:00
Mikhail Korobov
aa83e159c9
Bump version: 1.4.0 → 1.5.0
2017-12-30 02:09:52 +05:00
Daniel Graña
57d04aa960
Merge pull request #2767 from redapple/http-proxy-endpoint-key
...
[MRG+1] Use HTTP pool and proper endpoint key for ProxyAgent
2017-12-26 14:12:47 -03:00
Konstantin Lopuhin
a21b800419
Merge pull request #3011 from Jane222/master
...
[MRG+1] Issues a warning when user puts a URL into allowed_domains (#2250 )
2017-12-08 15:45:16 +03:00
Jana Cavojska
22c68baf99
url_pattern is now being compiled before entering the loop
2017-12-07 18:38:29 +01:00
Daniel Graña
3cf0332ec3
Merge pull request #2957 from ScrapingLab/add_meta_json_to_parse_command
...
[MRG+1] Scrapy Command: add --meta/-m to the "parse" command to pass additional meta data into the request
2017-11-29 16:26:48 -03:00
Jana Cavojska
454d5e5733
checking for subclass of URLWarning instead of checking error message text when URL in allowed_domains
2017-11-26 20:07:04 +01:00
Jana Cavojska
8ec3b476b0
triggering a warning when user puts URL in allowed_domains now covered by test
2017-11-26 16:36:15 +01:00
Jana Cavojska
91ff194d1e
looping over allowed_domains directly instead of via index
2017-11-20 21:23:31 +01:00
KosayJabre
5441cc18e4
Separated import statements
...
Just separated the import statements. Tiny change - testing GitHub!
2017-11-19 18:09:38 -04:00
Jana Cavojska
62a6261028
Issues a warning when user puts a URL into allowed_domains ( #2250 )
2017-11-18 20:03:59 +01:00