From 82d239f3b148d9ce69f67bd7a2cb00de7e934aa6 Mon Sep 17 00:00:00 2001 From: Anubhav Patel Date: Wed, 6 Mar 2019 12:08:09 +0530 Subject: [PATCH 001/387] docs for scrapy.logformatter --- docs/topics/logging.rst | 39 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 39 insertions(+) diff --git a/docs/topics/logging.rst b/docs/topics/logging.rst index 0986929ad..a5fecebba 100644 --- a/docs/topics/logging.rst +++ b/docs/topics/logging.rst @@ -193,6 +193,45 @@ to override some of the Scrapy settings regarding logging. Module `logging.handlers `_ Further documentation on available handlers +Custom Log Formats +------------------- + +Custom log format can be set for different actions by extending ``scrapy.logformatter.LogFormatter`` class. + +Each method of ``scrapy.logformatter.LogFormatter`` represents an action. All methods inherited from +``scrapy.logformatter.LogFormatter`` in your custom log formatting class must return a dictionary listing +the parameters ``level``, ``msg`` and ``args`` which are going to be used for constructing the log message. +Listed below is details of what each key represents : + +* ``level`` is the log level for that action, you can use those from the python logging library: + :setting:`logging.DEBUG`, :setting:`logging.INFO`, :setting:`logging.WARNING`, :setting:`logging.ERROR` + and :setting:`logging.CRITICAL`. + +* ``msg`` should be a string that can contain different formatting placeholders. This string, formatted + with the provided ``args``, is going to be the long message for that action. + +* ``args`` should be a tuple or dict with the formatting placeholders for `msg`. The final log message is + computed as ``msg % args``. + +.. note:: To use custom log formatting class, you must mention it in ``settings.py``, by adding a line + ``LOG_FORMATTER = '’`` + +.. class:: scrapy.logformatter.LogFormatter + + The default log formatting class in Scrapy. + + .. method:: crawled (request, response, spider) + + ``crawled`` is called to log message when the crawler finds a webpage. + + .. method:: scraped(item, response, spider) + + ``scraped`` is called to log message when an item scraped by a spider. + + .. method:: dropped(item, exception, response, spider) + + ``dropped`` is called to log message when an item is dropped while it is passing through the item pipeline. + Advanced customization ---------------------- From 924b67437b92f14601816d02c5d153e7281da6d4 Mon Sep 17 00:00:00 2001 From: Anubhav Patel Date: Thu, 7 Mar 2019 16:40:59 +0530 Subject: [PATCH 002/387] move api docs to source code --- docs/topics/logging.rst | 38 ++++---------------------------------- docs/topics/settings.rst | 9 +++++++++ scrapy/logformatter.py | 36 +++++++++++++++++++++++------------- 3 files changed, 36 insertions(+), 47 deletions(-) diff --git a/docs/topics/logging.rst b/docs/topics/logging.rst index a5fecebba..72f24bae6 100644 --- a/docs/topics/logging.rst +++ b/docs/topics/logging.rst @@ -196,41 +196,11 @@ to override some of the Scrapy settings regarding logging. Custom Log Formats ------------------- -Custom log format can be set for different actions by extending ``scrapy.logformatter.LogFormatter`` class. - -Each method of ``scrapy.logformatter.LogFormatter`` represents an action. All methods inherited from -``scrapy.logformatter.LogFormatter`` in your custom log formatting class must return a dictionary listing -the parameters ``level``, ``msg`` and ``args`` which are going to be used for constructing the log message. -Listed below is details of what each key represents : - -* ``level`` is the log level for that action, you can use those from the python logging library: - :setting:`logging.DEBUG`, :setting:`logging.INFO`, :setting:`logging.WARNING`, :setting:`logging.ERROR` - and :setting:`logging.CRITICAL`. - -* ``msg`` should be a string that can contain different formatting placeholders. This string, formatted - with the provided ``args``, is going to be the long message for that action. - -* ``args`` should be a tuple or dict with the formatting placeholders for `msg`. The final log message is - computed as ``msg % args``. - -.. note:: To use custom log formatting class, you must mention it in ``settings.py``, by adding a line - ``LOG_FORMATTER = '’`` +Custom log format can be set for different actions by extending :class:`~scrapy.logformatter.LogFormatter` class +and making :setting:`LOG_FORMATTER` inside ``settings.py`` point to your new class. -.. class:: scrapy.logformatter.LogFormatter - - The default log formatting class in Scrapy. - - .. method:: crawled (request, response, spider) - - ``crawled`` is called to log message when the crawler finds a webpage. - - .. method:: scraped(item, response, spider) - - ``scraped`` is called to log message when an item scraped by a spider. - - .. method:: dropped(item, exception, response, spider) - - ``dropped`` is called to log message when an item is dropped while it is passing through the item pipeline. +.. autoclass:: scrapy.logformatter.LogFormatter + :members: Advanced customization ---------------------- diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 0ac26a9bd..1dfb5b8aa 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -866,6 +866,15 @@ directives. .. _Python datetime documentation: https://docs.python.org/2/library/datetime.html#strftime-and-strptime-behavior +.. setting:: LOG_FORMATTER + +LOG_FORMATTER +------------- + +Default: ``scrapy.logformatter.LogFormatter`` + +The class to use for formatting log messages for different actions. + .. setting:: LOG_LEVEL LOG_LEVEL diff --git a/scrapy/logformatter.py b/scrapy/logformatter.py index 075a6d862..0bb8aee58 100644 --- a/scrapy/logformatter.py +++ b/scrapy/logformatter.py @@ -13,25 +13,29 @@ CRAWLEDMSG = u"Crawled (%(status)s) %(request)s%(request_flags)s (referer: %(ref class LogFormatter(object): """Class for generating log messages for different actions. - All methods must return a dictionary listing the parameters `level`, `msg` - and `args` which are going to be used for constructing the log message when - calling logging.log. + All methods must return a dictionary listing the parameters ``level``, ``msg`` + and ``args`` which are going to be used for constructing the log message when + calling ``logging.log``. Dictionary keys for the method outputs: - * `level` should be the log level for that action, you can use those - from the python logging library: logging.DEBUG, logging.INFO, - logging.WARNING, logging.ERROR and logging.CRITICAL. - * `msg` should be a string that can contain different formatting - placeholders. This string, formatted with the provided `args`, is going - to be the log message for that action. + * ``level`` is the log level for that action, you can use those from the + `python logging library `_ : + ``logging.DEBUG``, ``logging.INFO``, ``logging.WARNING``, ``logging.ERROR`` + and ``logging.CRITICAL``. + + * ``msg`` should be a string that can contain different formatting placeholders. This string, formatted + with the provided ``args``, is going to be the long message for that action. + + * ``args`` should be a tuple or dict with the formatting placeholders for ``msg``. The final log message is + computed as ``msg % args``. - * `args` should be a tuple or dict with the formatting placeholders for - `msg`. The final log message is computed as output['msg'] % - output['args']. """ def crawled(self, request, response, spider): + """ + ``crawled`` is called to log message when the crawler finds a webpage. + """ request_flags = ' %s' % str(request.flags) if request.flags else '' response_flags = ' %s' % str(response.flags) if response.flags else '' return { @@ -40,7 +44,7 @@ class LogFormatter(object): 'args': { 'status': response.status, 'request': request, - 'request_flags' : request_flags, + 'request_flags': request_flags, 'referer': referer_str(request), 'response_flags': response_flags, # backward compatibility with Scrapy logformatter below 1.4 version @@ -49,6 +53,9 @@ class LogFormatter(object): } def scraped(self, item, response, spider): + """ + ``scraped`` is called to log message when an item is scraped by a spider. + """ if isinstance(response, Failure): src = response.getErrorMessage() else: @@ -63,6 +70,9 @@ class LogFormatter(object): } def dropped(self, item, exception, response, spider): + """ + ``dropped`` is called to log message when an item is dropped while it is passing through the item pipeline. + """ return { 'level': logging.WARNING, 'msg': DROPPEDMSG, From 82049e9c41f878d84f0fe10f827c6fe2a33f7ba6 Mon Sep 17 00:00:00 2001 From: Anubhav Patel Date: Sun, 10 Mar 2019 20:14:55 +0530 Subject: [PATCH 003/387] make suggested changes. --- docs/topics/logging.rst | 10 ++++++---- docs/topics/settings.rst | 4 ++-- scrapy/logformatter.py | 26 +++++++++++++++++--------- 3 files changed, 25 insertions(+), 15 deletions(-) diff --git a/docs/topics/logging.rst b/docs/topics/logging.rst index 72f24bae6..006530a8c 100644 --- a/docs/topics/logging.rst +++ b/docs/topics/logging.rst @@ -193,11 +193,13 @@ to override some of the Scrapy settings regarding logging. Module `logging.handlers `_ Further documentation on available handlers -Custom Log Formats -------------------- +.. _custom-log-formats: -Custom log format can be set for different actions by extending :class:`~scrapy.logformatter.LogFormatter` class -and making :setting:`LOG_FORMATTER` inside ``settings.py`` point to your new class. +Custom Log Formats +------------------ + +A custom log format can be set for different actions by extending :class:`~scrapy.logformatter.LogFormatter` class +and making :setting:`LOG_FORMATTER` point to your new class. .. autoclass:: scrapy.logformatter.LogFormatter :members: diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 1dfb5b8aa..a36c0b34c 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -871,9 +871,9 @@ directives. LOG_FORMATTER ------------- -Default: ``scrapy.logformatter.LogFormatter`` +Default: :class:`scrapy.logformatter.LogFormatter` -The class to use for formatting log messages for different actions. +The class to use for :ref:`formatting log messages ` for different actions. .. setting:: LOG_LEVEL diff --git a/scrapy/logformatter.py b/scrapy/logformatter.py index 0bb8aee58..17c69cba8 100644 --- a/scrapy/logformatter.py +++ b/scrapy/logformatter.py @@ -30,12 +30,24 @@ class LogFormatter(object): * ``args`` should be a tuple or dict with the formatting placeholders for ``msg``. The final log message is computed as ``msg % args``. + Here is an example on how to create a custom log formatter to lower the severity level of the log message + when an item is dropped from the pipeline:: + + class PoliteLogFormatter(logformatter.LogFormatter): + def dropped(self, item, exception, response, spider): + return { + 'level': logging.INFO, # lowering the level from logging.WARNING + 'msg': u"Dropped: %(exception)s" + os.linesep + "%(item)s", + 'args': { + 'exception': exception, + 'item': item, + } + } + """ def crawled(self, request, response, spider): - """ - ``crawled`` is called to log message when the crawler finds a webpage. - """ + """Logs a message when the crawler finds a webpage.""" request_flags = ' %s' % str(request.flags) if request.flags else '' response_flags = ' %s' % str(response.flags) if response.flags else '' return { @@ -53,9 +65,7 @@ class LogFormatter(object): } def scraped(self, item, response, spider): - """ - ``scraped`` is called to log message when an item is scraped by a spider. - """ + """Logs a message when an item is scraped by a spider.""" if isinstance(response, Failure): src = response.getErrorMessage() else: @@ -70,9 +80,7 @@ class LogFormatter(object): } def dropped(self, item, exception, response, spider): - """ - ``dropped`` is called to log message when an item is dropped while it is passing through the item pipeline. - """ + """Logs a message when an item is dropped while it is passing through the item pipeline.""" return { 'level': logging.WARNING, 'msg': DROPPEDMSG, From e9cd4ee03aa41e27bea0408b10970ec5bedf35d3 Mon Sep 17 00:00:00 2001 From: Anubhav Patel Date: Sun, 10 Mar 2019 20:37:56 +0530 Subject: [PATCH 004/387] fix list alignment and line width --- scrapy/logformatter.py | 23 +++++++++++------------ 1 file changed, 11 insertions(+), 12 deletions(-) diff --git a/scrapy/logformatter.py b/scrapy/logformatter.py index 17c69cba8..717120242 100644 --- a/scrapy/logformatter.py +++ b/scrapy/logformatter.py @@ -19,19 +19,18 @@ class LogFormatter(object): Dictionary keys for the method outputs: - * ``level`` is the log level for that action, you can use those from the - `python logging library `_ : - ``logging.DEBUG``, ``logging.INFO``, ``logging.WARNING``, ``logging.ERROR`` - and ``logging.CRITICAL``. + * ``level`` is the log level for that action, you can use those from the + `python logging library `_ : + ``logging.DEBUG``, ``logging.INFO``, ``logging.WARNING``, ``logging.ERROR`` + and ``logging.CRITICAL``. + * ``msg`` should be a string that can contain different formatting placeholders. + This string, formatted with the provided ``args``, is going to be the long message + for that action. + * ``args`` should be a tuple or dict with the formatting placeholders for ``msg``. + The final log message is computed as ``msg % args``. - * ``msg`` should be a string that can contain different formatting placeholders. This string, formatted - with the provided ``args``, is going to be the long message for that action. - - * ``args`` should be a tuple or dict with the formatting placeholders for ``msg``. The final log message is - computed as ``msg % args``. - - Here is an example on how to create a custom log formatter to lower the severity level of the log message - when an item is dropped from the pipeline:: + Here is an example on how to create a custom log formatter to lower the severity level of + the log message when an item is dropped from the pipeline:: class PoliteLogFormatter(logformatter.LogFormatter): def dropped(self, item, exception, response, spider): From 69b1d5d3d7050b5c60a95cc547febaded1b3685f Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Fri, 5 Oct 2018 18:21:26 +0500 Subject: [PATCH 005/387] Log cipher, certificate and temp key info on establishing an SSL connection. --- scrapy/core/downloader/tls.py | 18 ++++++++++++- scrapy/utils/ssl.py | 50 +++++++++++++++++++++++++++++++++++ 2 files changed, 67 insertions(+), 1 deletion(-) create mode 100644 scrapy/utils/ssl.py diff --git a/scrapy/core/downloader/tls.py b/scrapy/core/downloader/tls.py index df8051182..2ba72593f 100644 --- a/scrapy/core/downloader/tls.py +++ b/scrapy/core/downloader/tls.py @@ -2,6 +2,7 @@ import logging from OpenSSL import SSL from scrapy import twisted_version +from scrapy.utils.ssl import x509name_to_string, get_temp_key_info logger = logging.getLogger(__name__) @@ -20,6 +21,7 @@ openssl_methods = { METHOD_TLSv12: getattr(SSL, 'TLSv1_2_METHOD', 6), # TLS 1.2 only } + if twisted_version >= (14, 0, 0): # ClientTLSOptions requires a recent-enough version of Twisted. # Not having ScrapyClientTLSOptions should not matter for older @@ -65,13 +67,27 @@ if twisted_version >= (14, 0, 0): Same as Twisted's private _sslverify.ClientTLSOptions, except that VerificationError, CertificateError and ValueError exceptions are caught, so that the connection is not closed, only - logging warnings. + logging warnings. Also, HTTPS connection parameters logging is added. """ def _identityVerifyingInfoCallback(self, connection, where, ret): if where & SSL_CB_HANDSHAKE_START: set_tlsext_host_name(connection, self._hostnameBytes) elif where & SSL_CB_HANDSHAKE_DONE: + logger.debug('SSL connection to %s using protocol %s, cipher %s', + self._hostnameASCII, + connection.get_protocol_version_name(), + connection.get_cipher_name(), + ) + server_cert = connection.get_peer_certificate() + logger.debug('SSL connection certificate: issuer "%s", subject "%s"', + x509name_to_string(server_cert.get_issuer()), + x509name_to_string(server_cert.get_subject()), + ) + key_info = get_temp_key_info(connection._ssl) + if key_info: + logger.debug('SSL temp key: %s', key_info) + try: verifyHostname(connection, self._hostnameASCII) except verification_errors as e: diff --git a/scrapy/utils/ssl.py b/scrapy/utils/ssl.py new file mode 100644 index 000000000..5db1608bf --- /dev/null +++ b/scrapy/utils/ssl.py @@ -0,0 +1,50 @@ +# -*- coding: utf-8 -*- + +import OpenSSL._util as pyOpenSSLutil + +from scrapy.utils.python import to_native_str + + +def ffi_buf_to_string(buf): + return to_native_str(pyOpenSSLutil.ffi.string(buf)) + + +def x509name_to_string(x509name): + # from OpenSSL.crypto.X509Name.__repr__ + result_buffer = pyOpenSSLutil.ffi.new("char[]", 512) + pyOpenSSLutil.lib.X509_NAME_oneline(x509name._name, result_buffer, len(result_buffer)) + + return ffi_buf_to_string(result_buffer) + + +def get_temp_key_info(ssl_object): + if not hasattr(pyOpenSSLutil.lib, 'SSL_get_server_tmp_key'): # requires OpenSSL 1.0.2 + return None + + # adapted from OpenSSL apps/s_cb.c::ssl_print_tmp_key() + temp_key_p = pyOpenSSLutil.ffi.new("EVP_PKEY **") + pyOpenSSLutil.lib.SSL_get_server_tmp_key(ssl_object, temp_key_p) + if temp_key_p == pyOpenSSLutil.ffi.NULL: + return None + + temp_key = temp_key_p[0] + pyOpenSSLutil.ffi.gc(temp_key, pyOpenSSLutil.lib.EVP_PKEY_free) + key_info = [] + key_type = pyOpenSSLutil.lib.EVP_PKEY_id(temp_key) + if key_type == pyOpenSSLutil.lib.EVP_PKEY_RSA: + key_info.append('RSA') + elif key_type == pyOpenSSLutil.lib.EVP_PKEY_DH: + key_info.append('DH') + elif key_type == pyOpenSSLutil.lib.EVP_PKEY_EC: + key_info.append('ECDH') + ec_key = pyOpenSSLutil.lib.EVP_PKEY_get1_EC_KEY(temp_key) + pyOpenSSLutil.ffi.gc(ec_key, pyOpenSSLutil.lib.EC_KEY_free) + nid = pyOpenSSLutil.lib.EC_GROUP_get_curve_name(pyOpenSSLutil.lib.EC_KEY_get0_group(ec_key)) + cname = pyOpenSSLutil.lib.EC_curve_nid2nist(nid) + if cname == pyOpenSSLutil.ffi.NULL: + cname = pyOpenSSLutil.lib.OBJ_nid2sn(nid) + key_info.append(ffi_buf_to_string(cname)) + else: + key_info.append(ffi_buf_to_string(pyOpenSSLutil.lib.OBJ_nid2sn(key_type))) + key_info.append('%s bits' % pyOpenSSLutil.lib.EVP_PKEY_bits(temp_key)) + return ', '.join(key_info) From 67a400092805ec0d643bd7de0481cc45d5ce8471 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Mon, 8 Jul 2019 10:31:52 +0500 Subject: [PATCH 006/387] Work around older pyOpenSSL not having get_cipher_name or get_protocol_version_name. --- scrapy/core/downloader/tls.py | 17 ++++++++++++----- 1 file changed, 12 insertions(+), 5 deletions(-) diff --git a/scrapy/core/downloader/tls.py b/scrapy/core/downloader/tls.py index 2ba72593f..7e5882663 100644 --- a/scrapy/core/downloader/tls.py +++ b/scrapy/core/downloader/tls.py @@ -74,11 +74,18 @@ if twisted_version >= (14, 0, 0): if where & SSL_CB_HANDSHAKE_START: set_tlsext_host_name(connection, self._hostnameBytes) elif where & SSL_CB_HANDSHAKE_DONE: - logger.debug('SSL connection to %s using protocol %s, cipher %s', - self._hostnameASCII, - connection.get_protocol_version_name(), - connection.get_cipher_name(), - ) + if hasattr(connection, 'get_cipher_name'): # requires pyOPenSSL 0.15 + if hasattr(connection, 'get_protocol_version_name'): # requires pyOPenSSL 16.0.0 + logger.debug('SSL connection to %s using protocol %s, cipher %s', + self._hostnameASCII, + connection.get_protocol_version_name(), + connection.get_cipher_name(), + ) + else: + logger.debug('SSL connection to %s using cipher %s', + self._hostnameASCII, + connection.get_cipher_name(), + ) server_cert = connection.get_peer_certificate() logger.debug('SSL connection certificate: issuer "%s", subject "%s"', x509name_to_string(server_cert.get_issuer()), From 0b9dce3a6c17d8dc827df57367383e0b82fa8b07 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Mon, 8 Jul 2019 17:40:56 +0500 Subject: [PATCH 007/387] Add DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING setting. --- docs/topics/settings.rst | 21 +++++++++-- scrapy/core/downloader/contextfactory.py | 14 ++++++- scrapy/core/downloader/handlers/http10.py | 7 ++-- scrapy/core/downloader/handlers/http11.py | 11 +++--- scrapy/core/downloader/tls.py | 45 +++++++++++++---------- scrapy/settings/default_settings.py | 1 + 6 files changed, 66 insertions(+), 33 deletions(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 371f21c72..5cc87bb64 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -438,9 +438,10 @@ or even enable client-side authentication (and various other things). which uses the platform's certificates to validate remote endpoints. **This is only available if you use Twisted>=14.0.** -If you do use a custom ContextFactory, make sure it accepts a ``method`` -parameter at init (this is the ``OpenSSL.SSL`` method mapping -:setting:`DOWNLOADER_CLIENT_TLS_METHOD`). +If you do use a custom ContextFactory, make sure its ``__init__` method accepts +a ``method`` parameter (this is the ``OpenSSL.SSL`` method mapping +:setting:`DOWNLOADER_CLIENT_TLS_METHOD`) and a ``settings`` parameter (this is +the Scrapy settings object). .. setting:: DOWNLOADER_CLIENT_TLS_METHOD @@ -468,6 +469,20 @@ This setting must be one of these string values: We recommend that you use PyOpenSSL>=0.13 and Twisted>=0.13 or above (Twisted>=14.0 if you can). +.. setting:: DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING + +DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING +------------------------------------- + +Default: ``False`` + +Setting this to ``True`` will enable DEBUG level messages about TLS connection +parameters after establishing HTTPS connections. The kind of information logged +depends on the versions of OpenSSL and pyOpenSSL. + +This setting is only used for the default +:setting:`DOWNLOADER_CLIENTCONTEXTFACTORY`. + .. setting:: DOWNLOADER_MIDDLEWARES DOWNLOADER_MIDDLEWARES diff --git a/scrapy/core/downloader/contextfactory.py b/scrapy/core/downloader/contextfactory.py index 783d4c383..80c784f5a 100644 --- a/scrapy/core/downloader/contextfactory.py +++ b/scrapy/core/downloader/contextfactory.py @@ -2,6 +2,7 @@ from OpenSSL import SSL from twisted.internet.ssl import ClientContextFactory from scrapy import twisted_version +from scrapy.utils.misc import create_instance if twisted_version >= (14, 0, 0): @@ -28,9 +29,17 @@ if twisted_version >= (14, 0, 0): understand the SSLv3, TLSv1, TLSv1.1 and TLSv1.2 protocols.' """ - def __init__(self, method=SSL.SSLv23_METHOD, *args, **kwargs): + def __init__(self, method=SSL.SSLv23_METHOD, settings=None, *args, **kwargs): super(ScrapyClientContextFactory, self).__init__(*args, **kwargs) self._ssl_method = method + if settings: + self.tls_verbose_logging = settings['DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING'] + else: + self.tls_verbose_logging = False + + @classmethod + def from_settings(cls, settings, method=SSL.SSLv23_METHOD, *args, **kwargs): + return cls(method=method, settings=settings, *args, **kwargs) def getCertificateOptions(self): # setting verify=True will require you to provide CAs @@ -56,7 +65,8 @@ if twisted_version >= (14, 0, 0): return self.getCertificateOptions().getContext() def creatorForNetloc(self, hostname, port): - return ScrapyClientTLSOptions(hostname.decode("ascii"), self.getContext()) + return ScrapyClientTLSOptions(hostname.decode("ascii"), self.getContext(), + verbose_logging=self.tls_verbose_logging) @implementer(IPolicyForHTTPS) diff --git a/scrapy/core/downloader/handlers/http10.py b/scrapy/core/downloader/handlers/http10.py index d875fb1e4..be7298531 100644 --- a/scrapy/core/downloader/handlers/http10.py +++ b/scrapy/core/downloader/handlers/http10.py @@ -1,7 +1,7 @@ """Download handlers for http and https schemes """ from twisted.internet import reactor -from scrapy.utils.misc import load_object +from scrapy.utils.misc import load_object, create_instance from scrapy.utils.python import to_unicode @@ -11,6 +11,7 @@ class HTTP10DownloadHandler(object): def __init__(self, settings): self.HTTPClientFactory = load_object(settings['DOWNLOADER_HTTPCLIENTFACTORY']) self.ClientContextFactory = load_object(settings['DOWNLOADER_CLIENTCONTEXTFACTORY']) + self._settings = settings def download_request(self, request, spider): """Return a deferred for the HTTP download""" @@ -21,7 +22,7 @@ class HTTP10DownloadHandler(object): def _connect(self, factory): host, port = to_unicode(factory.host), factory.port if factory.scheme == b'https': - return reactor.connectSSL(host, port, factory, - self.ClientContextFactory()) + client_context_factory = create_instance(self.ClientContextFactory, settings=self._settings, crawler=None) + return reactor.connectSSL(host, port, factory, client_context_factory) else: return reactor.connectTCP(host, port, factory) diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index 0673188a1..9b0c7977d 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -25,7 +25,7 @@ from scrapy.http import Headers from scrapy.responsetypes import responsetypes from scrapy.core.downloader.webclient import _parse from scrapy.core.downloader.tls import openssl_methods -from scrapy.utils.misc import load_object +from scrapy.utils.misc import load_object, create_instance from scrapy.utils.python import to_bytes, to_unicode from scrapy import twisted_version @@ -44,14 +44,15 @@ class HTTP11DownloadHandler(object): self._contextFactoryClass = load_object(settings['DOWNLOADER_CLIENTCONTEXTFACTORY']) # try method-aware context factory try: - self._contextFactory = self._contextFactoryClass(method=self._sslMethod) + self._contextFactory = create_instance(self._contextFactoryClass, settings=settings, crawler=None, + method=self._sslMethod) except TypeError: # use context factory defaults - self._contextFactory = self._contextFactoryClass() + self._contextFactory = create_instance(self._contextFactoryClass, settings=settings, crawler=None) msg = """ '%s' does not accept `method` argument (type OpenSSL.SSL method,\ - e.g. OpenSSL.SSL.SSLv23_METHOD).\ - Please upgrade your context factory class to handle it or ignore it.""" % ( + e.g. OpenSSL.SSL.SSLv23_METHOD) and/or `settings` argument.\ + Please upgrade your context factory class to handle them or ignore them.""" % ( settings['DOWNLOADER_CLIENTCONTEXTFACTORY'],) warnings.warn(msg) self._default_maxsize = settings.getint('DOWNLOAD_MAXSIZE') diff --git a/scrapy/core/downloader/tls.py b/scrapy/core/downloader/tls.py index 7e5882663..74be85d52 100644 --- a/scrapy/core/downloader/tls.py +++ b/scrapy/core/downloader/tls.py @@ -70,30 +70,35 @@ if twisted_version >= (14, 0, 0): logging warnings. Also, HTTPS connection parameters logging is added. """ + def __init__(self, hostname, ctx, verbose_logging=False): + super().__init__(hostname, ctx) + self.verbose_logging = verbose_logging + def _identityVerifyingInfoCallback(self, connection, where, ret): if where & SSL_CB_HANDSHAKE_START: set_tlsext_host_name(connection, self._hostnameBytes) elif where & SSL_CB_HANDSHAKE_DONE: - if hasattr(connection, 'get_cipher_name'): # requires pyOPenSSL 0.15 - if hasattr(connection, 'get_protocol_version_name'): # requires pyOPenSSL 16.0.0 - logger.debug('SSL connection to %s using protocol %s, cipher %s', - self._hostnameASCII, - connection.get_protocol_version_name(), - connection.get_cipher_name(), - ) - else: - logger.debug('SSL connection to %s using cipher %s', - self._hostnameASCII, - connection.get_cipher_name(), - ) - server_cert = connection.get_peer_certificate() - logger.debug('SSL connection certificate: issuer "%s", subject "%s"', - x509name_to_string(server_cert.get_issuer()), - x509name_to_string(server_cert.get_subject()), - ) - key_info = get_temp_key_info(connection._ssl) - if key_info: - logger.debug('SSL temp key: %s', key_info) + if self.verbose_logging: + if hasattr(connection, 'get_cipher_name'): # requires pyOPenSSL 0.15 + if hasattr(connection, 'get_protocol_version_name'): # requires pyOPenSSL 16.0.0 + logger.debug('SSL connection to %s using protocol %s, cipher %s', + self._hostnameASCII, + connection.get_protocol_version_name(), + connection.get_cipher_name(), + ) + else: + logger.debug('SSL connection to %s using cipher %s', + self._hostnameASCII, + connection.get_cipher_name(), + ) + server_cert = connection.get_peer_certificate() + logger.debug('SSL connection certificate: issuer "%s", subject "%s"', + x509name_to_string(server_cert.get_issuer()), + x509name_to_string(server_cert.get_subject()), + ) + key_info = get_temp_key_info(connection._ssl) + if key_info: + logger.debug('SSL temp key: %s', key_info) try: verifyHostname(connection, self._hostnameASCII) diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 10b6cf9bc..af8305b25 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -87,6 +87,7 @@ DOWNLOADER_HTTPCLIENTFACTORY = 'scrapy.core.downloader.webclient.ScrapyHTTPClien DOWNLOADER_CLIENTCONTEXTFACTORY = 'scrapy.core.downloader.contextfactory.ScrapyClientContextFactory' DOWNLOADER_CLIENT_TLS_METHOD = 'TLS' # Use highest TLS/SSL protocol version supported by the platform, # also allowing negotiation +DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING = False DOWNLOADER_MIDDLEWARES = {} From 0de6ffc8e1354ac7ffcdf7a75e3ec8d26ed23a01 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 11 Jul 2019 13:12:56 +0500 Subject: [PATCH 008/387] Fix super() call. --- scrapy/core/downloader/tls.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/core/downloader/tls.py b/scrapy/core/downloader/tls.py index 74be85d52..74afb3f10 100644 --- a/scrapy/core/downloader/tls.py +++ b/scrapy/core/downloader/tls.py @@ -71,7 +71,7 @@ if twisted_version >= (14, 0, 0): """ def __init__(self, hostname, ctx, verbose_logging=False): - super().__init__(hostname, ctx) + super(ScrapyClientTLSOptions, self).__init__(hostname, ctx) self.verbose_logging = verbose_logging def _identityVerifyingInfoCallback(self, connection, where, ret): From 98689b27a8839a21aed52ada44143456ead84e24 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 11 Jul 2019 14:02:35 +0500 Subject: [PATCH 009/387] Improve the DOWNLOADER_CLIENTCONTEXTFACTORY doc. --- docs/topics/settings.rst | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 5cc87bb64..53c624679 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -438,10 +438,10 @@ or even enable client-side authentication (and various other things). which uses the platform's certificates to validate remote endpoints. **This is only available if you use Twisted>=14.0.** -If you do use a custom ContextFactory, make sure its ``__init__` method accepts -a ``method`` parameter (this is the ``OpenSSL.SSL`` method mapping +If you do use a custom ContextFactory, make sure its ``__init__`` method +accepts a ``method`` parameter (this is the ``OpenSSL.SSL`` method mapping :setting:`DOWNLOADER_CLIENT_TLS_METHOD`) and a ``settings`` parameter (this is -the Scrapy settings object). +the Scrapy :class:`~scrapy.settings.Settings` object). .. setting:: DOWNLOADER_CLIENT_TLS_METHOD From a96a07bc762287a1f2056d1e142aff9f33e206fe Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Fri, 12 Jul 2019 18:44:45 +0500 Subject: [PATCH 010/387] Add a test for DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING. --- tests/test_downloader_handlers.py | 21 +++++++++++++++++++++ 1 file changed, 21 insertions(+) diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 81235a16f..8d0df6b5b 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -8,6 +8,7 @@ try: except ImportError: import mock +from testfixtures import LogCapture from twisted.trial import unittest from twisted.protocols.policies import WrappingFactory from twisted.python.filepath import FilePath @@ -503,6 +504,24 @@ class Http11TestCase(HttpTestCase): class Https11TestCase(Http11TestCase): scheme = 'https' + tls_log_message = 'SSL connection certificate: issuer "/C=IE/O=Scrapy/CN=localhost", subject "/C=IE/O=Scrapy/CN=localhost"' + + @defer.inlineCallbacks + def test_tls_logging(self): + download_handler = self.download_handler_cls(Settings({ + 'DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING': True, + })) + try: + with LogCapture() as log_capture: + request = Request(self.getURL('file')) + d = download_handler.download_request(request, Spider('foo')) + d.addCallback(lambda r: r.body) + d.addCallback(self.assertEqual, b"0123456789") + yield d + log_capture.check_present(('scrapy.core.downloader.tls', 'DEBUG', self.tls_log_message)) + finally: + yield download_handler.close() + class Https11WrongHostnameTestCase(Http11TestCase): scheme = 'https' @@ -523,6 +542,7 @@ class Https11InvalidDNSId(Https11TestCase): super(Https11InvalidDNSId, self).setUp() self.host = '127.0.0.1' + class Https11InvalidDNSPattern(Https11TestCase): """Connect to HTTPS hosts where the certificate are issued to an ip instead of a domain.""" @@ -534,6 +554,7 @@ class Https11InvalidDNSPattern(Https11TestCase): from service_identity.exceptions import CertificateError except ImportError: raise unittest.SkipTest("cryptography lib is too old") + self.tls_log_message = 'SSL connection certificate: issuer "/C=IE/O=Scrapy/CN=127.0.0.1", subject "/C=IE/O=Scrapy/CN=127.0.0.1"' super(Https11InvalidDNSPattern, self).setUp() From 42743fd9dd9d7116848fd3ad6b657453dd0b117d Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 18 Jul 2019 20:49:25 +0500 Subject: [PATCH 011/387] Move tls_verbose_logging extraction from __init__ to from_settings. --- docs/topics/settings.rst | 4 ++-- scrapy/core/downloader/contextfactory.py | 13 +++++++------ scrapy/core/downloader/handlers/http11.py | 2 +- 3 files changed, 10 insertions(+), 9 deletions(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 53c624679..8705a5249 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -440,8 +440,8 @@ or even enable client-side authentication (and various other things). If you do use a custom ContextFactory, make sure its ``__init__`` method accepts a ``method`` parameter (this is the ``OpenSSL.SSL`` method mapping -:setting:`DOWNLOADER_CLIENT_TLS_METHOD`) and a ``settings`` parameter (this is -the Scrapy :class:`~scrapy.settings.Settings` object). +:setting:`DOWNLOADER_CLIENT_TLS_METHOD`) and a ``tls_verbose_logging`` +parameter (``bool``). .. setting:: DOWNLOADER_CLIENT_TLS_METHOD diff --git a/scrapy/core/downloader/contextfactory.py b/scrapy/core/downloader/contextfactory.py index 80c784f5a..d5d238b9c 100644 --- a/scrapy/core/downloader/contextfactory.py +++ b/scrapy/core/downloader/contextfactory.py @@ -29,17 +29,18 @@ if twisted_version >= (14, 0, 0): understand the SSLv3, TLSv1, TLSv1.1 and TLSv1.2 protocols.' """ - def __init__(self, method=SSL.SSLv23_METHOD, settings=None, *args, **kwargs): + def __init__(self, method=SSL.SSLv23_METHOD, tls_verbose_logging=False, *args, **kwargs): super(ScrapyClientContextFactory, self).__init__(*args, **kwargs) self._ssl_method = method - if settings: - self.tls_verbose_logging = settings['DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING'] - else: - self.tls_verbose_logging = False + self.tls_verbose_logging = tls_verbose_logging @classmethod def from_settings(cls, settings, method=SSL.SSLv23_METHOD, *args, **kwargs): - return cls(method=method, settings=settings, *args, **kwargs) + if settings: + tls_verbose_logging = settings.getbool('DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING') + else: + tls_verbose_logging = False + return cls(method=method, tls_verbose_logging=tls_verbose_logging, *args, **kwargs) def getCertificateOptions(self): # setting verify=True will require you to provide CAs diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index 9b0c7977d..deb0f9d21 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -51,7 +51,7 @@ class HTTP11DownloadHandler(object): self._contextFactory = create_instance(self._contextFactoryClass, settings=settings, crawler=None) msg = """ '%s' does not accept `method` argument (type OpenSSL.SSL method,\ - e.g. OpenSSL.SSL.SSLv23_METHOD) and/or `settings` argument.\ + e.g. OpenSSL.SSL.SSLv23_METHOD) and/or `tls_verbose_logging` argument.\ Please upgrade your context factory class to handle them or ignore them.""" % ( settings['DOWNLOADER_CLIENTCONTEXTFACTORY'],) warnings.warn(msg) From 95dd2df7b5dc6836a784a2b373009b17ca2eb475 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 18 Jul 2019 20:51:26 +0500 Subject: [PATCH 012/387] Drop an unused import. --- scrapy/core/downloader/contextfactory.py | 1 - 1 file changed, 1 deletion(-) diff --git a/scrapy/core/downloader/contextfactory.py b/scrapy/core/downloader/contextfactory.py index d5d238b9c..188d9f917 100644 --- a/scrapy/core/downloader/contextfactory.py +++ b/scrapy/core/downloader/contextfactory.py @@ -2,7 +2,6 @@ from OpenSSL import SSL from twisted.internet.ssl import ClientContextFactory from scrapy import twisted_version -from scrapy.utils.misc import create_instance if twisted_version >= (14, 0, 0): From c6453800cd612297aa635636477c68a109a4a542 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 18 Jul 2019 22:17:39 +0500 Subject: [PATCH 013/387] Remove an unneeded if. --- scrapy/core/downloader/contextfactory.py | 5 +---- 1 file changed, 1 insertion(+), 4 deletions(-) diff --git a/scrapy/core/downloader/contextfactory.py b/scrapy/core/downloader/contextfactory.py index 188d9f917..5ac20c0bb 100644 --- a/scrapy/core/downloader/contextfactory.py +++ b/scrapy/core/downloader/contextfactory.py @@ -35,10 +35,7 @@ if twisted_version >= (14, 0, 0): @classmethod def from_settings(cls, settings, method=SSL.SSLv23_METHOD, *args, **kwargs): - if settings: - tls_verbose_logging = settings.getbool('DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING') - else: - tls_verbose_logging = False + tls_verbose_logging = settings.getbool('DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING') return cls(method=method, tls_verbose_logging=tls_verbose_logging, *args, **kwargs) def getCertificateOptions(self): From 43d5b5a524ff2cce6fd4620f8e2460489da39f42 Mon Sep 17 00:00:00 2001 From: Kristobal Junta Date: Mon, 22 Jul 2019 10:19:08 +0300 Subject: [PATCH 014/387] fix default RETRY_HTTP_CODES value in docs --- docs/topics/downloader-middleware.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 38a4fdb25..a3780a177 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -963,7 +963,7 @@ precedence over the :setting:`RETRY_TIMES` setting. RETRY_HTTP_CODES ^^^^^^^^^^^^^^^^ -Default: ``[500, 502, 503, 504, 522, 524, 408]`` +Default: ``[500, 502, 503, 504, 522, 524, 408, 429]`` Which HTTP response codes to retry. Other errors (DNS lookup issues, connections lost, etc) are always retried. From 7e622af4e5b0f49c88101ad370941b33e4833e1e Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Mon, 22 Jul 2019 14:53:17 -0300 Subject: [PATCH 015/387] Fix ConfigParser import in py2 --- scrapy/utils/conf.py | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/scrapy/utils/conf.py b/scrapy/utils/conf.py index 26d66eaf8..fb7ca3310 100644 --- a/scrapy/utils/conf.py +++ b/scrapy/utils/conf.py @@ -1,10 +1,13 @@ import os import sys import numbers -import configparser from operator import itemgetter import six +if six.PY2: + from ConfigParser import SafeConfigParser as ConfigParser +else: + from configparser import ConfigParser from scrapy.settings import BaseSettings from scrapy.utils.deprecate import update_classpath @@ -94,7 +97,7 @@ def init_env(project='default', set_syspath=True): def get_config(use_closest=True): """Get Scrapy config file as a ConfigParser""" sources = get_sources(use_closest) - cfg = configparser.ConfigParser() + cfg = ConfigParser() cfg.read(sources) return cfg From 7843101f9abad302e6c9c997f9d2a7adff98380b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Tue, 23 Jul 2019 12:04:26 +0200 Subject: [PATCH 016/387] Cover Scrapy 1.7.2 in the release notes --- docs/news.rst | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/news.rst b/docs/news.rst index a0f0c5697..d79844ed2 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -6,6 +6,12 @@ Release notes .. note:: Scrapy 1.x will be the last series supporting Python 2. Scrapy 2.0, planned for Q4 2019 or Q1 2020, will support **Python 3 only**. +Scrapy 1.7.2 (2019-07-23) +------------------------- + +Fix Python 2 support (:issue:`3889`, :issue:`3893`, :issue:`3896`). + + Scrapy 1.7.1 (2019-07-18) ------------------------- From 7551689c75a1f2b4dbed72184f1dabab2f6c3c4a Mon Sep 17 00:00:00 2001 From: Lucy Wang Date: Fri, 26 Jul 2019 09:07:29 +0800 Subject: [PATCH 017/387] s3 file store should accept all supported headers --- scrapy/pipelines/files.py | 13 +++++++++++++ 1 file changed, 13 insertions(+) diff --git a/scrapy/pipelines/files.py b/scrapy/pipelines/files.py index 2145e6d2b..ea06d2ae8 100644 --- a/scrapy/pipelines/files.py +++ b/scrapy/pipelines/files.py @@ -189,6 +189,19 @@ class S3FilesStore(object): 'X-Amz-Grant-Read': 'GrantRead', 'X-Amz-Grant-Read-ACP': 'GrantReadACP', 'X-Amz-Grant-Write-ACP': 'GrantWriteACP', + 'X-Amz-Object-Lock-Legal-Hold': 'ObjectLockLegalHoldStatus', + 'X-Amz-Object-Lock-Mode': 'ObjectLockMode', + 'X-Amz-Object-Lock-Retain-Until-Date': 'ObjectLockRetainUntilDate', + 'X-Amz-Request-Payer': 'RequestPayer', + 'X-Amz-Server-Side-Encryption': 'ServerSideEncryption', + 'X-Amz-Server-Side-Encryption-Aws-Kms-Key-Id': 'SSEKMSKeyId', + 'X-Amz-Server-Side-Encryption-Context': 'SSEKMSEncryptionContext', + 'X-Amz-Server-Side-Encryption-Customer-Algorithm': 'SSECustomerAlgorithm', + 'X-Amz-Server-Side-Encryption-Customer-Key': 'SSECustomerKey', + 'X-Amz-Server-Side-Encryption-Customer-Key-Md5': 'SSECustomerKeyMD5', + 'X-Amz-Storage-Class': 'StorageClass', + 'X-Amz-Tagging': 'Tagging', + 'X-Amz-Website-Redirect-Location': 'WebsiteRedirectLocation', }) extra = {} for key, value in six.iteritems(headers): From f21dc24a266a9e45088fc71daee97cbe9b11d4a4 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Tue, 30 Jul 2019 18:16:12 +0500 Subject: [PATCH 018/387] Fix memory handling and error handling in utils.ssl.get_temp_key_info. --- scrapy/utils/ssl.py | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/scrapy/utils/ssl.py b/scrapy/utils/ssl.py index 5db1608bf..e54232abd 100644 --- a/scrapy/utils/ssl.py +++ b/scrapy/utils/ssl.py @@ -23,12 +23,12 @@ def get_temp_key_info(ssl_object): # adapted from OpenSSL apps/s_cb.c::ssl_print_tmp_key() temp_key_p = pyOpenSSLutil.ffi.new("EVP_PKEY **") - pyOpenSSLutil.lib.SSL_get_server_tmp_key(ssl_object, temp_key_p) - if temp_key_p == pyOpenSSLutil.ffi.NULL: + if not pyOpenSSLutil.lib.SSL_get_server_tmp_key(ssl_object, temp_key_p): return None - temp_key = temp_key_p[0] - pyOpenSSLutil.ffi.gc(temp_key, pyOpenSSLutil.lib.EVP_PKEY_free) + if temp_key == pyOpenSSLutil.ffi.NULL: + return None + temp_key = pyOpenSSLutil.ffi.gc(temp_key, pyOpenSSLutil.lib.EVP_PKEY_free) key_info = [] key_type = pyOpenSSLutil.lib.EVP_PKEY_id(temp_key) if key_type == pyOpenSSLutil.lib.EVP_PKEY_RSA: @@ -38,7 +38,7 @@ def get_temp_key_info(ssl_object): elif key_type == pyOpenSSLutil.lib.EVP_PKEY_EC: key_info.append('ECDH') ec_key = pyOpenSSLutil.lib.EVP_PKEY_get1_EC_KEY(temp_key) - pyOpenSSLutil.ffi.gc(ec_key, pyOpenSSLutil.lib.EC_KEY_free) + ec_key = pyOpenSSLutil.ffi.gc(ec_key, pyOpenSSLutil.lib.EC_KEY_free) nid = pyOpenSSLutil.lib.EC_GROUP_get_curve_name(pyOpenSSLutil.lib.EC_KEY_get0_group(ec_key)) cname = pyOpenSSLutil.lib.EC_curve_nid2nist(nid) if cname == pyOpenSSLutil.ffi.NULL: From 7333fc02aa842ec4506ca00c345dcba72739f35c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Daniel=20Gra=C3=B1a?= Date: Tue, 30 Jul 2019 23:16:11 -0300 Subject: [PATCH 019/387] Pin Travis-ci build environment to previous default: Trusty Travis-ci changed the default build environment to Xenial as explained in https://blog.travis-ci.com/2019-04-15-xenial-default-build-environment This causes builds meant for Debian Jessie to break as noted by @wRAR in https://github.com/scrapy/scrapy/issues/3917#issuecomment-516426389 This change pins the environment to known working ubuntu trusty distribution prior to dropping Jessie support and upgrade to Xenial as base. Closes #1369 --- .travis.yml | 1 + 1 file changed, 1 insertion(+) diff --git a/.travis.yml b/.travis.yml index 08b0bf119..3116d9b48 100644 --- a/.travis.yml +++ b/.travis.yml @@ -1,4 +1,5 @@ language: python +dist: trusty branches: only: - master From a25e09ecdd6bb53779b713dead19637128259ee2 Mon Sep 17 00:00:00 2001 From: Renne Rocha Date: Mon, 29 Jul 2019 19:07:34 -0300 Subject: [PATCH 020/387] Added constrain on lxml version based on Python version --- requirements-py3.txt | 3 ++- setup.py | 3 ++- tests/constraints.txt | 3 +-- 3 files changed, 5 insertions(+), 4 deletions(-) diff --git a/requirements-py3.txt b/requirements-py3.txt index 5a5d4c95a..478ed0010 100644 --- a/requirements-py3.txt +++ b/requirements-py3.txt @@ -1,5 +1,6 @@ Twisted>=17.9.0 -lxml>=3.2.4 +lxml;python_version!="3.4" +lxml<=4.3.5;python_version=="3.4" pyOpenSSL>=0.13.1 cssselect>=0.9 queuelib>=1.1.1 diff --git a/setup.py b/setup.py index 4dc6d18c1..ee0aaabf0 100644 --- a/setup.py +++ b/setup.py @@ -69,7 +69,8 @@ setup( 'Twisted>=13.1.0,<=19.2.0;python_version=="3.4"', 'w3lib>=1.17.0', 'queuelib', - 'lxml', + 'lxml;python_version!="3.4"', + 'lxml<=4.3.5;python_version=="3.4"', 'pyOpenSSL', 'cssselect>=0.9', 'six>=1.5.2', diff --git a/tests/constraints.txt b/tests/constraints.txt index e59e68b3f..5655ac2d3 100644 --- a/tests/constraints.txt +++ b/tests/constraints.txt @@ -1,2 +1 @@ -Twisted!=18.4.0 -lxml!=4.2.2 \ No newline at end of file +Twisted!=18.4.0 \ No newline at end of file From 783d61d32aba53208b2b0bb9a1d82a56dbcbabd6 Mon Sep 17 00:00:00 2001 From: sbs2001 Date: Thu, 1 Aug 2019 14:11:27 +0530 Subject: [PATCH 021/387] [MRG+1] Update _monkeypatches.py (#3907) * Update _monkeypatches.py The workarounds are not required assuming the bugs regarding urlparse are absent in Python versions >2.7. We already exit the program if Python version<2.7 in the __init__.py(line 17).The monkeypatches are deployed after this check at line 27 in the __init__.py . * Update _monkeypatches.py Added the second workaround. * Update _monkeypatches.py * Update _monkeypatches.py --- scrapy/_monkeypatches.py | 7 +------ 1 file changed, 1 insertion(+), 6 deletions(-) diff --git a/scrapy/_monkeypatches.py b/scrapy/_monkeypatches.py index 935c4bfa3..b68099cad 100644 --- a/scrapy/_monkeypatches.py +++ b/scrapy/_monkeypatches.py @@ -4,12 +4,7 @@ from six.moves import copyreg if six.PY2: from urlparse import urlparse - - # workaround for https://bugs.python.org/issue7904 - Python < 2.7 - if urlparse('s3://bucket/key').netloc != 'bucket': - from urlparse import uses_netloc - uses_netloc.append('s3') - + # workaround for https://bugs.python.org/issue9374 - Python < 2.7.4 if urlparse('s3://bucket/key?key=value').query != 'key=value': from urlparse import uses_query From a12e8251e065182ab09b9b5a65a86c44a606c4a9 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 1 Aug 2019 17:06:53 +0200 Subject: [PATCH 022/387] Cover Scrapy 1.7.3 in the release notes --- docs/news.rst | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/docs/news.rst b/docs/news.rst index d79844ed2..ce5b8b406 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -6,6 +6,11 @@ Release notes .. note:: Scrapy 1.x will be the last series supporting Python 2. Scrapy 2.0, planned for Q4 2019 or Q1 2020, will support **Python 3 only**. +Scrapy 1.7.3 (2019-08-01) +------------------------- + +Enforce lxml 4.3.5 or lower for Python 3.4 (:issue:`3912`, :issue:`3918`). + Scrapy 1.7.2 (2019-07-23) ------------------------- From 8e813953bd4b25b60d30981a948e48fee43f8dad Mon Sep 17 00:00:00 2001 From: Anubhav Patel Date: Fri, 2 Aug 2019 13:13:29 +0530 Subject: [PATCH 023/387] [MRG+1] [GSoC 2019] Interface for robots.txt parsers (#3796) Make the robots.txt parser configurable through the new ROBOTSTXT_PARSER setting, support the Reppy and Robotexclusionrulesparser parsers, and allow implementing custom robots.txt parsers. --- .travis.yml | 6 + docs/topics/downloader-middleware.rst | 79 +++++++++++ docs/topics/settings.rst | 10 ++ scrapy/downloadermiddlewares/robotstxt.py | 41 ++---- scrapy/robotstxt.py | 112 +++++++++++++++ scrapy/settings/default_settings.py | 1 + tests/test_downloadermiddleware_robotstxt.py | 25 ++++ tests/test_robotstxt_interface.py | 142 +++++++++++++++++++ tox.ini | 14 ++ 9 files changed, 404 insertions(+), 26 deletions(-) create mode 100644 scrapy/robotstxt.py create mode 100644 tests/test_robotstxt_interface.py diff --git a/.travis.yml b/.travis.yml index 3116d9b48..138e81c64 100644 --- a/.travis.yml +++ b/.travis.yml @@ -27,6 +27,12 @@ matrix: sudo: true - python: 3.6 env: TOXENV=docs + - python: 3.7 + env: TOXENV=py37-extra-deps + dist: xenial + sudo: true + - python: 2.7 + env: TOXENV=py27-extra-deps install: - | if [ "$TOXENV" = "pypy" ]; then diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index a3780a177..616b56101 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -989,6 +989,17 @@ RobotsTxtMiddleware To make sure Scrapy respects robots.txt make sure the middleware is enabled and the :setting:`ROBOTSTXT_OBEY` setting is enabled. + This middleware has to be combined with a robots.txt_ parser. + + Scrapy ships with support for the following robots.txt_ parsers: + + * :ref:`RobotFileParser ` (default) + * :ref:`Reppy ` + * :ref:`Robotexclusionrulesparser ` + + You can change the robots.txt_ parser with the :setting:`ROBOTSTXT_PARSER` + setting. Or you can also :ref:`implement support for a new parser `. + .. reqmeta:: dont_obey_robotstxt If :attr:`Request.meta ` has @@ -996,6 +1007,74 @@ If :attr:`Request.meta ` has the request will be ignored by this middleware even if :setting:`ROBOTSTXT_OBEY` is enabled. +.. _python-robotfileparser: + +RobotFileParser +~~~~~~~~~~~~~~~ + +`RobotFileParser `_ is +Python's inbuilt ``robots.txt`` parser. The parser is fully compliant with `Martijn Koster's +1996 draft specification `_. It lacks +support for wildcard matching. Scrapy uses this parser by default. + +In order to use this parser, set: + +* :setting:`ROBOTSTXT_PARSER` to ``scrapy.robotstxt.PythonRobotParser`` + +.. _rerp-parser: + +Robotexclusionrulesparser +~~~~~~~~~~~~~~~~~~~~~~~~~ + +`Robotexclusionrulesparser `_ is fully compliant +with `Martijn Koster's 1996 draft specification `_, +with support for wildcard matching. + +In order to use this parser: + +* Install `Robotexclusionrulesparser `_ by running + ``pip install robotexclusionrulesparser`` + +* Set :setting:`ROBOTSTXT_PARSER` setting to + ``scrapy.robotstxt.RerpRobotParser`` + +.. _reppy-parser: + +Reppy parser +~~~~~~~~~~~~ + +`Reppy `_ is a Python wrapper around `Robots Exclusion +Protocol Parser for C++ `_. The parser is fully compliant +with `Martijn Koster's 1996 draft specification `_, +with support for wildcard matching. Unlike +`RobotFileParser `_ and +`Robotexclusionrulesparser `_, it uses the length based +rule, in particular for ``Allow`` and ``Disallow`` directives, where the most specific +rule based on the length of the path trumps the less specific (shorter) rule. + +In order to use this parser: + +* Install `Reppy `_ by running ``pip install reppy`` + +* Set :setting:`ROBOTSTXT_PARSER` setting to + ``scrapy.robotstxt.ReppyRobotParser`` + +.. _support-for-new-robots-parser: + +Implementing support for a new parser +------------------------------------- + +You can implement support for a new robots.txt_ parser by subclassing +the abstract base class :class:`~scrapy.robotstxt.RobotParser` and +implementing the methods described below. + +.. module:: scrapy.robotstxt + :synopsis: robots.txt parser interface and implementations + +.. autoclass:: RobotParser + :members: + +.. _robots.txt: http://www.robotstxt.org/ DownloaderStats --------------- diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 85ae2a305..12606fe47 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -1141,6 +1141,16 @@ If enabled, Scrapy will respect robots.txt policies. For more information see this option is enabled by default in settings.py file generated by ``scrapy startproject`` command. +.. setting:: ROBOTSTXT_PARSER + +ROBOTSTXT_PARSER +---------------- + +Default: ``'scrapy.robotstxt.PythonRobotParser'`` + +The parser backend to use for parsing ``robots.txt`` files. For more information see +:ref:`topics-dlmw-robots`. + .. setting:: SCHEDULER SCHEDULER diff --git a/scrapy/downloadermiddlewares/robotstxt.py b/scrapy/downloadermiddlewares/robotstxt.py index 200245210..c5a60d355 100644 --- a/scrapy/downloadermiddlewares/robotstxt.py +++ b/scrapy/downloadermiddlewares/robotstxt.py @@ -5,8 +5,8 @@ enable this middleware and enable the ROBOTSTXT_OBEY setting. """ import logging - -from six.moves.urllib import robotparser +import sys +import re from twisted.internet.defer import Deferred, maybeDeferred from scrapy.exceptions import NotConfigured, IgnoreRequest @@ -14,6 +14,7 @@ from scrapy.http import Request from scrapy.utils.httpobj import urlparse_cached from scrapy.utils.log import failure_to_exc_info from scrapy.utils.python import to_native_str +from scrapy.utils.misc import load_object logger = logging.getLogger(__name__) @@ -24,10 +25,13 @@ class RobotsTxtMiddleware(object): def __init__(self, crawler): if not crawler.settings.getbool('ROBOTSTXT_OBEY'): raise NotConfigured - + self._default_useragent = crawler.settings.get('USER_AGENT', 'Scrapy') self.crawler = crawler - self._useragent = crawler.settings.get('USER_AGENT') self._parsers = {} + self._parserimpl = load_object(crawler.settings.get('ROBOTSTXT_PARSER')) + + # check if parser dependencies are met, this should throw an error otherwise. + self._parserimpl.from_crawler(self.crawler, b'') @classmethod def from_crawler(cls, crawler): @@ -43,7 +47,8 @@ class RobotsTxtMiddleware(object): def process_request_2(self, rp, request, spider): if rp is None: return - if not rp.can_fetch(to_native_str(self._useragent), request.url): + useragent = request.headers.get(b'User-Agent', self._default_useragent) + if not rp.allowed(request.url, useragent): logger.debug("Forbidden by robots.txt: %(request)s", {'request': request}, extra={'spider': spider}) self.crawler.stats.inc_value('robotstxt/forbidden') @@ -62,13 +67,14 @@ class RobotsTxtMiddleware(object): meta={'dont_obey_robotstxt': True} ) dfd = self.crawler.engine.download(robotsreq, spider) - dfd.addCallback(self._parse_robots, netloc) + dfd.addCallback(self._parse_robots, netloc, spider) dfd.addErrback(self._logerror, robotsreq, spider) dfd.addErrback(self._robots_error, netloc) self.crawler.stats.inc_value('robotstxt/request_count') if isinstance(self._parsers[netloc], Deferred): d = Deferred() + def cb(result): d.callback(result) return result @@ -85,27 +91,10 @@ class RobotsTxtMiddleware(object): extra={'spider': spider}) return failure - def _parse_robots(self, response, netloc): + def _parse_robots(self, response, netloc, spider): self.crawler.stats.inc_value('robotstxt/response_count') - self.crawler.stats.inc_value( - 'robotstxt/response_status_count/{}'.format(response.status)) - rp = robotparser.RobotFileParser(response.url) - body = '' - if hasattr(response, 'text'): - body = response.text - else: # last effort try - try: - body = response.body.decode('utf-8') - except UnicodeDecodeError: - # If we found garbage, disregard it:, - # but keep the lookup cached (in self._parsers) - # Running rp.parse() will set rp state from - # 'disallow all' to 'allow any'. - self.crawler.stats.inc_value('robotstxt/unicode_error_count') - # stdlib's robotparser expects native 'str' ; - # with unicode input, non-ASCII encoded bytes decoding fails in Python2 - rp.parse(to_native_str(body).splitlines()) - + self.crawler.stats.inc_value('robotstxt/response_status_count/{}'.format(response.status)) + rp = self._parserimpl.from_crawler(self.crawler, response.body) rp_dfd = self._parsers[netloc] self._parsers[netloc] = rp rp_dfd.callback(rp) diff --git a/scrapy/robotstxt.py b/scrapy/robotstxt.py new file mode 100644 index 000000000..4bfb275fd --- /dev/null +++ b/scrapy/robotstxt.py @@ -0,0 +1,112 @@ +import sys +import logging +from abc import ABCMeta, abstractmethod +from six import with_metaclass + +from scrapy.utils.python import to_native_str, to_unicode + +logger = logging.getLogger(__name__) + + +class RobotParser(with_metaclass(ABCMeta)): + @classmethod + @abstractmethod + def from_crawler(cls, crawler, robotstxt_body): + """Parse the content of a robots.txt_ file as bytes. This must be a class method. + It must return a new instance of the parser backend. + + :param crawler: crawler which made the request + :type crawler: :class:`~scrapy.crawler.Crawler` instance + + :param robotstxt_body: content of a robots.txt_ file. + :type robotstxt_body: bytes + """ + pass + + @abstractmethod + def allowed(self, url, user_agent): + """Return ``True`` if ``user_agent`` is allowed to crawl ``url``, otherwise return ``False``. + + :param url: Absolute URL + :type url: string + + :param user_agent: User agent + :type user_agent: string + """ + pass + + +class PythonRobotParser(RobotParser): + def __init__(self, robotstxt_body, spider): + from six.moves.urllib_robotparser import RobotFileParser + self.spider = spider + try: + robotstxt_body = to_native_str(robotstxt_body) + except UnicodeDecodeError: + # If we found garbage or robots.txt in an encoding other than UTF-8, disregard it. + # Switch to 'allow all' state. + logger.warning("Failure while parsing robots.txt using %(parser)s." + " File either contains garbage or is in an encoding other than UTF-8, treating it as an empty file.", + {'parser': "RobotFileParser"}, + exc_info=sys.exc_info(), + extra={'spider': self.spider}) + robotstxt_body = '' + self.rp = RobotFileParser() + self.rp.parse(robotstxt_body.splitlines()) + + @classmethod + def from_crawler(cls, crawler, robotstxt_body): + spider = None if not crawler else crawler.spider + o = cls(robotstxt_body, spider) + return o + + def allowed(self, url, user_agent): + user_agent = to_native_str(user_agent) + url = to_native_str(url) + return self.rp.can_fetch(user_agent, url) + + +class ReppyRobotParser(RobotParser): + def __init__(self, robotstxt_body, spider): + from reppy.robots import Robots + self.spider = spider + self.rp = Robots.parse('', robotstxt_body) + + @classmethod + def from_crawler(cls, crawler, robotstxt_body): + spider = None if not crawler else crawler.spider + o = cls(robotstxt_body, spider) + return o + + def allowed(self, url, user_agent): + return self.rp.allowed(url, user_agent) + + +class RerpRobotParser(RobotParser): + def __init__(self, robotstxt_body, spider): + from robotexclusionrulesparser import RobotExclusionRulesParser + self.spider = spider + self.rp = RobotExclusionRulesParser() + try: + robotstxt_body = robotstxt_body.decode('utf-8') + except UnicodeDecodeError: + # If we found garbage or robots.txt in an encoding other than UTF-8, disregard it. + # Switch to 'allow all' state. + logger.warning("Failure while parsing robots.txt using %(parser)s." + " File either contains garbage or is in an encoding other than UTF-8, treating it as an empty file.", + {'parser': "RobotExclusionRulesParser"}, + exc_info=sys.exc_info(), + extra={'spider': self.spider}) + robotstxt_body = '' + self.rp.parse(robotstxt_body) + + @classmethod + def from_crawler(cls, crawler, robotstxt_body): + spider = None if not crawler else crawler.spider + o = cls(robotstxt_body, spider) + return o + + def allowed(self, url, user_agent): + user_agent = to_unicode(user_agent) + url = to_unicode(url) + return self.rp.is_allowed(user_agent, url) diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 086adf48e..81fee543f 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -245,6 +245,7 @@ RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408, 429] RETRY_PRIORITY_ADJUST = -1 ROBOTSTXT_OBEY = False +ROBOTSTXT_PARSER = 'scrapy.robotstxt.PythonRobotParser' SCHEDULER = 'scrapy.core.scheduler.Scheduler' SCHEDULER_DISK_QUEUE = 'scrapy.squeues.PickleLifoDiskQueue' diff --git a/tests/test_downloadermiddleware_robotstxt.py b/tests/test_downloadermiddleware_robotstxt.py index 2b3548bdd..79f172848 100644 --- a/tests/test_downloadermiddleware_robotstxt.py +++ b/tests/test_downloadermiddleware_robotstxt.py @@ -10,6 +10,7 @@ from scrapy.exceptions import IgnoreRequest, NotConfigured from scrapy.http import Request, Response, TextResponse from scrapy.settings import Settings from tests import mock +from tests.test_robotstxt_interface import rerp_available, reppy_available class RobotsTxtMiddlewareTest(unittest.TestCase): @@ -41,6 +42,7 @@ User-Agent: UnicödeBöt Disallow: /some/randome/page.html """.encode('utf-8') response = TextResponse('http://site.local/robots.txt', body=ROBOTS) + def return_response(request, spider): deferred = Deferred() reactor.callFromThread(deferred.callback, response) @@ -77,6 +79,7 @@ Disallow: /some/randome/page.html crawler = self.crawler crawler.settings.set('ROBOTSTXT_OBEY', True) response = Response('http://site.local/robots.txt', body=b'GIF89a\xd3\x00\xfe\x00\xa2') + def return_response(request, spider): deferred = Deferred() reactor.callFromThread(deferred.callback, response) @@ -99,6 +102,7 @@ Disallow: /some/randome/page.html crawler = self.crawler crawler.settings.set('ROBOTSTXT_OBEY', True) response = Response('http://site.local/robots.txt') + def return_response(request, spider): deferred = Deferred() reactor.callFromThread(deferred.callback, response) @@ -118,6 +122,7 @@ Disallow: /some/randome/page.html def test_robotstxt_error(self): self.crawler.settings.set('ROBOTSTXT_OBEY', True) err = error.DNSLookupError('Robotstxt address not found') + def return_failure(request, spider): deferred = Deferred() reactor.callFromThread(deferred.errback, failure.Failure(err)) @@ -133,6 +138,7 @@ Disallow: /some/randome/page.html def test_robotstxt_immediate_error(self): self.crawler.settings.set('ROBOTSTXT_OBEY', True) err = error.DNSLookupError('Robotstxt address not found') + def immediate_failure(request, spider): deferred = Deferred() deferred.errback(failure.Failure(err)) @@ -144,6 +150,7 @@ Disallow: /some/randome/page.html def test_ignore_robotstxt_request(self): self.crawler.settings.set('ROBOTSTXT_OBEY', True) + def ignore_request(request, spider): deferred = Deferred() reactor.callFromThread(deferred.errback, failure.Failure(IgnoreRequest())) @@ -167,3 +174,21 @@ Disallow: /some/randome/page.html spider = None # not actually used return self.assertFailure(maybeDeferred(middleware.process_request, request, spider), IgnoreRequest) + + +class RobotsTxtMiddlewareWithRerpTest(RobotsTxtMiddlewareTest): + if not rerp_available(): + skip = "Rerp parser is not installed" + + def setUp(self): + super(RobotsTxtMiddlewareWithRerpTest, self).setUp() + self.crawler.settings.set('ROBOTSTXT_PARSER', 'scrapy.robotstxt.RerpRobotParser') + + +class RobotsTxtMiddlewareWithReppyTest(RobotsTxtMiddlewareTest): + if not reppy_available(): + skip = "Reppy parser is not installed" + + def setUp(self): + super(RobotsTxtMiddlewareWithReppyTest, self).setUp() + self.crawler.settings.set('ROBOTSTXT_PARSER', 'scrapy.robotstxt.ReppyRobotParser') diff --git a/tests/test_robotstxt_interface.py b/tests/test_robotstxt_interface.py new file mode 100644 index 000000000..2819786b5 --- /dev/null +++ b/tests/test_robotstxt_interface.py @@ -0,0 +1,142 @@ +# coding=utf-8 +from twisted.trial import unittest +from scrapy.utils.python import to_native_str + + +def reppy_available(): + # check if reppy parser is installed + try: + from reppy.robots import Robots + except ImportError: + return False + return True + + +def rerp_available(): + # check if robotexclusionrulesparser is installed + try: + from robotexclusionrulesparser import RobotExclusionRulesParser + except ImportError: + return False + return True + + +class BaseRobotParserTest: + def _setUp(self, parser_cls): + self.parser_cls = parser_cls + + def test_allowed(self): + robotstxt_robotstxt_body = ("User-agent: * \n" + "Disallow: /disallowed \n" + "Allow: /allowed \n" + "Crawl-delay: 10".encode('utf-8')) + rp = self.parser_cls.from_crawler(crawler=None, robotstxt_body=robotstxt_robotstxt_body) + self.assertTrue(rp.allowed("https://www.site.local/allowed", "*")) + self.assertFalse(rp.allowed("https://www.site.local/disallowed", "*")) + + def test_allowed_wildcards(self): + robotstxt_robotstxt_body = """User-agent: first + Disallow: /disallowed/*/end$ + + User-agent: second + Allow: /*allowed + Disallow: / + """.encode('utf-8') + rp = self.parser_cls.from_crawler(crawler=None, robotstxt_body=robotstxt_robotstxt_body) + + self.assertTrue(rp.allowed("https://www.site.local/disallowed", "first")) + self.assertFalse(rp.allowed("https://www.site.local/disallowed/xyz/end", "first")) + self.assertFalse(rp.allowed("https://www.site.local/disallowed/abc/end", "first")) + self.assertTrue(rp.allowed("https://www.site.local/disallowed/xyz/endinglater", "first")) + + self.assertTrue(rp.allowed("https://www.site.local/allowed", "second")) + self.assertTrue(rp.allowed("https://www.site.local/is_still_allowed", "second")) + self.assertTrue(rp.allowed("https://www.site.local/is_allowed_too", "second")) + + def test_length_based_precedence(self): + robotstxt_robotstxt_body = ("User-agent: * \n" + "Disallow: / \n" + "Allow: /page".encode('utf-8')) + rp = self.parser_cls.from_crawler(crawler=None, robotstxt_body=robotstxt_robotstxt_body) + self.assertTrue(rp.allowed("https://www.site.local/page", "*")) + + def test_order_based_precedence(self): + robotstxt_robotstxt_body = ("User-agent: * \n" + "Disallow: / \n" + "Allow: /page".encode('utf-8')) + rp = self.parser_cls.from_crawler(crawler=None, robotstxt_body=robotstxt_robotstxt_body) + self.assertFalse(rp.allowed("https://www.site.local/page", "*")) + + def test_empty_response(self): + """empty response should equal 'allow all'""" + rp = self.parser_cls.from_crawler(crawler=None, robotstxt_body=b'') + self.assertTrue(rp.allowed("https://site.local/", "*")) + self.assertTrue(rp.allowed("https://site.local/", "chrome")) + self.assertTrue(rp.allowed("https://site.local/index.html", "*")) + self.assertTrue(rp.allowed("https://site.local/disallowed", "*")) + + def test_garbage_response(self): + """garbage response should be discarded, equal 'allow all'""" + robotstxt_robotstxt_body = b'GIF89a\xd3\x00\xfe\x00\xa2' + rp = self.parser_cls.from_crawler(crawler=None, robotstxt_body=robotstxt_robotstxt_body) + self.assertTrue(rp.allowed("https://site.local/", "*")) + self.assertTrue(rp.allowed("https://site.local/", "chrome")) + self.assertTrue(rp.allowed("https://site.local/index.html", "*")) + self.assertTrue(rp.allowed("https://site.local/disallowed", "*")) + + def test_unicode_url_and_useragent(self): + robotstxt_robotstxt_body = u""" + User-Agent: * + Disallow: /admin/ + Disallow: /static/ + # taken from https://en.wikipedia.org/robots.txt + Disallow: /wiki/K%C3%A4ytt%C3%A4j%C3%A4: + Disallow: /wiki/Käyttäjä: + + User-Agent: UnicödeBöt + Disallow: /some/randome/page.html""".encode('utf-8') + rp = self.parser_cls.from_crawler(crawler=None, robotstxt_body=robotstxt_robotstxt_body) + self.assertTrue(rp.allowed("https://site.local/", "*")) + self.assertFalse(rp.allowed("https://site.local/admin/", "*")) + self.assertFalse(rp.allowed("https://site.local/static/", "*")) + self.assertTrue(rp.allowed("https://site.local/admin/", u"UnicödeBöt")) + self.assertFalse(rp.allowed("https://site.local/wiki/K%C3%A4ytt%C3%A4j%C3%A4:", "*")) + self.assertFalse(rp.allowed(u"https://site.local/wiki/Käyttäjä:", "*")) + self.assertTrue(rp.allowed("https://site.local/some/randome/page.html", "*")) + self.assertFalse(rp.allowed("https://site.local/some/randome/page.html", u"UnicödeBöt")) + + +class PythonRobotParserTest(BaseRobotParserTest, unittest.TestCase): + def setUp(self): + from scrapy.robotstxt import PythonRobotParser + super(PythonRobotParserTest, self)._setUp(PythonRobotParser) + + def test_length_based_precedence(self): + raise unittest.SkipTest("RobotFileParser does not support length based directives precedence.") + + def test_allowed_wildcards(self): + raise unittest.SkipTest("RobotFileParser does not support wildcards.") + + +class ReppyRobotParserTest(BaseRobotParserTest, unittest.TestCase): + if not reppy_available(): + skip = "Reppy parser is not installed" + + def setUp(self): + from scrapy.robotstxt import ReppyRobotParser + super(ReppyRobotParserTest, self)._setUp(ReppyRobotParser) + + def test_order_based_precedence(self): + raise unittest.SkipTest("Rerp does not support order based directives precedence.") + + +class RerpRobotParserTest(BaseRobotParserTest, unittest.TestCase): + if not rerp_available(): + skip = "Rerp parser is not installed" + + def setUp(self): + from scrapy.robotstxt import RerpRobotParser + super(RerpRobotParserTest, self)._setUp(RerpRobotParser) + + def test_length_based_precedence(self): + raise unittest.SkipTest("Rerp does not support length based directives precedence.") diff --git a/tox.ini b/tox.ini index 157a8b3ed..c918731f4 100644 --- a/tox.ini +++ b/tox.ini @@ -116,3 +116,17 @@ changedir = {[docs]changedir} deps = {[docs]deps} commands = sphinx-build -W -b linkcheck . {envtmpdir}/linkcheck + +[testenv:py37-extra-deps] +basepython = python3.7 +deps = + {[testenv:py34]deps} + reppy + robotexclusionrulesparser + +[testenv:py27-extra-deps] +basepython = python2.7 +deps = + {[testenv]deps} + reppy + robotexclusionrulesparser From 9a4cd94244a01a7029e9f03959805a2e45a07874 Mon Sep 17 00:00:00 2001 From: Anubhav Patel Date: Sat, 3 Aug 2019 22:46:06 +0530 Subject: [PATCH 024/387] fixes typo --- scrapy/core/downloader/webclient.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/core/downloader/webclient.py b/scrapy/core/downloader/webclient.py index 1c89a0f9e..3a5890ed0 100644 --- a/scrapy/core/downloader/webclient.py +++ b/scrapy/core/downloader/webclient.py @@ -95,7 +95,7 @@ class ScrapyHTTPPageGetter(HTTPClient): class ScrapyHTTPClientFactory(HTTPClientFactory): """Scrapy implementation of the HTTPClientFactory overwriting the - serUrl method to make use of our Url object that cache the parse + setUrl method to make use of our Url object that cache the parse result. """ From 18d0affc015cd1be02cc9b227c4d89007dd8c816 Mon Sep 17 00:00:00 2001 From: Shivam Sandbhor Date: Mon, 5 Aug 2019 16:53:35 +0530 Subject: [PATCH 025/387] Update reactor.py, updated 'if' sequencing , possibly eliminating a bug if portrange=None This should be the proper ordering.This is the explanation. If 'not portrange' is True ,it is guaranteed that `not hasattr(portrange, '__iter__')` is also True the converse of this is not always true.(for example, consider portrange=None, for such case we were executing the logic for `not hasattr(portrange, '__iter__')` . ).Such case is eliminated by this PR. --- scrapy/utils/reactor.py | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/scrapy/utils/reactor.py b/scrapy/utils/reactor.py index 83186a372..eda7867e3 100644 --- a/scrapy/utils/reactor.py +++ b/scrapy/utils/reactor.py @@ -3,10 +3,10 @@ from twisted.internet import reactor, error def listen_tcp(portrange, host, factory): """Like reactor.listenTCP but tries different ports in a range.""" assert len(portrange) <= 2, "invalid portrange: %s" % portrange - if not hasattr(portrange, '__iter__'): - return reactor.listenTCP(portrange, factory, interface=host) if not portrange: return reactor.listenTCP(0, factory, interface=host) + if not hasattr(portrange, '__iter__'): + return reactor.listenTCP(portrange, factory, interface=host) if len(portrange) == 1: return reactor.listenTCP(portrange[0], factory, interface=host) for x in range(portrange[0], portrange[1]+1): From a8621bbc2929b904dfe14985dd1f94aeedc86648 Mon Sep 17 00:00:00 2001 From: tpeng Date: Thu, 26 Jun 2014 11:28:03 +0200 Subject: [PATCH 026/387] show all the missing field when scrapes contract fails --- scrapy/contracts/default.py | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/scrapy/contracts/default.py b/scrapy/contracts/default.py index 20582503d..0f6bdbad2 100644 --- a/scrapy/contracts/default.py +++ b/scrapy/contracts/default.py @@ -84,6 +84,6 @@ class ScrapesContract(Contract): def post_process(self, output): for x in output: if isinstance(x, (BaseItem, dict)): - for arg in self.args: - if not arg in x: - raise ContractFail("'%s' field is missing" % arg) + missing = [arg for arg in self.args if arg not in x] + if missing: + raise ContractFail("'%s' field is missing" % " ".join(missing)) From bff335cf7f90a0dc3aeafc199aaa81697d5f122b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 5 Aug 2019 15:47:58 +0200 Subject: [PATCH 027/387] Improve the error message in contract failures due to multiple missing fields --- scrapy/contracts/default.py | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/scrapy/contracts/default.py b/scrapy/contracts/default.py index 0f6bdbad2..7745959a7 100644 --- a/scrapy/contracts/default.py +++ b/scrapy/contracts/default.py @@ -86,4 +86,5 @@ class ScrapesContract(Contract): if isinstance(x, (BaseItem, dict)): missing = [arg for arg in self.args if arg not in x] if missing: - raise ContractFail("'%s' field is missing" % " ".join(missing)) + raise ContractFail( + "Missing fields: %s" % ", ".join(missing)) From 7b755a41a1046e031015130971c9b9cc8885176e Mon Sep 17 00:00:00 2001 From: Pengyu Chen Date: Tue, 6 Aug 2019 15:18:59 +0100 Subject: [PATCH 028/387] Added: Properly handling quoted passwords in FEED_URI for FTP --- scrapy/extensions/feedexport.py | 4 ++-- tests/test_feedexport.py | 9 ++++++++- 2 files changed, 10 insertions(+), 3 deletions(-) diff --git a/scrapy/extensions/feedexport.py b/scrapy/extensions/feedexport.py index d35551fdd..ce2846eba 100644 --- a/scrapy/extensions/feedexport.py +++ b/scrapy/extensions/feedexport.py @@ -11,7 +11,7 @@ import posixpath from tempfile import NamedTemporaryFile from datetime import datetime import six -from six.moves.urllib.parse import urlparse +from six.moves.urllib.parse import urlparse, unquote from ftplib import FTP from zope.interface import Interface, implementer @@ -162,7 +162,7 @@ class FTPFeedStorage(BlockingFeedStorage): self.host = u.hostname self.port = int(u.port or '21') self.username = u.username - self.password = u.password + self.password = unquote(u.password) self.path = u.path self.use_active_mode = use_active_mode diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index c5063253a..f32ac2a4b 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -6,7 +6,8 @@ import warnings from io import BytesIO import tempfile import shutil -from six.moves.urllib.parse import urljoin, urlparse +import string +from six.moves.urllib.parse import urljoin, urlparse, quote from six.moves.urllib.request import pathname2url from zope.interface.verify import verifyObject @@ -98,6 +99,12 @@ class FTPFeedStorageTest(unittest.TestCase): verifyObject(IFeedStorage, st) return self._assert_stores(st, path) + def test_uri_auth_quote(self): + # RFC3986: 3.2.1. User Information + pw_quoted = quote(string.punctuation, safe='') + st = FTPFeedStorage('ftp://foo:%s@example.com/some_path' % pw_quoted) + self.assertEqual(st.password, string.punctuation) + @defer.inlineCallbacks def _assert_stores(self, storage, path): spider = self.get_test_spider() From 5dbeece8da6686709eb00af9ffb91dd30a832a2e Mon Sep 17 00:00:00 2001 From: elacuesta Date: Wed, 7 Aug 2019 04:36:52 -0300 Subject: [PATCH 029/387] [MRG+1] Drop py34 support - Update CI envs (#3892) * Drop py34 support * Travis experiments * More Travis experiments * Bump Twisted version for py35+ (stretch) * Remove Debian build * Remove pinned lxml for Py34 * Fix merge error * Remove unused tox env * Add environment with pinned versions for py36 * Bump minimum Twisted version in py27; Envs with pinned versions for py27 and py35 * Add botocore as extra dep for py27 tests * Update requirements-py2.txt * Add botocore and Pillow as extra dependencies --- .travis.yml | 50 ++++++++++---------- README.rst | 2 +- docs/faq.rst | 2 +- docs/intro/install.rst | 2 +- requirements-py2.txt | 25 ++++++---- requirements-py3.txt | 26 +++++++---- setup.py | 26 +++++------ tests/requirements-py2.txt | 7 +-- tests/requirements-py3.txt | 9 ++-- tox.ini | 94 +++++++++++++++++++------------------- 10 files changed, 128 insertions(+), 115 deletions(-) diff --git a/.travis.yml b/.travis.yml index 138e81c64..0190a7f4d 100644 --- a/.travis.yml +++ b/.travis.yml @@ -1,5 +1,5 @@ language: python -dist: trusty +dist: xenial branches: only: - master @@ -7,32 +7,28 @@ branches: - /^\d\.\d+\.\d+(rc\d+|\.dev\d+)?$/ matrix: include: - - python: 2.7 - env: TOXENV=py27 - - python: 2.7 - env: TOXENV=jessie - - python: 2.7 - env: TOXENV=pypy - - python: 2.7 - env: TOXENV=pypy3 - - python: 3.4 - env: TOXENV=py34 - - python: 3.5 - env: TOXENV=py35 - - python: 3.6 - env: TOXENV=py36 - - python: 3.7 - env: TOXENV=py37 - dist: xenial - sudo: true - - python: 3.6 - env: TOXENV=docs - - python: 3.7 - env: TOXENV=py37-extra-deps - dist: xenial - sudo: true - - python: 2.7 - env: TOXENV=py27-extra-deps + - env: TOXENV=py27 + python: 2.7 + - env: TOXENV=py27-pinned + python: 2.7 + - env: TOXENV=py27-extra-deps + python: 2.7 + - env: TOXENV=pypy + python: 2.7 + - env: TOXENV=pypy3 + python: 3.5 + - env: TOXENV=py35 + python: 3.5 + - env: TOXENV=py35-pinned + python: 3.5 + - env: TOXENV=py36 + python: 3.6 + - env: TOXENV=py37 + python: 3.7 + - env: TOXENV=py37-extra-deps + python: 3.7 + - env: TOXENV=docs + python: 3.6 install: - | if [ "$TOXENV" = "pypy" ]; then diff --git a/README.rst b/README.rst index c28d217ff..bd82bff06 100644 --- a/README.rst +++ b/README.rst @@ -40,7 +40,7 @@ https://scrapy.org Requirements ============ -* Python 2.7 or Python 3.4+ +* Python 2.7 or Python 3.5+ * Works on Linux, Windows, Mac OSX, BSD Install diff --git a/docs/faq.rst b/docs/faq.rst index 44f2b97bd..9733471bf 100644 --- a/docs/faq.rst +++ b/docs/faq.rst @@ -69,7 +69,7 @@ Here's an example spider using BeautifulSoup API, with ``lxml`` as the HTML pars What Python versions does Scrapy support? ----------------------------------------- -Scrapy is supported under Python 2.7 and Python 3.4+ +Scrapy is supported under Python 2.7 and Python 3.5+ under CPython (default Python implementation) and PyPy (starting with PyPy 5.9). Python 2.6 support was dropped starting at Scrapy 0.20. Python 3 support was added in Scrapy 1.1. diff --git a/docs/intro/install.rst b/docs/intro/install.rst index daec7fcb7..2bf98dbdc 100644 --- a/docs/intro/install.rst +++ b/docs/intro/install.rst @@ -7,7 +7,7 @@ Installation guide Installing Scrapy ================= -Scrapy runs on Python 2.7 and Python 3.4 or above +Scrapy runs on Python 2.7 and Python 3.5 or above under CPython (default Python implementation) and PyPy (starting with PyPy 5.9). If you're using `Anaconda`_ or `Miniconda`_, you can install the package from diff --git a/requirements-py2.txt b/requirements-py2.txt index 0771aae3a..9e6944240 100644 --- a/requirements-py2.txt +++ b/requirements-py2.txt @@ -1,10 +1,17 @@ -Twisted>=13.1.0 -lxml -pyOpenSSL -cssselect>=0.9 -queuelib -w3lib>=1.17.0 -six>=1.5.2 +parsel>=1.5.0 PyDispatcher>=2.0.5 -parsel>=1.5 -service_identity +w3lib>=1.17.0 + +pyOpenSSL>=16.2.0 # Earlier versions fail with "AttributeError: module 'lib' has no attribute 'SSL_ST_INIT'" +queuelib>=1.4.2 # Earlier versions fail with "AttributeError: '...QueueTest' object has no attribute 'qpath'" +cryptography>=2.0 # Earlier versions would fail to install + +# Reference versions taken from +# https://packages.ubuntu.com/xenial/python/ +# https://packages.ubuntu.com/xenial/zope/ +cssselect>=0.9.1 +lxml>=3.5.0 +service_identity>=16.0.0 +six>=1.10.0 +Twisted>=16.0.0 +zope.interface>=4.1.3 diff --git a/requirements-py3.txt b/requirements-py3.txt index 478ed0010..cd183a525 100644 --- a/requirements-py3.txt +++ b/requirements-py3.txt @@ -1,11 +1,17 @@ -Twisted>=17.9.0 -lxml;python_version!="3.4" -lxml<=4.3.5;python_version=="3.4" -pyOpenSSL>=0.13.1 -cssselect>=0.9 -queuelib>=1.1.1 -w3lib>=1.17.0 -six>=1.5.2 +parsel>=1.5.0 PyDispatcher>=2.0.5 -parsel>=1.5 -service_identity +Twisted>=17.9.0 +w3lib>=1.17.0 + +pyOpenSSL>=16.2.0 # Earlier versions fail with "AttributeError: module 'lib' has no attribute 'SSL_ST_INIT'" +queuelib>=1.4.2 # Earlier versions fail with "AttributeError: '...QueueTest' object has no attribute 'qpath'" +cryptography>=2.0 # Earlier versions would fail to install + +# Reference versions taken from +# https://packages.ubuntu.com/xenial/python/ +# https://packages.ubuntu.com/xenial/zope/ +cssselect>=0.9.1 +lxml>=3.5.0 +service_identity>=16.0.0 +six>=1.10.0 +zope.interface>=4.1.3 diff --git a/setup.py b/setup.py index ee0aaabf0..37892cfbf 100644 --- a/setup.py +++ b/setup.py @@ -53,7 +53,6 @@ setup( 'Programming Language :: Python :: 2', 'Programming Language :: Python :: 2.7', 'Programming Language :: Python :: 3', - 'Programming Language :: Python :: 3.4', 'Programming Language :: Python :: 3.5', 'Programming Language :: Python :: 3.6', 'Programming Language :: Python :: 3.7', @@ -63,20 +62,21 @@ setup( 'Topic :: Software Development :: Libraries :: Application Frameworks', 'Topic :: Software Development :: Libraries :: Python Modules', ], - python_requires='>=2.7, !=3.0.*, !=3.1.*, !=3.2.*, !=3.3.*', + python_requires='>=2.7, !=3.0.*, !=3.1.*, !=3.2.*, !=3.3.*, !=3.4.*', install_requires=[ - 'Twisted>=13.1.0;python_version!="3.4"', - 'Twisted>=13.1.0,<=19.2.0;python_version=="3.4"', - 'w3lib>=1.17.0', - 'queuelib', - 'lxml;python_version!="3.4"', - 'lxml<=4.3.5;python_version=="3.4"', - 'pyOpenSSL', - 'cssselect>=0.9', - 'six>=1.5.2', - 'parsel>=1.5', + 'Twisted>=16.0.0;python_version=="2.7"', + 'Twisted>=17.9.0;python_version>="3.5"', + 'cryptography>=2.0', + 'cssselect>=0.9.1', + 'lxml>=3.5.0', + 'parsel>=1.5.0', 'PyDispatcher>=2.0.5', - 'service_identity', + 'pyOpenSSL>=16.2.0', + 'queuelib>=1.4.2', + 'service_identity>=16.0.0', + 'six>=1.10.0', + 'w3lib>=1.17.0', + 'zope.interface>=4.1.3', ], extras_require=extras_require, ) diff --git a/tests/requirements-py2.txt b/tests/requirements-py2.txt index be809b151..f621eb4eb 100644 --- a/tests/requirements-py2.txt +++ b/tests/requirements-py2.txt @@ -1,14 +1,15 @@ # Tests requirements -mock +brotlipy +jmespath mitmproxy==0.10.1 +mock netlib==0.10.1 pytest pytest-cov pytest-twisted pytest-xdist -jmespath -brotlipy testfixtures + # optional for shell wrapper tests bpython ipython<6.0 diff --git a/tests/requirements-py3.txt b/tests/requirements-py3.txt index ed7bf0be0..cb67bc40e 100644 --- a/tests/requirements-py3.txt +++ b/tests/requirements-py3.txt @@ -1,13 +1,14 @@ +# Tests requirements +jmespath +leveldb; sys_platform != "win32" pytest pytest-cov pytest-twisted pytest-xdist testfixtures -jmespath -leveldb; sys_platform != "win32" -botocore + # optional for shell wrapper tests bpython -ipython brotlipy +ipython pywin32; sys_platform == "win32" diff --git a/tox.ini b/tox.ini index c918731f4..c3502c2ca 100644 --- a/tox.ini +++ b/tox.ini @@ -11,10 +11,10 @@ deps = -ctests/constraints.txt -rrequirements-py2.txt # Extras - botocore + botocore>=1.3.23 google-cloud-storage - Pillow != 3.0.0 leveldb + Pillow>=3.4.2 -rtests/requirements-py2.txt passenv = S3_TEST_FILE_URI @@ -25,72 +25,74 @@ passenv = commands = py.test --cov=scrapy --cov-report= {posargs:scrapy tests} -[testenv:trusty] +[testenv:py27-pinned] basepython = python2.7 deps = - pyOpenSSL==0.13 - lxml==3.3.3 - Twisted==13.2.0 - boto==2.20.1 - Pillow==2.3.0 + -ctests/constraints.txt + cryptography==2.0 cssselect==0.9.1 - zope.interface==4.0.5 + lxml==3.5.0 + parsel==1.5.0 + PyDispatcher==2.0.5 + pyOpenSSL==16.2.0 + queuelib==1.4.2 + service_identity==16.0.0 + six==1.10.0 + Twisted==16.0.0 + w3lib==1.17.0 + zope.interface==4.1.3 -rtests/requirements-py2.txt - -[testenv:jessie] -# https://packages.debian.org/en/jessie/python/ -# https://packages.debian.org/en/jessie/zope/ -basepython = python2.7 -deps = - cryptography==0.6.1 - pyOpenSSL==0.14 - lxml==3.4.0 - Twisted==14.0.2 - boto==2.34.0 - Pillow==2.6.1 - cssselect==0.9.1 - zope.interface==4.1.1 - -rtests/requirements-py2.txt -# Not used directly but allows boto GCE plugins to load. -# https://github.com/GoogleCloudPlatform/compute-image-packages/issues/262 - google-compute-engine==2.8.12 - -[testenv:trunk] -basepython = python2.7 -commands = - pip install -U https://github.com/scrapy/w3lib/archive/master.zip#egg=w3lib - pip install -U https://github.com/scrapy/queuelib/archive/master.zip#egg=queuelib - py.test --cov=scrapy --cov-report= {posargs:scrapy tests} + # Extras + botocore==1.3.23 + Pillow==3.4.2 [testenv:pypy] basepython = pypy commands = py.test {posargs:scrapy tests} -[testenv:py34] -basepython = python3.4 +[testenv:py35] +basepython = python3.5 deps = -ctests/constraints.txt -rrequirements-py3.txt - # Extras - Pillow -rtests/requirements-py3.txt + # Extras + botocore>=1.3.23 + Pillow>=3.4.2 -[testenv:py35] +[testenv:py35-pinned] basepython = python3.5 -deps = {[testenv:py34]deps} +deps = + -ctests/constraints.txt + cryptography==2.0 + cssselect==0.9.1 + lxml==3.5.0 + parsel==1.5.0 + PyDispatcher==2.0.5 + pyOpenSSL==16.2.0 + queuelib==1.4.2 + service_identity==16.0.0 + six==1.10.0 + Twisted==17.9.0 + w3lib==1.17.0 + zope.interface==4.1.3 + -rtests/requirements-py3.txt + # Extras + botocore==1.3.23 + Pillow==3.4.2 [testenv:py36] basepython = python3.6 -deps = {[testenv:py34]deps} +deps = {[testenv:py35]deps} [testenv:py37] basepython = python3.7 -deps = {[testenv:py34]deps} +deps = {[testenv:py35]deps} [testenv:pypy3] basepython = pypy3 -deps = {[testenv:py34]deps} +deps = {[testenv:py35]deps} commands = py.test {posargs:scrapy tests} @@ -119,14 +121,14 @@ commands = [testenv:py37-extra-deps] basepython = python3.7 -deps = - {[testenv:py34]deps} +deps = + {[testenv:py35]deps} reppy robotexclusionrulesparser [testenv:py27-extra-deps] basepython = python2.7 -deps = +deps = {[testenv]deps} reppy robotexclusionrulesparser From 595c995ee640e360e5b574577594b273cc8af071 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Wed, 7 Aug 2019 15:38:04 -0300 Subject: [PATCH 030/387] Simplify version reporting --- scrapy/utils/versions.py | 26 +++++++------------------- 1 file changed, 7 insertions(+), 19 deletions(-) diff --git a/scrapy/utils/versions.py b/scrapy/utils/versions.py index 58c7aef85..3f8122154 100644 --- a/scrapy/utils/versions.py +++ b/scrapy/utils/versions.py @@ -1,8 +1,10 @@ import platform import sys +import cryptography import cssselect import lxml.etree +import OpenSSL import parsel import twisted import w3lib @@ -13,15 +15,6 @@ import scrapy def scrapy_components_versions(): lxml_version = ".".join(map(str, lxml.etree.LXML_VERSION)) libxml2_version = ".".join(map(str, lxml.etree.LIBXML_VERSION)) - try: - w3lib_version = w3lib.__version__ - except AttributeError: - w3lib_version = "<1.14.3" - try: - import cryptography - cryptography_version = cryptography.__version__ - except ImportError: - cryptography_version = "unknown" return [ ("Scrapy", scrapy.__version__), @@ -29,22 +22,17 @@ def scrapy_components_versions(): ("libxml2", libxml2_version), ("cssselect", cssselect.__version__), ("parsel", parsel.__version__), - ("w3lib", w3lib_version), + ("w3lib", w3lib.__version__), ("Twisted", twisted.version.short()), ("Python", sys.version.replace("\n", "- ")), ("pyOpenSSL", _get_openssl_version()), - ("cryptography", cryptography_version), + ("cryptography", cryptography.__version__), ("Platform", platform.platform()), ] def _get_openssl_version(): - try: - import OpenSSL - openssl = OpenSSL.SSL.SSLeay_version(OpenSSL.SSL.SSLEAY_VERSION)\ - .decode('ascii', errors='replace') - # pyOpenSSL 0.12 does not expose openssl version - except AttributeError: - openssl = 'Unknown OpenSSL version' - + openssl = OpenSSL.SSL.SSLeay_version( + OpenSSL.SSL.SSLEAY_VERSION + ).decode('ascii', errors='replace') return '{} ({})'.format(OpenSSL.version.__version__, openssl) From d76b6944c9081684e4961c15dc3426a4d4c67a3c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Marc=20Hern=C3=A1ndez?= Date: Thu, 8 Aug 2019 09:43:42 +0200 Subject: [PATCH 031/387] Create Request from curl command (#3862) --- docs/topics/developer-tools.rst | 30 ++++- docs/topics/dynamic-content.rst | 7 + docs/topics/request-response.rst | 2 + scrapy/http/request/__init__.py | 32 +++++ scrapy/utils/curl.py | 95 ++++++++++++++ tests/test_http_request.py | 76 +++++++++++ tests/test_utils_curl.py | 211 +++++++++++++++++++++++++++++++ 7 files changed, 450 insertions(+), 3 deletions(-) create mode 100644 scrapy/utils/curl.py create mode 100644 tests/test_utils_curl.py diff --git a/docs/topics/developer-tools.rst b/docs/topics/developer-tools.rst index 82857c9da..dcf8af365 100644 --- a/docs/topics/developer-tools.rst +++ b/docs/topics/developer-tools.rst @@ -252,9 +252,33 @@ If the handy ``has_next`` element is ``true`` (try loading `quotes.toscrape.com/api/quotes?page=10`_ in your browser or a page-number greater than 10), we increment the ``page`` attribute and ``yield`` a new request, inserting the incremented page-number -into our ``url``. +into our ``url``. -You can see that with a few inspections in the `Network`-tool we +.. _requests-from-curl: + +In more complex websites, it could be difficult to easily reproduce the +requests, as we could need to add ``headers`` or ``cookies`` to make it work. +In those cases you can export the requests in `cURL `_ +format, by right-clicking on each of them in the network tool and using the +:meth:`~scrapy.http.Request.from_curl()` method to generate an equivalent +request:: + + from scrapy import Request + + request = Request.from_curl( + "curl 'http://quotes.toscrape.com/api/quotes?page=1' -H 'User-Agent: Mozil" + "la/5.0 (X11; Linux x86_64; rv:67.0) Gecko/20100101 Firefox/67.0' -H 'Acce" + "pt: */*' -H 'Accept-Language: ca,en-US;q=0.7,en;q=0.3' --compressed -H 'X" + "-Requested-With: XMLHttpRequest' -H 'Proxy-Authorization: Basic QFRLLTAzM" + "zEwZTAxLTk5MWUtNDFiNC1iZWRmLTJjNGI4M2ZiNDBmNDpAVEstMDMzMTBlMDEtOTkxZS00MW" + "I0LWJlZGYtMmM0YjgzZmI0MGY0' -H 'Connection: keep-alive' -H 'Referer: http" + "://quotes.toscrape.com/scroll' -H 'Cache-Control: max-age=0'") + +Alternatively, if you want to know the arguments needed to recreate that +request you can use the :func:`scrapy.utils.curl.curl_to_request_kwargs` +function to get a dictionary with the equivalent arguments. + +As you can see, with a few inspections in the `Network`-tool we were able to easily replicate the dynamic requests of the scrolling functionality of the page. Crawling dynamic pages can be quite daunting and pages can be very complex, but it (mostly) boils down @@ -262,7 +286,7 @@ to identifying the correct request and replicating it in your spider. .. _Developer Tools: https://en.wikipedia.org/wiki/Web_development_tools .. _quotes.toscrape.com: http://quotes.toscrape.com -.. _quotes.toscrape.com/scroll: quotes.toscrape.com/scroll/ +.. _quotes.toscrape.com/scroll: http://quotes.toscrape.com/scroll .. _quotes.toscrape.com/api/quotes?page=10: http://quotes.toscrape.com/api/quotes?page=10 .. _has-class-extension: https://parsel.readthedocs.io/en/latest/usage.html#other-xpath-extensions diff --git a/docs/topics/dynamic-content.rst b/docs/topics/dynamic-content.rst index 8b5dacf56..8334ddcec 100644 --- a/docs/topics/dynamic-content.rst +++ b/docs/topics/dynamic-content.rst @@ -85,6 +85,13 @@ It might be enough to yield a :class:`~scrapy.http.Request` with the same HTTP method and URL. However, you may also need to reproduce the body, headers and form parameters (see :class:`~scrapy.http.FormRequest`) of that request. +As all major browsers allow to export the requests in `cURL +`_ format, Scrapy incorporates the method +:meth:`~scrapy.http.Request.from_curl()` to generate an equivalent +:class:`~scrapy.http.Request` from a cURL command. To get more information +visit :ref:`request from curl ` inside the network +tool section. + Once you get the expected response, you can :ref:`extract the desired data from it `. diff --git a/docs/topics/request-response.rst b/docs/topics/request-response.rst index 4e81ce878..9a5c65b0d 100644 --- a/docs/topics/request-response.rst +++ b/docs/topics/request-response.rst @@ -194,6 +194,8 @@ Request objects copied by default (unless new values are given as arguments). See also :ref:`topics-request-response-ref-request-callback-arguments`. + .. automethod:: from_curl + .. _topics-request-response-ref-request-callback-arguments: Passing additional data to callback functions diff --git a/scrapy/http/request/__init__.py b/scrapy/http/request/__init__.py index f5935c4ef..d09eaf849 100644 --- a/scrapy/http/request/__init__.py +++ b/scrapy/http/request/__init__.py @@ -12,6 +12,7 @@ from scrapy.utils.python import to_bytes from scrapy.utils.trackref import object_ref from scrapy.utils.url import escape_ajax from scrapy.http.common import obsolete_setter +from scrapy.utils.curl import curl_to_request_kwargs class Request(object_ref): @@ -103,3 +104,34 @@ class Request(object_ref): kwargs.setdefault(x, getattr(self, x)) cls = kwargs.pop('cls', self.__class__) return cls(*args, **kwargs) + + @classmethod + def from_curl(cls, curl_command, ignore_unknown_options=True, **kwargs): + """Create a Request object from a string containing a `cURL + `_ command. It populates the HTTP method, the + URL, the headers, the cookies and the body. It accepts the same + arguments as the :class:`Request` class, taking preference and + overriding the values of the same arguments contained in the cURL + command. + + Unrecognized options are ignored by default. To raise an error when + finding unknown options call this method by passing + ``ignore_unknown_options=False``. + + .. caution:: Using :meth:`from_curl` from :class:`~scrapy.http.Request` + subclasses, such as :class:`~scrapy.http.JSONRequest`, or + :class:`~scrapy.http.XmlRpcRequest`, as well as having + :ref:`downloader middlewares ` + and + :ref:`spider middlewares ` + enabled, such as + :class:`~scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware`, + :class:`~scrapy.downloadermiddlewares.useragent.UserAgentMiddleware`, + or + :class:`~scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware`, + may modify the :class:`~scrapy.http.Request` object. + + """ + request_kwargs = curl_to_request_kwargs(curl_command, ignore_unknown_options) + request_kwargs.update(kwargs) + return cls(**request_kwargs) diff --git a/scrapy/utils/curl.py b/scrapy/utils/curl.py new file mode 100644 index 000000000..b3fd0a497 --- /dev/null +++ b/scrapy/utils/curl.py @@ -0,0 +1,95 @@ +import argparse +import warnings +from shlex import split + +from six.moves.http_cookies import SimpleCookie +from six.moves.urllib.parse import urlparse +from six import string_types, iteritems +from w3lib.http import basic_auth_header + + +class CurlParser(argparse.ArgumentParser): + def error(self, message): + error_msg = \ + 'There was an error parsing the curl command: {}'.format(message) + raise ValueError(error_msg) + + +curl_parser = CurlParser() +curl_parser.add_argument('url') +curl_parser.add_argument('-H', '--header', dest='headers', action='append') +curl_parser.add_argument('-X', '--request', dest='method', default='get') +curl_parser.add_argument('-d', '--data', dest='data') +curl_parser.add_argument('-u', '--user', dest='auth') + + +safe_to_ignore_arguments = [ + ['--compressed'], + # `--compressed` argument is not safe to ignore, but it's included here + # because the `HttpCompressionMiddleware` is enabled by default + ['-s', '--silent'], + ['-v', '--verbose'], + ['-#', '--progress-bar'] +] + +for argument in safe_to_ignore_arguments: + curl_parser.add_argument(*argument, action='store_true') + + +def curl_to_request_kwargs(curl_command, ignore_unknown_options=True): + """Convert a cURL command syntax to Request kwargs. + + :param str curl_command: string containing the curl command + :param bool ignore_unknown_options: If true, only a warning is emitted when + cURL options are unknown. Otherwise raises an error. (default: True) + :return: dictionary of Request kwargs + """ + + curl_args = split(curl_command) + + if curl_args[0] != 'curl': + raise ValueError('A curl command must start with "curl"') + + parsed_args, argv = curl_parser.parse_known_args(curl_args[1:]) + + if argv: + msg = 'Unrecognized options: {}'.format(', '.join(argv)) + if ignore_unknown_options: + warnings.warn(msg) + else: + raise ValueError(msg) + + url = parsed_args.url + + # curl automatically prepends 'http' if the scheme is missing, but Request + # needs the scheme to work + parsed_url = urlparse(url) + if not parsed_url.scheme: + url = 'http://' + url + + result = {'method': parsed_args.method.upper(), 'url': url} + + headers = [] + cookies = {} + for header in parsed_args.headers or (): + name, val = header.split(':', 1) + name = name.strip() + val = val.strip() + if name.title() == 'Cookie': + for name, morsel in iteritems(SimpleCookie(val)): + cookies[name] = morsel.value + else: + headers.append((name, val)) + + if parsed_args.auth: + user, password = parsed_args.auth.split(':', 1) + headers.append(('Authorization', basic_auth_header(user, password))) + + if headers: + result['headers'] = headers + if cookies: + result['cookies'] = cookies + if parsed_args.data: + result['body'] = parsed_args.data + + return result diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 952e208de..60494d792 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -269,6 +269,82 @@ class RequestTest(unittest.TestCase): with self.assertRaises(TypeError): self.request_class('http://example.com', a_function, errback='a_function') + def test_from_curl(self): + # Note: more curated tests regarding curl conversion are in + # `test_utils_curl.py` + curl_command = ( + "curl 'http://httpbin.org/post' -X POST -H 'Cookie: _gauges_unique" + "_year=1; _gauges_unique=1; _gauges_unique_month=1; _gauges_unique" + "_hour=1; _gauges_unique_day=1' -H 'Origin: http://httpbin.org' -H" + " 'Accept-Encoding: gzip, deflate' -H 'Accept-Language: en-US,en;q" + "=0.9,ru;q=0.8,es;q=0.7' -H 'Upgrade-Insecure-Requests: 1' -H 'Use" + "r-Agent: Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTM" + "L, like Gecko) Ubuntu Chromium/62.0.3202.75 Chrome/62.0.3202.75 S" + "afari/537.36' -H 'Content-Type: application /x-www-form-urlencode" + "d' -H 'Accept: text/html,application/xhtml+xml,application/xml;q=" + "0.9,image/webp,image/apng,*/*;q=0.8' -H 'Cache-Control: max-age=0" + "' -H 'Referer: http://httpbin.org/forms/post' -H 'Connection: kee" + "p-alive' --data 'custname=John+Smith&custtel=500&custemail=jsmith" + "%40example.org&size=small&topping=cheese&topping=onion&delivery=1" + "2%3A15&comments=' --compressed" + ) + r = self.request_class.from_curl(curl_command) + self.assertEqual(r.method, "POST") + self.assertEqual(r.url, "http://httpbin.org/post") + self.assertEqual(r.body, + b"custname=John+Smith&custtel=500&custemail=jsmith%40" + b"example.org&size=small&topping=cheese&topping=onion" + b"&delivery=12%3A15&comments=") + self.assertEqual(r.cookies, { + '_gauges_unique_year': '1', + '_gauges_unique': '1', + '_gauges_unique_month': '1', + '_gauges_unique_hour': '1', + '_gauges_unique_day': '1' + }) + self.assertEqual(r.headers, { + b'Origin': [b'http://httpbin.org'], + b'Accept-Encoding': [b'gzip, deflate'], + b'Accept-Language': [b'en-US,en;q=0.9,ru;q=0.8,es;q=0.7'], + b'Upgrade-Insecure-Requests': [b'1'], + b'User-Agent': [b'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.' + b'36 (KHTML, like Gecko) Ubuntu Chromium/62.0.3202' + b'.75 Chrome/62.0.3202.75 Safari/537.36'], + b'Content-Type': [b'application /x-www-form-urlencoded'], + b'Accept': [b'text/html,application/xhtml+xml,application/xml;q=0.' + b'9,image/webp,image/apng,*/*;q=0.8'], + b'Cache-Control': [b'max-age=0'], + b'Referer': [b'http://httpbin.org/forms/post'], + b'Connection': [b'keep-alive']}) + + def test_from_curl_with_kwargs(self): + r = self.request_class.from_curl( + 'curl -X PATCH "http://example.org"', + method="POST", + meta={'key': 'value'} + ) + self.assertEqual(r.method, "POST") + self.assertEqual(r.meta, {"key": "value"}) + + def test_from_curl_ignore_unknown_options(self): + # By default: it works and ignores the unknown options: --foo and -z + with warnings.catch_warnings(): # avoid warning when executing tests + warnings.simplefilter('ignore') + r = self.request_class.from_curl( + 'curl -X DELETE "http://example.org" --foo -z', + ) + self.assertEqual(r.method, "DELETE") + + # If `ignore_unknon_options` is set to `False` it raises an error with + # the unknown options: --foo and -z + self.assertRaises( + ValueError, + lambda: self.request_class.from_curl( + 'curl -X PATCH "http://example.org" --foo -z', + ignore_unknown_options=False, + ), + ) + class FormRequestTest(RequestTest): diff --git a/tests/test_utils_curl.py b/tests/test_utils_curl.py new file mode 100644 index 000000000..c5655df7e --- /dev/null +++ b/tests/test_utils_curl.py @@ -0,0 +1,211 @@ +import unittest +import warnings + +from six import assertRaisesRegex +from w3lib.http import basic_auth_header + +from scrapy import Request +from scrapy.utils.curl import curl_to_request_kwargs + + +class CurlToRequestKwargsTest(unittest.TestCase): + maxDiff = 5000 + + def _test_command(self, curl_command, expected_result): + result = curl_to_request_kwargs(curl_command) + self.assertEqual(result, expected_result) + try: + Request(**result) + except TypeError as e: + self.fail("Request kwargs are not correct {}".format(e)) + + def test_get(self): + curl_command = "curl http://example.org/" + expected_result = {"method": "GET", "url": "http://example.org/"} + self._test_command(curl_command, expected_result) + + def test_get_without_scheme(self): + curl_command = "curl www.example.org" + expected_result = {"method": "GET", "url": "http://www.example.org"} + self._test_command(curl_command, expected_result) + + def test_get_basic_auth(self): + curl_command = 'curl "https://api.test.com/" -u ' \ + '"some_username:some_password"' + expected_result = { + "method": "GET", + "url": "https://api.test.com/", + "headers": [ + ( + "Authorization", + basic_auth_header("some_username", "some_password") + ) + ], + } + self._test_command(curl_command, expected_result) + + def test_get_complex(self): + curl_command = ( + "curl 'http://httpbin.org/get' -H 'Accept-Encoding: gzip, deflate'" + " -H 'Accept-Language: en-US,en;q=0.9,ru;q=0.8,es;q=0.7' -H 'Upgra" + "de-Insecure-Requests: 1' -H 'User-Agent: Mozilla/5.0 (X11; Linux " + "x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Ubuntu Chromium/62" + ".0.3202.75 Chrome/62.0.3202.75 Safari/537.36' -H 'Accept: text/ht" + "ml,application/xhtml+xml,application/xml;q=0.9,image/webp,image/a" + "png,*/*;q=0.8' -H 'Referer: http://httpbin.org/' -H 'Cookie: _gau" + "ges_unique_year=1; _gauges_unique=1; _gauges_unique_month=1; _gau" + "ges_unique_hour=1; _gauges_unique_day=1' -H 'Connection: keep-ali" + "ve' --compressed" + ) + expected_result = { + "method": "GET", + "url": "http://httpbin.org/get", + "headers": [ + ("Accept-Encoding", "gzip, deflate"), + ("Accept-Language", "en-US,en;q=0.9,ru;q=0.8,es;q=0.7"), + ("Upgrade-Insecure-Requests", "1"), + ( + "User-Agent", + "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML" + ", like Gecko) Ubuntu Chromium/62.0.3202.75 Chrome/62.0.32" + "02.75 Safari/537.36", + ), + ( + "Accept", + "text/html,application/xhtml+xml,application/xml;q=0.9,ima" + "ge/webp,image/apng,*/*;q=0.8", + ), + ("Referer", "http://httpbin.org/"), + ("Connection", "keep-alive"), + ], + "cookies": { + '_gauges_unique_year': '1', + '_gauges_unique_hour': '1', + '_gauges_unique_day': '1', + '_gauges_unique': '1', + '_gauges_unique_month': '1' + }, + } + self._test_command(curl_command, expected_result) + + def test_post(self): + curl_command = ( + "curl 'http://httpbin.org/post' -X POST -H 'Cookie: _gauges_unique" + "_year=1; _gauges_unique=1; _gauges_unique_month=1; _gauges_unique" + "_hour=1; _gauges_unique_day=1' -H 'Origin: http://httpbin.org' -H" + " 'Accept-Encoding: gzip, deflate' -H 'Accept-Language: en-US,en;q" + "=0.9,ru;q=0.8,es;q=0.7' -H 'Upgrade-Insecure-Requests: 1' -H 'Use" + "r-Agent: Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTM" + "L, like Gecko) Ubuntu Chromium/62.0.3202.75 Chrome/62.0.3202.75 S" + "afari/537.36' -H 'Content-Type: application/x-www-form-urlencoded" + "' -H 'Accept: text/html,application/xhtml+xml,application/xml;q=0" + ".9,image/webp,image/apng,*/*;q=0.8' -H 'Cache-Control: max-age=0'" + " -H 'Referer: http://httpbin.org/forms/post' -H 'Connection: keep" + "-alive' --data 'custname=John+Smith&custtel=500&custemail=jsmith%" + "40example.org&size=small&topping=cheese&topping=onion&delivery=12" + "%3A15&comments=' --compressed" + ) + expected_result = { + "method": "POST", + "url": "http://httpbin.org/post", + "body": "custname=John+Smith&custtel=500&custemail=jsmith%40exampl" + "e.org&size=small&topping=cheese&topping=onion&delivery=12" + "%3A15&comments=", + "cookies": { + '_gauges_unique_year': '1', + '_gauges_unique_hour': '1', + '_gauges_unique_day': '1', + '_gauges_unique': '1', + '_gauges_unique_month': '1' + }, + "headers": [ + ("Origin", "http://httpbin.org"), + ("Accept-Encoding", "gzip, deflate"), + ("Accept-Language", "en-US,en;q=0.9,ru;q=0.8,es;q=0.7"), + ("Upgrade-Insecure-Requests", "1"), + ( + "User-Agent", + "Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML" + ", like Gecko) Ubuntu Chromium/62.0.3202.75 Chrome/62.0.32" + "02.75 Safari/537.36", + ), + ("Content-Type", "application/x-www-form-urlencoded"), + ( + "Accept", + "text/html,application/xhtml+xml,application/xml;q=0.9,ima" + "ge/webp,image/apng,*/*;q=0.8", + ), + ("Cache-Control", "max-age=0"), + ("Referer", "http://httpbin.org/forms/post"), + ("Connection", "keep-alive"), + ], + } + self._test_command(curl_command, expected_result) + + def test_patch(self): + curl_command = ( + 'curl "https://example.com/api/fake" -u "username:password" -H "Ac' + 'cept: application/vnd.go.cd.v4+json" -H "Content-Type: applicatio' + 'n/json" -X PATCH -d \'{"hostname": "agent02.example.com", "agent' + '_config_state": "Enabled", "resources": ["Java","Linux"], "enviro' + 'nments": ["Dev"]}\'' + ) + expected_result = { + "method": "PATCH", + "url": "https://example.com/api/fake", + "headers": [ + ("Accept", "application/vnd.go.cd.v4+json"), + ("Content-Type", "application/json"), + ("Authorization", basic_auth_header("username", "password")), + ], + "body": '{"hostname": "agent02.example.com", "agent_config_state"' + ': "Enabled", "resources": ["Java","Linux"], "environments' + '": ["Dev"]}', + } + self._test_command(curl_command, expected_result) + + def test_delete(self): + curl_command = 'curl -X "DELETE" https://www.url.com/page' + expected_result = { + "method": "DELETE", "url": "https://www.url.com/page" + } + self._test_command(curl_command, expected_result) + + def test_get_silent(self): + curl_command = 'curl --silent "www.example.com"' + expected_result = {"method": "GET", "url": "http://www.example.com"} + self.assertEqual(curl_to_request_kwargs(curl_command), expected_result) + + def test_too_few_arguments_error(self): + assertRaisesRegex( + self, + ValueError, + r"too few arguments|the following arguments are required:\s*url", + lambda: curl_to_request_kwargs("curl"), + ) + + def test_ignore_unknown_options(self): + # case 1: ignore_unknown_options=True: + with warnings.catch_warnings(): # avoid warning when executing tests + warnings.simplefilter('ignore') + curl_command = 'curl --bar --baz http://www.example.com' + expected_result = \ + {"method": "GET", "url": "http://www.example.com"} + self.assertEqual(curl_to_request_kwargs(curl_command), expected_result) + + # case 2: ignore_unknown_options=False (raise exception): + assertRaisesRegex( + self, + ValueError, + "Unrecognized options:.*--bar.*--baz", + lambda: curl_to_request_kwargs( + "curl --bar --baz http://www.example.com", + ignore_unknown_options=False + ), + ) + + def test_must_start_with_curl_error(self): + self.assertRaises( + ValueError, + lambda: curl_to_request_kwargs("carl -X POST http://example.org") + ) From 9119798a5ce10aaf015d1af647c5b85e312d2386 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 5 Aug 2019 15:49:07 +0200 Subject: [PATCH 032/387] Add test coverage for contract failures involving multiple missing fields --- tests/test_contracts.py | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/tests/test_contracts.py b/tests/test_contracts.py index a06bb2cc3..a728099c0 100644 --- a/tests/test_contracts.py +++ b/tests/test_contracts.py @@ -123,6 +123,14 @@ class TestSpider(Spider): """ return {'url': response.url} + def scrapes_multiple_missing_fields(self, response): + """ returns item with no name + @url http://scrapy.org + @returns items 1 1 + @scrapes name url + """ + return {} + def parse_no_url(self, response): """ method with no url @returns items 1 1 @@ -256,6 +264,13 @@ class ContractsManagerTest(unittest.TestCase): request.callback(response) self.should_fail() + # scrapes_multiple_missing_fields + request = self.conman.from_method(spider.scrapes_multiple_missing_fields, self.results) + request.callback(response) + self.should_fail() + message = 'ContractFail: Missing fields: name, url' + assert message in self.results.failures[-1][-1] + def test_custom_contracts(self): self.conman.from_spider(CustomContractSuccessSpider(), self.results) self.should_succeed() From 3040f77468d7dd997afd2f8a35590cafaff454af Mon Sep 17 00:00:00 2001 From: Shivam Sandbhor Date: Thu, 8 Aug 2019 17:28:22 +0530 Subject: [PATCH 033/387] [MRG+1] Update project.py removed one 'hack', seems irrelevant. (#3910) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * Update project.py removed one 'hack', seems irrelevant. As mentioned by @Gallaecio in issue #3871, the 'hack' is cleared. I also double checked whether the environment variable "SCRAPY_PICKLED_SETTINGS_TO_OVERRIDE" was ever set in our codebase and it turns out we didn't set it or used it anywhere else.So I guess the 'hack' was not used in the current version. Also the name of this environment variable rather doesn't suggest it was a boolean(it is used in an 'if' condition which has perplexed me ) * Update project.py * Update project.py How about this? * Update project.py * Update project.py * Update scrapy/utils/project.py Co-Authored-By: Adrián Chaves * Update scrapy/utils/project.py Co-Authored-By: Adrián Chaves * Update project.py --- scrapy/utils/project.py | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/scrapy/utils/project.py b/scrapy/utils/project.py index 95c6a8035..1cbda141a 100644 --- a/scrapy/utils/project.py +++ b/scrapy/utils/project.py @@ -8,6 +8,7 @@ from os.path import join, dirname, abspath, isabs, exists from scrapy.utils.conf import closest_scrapy_cfg, get_config, init_env from scrapy.settings import Settings from scrapy.exceptions import NotConfigured +from scrapy.exceptions import ScrapyDeprecationWarning ENVVAR = 'SCRAPY_SETTINGS_MODULE' DATADIR_CFG_SECTION = 'datadir' @@ -70,6 +71,9 @@ def get_project_settings(): # XXX: remove this hack pickled_settings = os.environ.get("SCRAPY_PICKLED_SETTINGS_TO_OVERRIDE") if pickled_settings: + warnings.warn("Use of environment variable " + "'SCRAPY_PICKLED_SETTINGS_TO_OVERRIDE' " + "is deprecated.", ScrapyDeprecationWarning) settings.setdict(pickle.loads(pickled_settings), priority='project') # XXX: deprecate and remove this functionality From da385b56b10aaf3393cc10d13421d45c5a0edbdb Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 8 Aug 2019 12:44:23 -0300 Subject: [PATCH 034/387] Move get_openssl_version function to scrapy.utils.ssl --- scrapy/utils/ssl.py | 8 ++++++++ scrapy/utils/versions.py | 11 ++--------- 2 files changed, 10 insertions(+), 9 deletions(-) diff --git a/scrapy/utils/ssl.py b/scrapy/utils/ssl.py index e54232abd..632827471 100644 --- a/scrapy/utils/ssl.py +++ b/scrapy/utils/ssl.py @@ -1,5 +1,6 @@ # -*- coding: utf-8 -*- +import OpenSSL import OpenSSL._util as pyOpenSSLutil from scrapy.utils.python import to_native_str @@ -48,3 +49,10 @@ def get_temp_key_info(ssl_object): key_info.append(ffi_buf_to_string(pyOpenSSLutil.lib.OBJ_nid2sn(key_type))) key_info.append('%s bits' % pyOpenSSLutil.lib.EVP_PKEY_bits(temp_key)) return ', '.join(key_info) + + +def get_openssl_version(): + system_openssl = OpenSSL.SSL.SSLeay_version( + OpenSSL.SSL.SSLEAY_VERSION + ).decode('ascii', errors='replace') + return '{} ({})'.format(OpenSSL.version.__version__, system_openssl) diff --git a/scrapy/utils/versions.py b/scrapy/utils/versions.py index 3f8122154..48484b303 100644 --- a/scrapy/utils/versions.py +++ b/scrapy/utils/versions.py @@ -4,12 +4,12 @@ import sys import cryptography import cssselect import lxml.etree -import OpenSSL import parsel import twisted import w3lib import scrapy +from scrapy.utils.ssl import get_openssl_version def scrapy_components_versions(): @@ -25,14 +25,7 @@ def scrapy_components_versions(): ("w3lib", w3lib.__version__), ("Twisted", twisted.version.short()), ("Python", sys.version.replace("\n", "- ")), - ("pyOpenSSL", _get_openssl_version()), + ("pyOpenSSL", get_openssl_version()), ("cryptography", cryptography.__version__), ("Platform", platform.platform()), ] - - -def _get_openssl_version(): - openssl = OpenSSL.SSL.SSLeay_version( - OpenSSL.SSL.SSLEAY_VERSION - ).decode('ascii', errors='replace') - return '{} ({})'.format(OpenSSL.version.__version__, openssl) From fa9a9033f0a73e5c3c1b3bd54408e3bab7ea1b91 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 8 Aug 2019 22:54:51 -0300 Subject: [PATCH 035/387] Remove check for Twisted>=14.0.0 16.0.0 is currently the minimum supported version --- scrapy/core/downloader/tls.py | 158 +++++++++++++++++----------------- 1 file changed, 78 insertions(+), 80 deletions(-) diff --git a/scrapy/core/downloader/tls.py b/scrapy/core/downloader/tls.py index 74afb3f10..2e218dfb4 100644 --- a/scrapy/core/downloader/tls.py +++ b/scrapy/core/downloader/tls.py @@ -1,5 +1,8 @@ import logging + from OpenSSL import SSL +from twisted.internet.ssl import AcceptableCiphers +from twisted.internet._sslverify import ClientTLSOptions, verifyHostname, VerificationError from scrapy import twisted_version from scrapy.utils.ssl import x509name_to_string, get_temp_key_info @@ -22,95 +25,90 @@ openssl_methods = { } -if twisted_version >= (14, 0, 0): - # ClientTLSOptions requires a recent-enough version of Twisted. - # Not having ScrapyClientTLSOptions should not matter for older - # Twisted versions because it is not used in the fallback - # ScrapyClientContextFactory. +# ClientTLSOptions requires a recent-enough version of Twisted (14.0.0+) +# Not having ScrapyClientTLSOptions should not matter for older +# Twisted versions because it is not used in the fallback +# ScrapyClientContextFactory. - # taken from twisted/twisted/internet/_sslverify.py +# taken from twisted/twisted/internet/_sslverify.py - try: - # XXX: this try-except is not needed in Twisted 17.0.0+ because - # it requires pyOpenSSL 0.16+. - from OpenSSL.SSL import SSL_CB_HANDSHAKE_DONE, SSL_CB_HANDSHAKE_START - except ImportError: - SSL_CB_HANDSHAKE_START = 0x10 - SSL_CB_HANDSHAKE_DONE = 0x20 +try: + # XXX: this try-except is not needed in Twisted 17.0.0+ because + # it requires pyOpenSSL 0.16+. + from OpenSSL.SSL import SSL_CB_HANDSHAKE_DONE, SSL_CB_HANDSHAKE_START +except ImportError: + SSL_CB_HANDSHAKE_START = 0x10 + SSL_CB_HANDSHAKE_DONE = 0x20 - from twisted.internet.ssl import AcceptableCiphers - from twisted.internet._sslverify import (ClientTLSOptions, - verifyHostname, - VerificationError) - try: - # XXX: this import would fail on Debian jessie with system installed - # service_identity library, due to lack of cryptography.x509 dependency - # See https://github.com/pyca/service_identity/issues/21 - from service_identity.exceptions import CertificateError - verification_errors = (CertificateError, VerificationError) - except ImportError: - verification_errors = VerificationError +try: + # XXX: this import would fail on Debian jessie with system installed + # service_identity library, due to lack of cryptography.x509 dependency + # See https://github.com/pyca/service_identity/issues/21 + from service_identity.exceptions import CertificateError + verification_errors = (CertificateError, VerificationError) +except ImportError: + verification_errors = VerificationError - if twisted_version < (17, 0, 0): - from twisted.internet._sslverify import _maybeSetHostNameIndication - set_tlsext_host_name = _maybeSetHostNameIndication - else: - def set_tlsext_host_name(connection, hostNameBytes): - connection.set_tlsext_host_name(hostNameBytes) +if twisted_version < (17, 0, 0): + from twisted.internet._sslverify import _maybeSetHostNameIndication + set_tlsext_host_name = _maybeSetHostNameIndication +else: + def set_tlsext_host_name(connection, hostNameBytes): + connection.set_tlsext_host_name(hostNameBytes) - class ScrapyClientTLSOptions(ClientTLSOptions): - """ - SSL Client connection creator ignoring certificate verification errors - (for genuinely invalid certificates or bugs in verification code). +class ScrapyClientTLSOptions(ClientTLSOptions): + """ + SSL Client connection creator ignoring certificate verification errors + (for genuinely invalid certificates or bugs in verification code). - Same as Twisted's private _sslverify.ClientTLSOptions, - except that VerificationError, CertificateError and ValueError - exceptions are caught, so that the connection is not closed, only - logging warnings. Also, HTTPS connection parameters logging is added. - """ + Same as Twisted's private _sslverify.ClientTLSOptions, + except that VerificationError, CertificateError and ValueError + exceptions are caught, so that the connection is not closed, only + logging warnings. Also, HTTPS connection parameters logging is added. + """ - def __init__(self, hostname, ctx, verbose_logging=False): - super(ScrapyClientTLSOptions, self).__init__(hostname, ctx) - self.verbose_logging = verbose_logging + def __init__(self, hostname, ctx, verbose_logging=False): + super(ScrapyClientTLSOptions, self).__init__(hostname, ctx) + self.verbose_logging = verbose_logging - def _identityVerifyingInfoCallback(self, connection, where, ret): - if where & SSL_CB_HANDSHAKE_START: - set_tlsext_host_name(connection, self._hostnameBytes) - elif where & SSL_CB_HANDSHAKE_DONE: - if self.verbose_logging: - if hasattr(connection, 'get_cipher_name'): # requires pyOPenSSL 0.15 - if hasattr(connection, 'get_protocol_version_name'): # requires pyOPenSSL 16.0.0 - logger.debug('SSL connection to %s using protocol %s, cipher %s', - self._hostnameASCII, - connection.get_protocol_version_name(), - connection.get_cipher_name(), - ) - else: - logger.debug('SSL connection to %s using cipher %s', - self._hostnameASCII, - connection.get_cipher_name(), - ) - server_cert = connection.get_peer_certificate() - logger.debug('SSL connection certificate: issuer "%s", subject "%s"', - x509name_to_string(server_cert.get_issuer()), - x509name_to_string(server_cert.get_subject()), - ) - key_info = get_temp_key_info(connection._ssl) - if key_info: - logger.debug('SSL temp key: %s', key_info) + def _identityVerifyingInfoCallback(self, connection, where, ret): + if where & SSL_CB_HANDSHAKE_START: + set_tlsext_host_name(connection, self._hostnameBytes) + elif where & SSL_CB_HANDSHAKE_DONE: + if self.verbose_logging: + if hasattr(connection, 'get_cipher_name'): # requires pyOPenSSL 0.15 + if hasattr(connection, 'get_protocol_version_name'): # requires pyOPenSSL 16.0.0 + logger.debug('SSL connection to %s using protocol %s, cipher %s', + self._hostnameASCII, + connection.get_protocol_version_name(), + connection.get_cipher_name(), + ) + else: + logger.debug('SSL connection to %s using cipher %s', + self._hostnameASCII, + connection.get_cipher_name(), + ) + server_cert = connection.get_peer_certificate() + logger.debug('SSL connection certificate: issuer "%s", subject "%s"', + x509name_to_string(server_cert.get_issuer()), + x509name_to_string(server_cert.get_subject()), + ) + key_info = get_temp_key_info(connection._ssl) + if key_info: + logger.debug('SSL temp key: %s', key_info) - try: - verifyHostname(connection, self._hostnameASCII) - except verification_errors as e: - logger.warning( - 'Remote certificate is not valid for hostname "{}"; {}'.format( - self._hostnameASCII, e)) + try: + verifyHostname(connection, self._hostnameASCII) + except verification_errors as e: + logger.warning( + 'Remote certificate is not valid for hostname "{}"; {}'.format( + self._hostnameASCII, e)) - except ValueError as e: - logger.warning( - 'Ignoring error while verifying certificate ' - 'from host "{}" (exception: {})'.format( - self._hostnameASCII, repr(e))) + except ValueError as e: + logger.warning( + 'Ignoring error while verifying certificate ' + 'from host "{}" (exception: {})'.format( + self._hostnameASCII, repr(e))) - DEFAULT_CIPHERS = AcceptableCiphers.fromOpenSSLCipherString('DEFAULT') +DEFAULT_CIPHERS = AcceptableCiphers.fromOpenSSLCipherString('DEFAULT') From a940a80f5886c631425c4790ac92928926b1cc92 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 8 Aug 2019 23:05:52 -0300 Subject: [PATCH 036/387] Remove check for pyOpenSSL>=0.16 16.2.0 is currently the minimum supported version --- scrapy/core/downloader/tls.py | 14 ++------------ 1 file changed, 2 insertions(+), 12 deletions(-) diff --git a/scrapy/core/downloader/tls.py b/scrapy/core/downloader/tls.py index 2e218dfb4..d6b7967da 100644 --- a/scrapy/core/downloader/tls.py +++ b/scrapy/core/downloader/tls.py @@ -30,16 +30,6 @@ openssl_methods = { # Twisted versions because it is not used in the fallback # ScrapyClientContextFactory. -# taken from twisted/twisted/internet/_sslverify.py - -try: - # XXX: this try-except is not needed in Twisted 17.0.0+ because - # it requires pyOpenSSL 0.16+. - from OpenSSL.SSL import SSL_CB_HANDSHAKE_DONE, SSL_CB_HANDSHAKE_START -except ImportError: - SSL_CB_HANDSHAKE_START = 0x10 - SSL_CB_HANDSHAKE_DONE = 0x20 - try: # XXX: this import would fail on Debian jessie with system installed # service_identity library, due to lack of cryptography.x509 dependency @@ -73,9 +63,9 @@ class ScrapyClientTLSOptions(ClientTLSOptions): self.verbose_logging = verbose_logging def _identityVerifyingInfoCallback(self, connection, where, ret): - if where & SSL_CB_HANDSHAKE_START: + if where & SSL.SSL_CB_HANDSHAKE_START: set_tlsext_host_name(connection, self._hostnameBytes) - elif where & SSL_CB_HANDSHAKE_DONE: + elif where & SSL.SSL_CB_HANDSHAKE_DONE: if self.verbose_logging: if hasattr(connection, 'get_cipher_name'): # requires pyOPenSSL 0.15 if hasattr(connection, 'get_protocol_version_name'): # requires pyOPenSSL 16.0.0 From 3164543ed1659a11cea1f6a04d8545f9cfebe367 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 8 Aug 2019 23:15:03 -0300 Subject: [PATCH 037/387] Remove fallback ScrapyClientContextFactory class (used in Twisted < 14.0.0) 16.0.0 is currently the minimum supported version --- scrapy/core/downloader/contextfactory.py | 161 ++++++++++------------- scrapy/core/downloader/tls.py | 7 +- 2 files changed, 69 insertions(+), 99 deletions(-) diff --git a/scrapy/core/downloader/contextfactory.py b/scrapy/core/downloader/contextfactory.py index 5ac20c0bb..127a246f5 100644 --- a/scrapy/core/downloader/contextfactory.py +++ b/scrapy/core/downloader/contextfactory.py @@ -1,112 +1,85 @@ from OpenSSL import SSL -from twisted.internet.ssl import ClientContextFactory +from twisted.internet.ssl import optionsForClientTLS, CertificateOptions, platformTrust +from twisted.web.client import BrowserLikePolicyForHTTPS +from twisted.web.iweb import IPolicyForHTTPS +from zope.interface.declarations import implementer -from scrapy import twisted_version - -if twisted_version >= (14, 0, 0): - - from zope.interface.declarations import implementer - - from twisted.internet.ssl import (optionsForClientTLS, - CertificateOptions, - platformTrust) - from twisted.web.client import BrowserLikePolicyForHTTPS - from twisted.web.iweb import IPolicyForHTTPS - - from scrapy.core.downloader.tls import ScrapyClientTLSOptions, DEFAULT_CIPHERS +from scrapy.core.downloader.tls import ScrapyClientTLSOptions, DEFAULT_CIPHERS - @implementer(IPolicyForHTTPS) - class ScrapyClientContextFactory(BrowserLikePolicyForHTTPS): - """ - Non-peer-certificate verifying HTTPS context factory +@implementer(IPolicyForHTTPS) +class ScrapyClientContextFactory(BrowserLikePolicyForHTTPS): + """ + Non-peer-certificate verifying HTTPS context factory - Default OpenSSL method is TLS_METHOD (also called SSLv23_METHOD) - which allows TLS protocol negotiation + Default OpenSSL method is TLS_METHOD (also called SSLv23_METHOD) + which allows TLS protocol negotiation - 'A TLS/SSL connection established with [this method] may - understand the SSLv3, TLSv1, TLSv1.1 and TLSv1.2 protocols.' - """ + 'A TLS/SSL connection established with [this method] may + understand the SSLv3, TLSv1, TLSv1.1 and TLSv1.2 protocols.' + """ - def __init__(self, method=SSL.SSLv23_METHOD, tls_verbose_logging=False, *args, **kwargs): - super(ScrapyClientContextFactory, self).__init__(*args, **kwargs) - self._ssl_method = method - self.tls_verbose_logging = tls_verbose_logging + def __init__(self, method=SSL.SSLv23_METHOD, tls_verbose_logging=False, *args, **kwargs): + super(ScrapyClientContextFactory, self).__init__(*args, **kwargs) + self._ssl_method = method + self.tls_verbose_logging = tls_verbose_logging - @classmethod - def from_settings(cls, settings, method=SSL.SSLv23_METHOD, *args, **kwargs): - tls_verbose_logging = settings.getbool('DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING') - return cls(method=method, tls_verbose_logging=tls_verbose_logging, *args, **kwargs) + @classmethod + def from_settings(cls, settings, method=SSL.SSLv23_METHOD, *args, **kwargs): + tls_verbose_logging = settings.getbool('DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING') + return cls(method=method, tls_verbose_logging=tls_verbose_logging, *args, **kwargs) - def getCertificateOptions(self): - # setting verify=True will require you to provide CAs - # to verify against; in other words: it's not that simple + def getCertificateOptions(self): + # setting verify=True will require you to provide CAs + # to verify against; in other words: it's not that simple - # backward-compatible SSL/TLS method: - # - # * this will respect `method` attribute in often recommended - # `ScrapyClientContextFactory` subclass - # (https://github.com/scrapy/scrapy/issues/1429#issuecomment-131782133) - # - # * getattr() for `_ssl_method` attribute for context factories - # not calling super(..., self).__init__ - return CertificateOptions(verify=False, - method=getattr(self, 'method', - getattr(self, '_ssl_method', None)), - fixBrokenPeers=True, - acceptableCiphers=DEFAULT_CIPHERS) + # backward-compatible SSL/TLS method: + # + # * this will respect `method` attribute in often recommended + # `ScrapyClientContextFactory` subclass + # (https://github.com/scrapy/scrapy/issues/1429#issuecomment-131782133) + # + # * getattr() for `_ssl_method` attribute for context factories + # not calling super(..., self).__init__ + return CertificateOptions(verify=False, + method=getattr(self, 'method', + getattr(self, '_ssl_method', None)), + fixBrokenPeers=True, + acceptableCiphers=DEFAULT_CIPHERS) - # kept for old-style HTTP/1.0 downloader context twisted calls, - # e.g. connectSSL() - def getContext(self, hostname=None, port=None): - return self.getCertificateOptions().getContext() + # kept for old-style HTTP/1.0 downloader context twisted calls, + # e.g. connectSSL() + def getContext(self, hostname=None, port=None): + return self.getCertificateOptions().getContext() - def creatorForNetloc(self, hostname, port): - return ScrapyClientTLSOptions(hostname.decode("ascii"), self.getContext(), - verbose_logging=self.tls_verbose_logging) + def creatorForNetloc(self, hostname, port): + return ScrapyClientTLSOptions(hostname.decode("ascii"), self.getContext(), + verbose_logging=self.tls_verbose_logging) - @implementer(IPolicyForHTTPS) - class BrowserLikeContextFactory(ScrapyClientContextFactory): - """ - Twisted-recommended context factory for web clients. +@implementer(IPolicyForHTTPS) +class BrowserLikeContextFactory(ScrapyClientContextFactory): + """ + Twisted-recommended context factory for web clients. - Quoting https://twistedmatrix.com/documents/current/api/twisted.web.client.Agent.html: - "The default is to use a BrowserLikePolicyForHTTPS, - so unless you have special requirements you can leave this as-is." + Quoting https://twistedmatrix.com/documents/current/api/twisted.web.client.Agent.html: + "The default is to use a BrowserLikePolicyForHTTPS, + so unless you have special requirements you can leave this as-is." - creatorForNetloc() is the same as BrowserLikePolicyForHTTPS - except this context factory allows setting the TLS/SSL method to use. + creatorForNetloc() is the same as BrowserLikePolicyForHTTPS + except this context factory allows setting the TLS/SSL method to use. - Default OpenSSL method is TLS_METHOD (also called SSLv23_METHOD) - which allows TLS protocol negotiation. - """ - def creatorForNetloc(self, hostname, port): + Default OpenSSL method is TLS_METHOD (also called SSLv23_METHOD) + which allows TLS protocol negotiation. + """ + def creatorForNetloc(self, hostname, port): - # trustRoot set to platformTrust() will use the platform's root CAs. - # - # This means that a website like https://www.cacert.org will be rejected - # by default, since CAcert.org CA certificate is seldom shipped. - return optionsForClientTLS(hostname.decode("ascii"), - trustRoot=platformTrust(), - extraCertificateOptions={ - 'method': self._ssl_method, - }) - -else: - - class ScrapyClientContextFactory(ClientContextFactory): - "A SSL context factory which is more permissive against SSL bugs." - # see https://github.com/scrapy/scrapy/issues/82 - # and https://github.com/scrapy/scrapy/issues/26 - # and https://github.com/scrapy/scrapy/issues/981 - - def __init__(self, method=SSL.SSLv23_METHOD): - self.method = method - - def getContext(self, hostname=None, port=None): - ctx = ClientContextFactory.getContext(self) - # Enable all workarounds to SSL bugs as documented by - # https://www.openssl.org/docs/manmaster/man3/SSL_CTX_set_options.html - ctx.set_options(SSL.OP_ALL) - return ctx + # trustRoot set to platformTrust() will use the platform's root CAs. + # + # This means that a website like https://www.cacert.org will be rejected + # by default, since CAcert.org CA certificate is seldom shipped. + return optionsForClientTLS(hostname.decode("ascii"), + trustRoot=platformTrust(), + extraCertificateOptions={ + 'method': self._ssl_method, + }) diff --git a/scrapy/core/downloader/tls.py b/scrapy/core/downloader/tls.py index d6b7967da..995f8cbba 100644 --- a/scrapy/core/downloader/tls.py +++ b/scrapy/core/downloader/tls.py @@ -10,12 +10,14 @@ from scrapy.utils.ssl import x509name_to_string, get_temp_key_info logger = logging.getLogger(__name__) + METHOD_SSLv3 = 'SSLv3' METHOD_TLS = 'TLS' METHOD_TLSv10 = 'TLSv1.0' METHOD_TLSv11 = 'TLSv1.1' METHOD_TLSv12 = 'TLSv1.2' + openssl_methods = { METHOD_TLS: SSL.SSLv23_METHOD, # protocol negotiation (recommended) METHOD_SSLv3: SSL.SSLv3_METHOD, # SSL 3 (NOT recommended) @@ -25,11 +27,6 @@ openssl_methods = { } -# ClientTLSOptions requires a recent-enough version of Twisted (14.0.0+) -# Not having ScrapyClientTLSOptions should not matter for older -# Twisted versions because it is not used in the fallback -# ScrapyClientContextFactory. - try: # XXX: this import would fail on Debian jessie with system installed # service_identity library, due to lack of cryptography.x509 dependency From b404941e0de86ddf2e108c42a4847f0341d53d14 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 8 Aug 2019 23:18:47 -0300 Subject: [PATCH 038/387] Remove import check for service_identity service_identity.exceptions.CertificateError is available in the current minimum version (16.0.0) --- scrapy/core/downloader/tls.py | 14 +++----------- 1 file changed, 3 insertions(+), 11 deletions(-) diff --git a/scrapy/core/downloader/tls.py b/scrapy/core/downloader/tls.py index 995f8cbba..1c3d94b29 100644 --- a/scrapy/core/downloader/tls.py +++ b/scrapy/core/downloader/tls.py @@ -1,8 +1,9 @@ import logging from OpenSSL import SSL -from twisted.internet.ssl import AcceptableCiphers +from service_identity.exceptions import CertificateError from twisted.internet._sslverify import ClientTLSOptions, verifyHostname, VerificationError +from twisted.internet.ssl import AcceptableCiphers from scrapy import twisted_version from scrapy.utils.ssl import x509name_to_string, get_temp_key_info @@ -27,15 +28,6 @@ openssl_methods = { } -try: - # XXX: this import would fail on Debian jessie with system installed - # service_identity library, due to lack of cryptography.x509 dependency - # See https://github.com/pyca/service_identity/issues/21 - from service_identity.exceptions import CertificateError - verification_errors = (CertificateError, VerificationError) -except ImportError: - verification_errors = VerificationError - if twisted_version < (17, 0, 0): from twisted.internet._sslverify import _maybeSetHostNameIndication set_tlsext_host_name = _maybeSetHostNameIndication @@ -87,7 +79,7 @@ class ScrapyClientTLSOptions(ClientTLSOptions): try: verifyHostname(connection, self._hostnameASCII) - except verification_errors as e: + except (CertificateError, VerificationError) as e: logger.warning( 'Remote certificate is not valid for hostname "{}"; {}'.format( self._hostnameASCII, e)) From d92f1b18580940c27f303f1dfcf47bb66e508ce4 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 8 Aug 2019 23:53:35 -0300 Subject: [PATCH 039/387] Simplify import + assignment --- scrapy/core/downloader/tls.py | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/scrapy/core/downloader/tls.py b/scrapy/core/downloader/tls.py index 1c3d94b29..4ed482058 100644 --- a/scrapy/core/downloader/tls.py +++ b/scrapy/core/downloader/tls.py @@ -29,8 +29,7 @@ openssl_methods = { if twisted_version < (17, 0, 0): - from twisted.internet._sslverify import _maybeSetHostNameIndication - set_tlsext_host_name = _maybeSetHostNameIndication + from twisted.internet._sslverify import _maybeSetHostNameIndication as set_tlsext_host_name else: def set_tlsext_host_name(connection, hostNameBytes): connection.set_tlsext_host_name(hostNameBytes) From e17c9a48fdd41e838c235be75038fab1ab9bb8e2 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 8 Aug 2019 23:59:17 -0300 Subject: [PATCH 040/387] Remove check for Twisted>=15.0.0 16.0.0 is currently the minimum supported version --- scrapy/core/downloader/handlers/http11.py | 28 +++++++---------------- 1 file changed, 8 insertions(+), 20 deletions(-) diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index 2ccab2614..8b4ae6b24 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -27,7 +27,7 @@ from scrapy.core.downloader.webclient import _parse from scrapy.core.downloader.tls import openssl_methods from scrapy.utils.misc import load_object, create_instance from scrapy.utils.python import to_bytes, to_unicode -from scrapy import twisted_version + logger = logging.getLogger(__name__) @@ -211,21 +211,14 @@ class TunnelingAgent(Agent): self._proxyConf = proxyConf self._contextFactory = contextFactory - if twisted_version >= (15, 0, 0): - def _getEndpoint(self, uri): - return TunnelingTCP4ClientEndpoint( - self._reactor, uri.host, uri.port, self._proxyConf, - self._contextFactory, self._endpointFactory._connectTimeout, - self._endpointFactory._bindAddress) - else: - def _getEndpoint(self, scheme, host, port): - return TunnelingTCP4ClientEndpoint( - self._reactor, host, port, self._proxyConf, - self._contextFactory, self._connectTimeout, - self._bindAddress) + def _getEndpoint(self, uri): + return TunnelingTCP4ClientEndpoint( + self._reactor, uri.host, uri.port, self._proxyConf, + self._contextFactory, self._endpointFactory._connectTimeout, + self._endpointFactory._bindAddress) def _requestWithEndpoint(self, key, endpoint, method, parsedURI, - headers, bodyProducer, requestPath): + headers, bodyProducer, requestPath): # proxy host and port are required for HTTP pool `key` # otherwise, same remote host connection request could reuse # a cached tunneled connection to a different proxy @@ -250,12 +243,7 @@ class ScrapyProxyAgent(Agent): """ # Cache *all* connections under the same key, since we are only # connecting to a single destination, the proxy: - if twisted_version >= (15, 0, 0): - proxyEndpoint = self._getEndpoint(self._proxyURI) - else: - proxyEndpoint = self._getEndpoint(self._proxyURI.scheme, - self._proxyURI.host, - self._proxyURI.port) + proxyEndpoint = self._getEndpoint(self._proxyURI) key = ("http-proxy", self._proxyURI.host, self._proxyURI.port) return self._requestWithEndpoint(key, proxyEndpoint, method, URI.fromBytes(uri), headers, From d3737d869b1b49497f53376283b02ac604543e2c Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Fri, 9 Aug 2019 00:21:43 -0300 Subject: [PATCH 041/387] Remove check for Twisted>=14.0 --- scrapy/core/downloader/handlers/http11.py | 10 ++-------- 1 file changed, 2 insertions(+), 8 deletions(-) diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index 8b4ae6b24..f0ed1a4af 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -141,14 +141,8 @@ class TunnelingTCP4ClientEndpoint(TCP4ClientEndpoint): self._protocol.dataReceived = self._protocolDataReceived respm = TunnelingTCP4ClientEndpoint._responseMatcher.match(self._connectBuffer) if respm and int(respm.group('status')) == 200: - try: - # this sets proper Server Name Indication extension - # but is only available for Twisted>=14.0 - sslOptions = self._contextFactory.creatorForNetloc( - self._tunneledHost, self._tunneledPort) - except AttributeError: - # fall back to non-SNI SSL context factory - sslOptions = self._contextFactory + # set proper Server Name Indication extension + sslOptions = self._contextFactory.creatorForNetloc(self._tunneledHost, self._tunneledPort) self._protocol.transport.startTLS(sslOptions, self._protocolFactory) self._tunnelReadyDeferred.callback(self._protocol) From d5dcc5eaef80ef383dff90f19349c0e06f1836a6 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Fri, 9 Aug 2019 00:30:58 -0300 Subject: [PATCH 042/387] Import twisted.web.client.URI directly --- scrapy/core/downloader/handlers/http11.py | 8 ++------ 1 file changed, 2 insertions(+), 6 deletions(-) diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index f0ed1a4af..9da20e032 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -13,12 +13,8 @@ from twisted.web.http_headers import Headers as TxHeaders from twisted.web.iweb import IBodyProducer, UNKNOWN_LENGTH from twisted.internet.error import TimeoutError from twisted.web.http import _DataLoss, PotentialDataLoss -from twisted.web.client import Agent, ProxyAgent, ResponseDone, \ - HTTPConnectionPool, ResponseFailed -try: - from twisted.web.client import URI -except ImportError: - from twisted.web.client import _URI as URI +from twisted.web.client import (Agent, ProxyAgent, ResponseDone, + HTTPConnectionPool, ResponseFailed, URI) from twisted.internet.endpoints import TCP4ClientEndpoint from scrapy.http import Headers From 26fb28b20f661e6d7cfd02c904f4b1dc2a79e1dc Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Fri, 9 Aug 2019 00:49:46 -0300 Subject: [PATCH 043/387] PEP8-ify HTTP/1.1 downloader handler Signed-off-by: Eugenio Lacuesta --- scrapy/core/downloader/handlers/http11.py | 150 +++++++++++++--------- 1 file changed, 91 insertions(+), 59 deletions(-) diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index 9da20e032..e72052afc 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -13,8 +13,7 @@ from twisted.web.http_headers import Headers as TxHeaders from twisted.web.iweb import IBodyProducer, UNKNOWN_LENGTH from twisted.internet.error import TimeoutError from twisted.web.http import _DataLoss, PotentialDataLoss -from twisted.web.client import (Agent, ProxyAgent, ResponseDone, - HTTPConnectionPool, ResponseFailed, URI) +from twisted.web.client import Agent, ResponseDone, HTTPConnectionPool, ResponseFailed, URI from twisted.internet.endpoints import TCP4ClientEndpoint from scrapy.http import Headers @@ -40,11 +39,19 @@ class HTTP11DownloadHandler(object): self._contextFactoryClass = load_object(settings['DOWNLOADER_CLIENTCONTEXTFACTORY']) # try method-aware context factory try: - self._contextFactory = create_instance(self._contextFactoryClass, settings=settings, crawler=None, - method=self._sslMethod) + self._contextFactory = create_instance( + self._contextFactoryClass, + settings=settings, + crawler=None, + method=self._sslMethod, + ) except TypeError: # use context factory defaults - self._contextFactory = create_instance(self._contextFactoryClass, settings=settings, crawler=None) + self._contextFactory = create_instance( + self._contextFactoryClass, + settings=settings, + crawler=None, + ) msg = """ '%s' does not accept `method` argument (type OpenSSL.SSL method,\ e.g. OpenSSL.SSL.SSLv23_METHOD) and/or `tls_verbose_logging` argument.\ @@ -58,10 +65,13 @@ class HTTP11DownloadHandler(object): def download_request(self, request, spider): """Return a deferred for the HTTP download""" - agent = ScrapyAgent(contextFactory=self._contextFactory, pool=self._pool, + agent = ScrapyAgent( + contextFactory=self._contextFactory, + pool=self._pool, maxsize=getattr(spider, 'download_maxsize', self._default_maxsize), warnsize=getattr(spider, 'download_warnsize', self._default_warnsize), - fail_on_dataloss=self._fail_on_dataloss) + fail_on_dataloss=self._fail_on_dataloss, + ) return agent.download_request(request) def close(self): @@ -100,11 +110,9 @@ class TunnelingTCP4ClientEndpoint(TCP4ClientEndpoint): _responseMatcher = re.compile(br'HTTP/1\.. (?P\d{3})(?P.{,32})') - def __init__(self, reactor, host, port, proxyConf, contextFactory, - timeout=30, bindAddress=None): + def __init__(self, reactor, host, port, proxyConf, contextFactory, timeout=30, bindAddress=None): proxyHost, proxyPort, self._proxyAuthHeader = proxyConf - super(TunnelingTCP4ClientEndpoint, self).__init__(reactor, proxyHost, - proxyPort, timeout, bindAddress) + super(TunnelingTCP4ClientEndpoint, self).__init__(reactor, proxyHost, proxyPort, timeout, bindAddress) self._tunnelReadyDeferred = defer.Deferred() self._tunneledHost = host self._tunneledPort = port @@ -113,8 +121,7 @@ class TunnelingTCP4ClientEndpoint(TCP4ClientEndpoint): def requestTunnel(self, protocol): """Asks the proxy to open a tunnel.""" - tunnelReq = tunnel_request_data(self._tunneledHost, self._tunneledPort, - self._proxyAuthHeader) + tunnelReq = tunnel_request_data(self._tunneledHost, self._tunneledPort, self._proxyAuthHeader) protocol.transport.write(tunnelReq) self._protocolDataReceived = protocol.dataReceived protocol.dataReceived = self.processProxyResponse @@ -139,8 +146,7 @@ class TunnelingTCP4ClientEndpoint(TCP4ClientEndpoint): if respm and int(respm.group('status')) == 200: # set proper Server Name Indication extension sslOptions = self._contextFactory.creatorForNetloc(self._tunneledHost, self._tunneledPort) - self._protocol.transport.startTLS(sslOptions, - self._protocolFactory) + self._protocol.transport.startTLS(sslOptions, self._protocolFactory) self._tunnelReadyDeferred.callback(self._protocol) else: if respm: @@ -158,8 +164,7 @@ class TunnelingTCP4ClientEndpoint(TCP4ClientEndpoint): def connect(self, protocolFactory): self._protocolFactory = protocolFactory - connectDeferred = super(TunnelingTCP4ClientEndpoint, - self).connect(protocolFactory) + connectDeferred = super(TunnelingTCP4ClientEndpoint, self).connect(protocolFactory) connectDeferred.addCallback(self.requestTunnel) connectDeferred.addErrback(self.connectFailed) return self._tunnelReadyDeferred @@ -196,35 +201,46 @@ class TunnelingAgent(Agent): def __init__(self, reactor, proxyConf, contextFactory=None, connectTimeout=None, bindAddress=None, pool=None): - super(TunnelingAgent, self).__init__(reactor, contextFactory, - connectTimeout, bindAddress, pool) + super(TunnelingAgent, self).__init__(reactor, contextFactory, connectTimeout, bindAddress, pool) self._proxyConf = proxyConf self._contextFactory = contextFactory def _getEndpoint(self, uri): return TunnelingTCP4ClientEndpoint( - self._reactor, uri.host, uri.port, self._proxyConf, - self._contextFactory, self._endpointFactory._connectTimeout, - self._endpointFactory._bindAddress) + reactor=self._reactor, + host=uri.host, + port=uri.port, + proxyConf=self._proxyConf, + contextFactory=self._contextFactory, + timeout=self._endpointFactory._connectTimeout, + bindAddress=self._endpointFactory._bindAddress, + ) - def _requestWithEndpoint(self, key, endpoint, method, parsedURI, - headers, bodyProducer, requestPath): + def _requestWithEndpoint(self, key, endpoint, method, parsedURI, headers, bodyProducer, requestPath): # proxy host and port are required for HTTP pool `key` # otherwise, same remote host connection request could reuse # a cached tunneled connection to a different proxy key = key + self._proxyConf - return super(TunnelingAgent, self)._requestWithEndpoint(key, endpoint, method, parsedURI, - headers, bodyProducer, requestPath) + return super(TunnelingAgent, self)._requestWithEndpoint( + key=key, + endpoint=endpoint, + method=method, + parsedURI=parsedURI, + headers=headers, + bodyProducer=bodyProducer, + requestPath=requestPath, + ) class ScrapyProxyAgent(Agent): - def __init__(self, reactor, proxyURI, - connectTimeout=None, bindAddress=None, pool=None): - super(ScrapyProxyAgent, self).__init__(reactor, - connectTimeout=connectTimeout, - bindAddress=bindAddress, - pool=pool) + def __init__(self, reactor, proxyURI, connectTimeout=None, bindAddress=None, pool=None): + super(ScrapyProxyAgent, self).__init__( + reactor=reactor, + connectTimeout=connectTimeout, + bindAddress=bindAddress, + pool=pool, + ) self._proxyURI = URI.fromBytes(proxyURI) def request(self, method, uri, headers=None, bodyProducer=None): @@ -233,11 +249,15 @@ class ScrapyProxyAgent(Agent): """ # Cache *all* connections under the same key, since we are only # connecting to a single destination, the proxy: - proxyEndpoint = self._getEndpoint(self._proxyURI) - key = ("http-proxy", self._proxyURI.host, self._proxyURI.port) - return self._requestWithEndpoint(key, proxyEndpoint, method, - URI.fromBytes(uri), headers, - bodyProducer, uri) + return self._requestWithEndpoint( + key=("http-proxy", self._proxyURI.host, self._proxyURI.port), + endpoint=self._getEndpoint(self._proxyURI), + method=method, + parsedURI=URI.fromBytes(uri), + headers=headers, + bodyProducer=bodyProducer, + requestPath=uri, + ) class ScrapyAgent(object): @@ -265,18 +285,33 @@ class ScrapyAgent(object): scheme = _parse(request.url)[0] proxyHost = to_unicode(proxyHost) omitConnectTunnel = b'noconnect' in proxyParams - if scheme == b'https' and not omitConnectTunnel: - proxyConf = (proxyHost, proxyPort, - request.headers.get(b'Proxy-Authorization', None)) - return self._TunnelingAgent(reactor, proxyConf, - contextFactory=self._contextFactory, connectTimeout=timeout, - bindAddress=bindaddress, pool=self._pool) + if scheme == b'https' and not omitConnectTunnel: + proxyAuth = request.headers.get(b'Proxy-Authorization', None) + proxyConf = (proxyHost, proxyPort, proxyAuth) + return self._TunnelingAgent( + reactor=reactor, + proxyConf=proxyConf, + contextFactory=self._contextFactory, + connectTimeout=timeout, + bindAddress=bindaddress, + pool=self._pool, + ) else: - return self._ProxyAgent(reactor, proxyURI=to_bytes(proxy, encoding='ascii'), - connectTimeout=timeout, bindAddress=bindaddress, pool=self._pool) + return self._ProxyAgent( + reactor=reactor, + proxyURI=to_bytes(proxy, encoding='ascii'), + connectTimeout=timeout, + bindAddress=bindaddress, + pool=self._pool, + ) - return self._Agent(reactor, contextFactory=self._contextFactory, - connectTimeout=timeout, bindAddress=bindaddress, pool=self._pool) + return self._Agent( + reactor=reactor, + contextFactory=self._contextFactory, + connectTimeout=timeout, + bindAddress=bindaddress, + pool=self._pool, + ) def download_request(self, request): timeout = request.meta.get('download_timeout') or self._connectTimeout @@ -307,8 +342,7 @@ class ScrapyAgent(object): else: bodyproducer = None start_time = time() - d = agent.request( - method, to_bytes(url, encoding='ascii'), headers, bodyproducer) + d = agent.request(method, to_bytes(url, encoding='ascii'), headers, bodyproducer) # set download latency d.addCallback(self._cb_latency, request, start_time) # response body is ready to be consumed @@ -363,8 +397,9 @@ class ScrapyAgent(object): txresponse._transport._producer.abortConnection() d = defer.Deferred(_cancel) - txresponse.deliverBody(_ResponseReader( - d, txresponse, request, maxsize, warnsize, fail_on_dataloss)) + txresponse.deliverBody( + _ResponseReader(d, txresponse, request, maxsize, warnsize, fail_on_dataloss) + ) # save response for timeouts self._txresponse = txresponse @@ -399,22 +434,20 @@ class _RequestBodyProducer(object): class _ResponseReader(protocol.Protocol): - def __init__(self, finished, txresponse, request, maxsize, warnsize, - fail_on_dataloss): + def __init__(self, finished, txresponse, request, maxsize, warnsize, fail_on_dataloss): self._finished = finished self._txresponse = txresponse self._request = request self._bodybuf = BytesIO() - self._maxsize = maxsize - self._warnsize = warnsize + self._maxsize = maxsize + self._warnsize = warnsize self._fail_on_dataloss = fail_on_dataloss self._fail_on_dataloss_warned = False self._reached_warnsize = False self._bytes_received = 0 def dataReceived(self, bodyBytes): - # This maybe called several times after cancel was called with buffered - # data. + # This maybe called several times after cancel was called with buffered data. if self._finished.called: return @@ -427,8 +460,7 @@ class _ResponseReader(protocol.Protocol): {'bytes': self._bytes_received, 'maxsize': self._maxsize, 'request': self._request}) - # Clear buffer earlier to avoid keeping data in memory for a long - # time. + # Clear buffer earlier to avoid keeping data in memory for a long time. self._bodybuf.truncate(0) self._finished.cancel() From 3384db92b4fb2bce66d01ddf5365478a695caad6 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 27 Sep 2018 19:51:27 +0500 Subject: [PATCH 044/387] Add support for setting SSL ciphers. --- scrapy/core/downloader/contextfactory.py | 13 +++++++++---- scrapy/settings/default_settings.py | 1 + 2 files changed, 10 insertions(+), 4 deletions(-) diff --git a/scrapy/core/downloader/contextfactory.py b/scrapy/core/downloader/contextfactory.py index 127a246f5..89d2776ae 100644 --- a/scrapy/core/downloader/contextfactory.py +++ b/scrapy/core/downloader/contextfactory.py @@ -1,5 +1,5 @@ from OpenSSL import SSL -from twisted.internet.ssl import optionsForClientTLS, CertificateOptions, platformTrust +from twisted.internet.ssl import optionsForClientTLS, CertificateOptions, platformTrust, AcceptableCiphers from twisted.web.client import BrowserLikePolicyForHTTPS from twisted.web.iweb import IPolicyForHTTPS from zope.interface.declarations import implementer @@ -19,15 +19,20 @@ class ScrapyClientContextFactory(BrowserLikePolicyForHTTPS): understand the SSLv3, TLSv1, TLSv1.1 and TLSv1.2 protocols.' """ - def __init__(self, method=SSL.SSLv23_METHOD, tls_verbose_logging=False, *args, **kwargs): + def __init__(self, method=SSL.SSLv23_METHOD, tls_verbose_logging=False, tls_ciphers=None, *args, **kwargs): super(ScrapyClientContextFactory, self).__init__(*args, **kwargs) self._ssl_method = method self.tls_verbose_logging = tls_verbose_logging + if tls_ciphers: + self.tls_ciphers = AcceptableCiphers.fromOpenSSLCipherString(tls_ciphers) + else: + self.tls_ciphers = DEFAULT_CIPHERS @classmethod def from_settings(cls, settings, method=SSL.SSLv23_METHOD, *args, **kwargs): tls_verbose_logging = settings.getbool('DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING') - return cls(method=method, tls_verbose_logging=tls_verbose_logging, *args, **kwargs) + tls_ciphers = settings['DOWNLOADER_CLIENT_TLS_CIPHERS'] + return cls(method=method, tls_verbose_logging=tls_verbose_logging, tls_ciphers=tls_ciphers, *args, **kwargs) def getCertificateOptions(self): # setting verify=True will require you to provide CAs @@ -45,7 +50,7 @@ class ScrapyClientContextFactory(BrowserLikePolicyForHTTPS): method=getattr(self, 'method', getattr(self, '_ssl_method', None)), fixBrokenPeers=True, - acceptableCiphers=DEFAULT_CIPHERS) + acceptableCiphers=self.tls_ciphers) # kept for old-style HTTP/1.0 downloader context twisted calls, # e.g. connectSSL() diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 81fee543f..742c8e8a1 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -85,6 +85,7 @@ DOWNLOADER = 'scrapy.core.downloader.Downloader' DOWNLOADER_HTTPCLIENTFACTORY = 'scrapy.core.downloader.webclient.ScrapyHTTPClientFactory' DOWNLOADER_CLIENTCONTEXTFACTORY = 'scrapy.core.downloader.contextfactory.ScrapyClientContextFactory' +DOWNLOADER_CLIENT_TLS_CIPHERS = 'DEFAULT' DOWNLOADER_CLIENT_TLS_METHOD = 'TLS' # Use highest TLS/SSL protocol version supported by the platform, # also allowing negotiation DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING = False From ce281d890dfbd905922e16552fdcee34eb0d6c42 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Fri, 28 Sep 2018 20:17:12 +0500 Subject: [PATCH 045/387] Documentation for DOWNLOADER_CLIENT_TLS_CIPHERS. --- docs/topics/settings.rst | 18 ++++++++++++++++++ 1 file changed, 18 insertions(+) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 12606fe47..c042d3f43 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -445,6 +445,24 @@ accepts a ``method`` parameter (this is the ``OpenSSL.SSL`` method mapping :setting:`DOWNLOADER_CLIENT_TLS_METHOD`) and a ``tls_verbose_logging`` parameter (``bool``). +.. setting:: DOWNLOADER_CLIENT_TLS_CIPHERS + +DOWNLOADER_CLIENT_TLS_CIPHERS +----------------------------- + +Default: ``'DEFAULT'`` + +Use this setting to customize the TLS/SSL ciphers used by the default +HTTP/1.1 downloader. + +The setting should contain a string in the `OpenSSL cipher list format`_, +these ciphers will be used as client ciphers. Changing this setting may be +necessary to access certain HTTPS websites: for example, you may need to use +``'DEFAULT:!DH'`` for a website with weak DH parameters or enable a +specific cipher that is not included in ``DEFAULT`` if a website requires it. + +.. _OpenSSL cipher list format: https://www.openssl.org/docs/manmaster/man1/ciphers.html#CIPHER-LIST-FORMAT + .. setting:: DOWNLOADER_CLIENT_TLS_METHOD DOWNLOADER_CLIENT_TLS_METHOD From 9a8edf2bf1172e8eec282a70ad46a9dacff76d62 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 27 Sep 2018 19:52:29 +0500 Subject: [PATCH 046/387] Tests for setting SSL ciphers. --- tests/mockserver.py | 10 +++++++ tests/test_downloader_handlers.py | 43 +++++++++++++++++++++++++- tests/test_webclient.py | 50 ++++++++++++++++++++++++++++++- 3 files changed, 101 insertions(+), 2 deletions(-) diff --git a/tests/mockserver.py b/tests/mockserver.py index 3fa4bc0f0..8be8a36bb 100644 --- a/tests/mockserver.py +++ b/tests/mockserver.py @@ -3,6 +3,8 @@ import sys, time, random, os, json from six.moves.urllib.parse import urlencode from subprocess import Popen, PIPE +from OpenSSL import SSL + from twisted.web.server import Site, NOT_DONE_YET from twisted.web.resource import Resource from twisted.web.static import File @@ -222,6 +224,14 @@ def ssl_context_factory(keyfile='keys/localhost.key', certfile='keys/localhost.c ) +def broken_ssl_context_factory(keyfile='keys/localhost.key', certfile='keys/localhost.crt', cipher_string='DEFAULT'): + factory = ssl_context_factory(keyfile, certfile) + ctx = factory.getContext() + ctx.set_options(SSL.OP_CIPHER_SERVER_PREFERENCE | SSL.OP_NO_TLSv1_2) + ctx.set_cipher_list(cipher_string) + return factory + + if __name__ == "__main__": root = Root() factory = Site(root) diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index efef4192c..5df2ffe45 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -39,7 +39,7 @@ from scrapy.utils.test import get_crawler, skip_if_no_boto from scrapy.utils.python import to_bytes from scrapy.exceptions import NotConfigured -from tests.mockserver import MockServer, ssl_context_factory, Echo +from tests.mockserver import MockServer, ssl_context_factory, Echo, broken_ssl_context_factory from tests.spiders import SingleRequestSpider @@ -553,6 +553,47 @@ class Https11InvalidDNSPattern(Https11TestCase): super(Https11InvalidDNSPattern, self).setUp() +class Https11BadCiphers(unittest.TestCase): + scheme = 'https' + download_handler_cls = HTTP11DownloadHandler + + keyfile = 'keys/localhost.key' + certfile = 'keys/localhost.crt' + + def setUp(self): + self.tmpname = self.mktemp() + os.mkdir(self.tmpname) + FilePath(self.tmpname).child("file").setContent(b"0123456789") + r = static.File(self.tmpname) + self.site = server.Site(r, timeout=None) + self.wrapper = WrappingFactory(self.site) + self.host = 'localhost' + self.port = reactor.listenSSL( + 0, self.wrapper, broken_ssl_context_factory(self.keyfile, self.certfile, cipher_string='CAMELLIA256-SHA'), + interface=self.host) + self.portno = self.port.getHost().port + self.download_handler = self.download_handler_cls( + Settings({'DOWNLOADER_CLIENT_TLS_CIPHERS': 'CAMELLIA256-SHA'})) + self.download_request = self.download_handler.download_request + + @defer.inlineCallbacks + def tearDown(self): + yield self.port.stopListening() + if hasattr(self.download_handler, 'close'): + yield self.download_handler.close() + shutil.rmtree(self.tmpname) + + def getURL(self, path): + return "%s://%s:%d/%s" % (self.scheme, self.host, self.portno, path) + + def test_download(self): + request = Request(self.getURL('file')) + d = self.download_request(request, Spider('foo')) + d.addCallback(lambda r: r.body) + d.addCallback(self.assertEqual, b"0123456789") + return d + + class Http11MockServerTestCase(unittest.TestCase): """HTTP 1.1 test case with MockServer""" diff --git a/tests/test_webclient.py b/tests/test_webclient.py index 766329b57..2ebe075ab 100644 --- a/tests/test_webclient.py +++ b/tests/test_webclient.py @@ -8,15 +8,18 @@ import shutil from twisted.trial import unittest from twisted.web import server, static, util, resource -from twisted.internet import reactor, defer +from twisted.internet import reactor, defer, ssl from twisted.test.proto_helpers import StringTransport from twisted.python.filepath import FilePath from twisted.protocols.policies import WrappingFactory from twisted.internet.defer import inlineCallbacks from scrapy.core.downloader import webclient as client +from scrapy.core.downloader.contextfactory import ScrapyClientContextFactory from scrapy.http import Request, Headers +from scrapy.settings import Settings from scrapy.utils.python import to_bytes, to_unicode +from tests.mockserver import ssl_context_factory, broken_ssl_context_factory def getPage(url, contextFactory=None, response_transform=None, *args, **kwargs): @@ -363,3 +366,48 @@ class WebClientTestCase(unittest.TestCase): self.assertEqual(content_encoding, EncodingResource.out_encoding) self.assertEqual( response.body.decode(content_encoding), to_unicode(original_body)) + + +class WebClientSSLTestCase(unittest.TestCase): + context_factory = None + + def _listen(self, site): + return reactor.listenSSL( + 0, site, + contextFactory=self.context_factory or ssl_context_factory(), + interface="127.0.0.1") + + def getURL(self, path): + return "https://127.0.0.1:%d/%s" % (self.portno, path) + + def setUp(self): + self.tmpname = self.mktemp() + os.mkdir(self.tmpname) + FilePath(self.tmpname).child("file").setContent(b"0123456789") + r = static.File(self.tmpname) + r.putChild(b"payload", PayloadResource()) + self.site = server.Site(r, timeout=None) + self.wrapper = WrappingFactory(self.site) + self.port = self._listen(self.wrapper) + self.portno = self.port.getHost().port + + @inlineCallbacks + def tearDown(self): + yield self.port.stopListening() + shutil.rmtree(self.tmpname) + + def testPayload(self): + s = "0123456789" * 10 + return getPage(self.getURL("payload"), body=s).addCallback( + self.assertEqual, to_bytes(s)) + + +class WebClientBrokenSSLTestCase(WebClientSSLTestCase): + context_factory = broken_ssl_context_factory(cipher_string='CAMELLIA256-SHA') + + def testPayload(self): + s = "0123456789" * 10 + settings = Settings({'DOWNLOADER_CLIENT_TLS_CIPHERS': 'CAMELLIA256-SHA'}) + return getPage(self.getURL("payload"), body=s, + contextFactory=ScrapyClientContextFactory(settings=settings)).addCallback(self.assertEqual, + to_bytes(s)) From aaa5229e5db4f1c014ae69e0fb9e1a933f4952b0 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Mon, 15 Jul 2019 17:47:06 +0500 Subject: [PATCH 047/387] Fixes and improvements for DOWNLOADER_CLIENT_TLS_CIPHERS. --- docs/topics/settings.rst | 5 +++-- scrapy/core/downloader/handlers/http11.py | 2 +- scrapy/utils/ssl.py | 5 +++++ tests/mockserver.py | 19 ++++++++----------- tests/test_downloader_handlers.py | 6 +++--- tests/test_webclient.py | 23 ++++++++++++++++------- 6 files changed, 36 insertions(+), 24 deletions(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index c042d3f43..0cb81c43e 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -442,8 +442,9 @@ or even enable client-side authentication (and various other things). If you do use a custom ContextFactory, make sure its ``__init__`` method accepts a ``method`` parameter (this is the ``OpenSSL.SSL`` method mapping -:setting:`DOWNLOADER_CLIENT_TLS_METHOD`) and a ``tls_verbose_logging`` -parameter (``bool``). +:setting:`DOWNLOADER_CLIENT_TLS_METHOD`), a ``tls_verbose_logging`` +parameter (``bool``) and a ``tls_ciphers`` parameter (see +:setting:`DOWNLOADER_CLIENT_TLS_CIPHERS`). .. setting:: DOWNLOADER_CLIENT_TLS_CIPHERS diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index e72052afc..91b45a8fc 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -54,7 +54,7 @@ class HTTP11DownloadHandler(object): ) msg = """ '%s' does not accept `method` argument (type OpenSSL.SSL method,\ - e.g. OpenSSL.SSL.SSLv23_METHOD) and/or `tls_verbose_logging` argument.\ + e.g. OpenSSL.SSL.SSLv23_METHOD) and/or `tls_verbose_logging` argument and/or `tls_ciphers` argument.\ Please upgrade your context factory class to handle them or ignore them.""" % ( settings['DOWNLOADER_CLIENTCONTEXTFACTORY'],) warnings.warn(msg) diff --git a/scrapy/utils/ssl.py b/scrapy/utils/ssl.py index 632827471..02aed60ee 100644 --- a/scrapy/utils/ssl.py +++ b/scrapy/utils/ssl.py @@ -6,6 +6,11 @@ import OpenSSL._util as pyOpenSSLutil from scrapy.utils.python import to_native_str +# The OpenSSL symbol is present since 1.1.1 but it's not currently supported in any version of pyOpenSSL. +# Using the binding directly, as this code does, requires cryptography 2.4. +SSL_OP_NO_TLSv1_3 = getattr(pyOpenSSLutil.lib, 'SSL_OP_NO_TLSv1_3', 0) + + def ffi_buf_to_string(buf): return to_native_str(pyOpenSSLutil.ffi.string(buf)) diff --git a/tests/mockserver.py b/tests/mockserver.py index 8be8a36bb..77908284b 100644 --- a/tests/mockserver.py +++ b/tests/mockserver.py @@ -4,7 +4,6 @@ from six.moves.urllib.parse import urlencode from subprocess import Popen, PIPE from OpenSSL import SSL - from twisted.web.server import Site, NOT_DONE_YET from twisted.web.resource import Resource from twisted.web.static import File @@ -15,8 +14,8 @@ from twisted.web.util import redirectTo from twisted.internet import reactor, ssl from twisted.internet.task import deferLater - from scrapy.utils.python import to_bytes, to_unicode +from scrapy.utils.ssl import SSL_OP_NO_TLSv1_3 def getarg(request, name, default=None, type=None): @@ -217,18 +216,16 @@ class MockServer(): return host + path -def ssl_context_factory(keyfile='keys/localhost.key', certfile='keys/localhost.crt'): - return ssl.DefaultOpenSSLContextFactory( +def ssl_context_factory(keyfile='keys/localhost.key', certfile='keys/localhost.crt', cipher_string=None): + factory = ssl.DefaultOpenSSLContextFactory( os.path.join(os.path.dirname(__file__), keyfile), os.path.join(os.path.dirname(__file__), certfile), ) - - -def broken_ssl_context_factory(keyfile='keys/localhost.key', certfile='keys/localhost.crt', cipher_string='DEFAULT'): - factory = ssl_context_factory(keyfile, certfile) - ctx = factory.getContext() - ctx.set_options(SSL.OP_CIPHER_SERVER_PREFERENCE | SSL.OP_NO_TLSv1_2) - ctx.set_cipher_list(cipher_string) + if cipher_string: + ctx = factory.getContext() + # disabling TLS1.2+ because it unconditionally enables some strong ciphers + ctx.set_options(SSL.OP_CIPHER_SERVER_PREFERENCE | SSL.OP_NO_TLSv1_2 | SSL_OP_NO_TLSv1_3) + ctx.set_cipher_list(to_bytes(cipher_string)) return factory diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 5df2ffe45..109469503 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -39,7 +39,7 @@ from scrapy.utils.test import get_crawler, skip_if_no_boto from scrapy.utils.python import to_bytes from scrapy.exceptions import NotConfigured -from tests.mockserver import MockServer, ssl_context_factory, Echo, broken_ssl_context_factory +from tests.mockserver import MockServer, ssl_context_factory, Echo from tests.spiders import SingleRequestSpider @@ -553,7 +553,7 @@ class Https11InvalidDNSPattern(Https11TestCase): super(Https11InvalidDNSPattern, self).setUp() -class Https11BadCiphers(unittest.TestCase): +class Https11CustomCiphers(unittest.TestCase): scheme = 'https' download_handler_cls = HTTP11DownloadHandler @@ -569,7 +569,7 @@ class Https11BadCiphers(unittest.TestCase): self.wrapper = WrappingFactory(self.site) self.host = 'localhost' self.port = reactor.listenSSL( - 0, self.wrapper, broken_ssl_context_factory(self.keyfile, self.certfile, cipher_string='CAMELLIA256-SHA'), + 0, self.wrapper, ssl_context_factory(self.keyfile, self.certfile, cipher_string='CAMELLIA256-SHA'), interface=self.host) self.portno = self.port.getHost().port self.download_handler = self.download_handler_cls( diff --git a/tests/test_webclient.py b/tests/test_webclient.py index 2ebe075ab..a81946490 100644 --- a/tests/test_webclient.py +++ b/tests/test_webclient.py @@ -6,9 +6,10 @@ import os import six import shutil +import OpenSSL.SSL from twisted.trial import unittest from twisted.web import server, static, util, resource -from twisted.internet import reactor, defer, ssl +from twisted.internet import reactor, defer from twisted.test.proto_helpers import StringTransport from twisted.python.filepath import FilePath from twisted.protocols.policies import WrappingFactory @@ -18,8 +19,9 @@ from scrapy.core.downloader import webclient as client from scrapy.core.downloader.contextfactory import ScrapyClientContextFactory from scrapy.http import Request, Headers from scrapy.settings import Settings +from scrapy.utils.misc import create_instance from scrapy.utils.python import to_bytes, to_unicode -from tests.mockserver import ssl_context_factory, broken_ssl_context_factory +from tests.mockserver import ssl_context_factory def getPage(url, contextFactory=None, response_transform=None, *args, **kwargs): @@ -402,12 +404,19 @@ class WebClientSSLTestCase(unittest.TestCase): self.assertEqual, to_bytes(s)) -class WebClientBrokenSSLTestCase(WebClientSSLTestCase): - context_factory = broken_ssl_context_factory(cipher_string='CAMELLIA256-SHA') +class WebClientCustomCiphersSSLTestCase(WebClientSSLTestCase): + # we try to use a cipher that is not enabled by default in OpenSSL + custom_ciphers = 'CAMELLIA256-SHA' + context_factory = ssl_context_factory(cipher_string=custom_ciphers) def testPayload(self): s = "0123456789" * 10 - settings = Settings({'DOWNLOADER_CLIENT_TLS_CIPHERS': 'CAMELLIA256-SHA'}) + settings = Settings({'DOWNLOADER_CLIENT_TLS_CIPHERS': self.custom_ciphers}) + client_context_factory = create_instance(ScrapyClientContextFactory, settings=settings, crawler=None) return getPage(self.getURL("payload"), body=s, - contextFactory=ScrapyClientContextFactory(settings=settings)).addCallback(self.assertEqual, - to_bytes(s)) + contextFactory=client_context_factory).addCallback(self.assertEqual, to_bytes(s)) + + def testPayloadDefaultCiphers(self): + s = "0123456789" * 10 + d = getPage(self.getURL("payload"), body=s, contextFactory=ScrapyClientContextFactory()) + return self.assertFailure(d, OpenSSL.SSL.Error) From 50c4cafe0ccc53064f38654fd7426ded876469f7 Mon Sep 17 00:00:00 2001 From: Tobias Hernstig Date: Sun, 11 Mar 2018 11:46:07 +0100 Subject: [PATCH 048/387] Update documentation for logging manually Usage of basicConfig() together with crawlerRunner is not recommended. Update documentation to highlight this fact. Closes #2149, #2352, #3146 --- docs/topics/logging.rst | 14 +++++++------- 1 file changed, 7 insertions(+), 7 deletions(-) diff --git a/docs/topics/logging.rst b/docs/topics/logging.rst index 87ea43c7d..2fd85196d 100644 --- a/docs/topics/logging.rst +++ b/docs/topics/logging.rst @@ -254,18 +254,18 @@ scrapy.utils.log module when running custom scripts using :class:`~scrapy.crawler.CrawlerRunner`. In that case, its usage is not required but it's recommended. - If you plan on configuring the handlers yourself is still recommended you - call this function, passing ``install_root_handler=False``. Bear in mind - there won't be any log output set by default in that case. + Another option when running custom scripts is to manually configure the logging. + To do this you can use `logging.basicConfig()`_ to set a basic root handler. - To get you started on manually configuring logging's output, you can use - `logging.basicConfig()`_ to set a basic root handler. This is an example - on how to redirect ``INFO`` or higher messages to a file:: + Note that ``scrapy.crawler.CrawlerProcess`` automatically calls ``configure_logging``, + so it is recommended to only use `logging.basicConfig()`_ together with + ``scrapy.crawler.CrawlerRunner`` + + This is an example on how to redirect ``INFO`` or higher messages to a file:: import logging from scrapy.utils.log import configure_logging - configure_logging(install_root_handler=False) logging.basicConfig( filename='log.txt', format='%(levelname)s: %(message)s', From 2b0de0606c3b1a379722e916e9ce350e8de08a20 Mon Sep 17 00:00:00 2001 From: Tobias Hernstig Date: Thu, 15 Aug 2019 18:54:28 +0200 Subject: [PATCH 049/387] Fix merge conflicts --- docs/topics/logging.rst | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/topics/logging.rst b/docs/topics/logging.rst index 2fd85196d..87ab6e19a 100644 --- a/docs/topics/logging.rst +++ b/docs/topics/logging.rst @@ -257,9 +257,9 @@ scrapy.utils.log module Another option when running custom scripts is to manually configure the logging. To do this you can use `logging.basicConfig()`_ to set a basic root handler. - Note that ``scrapy.crawler.CrawlerProcess`` automatically calls ``configure_logging``, + Note that :class:`~scrapy.crawler.CrawlerProcess` automatically calls ``configure_logging``, so it is recommended to only use `logging.basicConfig()`_ together with - ``scrapy.crawler.CrawlerRunner`` + :class:`~scrapy.crawler.CrawlerRunner`. This is an example on how to redirect ``INFO`` or higher messages to a file:: From 00fe05e53611d5c61fb23cdd6674fd89c99fd69e Mon Sep 17 00:00:00 2001 From: Anubhav Patel Date: Mon, 19 Aug 2019 09:19:06 +0530 Subject: [PATCH 050/387] adds ROBOTSTXT_USER_AGENT setting --- docs/topics/downloader-middleware.rst | 15 +++++++++++++++ docs/topics/settings.rst | 16 +++++++++++++++- scrapy/downloadermiddlewares/robotstxt.py | 6 +++++- scrapy/settings/default_settings.py | 1 + tests/test_downloadermiddleware_robotstxt.py | 9 +++++++++ 5 files changed, 45 insertions(+), 2 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 616b56101..ae413dc84 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -1074,6 +1074,21 @@ implementing the methods described below. .. autoclass:: RobotParser :members: +RobotsTxtMiddleware Settings +~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +.. setting:: ROBOTSTXT_USER_AGENT + +ROBOTSTXT_USER_AGENT +^^^^^^^^^^^^^^^^^^^^ + +Default: ``None`` + +The user agent string to use for matching in the robots.txt_ file. If ``None``, +the User-Agent header you are sending with the request or the +:setting:`USER_AGENT` setting (in that order) will be used for determining +the user agent to use in the robots.txt_ file. + .. _robots.txt: http://www.robotstxt.org/ DownloaderStats diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 12606fe47..c4b55cc7b 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -1151,6 +1151,18 @@ Default: ``'scrapy.robotstxt.PythonRobotParser'`` The parser backend to use for parsing ``robots.txt`` files. For more information see :ref:`topics-dlmw-robots`. +.. setting:: ROBOTSTXT_USER_AGENT + +ROBOTSTXT_USER_AGENT +^^^^^^^^^^^^^^^^^^^^ + +Default: ``None`` + +The user agent string to use for matching in the robots.txt file. If ``None``, +the User-Agent header you are sending with the request or the +:setting:`USER_AGENT` setting (in that order) will be used for determining +the user agent to use in the robots.txt file. + .. setting:: SCHEDULER SCHEDULER @@ -1409,7 +1421,9 @@ USER_AGENT Default: ``"Scrapy/VERSION (+https://scrapy.org)"`` -The default User-Agent to use when crawling, unless overridden. +The default User-Agent to use when crawling, unless overridden. This user agent is +also used in robots.txt if :setting:`ROBOTSTXT_USER_AGENT` setting is ``None`` and +there is no overridding User-Agent header specified for the request. Settings documented elsewhere: diff --git a/scrapy/downloadermiddlewares/robotstxt.py b/scrapy/downloadermiddlewares/robotstxt.py index c5a60d355..6a5dfb79c 100644 --- a/scrapy/downloadermiddlewares/robotstxt.py +++ b/scrapy/downloadermiddlewares/robotstxt.py @@ -26,6 +26,7 @@ class RobotsTxtMiddleware(object): if not crawler.settings.getbool('ROBOTSTXT_OBEY'): raise NotConfigured self._default_useragent = crawler.settings.get('USER_AGENT', 'Scrapy') + self._robotstxt_useragent = crawler.settings.get('ROBOTSTXT_USER_AGENT', None) self.crawler = crawler self._parsers = {} self._parserimpl = load_object(crawler.settings.get('ROBOTSTXT_PARSER')) @@ -47,7 +48,10 @@ class RobotsTxtMiddleware(object): def process_request_2(self, rp, request, spider): if rp is None: return - useragent = request.headers.get(b'User-Agent', self._default_useragent) + + useragent = self._robotstxt_useragent + if not useragent: + useragent = request.headers.get(b'User-Agent', self._default_useragent) if not rp.allowed(request.url, useragent): logger.debug("Forbidden by robots.txt: %(request)s", {'request': request}, extra={'spider': spider}) diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 81fee543f..efb257e2b 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -246,6 +246,7 @@ RETRY_PRIORITY_ADJUST = -1 ROBOTSTXT_OBEY = False ROBOTSTXT_PARSER = 'scrapy.robotstxt.PythonRobotParser' +ROBOTSTXT_USER_AGENT = None SCHEDULER = 'scrapy.core.scheduler.Scheduler' SCHEDULER_DISK_QUEUE = 'scrapy.squeues.PickleLifoDiskQueue' diff --git a/tests/test_downloadermiddleware_robotstxt.py b/tests/test_downloadermiddleware_robotstxt.py index 79f172848..fbc46cba4 100644 --- a/tests/test_downloadermiddleware_robotstxt.py +++ b/tests/test_downloadermiddleware_robotstxt.py @@ -164,6 +164,15 @@ Disallow: /some/randome/page.html d.addCallback(lambda _: self.assertFalse(mw_module_logger.error.called)) return d + def test_robotstxt_user_agent_setting(self): + crawler = self._get_successful_crawler() + crawler.settings.set('ROBOTSTXT_USER_AGENT', 'Examplebot') + crawler.settings.set('USER_AGENT', 'Mozilla/5.0 (X11; Linux x86_64)') + middleware = RobotsTxtMiddleware(crawler) + rp = mock.MagicMock(return_value=True) + middleware.process_request_2(rp, Request('http://site.local/allowed'), None) + rp.allowed.assert_called_once_with('http://site.local/allowed', 'Examplebot') + def assertNotIgnored(self, request, middleware): spider = None # not actually used dfd = maybeDeferred(middleware.process_request, request, spider) From 0fa384e80defaaac88b2f0d5449072f832611431 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 22 Aug 2019 20:10:42 +0200 Subject: [PATCH 051/387] Provide complete API documentation coverage of scrapy.exporters --- docs/conf.py | 4 ++++ docs/news.rst | 8 ++++---- docs/topics/exporters.rst | 15 +++++++++++++-- scrapy/exporters.py | 18 ++++++++++++++---- 4 files changed, 35 insertions(+), 10 deletions(-) diff --git a/docs/conf.py b/docs/conf.py index f49f79cd5..eba416cd6 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -252,6 +252,10 @@ coverage_ignore_pyobjects = [ # Private exception used by the command-line interface implementation. r'^scrapy\.exceptions\.UsageError', + + # Methods of BaseItemExporter subclasses are only documented in + # BaseItemExporter. + r'^scrapy\.exporters\.(?!BaseItemExporter\b)\w*?\.', ] diff --git a/docs/news.rst b/docs/news.rst index ce5b8b406..5915ba742 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -1269,8 +1269,8 @@ This 1.1 release brings a lot of interesting features and bug fixes: this behavior, update :setting:`ROBOTSTXT_OBEY` in ``settings.py`` file after creating a new project. - Exporters now work on unicode, instead of bytes by default (:issue:`1080`). - If you use ``PythonItemExporter``, you may want to update your code to - disable binary mode which is now deprecated. + If you use :class:`~scrapy.exporters.PythonItemExporter`, you may want to + update your code to disable binary mode which is now deprecated. - Accept XML node names containing dots as valid (:issue:`1533`). - When uploading files or images to S3 (with ``FilesPipeline`` or ``ImagesPipeline``), the default ACL policy is now "private" instead @@ -1408,8 +1408,8 @@ Bugfixes - Fixed bug on ``XMLItemExporter`` with non-string fields in items (:issue:`1738`). - Fixed startproject command in OS X (:issue:`1635`). -- Fixed PythonItemExporter and CSVExporter for non-string item - types (:issue:`1737`). +- Fixed :class:`~scrapy.exporters.PythonItemExporter` and CSVExporter for + non-string item types (:issue:`1737`). - Various logging related fixes (:issue:`1294`, :issue:`1419`, :issue:`1263`, :issue:`1624`, :issue:`1654`, :issue:`1722`, :issue:`1726` and :issue:`1303`). - Fixed bug in ``utils.template.render_templatefile()`` (:issue:`1212`). diff --git a/docs/topics/exporters.rst b/docs/topics/exporters.rst index f5048d2da..11b3045ec 100644 --- a/docs/topics/exporters.rst +++ b/docs/topics/exporters.rst @@ -165,9 +165,9 @@ BaseItemExporter value unchanged except for ``unicode`` values which are encoded to ``str`` using the encoding declared in the :attr:`encoding` attribute. - :param field: the field being serialized. If a raw dict is being + :param field: the field being serialized. If a raw dict is being exported (not :class:`~.Item`) *field* value is an empty dict. - :type field: :class:`~scrapy.item.Field` object or an empty dict + :type field: :class:`~scrapy.item.Field` object or an empty dict :param name: the name of the field being serialized :type name: str @@ -223,6 +223,12 @@ BaseItemExporter * ``indent<=0`` each item on its own line, no indentation * ``indent>0`` each item on its own line, indented with the provided numeric value +PythonItemExporter +------------------ + +.. autoclass:: PythonItemExporter + + .. highlight:: none XmlItemExporter @@ -410,3 +416,8 @@ JsonLinesItemExporter this exporter is well suited for serializing large amounts of data. .. _JSONEncoder: https://docs.python.org/2/library/json.html#json.JSONEncoder + +MarshalItemExporter +------------------- + +.. autoclass:: MarshalItemExporter diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 695c74fec..6fc87ed18 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -276,6 +276,13 @@ class PickleItemExporter(BaseItemExporter): class MarshalItemExporter(BaseItemExporter): + """Exports items in a Python-specific binary format (see + :mod:`marshal`). + + :param file: The file-like object to use for exporting the data. Its + ``write`` method should accept :class:`bytes` (a disk file + opened in binary mode, a :class:`~io.BytesIO` object, etc) + """ def __init__(self, file, **kwargs): self._configure(kwargs) @@ -297,10 +304,13 @@ class PprintItemExporter(BaseItemExporter): class PythonItemExporter(BaseItemExporter): - """The idea behind this exporter is to have a mechanism to serialize items - to built-in python types so any serialization library (like - json, msgpack, binc, etc) can be used on top of it. Its main goal is to - seamless support what BaseItemExporter does plus nested items. + """This is a base class for item exporters that extends + :class:`BaseItemExporter` with support for nested items. + + It serializes items to built-in Python types, so that any serialization + library (e.g. :mod:`json` or msgpack_) can be used on top of it. + + .. _msgpack: https://pypi.org/project/msgpack/ """ def _configure(self, options, dont_fail=False): self.binary = options.pop('binary', True) From 3abe7e6e6da38b1586b720749f53955b8d00390e Mon Sep 17 00:00:00 2001 From: elacuesta Date: Mon, 26 Aug 2019 04:35:44 -0300 Subject: [PATCH 052/387] Add Bug report and Feature request templates (#3471) --- .github/ISSUE_TEMPLATE/bug_report.md | 41 +++++++++++++++++++++++ .github/ISSUE_TEMPLATE/feature_request.md | 33 ++++++++++++++++++ 2 files changed, 74 insertions(+) create mode 100644 .github/ISSUE_TEMPLATE/bug_report.md create mode 100644 .github/ISSUE_TEMPLATE/feature_request.md diff --git a/.github/ISSUE_TEMPLATE/bug_report.md b/.github/ISSUE_TEMPLATE/bug_report.md new file mode 100644 index 000000000..66821171f --- /dev/null +++ b/.github/ISSUE_TEMPLATE/bug_report.md @@ -0,0 +1,41 @@ +--- +name: Bug report +about: Report a problem to help us improve +--- + + + +### Description + +[Description of the issue] + +### Steps to Reproduce + +1. [First Step] +2. [Second Step] +3. [and so on...] + +**Expected behavior:** [What you expect to happen] + +**Actual behavior:** [What actually happens] + +**Reproduces how often:** [What percentage of the time does it reproduce?] + +### Versions + +Please paste here the output of executing `scrapy version --verbose` in the command line. + +### Additional context + +Any additional information, configuration, data or output from commands that might be necessary to reproduce or understand the issue. Please try not to include screenshots of code or the command line, paste the contents as text instead. You can use [GitHub Flavored Markdown](https://help.github.com/en/articles/creating-and-highlighting-code-blocks) to make the text look better. diff --git a/.github/ISSUE_TEMPLATE/feature_request.md b/.github/ISSUE_TEMPLATE/feature_request.md new file mode 100644 index 000000000..df5127b4c --- /dev/null +++ b/.github/ISSUE_TEMPLATE/feature_request.md @@ -0,0 +1,33 @@ +--- +name: Feature request +about: Suggest an idea for an enhancement or new feature +--- + + + +## Summary + +One paragraph explanation of the feature. + +## Motivation + +Why are we doing this? What use cases does it support? What is the expected outcome? + +## Describe alternatives you've considered + +A clear and concise description of the alternative solutions you've considered. Be sure to explain why Scrapy's existing customizability isn't suitable for this feature. + +## Additional context + +Any additional information about the feature request here. From 3a7b949d6d8058cdc342dba683bc054d95ed6ccd Mon Sep 17 00:00:00 2001 From: Anubhav Patel Date: Tue, 27 Aug 2019 13:11:31 +0530 Subject: [PATCH 053/387] Adds integration with Protego robots.txt parser (#3935) --- docs/topics/downloader-middleware.rst | 20 ++++++++- scrapy/robotstxt.py | 58 +++++++++++++++++---------- tests/test_robotstxt_interface.py | 21 +++++++++- tox.ini | 2 + 4 files changed, 77 insertions(+), 24 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 616b56101..6aa714fb2 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -996,6 +996,7 @@ RobotsTxtMiddleware * :ref:`RobotFileParser ` (default) * :ref:`Reppy ` * :ref:`Robotexclusionrulesparser ` + * :ref:`Protego ` You can change the robots.txt_ parser with the :setting:`ROBOTSTXT_PARSER` setting. Or you can also :ref:`implement support for a new parser `. @@ -1013,7 +1014,7 @@ RobotFileParser ~~~~~~~~~~~~~~~ `RobotFileParser `_ is -Python's inbuilt ``robots.txt`` parser. The parser is fully compliant with `Martijn Koster's +Python's inbuilt robots.txt_ parser. The parser is fully compliant with `Martijn Koster's 1996 draft specification `_. It lacks support for wildcard matching. Scrapy uses this parser by default. @@ -1059,6 +1060,23 @@ In order to use this parser: * Set :setting:`ROBOTSTXT_PARSER` setting to ``scrapy.robotstxt.ReppyRobotParser`` +.. _protego-parser: + +Protego parser +~~~~~~~~~~~~~~ + +`Protego `_ is a pure-Python robots.txt_ parser. +The parser is fully compliant with `Google's Robots.txt Specification +`_ hence supports wildcard +matching, and uses the length based rule similar to `Reppy `_. + +In order to use this parser: + +* Install `Protego `_ by running ``pip install protego`` + +* Set :setting:`ROBOTSTXT_PARSER` setting to + ``scrapy.robotstxt.ProtegoRobotParser`` + .. _support-for-new-robots-parser: Implementing support for a new parser diff --git a/scrapy/robotstxt.py b/scrapy/robotstxt.py index 4bfb275fd..189f165d1 100644 --- a/scrapy/robotstxt.py +++ b/scrapy/robotstxt.py @@ -7,6 +7,21 @@ from scrapy.utils.python import to_native_str, to_unicode logger = logging.getLogger(__name__) +def decode_robotstxt(robotstxt_body, spider, to_native_str_type=False): + try: + if to_native_str_type: + robotstxt_body = to_native_str(robotstxt_body) + else: + robotstxt_body = robotstxt_body.decode('utf-8') + except UnicodeDecodeError: + # If we found garbage or robots.txt in an encoding other than UTF-8, disregard it. + # Switch to 'allow all' state. + logger.warning("Failure while parsing robots.txt. " + "File either contains garbage or is in an encoding other than UTF-8, treating it as an empty file.", + exc_info=sys.exc_info(), + extra={'spider': spider}) + robotstxt_body = '' + return robotstxt_body class RobotParser(with_metaclass(ABCMeta)): @classmethod @@ -40,17 +55,7 @@ class PythonRobotParser(RobotParser): def __init__(self, robotstxt_body, spider): from six.moves.urllib_robotparser import RobotFileParser self.spider = spider - try: - robotstxt_body = to_native_str(robotstxt_body) - except UnicodeDecodeError: - # If we found garbage or robots.txt in an encoding other than UTF-8, disregard it. - # Switch to 'allow all' state. - logger.warning("Failure while parsing robots.txt using %(parser)s." - " File either contains garbage or is in an encoding other than UTF-8, treating it as an empty file.", - {'parser': "RobotFileParser"}, - exc_info=sys.exc_info(), - extra={'spider': self.spider}) - robotstxt_body = '' + robotstxt_body = decode_robotstxt(robotstxt_body, spider, to_native_str_type=True) self.rp = RobotFileParser() self.rp.parse(robotstxt_body.splitlines()) @@ -87,17 +92,7 @@ class RerpRobotParser(RobotParser): from robotexclusionrulesparser import RobotExclusionRulesParser self.spider = spider self.rp = RobotExclusionRulesParser() - try: - robotstxt_body = robotstxt_body.decode('utf-8') - except UnicodeDecodeError: - # If we found garbage or robots.txt in an encoding other than UTF-8, disregard it. - # Switch to 'allow all' state. - logger.warning("Failure while parsing robots.txt using %(parser)s." - " File either contains garbage or is in an encoding other than UTF-8, treating it as an empty file.", - {'parser': "RobotExclusionRulesParser"}, - exc_info=sys.exc_info(), - extra={'spider': self.spider}) - robotstxt_body = '' + robotstxt_body = decode_robotstxt(robotstxt_body, spider) self.rp.parse(robotstxt_body) @classmethod @@ -110,3 +105,22 @@ class RerpRobotParser(RobotParser): user_agent = to_unicode(user_agent) url = to_unicode(url) return self.rp.is_allowed(user_agent, url) + + +class ProtegoRobotParser(RobotParser): + def __init__(self, robotstxt_body, spider): + from protego import Protego + self.spider = spider + robotstxt_body = decode_robotstxt(robotstxt_body, spider) + self.rp = Protego.parse(robotstxt_body) + + @classmethod + def from_crawler(cls, crawler, robotstxt_body): + spider = None if not crawler else crawler.spider + o = cls(robotstxt_body, spider) + return o + + def allowed(self, url, user_agent): + user_agent = to_unicode(user_agent) + url = to_unicode(url) + return self.rp.can_fetch(url, user_agent) diff --git a/tests/test_robotstxt_interface.py b/tests/test_robotstxt_interface.py index 2819786b5..9aaab560a 100644 --- a/tests/test_robotstxt_interface.py +++ b/tests/test_robotstxt_interface.py @@ -20,6 +20,13 @@ def rerp_available(): return False return True +def protego_available(): + # check if protego parser is installed + try: + from protego import Protego + except ImportError: + return False + return True class BaseRobotParserTest: def _setUp(self, parser_cls): @@ -127,7 +134,7 @@ class ReppyRobotParserTest(BaseRobotParserTest, unittest.TestCase): super(ReppyRobotParserTest, self)._setUp(ReppyRobotParser) def test_order_based_precedence(self): - raise unittest.SkipTest("Rerp does not support order based directives precedence.") + raise unittest.SkipTest("Reppy does not support order based directives precedence.") class RerpRobotParserTest(BaseRobotParserTest, unittest.TestCase): @@ -140,3 +147,15 @@ class RerpRobotParserTest(BaseRobotParserTest, unittest.TestCase): def test_length_based_precedence(self): raise unittest.SkipTest("Rerp does not support length based directives precedence.") + + +class ProtegoRobotParserTest(BaseRobotParserTest, unittest.TestCase): + if not protego_available(): + skip = "Protego parser is not installed" + + def setUp(self): + from scrapy.robotstxt import ProtegoRobotParser + super(ProtegoRobotParserTest, self)._setUp(ProtegoRobotParser) + + def test_order_based_precedence(self): + raise unittest.SkipTest("Protego does not support order based directives precedence.") diff --git a/tox.ini b/tox.ini index c3502c2ca..cc845faf1 100644 --- a/tox.ini +++ b/tox.ini @@ -125,6 +125,7 @@ deps = {[testenv:py35]deps} reppy robotexclusionrulesparser + protego [testenv:py27-extra-deps] basepython = python2.7 @@ -132,3 +133,4 @@ deps = {[testenv]deps} reppy robotexclusionrulesparser + protego \ No newline at end of file From ad824a264bfc69ec600c90ca380745b45da3dfe6 Mon Sep 17 00:00:00 2001 From: Anubhav Patel Date: Tue, 27 Aug 2019 18:30:11 +0530 Subject: [PATCH 054/387] fixes a link in doc --- docs/topics/settings.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 0cb81c43e..4cb76412e 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -952,7 +952,7 @@ LOGSTATS_INTERVAL Default: ``60.0`` The interval (in seconds) between each logging printout of the stats -by :class:`~extensions.logstats.LogStats`. +by :class:`~scrapy.extensions.logstats.LogStats`. .. setting:: MEMDEBUG_ENABLED From 77c8ab2e62f25496acfd812381d84aac369fd3b8 Mon Sep 17 00:00:00 2001 From: Anubhav Patel Date: Tue, 27 Aug 2019 18:44:08 +0530 Subject: [PATCH 055/387] makes suggested changes --- docs/topics/downloader-middleware.rst | 21 ++++++--------------- docs/topics/settings.rst | 3 ++- 2 files changed, 8 insertions(+), 16 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index ae413dc84..93c91f18a 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -989,6 +989,12 @@ RobotsTxtMiddleware To make sure Scrapy respects robots.txt make sure the middleware is enabled and the :setting:`ROBOTSTXT_OBEY` setting is enabled. + The :setting:`ROBOTSTXT_USER_AGENT` setting can be used to specify the + user agent string to use for matching in the robots.txt_ file. If it + is ``None``, the User-Agent header you are sending with the request or the + :setting:`USER_AGENT` setting (in that order) will be used for determining + the user agent to use in the robots.txt_ file. + This middleware has to be combined with a robots.txt_ parser. Scrapy ships with support for the following robots.txt_ parsers: @@ -1074,21 +1080,6 @@ implementing the methods described below. .. autoclass:: RobotParser :members: -RobotsTxtMiddleware Settings -~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - -.. setting:: ROBOTSTXT_USER_AGENT - -ROBOTSTXT_USER_AGENT -^^^^^^^^^^^^^^^^^^^^ - -Default: ``None`` - -The user agent string to use for matching in the robots.txt_ file. If ``None``, -the User-Agent header you are sending with the request or the -:setting:`USER_AGENT` setting (in that order) will be used for determining -the user agent to use in the robots.txt_ file. - .. _robots.txt: http://www.robotstxt.org/ DownloaderStats diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index c4b55cc7b..d3ad777f2 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -1422,7 +1422,8 @@ USER_AGENT Default: ``"Scrapy/VERSION (+https://scrapy.org)"`` The default User-Agent to use when crawling, unless overridden. This user agent is -also used in robots.txt if :setting:`ROBOTSTXT_USER_AGENT` setting is ``None`` and +also used by :class:`~scrapy.downloadermiddlewares.robotstxt.RobotsTxtMiddleware` +if :setting:`ROBOTSTXT_USER_AGENT` setting is ``None`` and there is no overridding User-Agent header specified for the request. From b6b76df0574d65e053eb2e0aba5dd757a4907c24 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Wed, 28 Aug 2019 18:28:31 -0300 Subject: [PATCH 056/387] CallbackKeywordArgumentsContract --- scrapy/contracts/__init__.py | 6 +++--- scrapy/contracts/default.py | 16 ++++++++++++++++ scrapy/settings/default_settings.py | 5 +++-- 3 files changed, 22 insertions(+), 5 deletions(-) diff --git a/scrapy/contracts/__init__.py b/scrapy/contracts/__init__.py index 536bbdafb..b3f02a291 100644 --- a/scrapy/contracts/__init__.py +++ b/scrapy/contracts/__init__.py @@ -92,7 +92,7 @@ class ContractsManager(object): @wraps(cb) def cb_wrapper(response): try: - output = cb(response) + output = cb(response, **request.cb_kwargs) output = list(iterate_spider_output(output)) except Exception: case = _create_testcase(method, 'callback') @@ -133,7 +133,7 @@ class Contract(object): else: results.addSuccess(self.testcase_pre) finally: - return list(iterate_spider_output(cb(response))) + return list(iterate_spider_output(cb(response, **request.cb_kwargs))) request.callback = wrapper @@ -145,7 +145,7 @@ class Contract(object): @wraps(cb) def wrapper(response): - output = list(iterate_spider_output(cb(response))) + output = list(iterate_spider_output(cb(response, **request.cb_kwargs))) try: results.startTest(self.testcase_post) self.post_process(output) diff --git a/scrapy/contracts/default.py b/scrapy/contracts/default.py index 7745959a7..24f6c2e77 100644 --- a/scrapy/contracts/default.py +++ b/scrapy/contracts/default.py @@ -1,3 +1,5 @@ +import json + from scrapy.item import BaseItem from scrapy.http import Request from scrapy.exceptions import ContractFail @@ -18,6 +20,20 @@ class UrlContract(Contract): return args +class CallbackKeywordArgumentsContract(Contract): + """ Contract to set the keyword arguments for the request. + The value should be a JSON-encoded dictionary, e.g.: + + @cb_kwargs {"arg1": "some value"} + """ + + name = 'cb_kwargs' + + def adjust_request_args(self, args): + args['cb_kwargs'] = json.loads(' '.join(self.args)) + return args + + class ReturnsContract(Contract): """ Contract to check the output of a callback diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 5884dfc60..52a701edf 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -291,6 +291,7 @@ TELNETCONSOLE_PASSWORD = None SPIDER_CONTRACTS = {} SPIDER_CONTRACTS_BASE = { 'scrapy.contracts.default.UrlContract': 1, - 'scrapy.contracts.default.ReturnsContract': 2, - 'scrapy.contracts.default.ScrapesContract': 3, + 'scrapy.contracts.default.CallbackKeywordArgumentsContract': 2, + 'scrapy.contracts.default.ReturnsContract': 3, + 'scrapy.contracts.default.ScrapesContract': 4, } From 97a7d775f7021e485a88d4becb363c6e49ca3bb4 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 29 Aug 2019 10:51:16 -0300 Subject: [PATCH 057/387] Aplly suggestions by @victor-torres --- scrapy/contracts/__init__.py | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/scrapy/contracts/__init__.py b/scrapy/contracts/__init__.py index b3f02a291..7b6591d86 100644 --- a/scrapy/contracts/__init__.py +++ b/scrapy/contracts/__init__.py @@ -90,9 +90,9 @@ class ContractsManager(object): cb = request.callback @wraps(cb) - def cb_wrapper(response): + def cb_wrapper(response, **cb_kwargs): try: - output = cb(response, **request.cb_kwargs) + output = cb(response, **cb_kwargs) output = list(iterate_spider_output(output)) except Exception: case = _create_testcase(method, 'callback') @@ -121,7 +121,7 @@ class Contract(object): cb = request.callback @wraps(cb) - def wrapper(response): + def wrapper(response, **cb_kwargs): try: results.startTest(self.testcase_pre) self.pre_process(response) @@ -133,7 +133,7 @@ class Contract(object): else: results.addSuccess(self.testcase_pre) finally: - return list(iterate_spider_output(cb(response, **request.cb_kwargs))) + return list(iterate_spider_output(cb(response, **cb_kwargs))) request.callback = wrapper @@ -144,8 +144,8 @@ class Contract(object): cb = request.callback @wraps(cb) - def wrapper(response): - output = list(iterate_spider_output(cb(response, **request.cb_kwargs))) + def wrapper(response, **cb_kwargs): + output = list(iterate_spider_output(cb(response, **cb_kwargs))) try: results.startTest(self.testcase_post) self.post_process(output) From eb0bd2daef8d91e6386390975b2e75d1a93b3161 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 29 Aug 2019 14:01:13 -0300 Subject: [PATCH 058/387] Revert backward-incompatible change (contract priorities) --- scrapy/settings/default_settings.py | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 52a701edf..05ab4b628 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -291,7 +291,7 @@ TELNETCONSOLE_PASSWORD = None SPIDER_CONTRACTS = {} SPIDER_CONTRACTS_BASE = { 'scrapy.contracts.default.UrlContract': 1, - 'scrapy.contracts.default.CallbackKeywordArgumentsContract': 2, - 'scrapy.contracts.default.ReturnsContract': 3, - 'scrapy.contracts.default.ScrapesContract': 4, + 'scrapy.contracts.default.CallbackKeywordArgumentsContract': 1, + 'scrapy.contracts.default.ReturnsContract': 2, + 'scrapy.contracts.default.ScrapesContract': 3, } From ace2df3d140ac783c622ea5403876fe76af2bb72 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Marc=20Hern=C3=A1ndez?= Date: Fri, 30 Aug 2019 11:03:44 +0200 Subject: [PATCH 059/387] Fix JSONRequest naming (#3982) --- docs/news.rst | 4 ++-- docs/topics/request-response.rst | 14 +++++++------- scrapy/http/__init__.py | 2 +- scrapy/http/request/json_request.py | 12 ++++++++---- tests/test_http_request.py | 10 +++++----- 5 files changed, 23 insertions(+), 19 deletions(-) diff --git a/docs/news.rst b/docs/news.rst index 5915ba742..aac750601 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -80,8 +80,8 @@ New features provides a cleaner way to pass keyword arguments to callback methods (:issue:`1138`, :issue:`3563`) -* A new :class:`~scrapy.http.JSONRequest` class offers a more convenient way - to build JSON requests (:issue:`3504`, :issue:`3505`) +* A new :class:`JSONRequest ` class offers a more + convenient way to build JSON requests (:issue:`3504`, :issue:`3505`) * A ``process_request`` callback passed to the :class:`~scrapy.spiders.Rule` constructor now receives the :class:`~scrapy.http.Response` object that diff --git a/docs/topics/request-response.rst b/docs/topics/request-response.rst index 9a5c65b0d..ad1b9af10 100644 --- a/docs/topics/request-response.rst +++ b/docs/topics/request-response.rst @@ -535,19 +535,19 @@ method for this job. Here's an example spider which uses it:: # continue scraping with authenticated session... -JSONRequest +JsonRequest ----------- -The JSONRequest class extends the base :class:`Request` class with functionality for +The JsonRequest class extends the base :class:`Request` class with functionality for dealing with JSON requests. -.. class:: JSONRequest(url, [... data, dumps_kwargs]) +.. class:: JsonRequest(url, [... data, dumps_kwargs]) - The :class:`JSONRequest` class adds two new argument to the constructor. The + The :class:`JsonRequest` class adds two new argument to the constructor. The remaining arguments are the same as for the :class:`Request` class and are not documented here. - Using the :class:`JSONRequest` will set the ``Content-Type`` header to ``application/json`` + Using the :class:`JsonRequest` will set the ``Content-Type`` header to ``application/json`` and ``Accept`` header to ``application/json, text/javascript, */*; q=0.01`` :param data: is any JSON serializable object that needs to be JSON encoded and assigned to body. @@ -562,7 +562,7 @@ dealing with JSON requests. .. _json.dumps: https://docs.python.org/3/library/json.html#json.dumps -JSONRequest usage example +JsonRequest usage example ------------------------- Sending a JSON POST request with a JSON payload:: @@ -571,7 +571,7 @@ Sending a JSON POST request with a JSON payload:: 'name1': 'value1', 'name2': 'value2', } - yield JSONRequest(url='http://www.example.com/post/action', data=data) + yield JsonRequest(url='http://www.example.com/post/action', data=data) Response objects diff --git a/scrapy/http/__init__.py b/scrapy/http/__init__.py index 4b2f7b33f..e6c58e1f1 100644 --- a/scrapy/http/__init__.py +++ b/scrapy/http/__init__.py @@ -10,7 +10,7 @@ from scrapy.http.headers import Headers from scrapy.http.request import Request from scrapy.http.request.form import FormRequest from scrapy.http.request.rpc import XmlRpcRequest -from scrapy.http.request.json_request import JSONRequest +from scrapy.http.request.json_request import JsonRequest from scrapy.http.response import Response from scrapy.http.response.html import HtmlResponse diff --git a/scrapy/http/request/json_request.py b/scrapy/http/request/json_request.py index 8f7a61a6d..f08b25280 100644 --- a/scrapy/http/request/json_request.py +++ b/scrapy/http/request/json_request.py @@ -1,5 +1,5 @@ """ -This module implements the JSONRequest class which is a more convenient class +This module implements the JsonRequest class which is a more convenient class (than Request) to generate JSON Requests. See documentation in docs/topics/request-response.rst @@ -10,9 +10,10 @@ import json import warnings from scrapy.http.request import Request +from scrapy.utils.deprecate import create_deprecated_class -class JSONRequest(Request): +class JsonRequest(Request): def __init__(self, *args, **kwargs): dumps_kwargs = copy.deepcopy(kwargs.pop('dumps_kwargs', {})) dumps_kwargs.setdefault('sort_keys', True) @@ -31,7 +32,7 @@ class JSONRequest(Request): if 'method' not in kwargs: kwargs['method'] = 'POST' - super(JSONRequest, self).__init__(*args, **kwargs) + super(JsonRequest, self).__init__(*args, **kwargs) self.headers.setdefault('Content-Type', 'application/json') self.headers.setdefault('Accept', 'application/json, text/javascript, */*; q=0.01') @@ -46,8 +47,11 @@ class JSONRequest(Request): elif not body_passed and data_passed: kwargs['body'] = self._dumps(data) - return super(JSONRequest, self).replace(*args, **kwargs) + return super(JsonRequest, self).replace(*args, **kwargs) def _dumps(self, data): """Convert to JSON """ return json.dumps(data, **self._dumps_kwargs) + + +JSONRequest = create_deprecated_class("JSONRequest", JsonRequest) diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 60494d792..16d7a1cb8 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -11,7 +11,7 @@ from six.moves.urllib.parse import urlparse, parse_qs, unquote if six.PY3: from urllib.parse import unquote_to_bytes -from scrapy.http import Request, FormRequest, XmlRpcRequest, JSONRequest, Headers, HtmlResponse +from scrapy.http import Request, FormRequest, XmlRpcRequest, JsonRequest, Headers, HtmlResponse from scrapy.utils.python import to_bytes, to_native_str from tests import mock @@ -1246,14 +1246,14 @@ class XmlRpcRequestTest(RequestTest): self._test_request(params=(u'pas£',), encoding='latin1') -class JSONRequestTest(RequestTest): - request_class = JSONRequest +class JsonRequestTest(RequestTest): + request_class = JsonRequest default_method = 'GET' default_headers = {b'Content-Type': [b'application/json'], b'Accept': [b'application/json, text/javascript, */*; q=0.01']} def setUp(self): warnings.simplefilter("always") - super(JSONRequestTest, self).setUp() + super(JsonRequestTest, self).setUp() def test_data(self): r1 = self.request_class(url="http://www.example.com/") @@ -1407,7 +1407,7 @@ class JSONRequestTest(RequestTest): def tearDown(self): warnings.resetwarnings() - super(JSONRequestTest, self).tearDown() + super(JsonRequestTest, self).tearDown() if __name__ == "__main__": From 2828cb769f0618d3986cd12c2ef0f8143f31bfba Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Fri, 30 Aug 2019 14:29:15 +0200 Subject: [PATCH 060/387] Provide complete API documentation coverage of scrapy.extensions --- docs/conf.py | 6 + docs/topics/downloader-middleware.rst | 153 +++++++++++++------------- docs/topics/extensions.rst | 14 +-- 3 files changed, 90 insertions(+), 83 deletions(-) diff --git a/docs/conf.py b/docs/conf.py index eba416cd6..fa257dead 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -256,6 +256,12 @@ coverage_ignore_pyobjects = [ # Methods of BaseItemExporter subclasses are only documented in # BaseItemExporter. r'^scrapy\.exporters\.(?!BaseItemExporter\b)\w*?\.', + + # Extension behavior is only modified through settings. Methods of + # extension classes, as well as helper functions, are implementation + # details that are not documented. + r'^scrapy\.extensions\.[a-z]\w*?\.[A-Z]\w*?\.', # methods + r'^scrapy\.extensions\.[a-z]\w*?\.[a-z]', # helper functions ] diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index c25e12bda..0845ef6e4 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -365,24 +365,25 @@ HttpCacheMiddleware You can also avoid caching a response on every policy using :reqmeta:`dont_cache` meta key equals ``True``. +.. module:: scrapy.extensions.httpcache + :noindex: + .. _httpcache-policy-dummy: Dummy policy (default) ~~~~~~~~~~~~~~~~~~~~~~ -This policy has no awareness of any HTTP Cache-Control directives. -Every request and its corresponding response are cached. When the same -request is seen again, the response is returned without transferring -anything from the Internet. +.. class:: DummyPolicy -The Dummy policy is useful for testing spiders faster (without having -to wait for downloads every time) and for trying your spider offline, -when an Internet connection is not available. The goal is to be able to -"replay" a spider run *exactly as it ran before*. + This policy has no awareness of any HTTP Cache-Control directives. + Every request and its corresponding response are cached. When the same + request is seen again, the response is returned without transferring + anything from the Internet. -In order to use this policy, set: - -* :setting:`HTTPCACHE_POLICY` to ``scrapy.extensions.httpcache.DummyPolicy`` + The Dummy policy is useful for testing spiders faster (without having + to wait for downloads every time) and for trying your spider offline, + when an Internet connection is not available. The goal is to be able to + "replay" a spider run *exactly as it ran before*. .. _httpcache-policy-rfc2616: @@ -390,45 +391,44 @@ In order to use this policy, set: RFC2616 policy ~~~~~~~~~~~~~~ -This policy provides a RFC2616 compliant HTTP cache, i.e. with HTTP -Cache-Control awareness, aimed at production and used in continuous -runs to avoid downloading unmodified data (to save bandwidth and speed up crawls). +.. class:: RFC2616Policy -what is implemented: + This policy provides a RFC2616 compliant HTTP cache, i.e. with HTTP + Cache-Control awareness, aimed at production and used in continuous + runs to avoid downloading unmodified data (to save bandwidth and speed up + crawls). -* Do not attempt to store responses/requests with ``no-store`` cache-control directive set -* Do not serve responses from cache if ``no-cache`` cache-control directive is set even for fresh responses -* Compute freshness lifetime from ``max-age`` cache-control directive -* Compute freshness lifetime from ``Expires`` response header -* Compute freshness lifetime from ``Last-Modified`` response header (heuristic used by Firefox) -* Compute current age from ``Age`` response header -* Compute current age from ``Date`` header -* Revalidate stale responses based on ``Last-Modified`` response header -* Revalidate stale responses based on ``ETag`` response header -* Set ``Date`` header for any received response missing it -* Support ``max-stale`` cache-control directive in requests + What is implemented: - This allows spiders to be configured with the full RFC2616 cache policy, - but avoid revalidation on a request-by-request basis, while remaining - conformant with the HTTP spec. + * Do not attempt to store responses/requests with ``no-store`` cache-control directive set + * Do not serve responses from cache if ``no-cache`` cache-control directive is set even for fresh responses + * Compute freshness lifetime from ``max-age`` cache-control directive + * Compute freshness lifetime from ``Expires`` response header + * Compute freshness lifetime from ``Last-Modified`` response header (heuristic used by Firefox) + * Compute current age from ``Age`` response header + * Compute current age from ``Date`` header + * Revalidate stale responses based on ``Last-Modified`` response header + * Revalidate stale responses based on ``ETag`` response header + * Set ``Date`` header for any received response missing it + * Support ``max-stale`` cache-control directive in requests - Example: + This allows spiders to be configured with the full RFC2616 cache policy, + but avoid revalidation on a request-by-request basis, while remaining + conformant with the HTTP spec. - Add ``Cache-Control: max-stale=600`` to Request headers to accept responses that - have exceeded their expiration time by no more than 600 seconds. + Example: - See also: RFC2616, 14.9.3 + Add ``Cache-Control: max-stale=600`` to Request headers to accept responses that + have exceeded their expiration time by no more than 600 seconds. -what is missing: + See also: RFC2616, 14.9.3 -* ``Pragma: no-cache`` support https://www.w3.org/Protocols/rfc2616/rfc2616-sec14.html#sec14.9.1 -* ``Vary`` header support https://www.w3.org/Protocols/rfc2616/rfc2616-sec13.html#sec13.6 -* Invalidation after updates or deletes https://www.w3.org/Protocols/rfc2616/rfc2616-sec13.html#sec13.10 -* ... probably others .. + What is missing: -In order to use this policy, set: - -* :setting:`HTTPCACHE_POLICY` to ``scrapy.extensions.httpcache.RFC2616Policy`` + * ``Pragma: no-cache`` support https://www.w3.org/Protocols/rfc2616/rfc2616-sec14.html#sec14.9.1 + * ``Vary`` header support https://www.w3.org/Protocols/rfc2616/rfc2616-sec13.html#sec13.6 + * Invalidation after updates or deletes https://www.w3.org/Protocols/rfc2616/rfc2616-sec13.html#sec13.10 + * ... probably others .. .. _httpcache-storage-fs: @@ -436,67 +436,68 @@ In order to use this policy, set: Filesystem storage backend (default) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -File system storage backend is available for the HTTP cache middleware. +.. class:: FilesystemCacheStorage -In order to use this storage backend, set: + File system storage backend is available for the HTTP cache middleware. -* :setting:`HTTPCACHE_STORAGE` to ``scrapy.extensions.httpcache.FilesystemCacheStorage`` + Each request/response pair is stored in a different directory containing + the following files: -Each request/response pair is stored in a different directory containing -the following files: + * ``request_body`` - the plain request body - * ``request_body`` - the plain request body - * ``request_headers`` - the request headers (in raw HTTP format) - * ``response_body`` - the plain response body - * ``response_headers`` - the request headers (in raw HTTP format) - * ``meta`` - some metadata of this cache resource in Python ``repr()`` format - (grep-friendly format) - * ``pickled_meta`` - the same metadata in ``meta`` but pickled for more - efficient deserialization + * ``request_headers`` - the request headers (in raw HTTP format) -The directory name is made from the request fingerprint (see -``scrapy.utils.request.fingerprint``), and one level of subdirectories is -used to avoid creating too many files into the same directory (which is -inefficient in many file systems). An example directory could be:: + * ``response_body`` - the plain response body - /path/to/cache/dir/example.com/72/72811f648e718090f041317756c03adb0ada46c7 + * ``response_headers`` - the request headers (in raw HTTP format) + + * ``meta`` - some metadata of this cache resource in Python ``repr()`` + format (grep-friendly format) + + * ``pickled_meta`` - the same metadata in ``meta`` but pickled for more + efficient deserialization + + The directory name is made from the request fingerprint (see + ``scrapy.utils.request.fingerprint``), and one level of subdirectories is + used to avoid creating too many files into the same directory (which is + inefficient in many file systems). An example directory could be:: + + /path/to/cache/dir/example.com/72/72811f648e718090f041317756c03adb0ada46c7 .. _httpcache-storage-dbm: DBM storage backend ~~~~~~~~~~~~~~~~~~~ -.. versionadded:: 0.13 +.. class:: DbmCacheStorage -A DBM_ storage backend is also available for the HTTP cache middleware. + .. versionadded:: 0.13 -By default, it uses the anydbm_ module, but you can change it with the -:setting:`HTTPCACHE_DBM_MODULE` setting. + A DBM_ storage backend is also available for the HTTP cache middleware. -In order to use this storage backend, set: - -* :setting:`HTTPCACHE_STORAGE` to ``scrapy.extensions.httpcache.DbmCacheStorage`` + By default, it uses the anydbm_ module, but you can change it with the + :setting:`HTTPCACHE_DBM_MODULE` setting. .. _httpcache-storage-leveldb: LevelDB storage backend ~~~~~~~~~~~~~~~~~~~~~~~ -.. versionadded:: 0.23 +.. class:: LeveldbCacheStorage -A LevelDB_ storage backend is also available for the HTTP cache middleware. + .. versionadded:: 0.23 -This backend is not recommended for development because only one process can -access LevelDB databases at the same time, so you can't run a crawl and open -the scrapy shell in parallel for the same spider. + A LevelDB_ storage backend is also available for the HTTP cache middleware. -In order to use this storage backend: + This backend is not recommended for development because only one process + can access LevelDB databases at the same time, so you can't run a crawl and + open the scrapy shell in parallel for the same spider. -* set :setting:`HTTPCACHE_STORAGE` to ``scrapy.extensions.httpcache.LeveldbCacheStorage`` -* install `LevelDB python bindings`_ like ``pip install leveldb`` + In order to use this storage backend, install the `LevelDB python + bindings`_ (e.g. ``pip install leveldb``). -.. _LevelDB: https://github.com/google/leveldb -.. _leveldb python bindings: https://pypi.python.org/pypi/leveldb + .. _LevelDB: https://github.com/google/leveldb + .. _leveldb python bindings: https://pypi.python.org/pypi/leveldb .. _httpcache-storage-custom: diff --git a/docs/topics/extensions.rst b/docs/topics/extensions.rst index d6e7452a1..72c2290b5 100644 --- a/docs/topics/extensions.rst +++ b/docs/topics/extensions.rst @@ -183,7 +183,7 @@ Telnet console extension .. module:: scrapy.extensions.telnet :synopsis: Telnet console -.. class:: scrapy.extensions.telnet.TelnetConsole +.. class:: TelnetConsole Provides a telnet console for getting into a Python interpreter inside the currently running Scrapy process, which can be very useful for debugging. @@ -200,7 +200,7 @@ Memory usage extension .. module:: scrapy.extensions.memusage :synopsis: Memory usage extension -.. class:: scrapy.extensions.memusage.MemoryUsage +.. class:: MemoryUsage .. note:: This extension does not work in Windows. @@ -228,7 +228,7 @@ Memory debugger extension .. module:: scrapy.extensions.memdebug :synopsis: Memory debugger extension -.. class:: scrapy.extensions.memdebug.MemoryDebugger +.. class:: MemoryDebugger An extension for debugging memory usage. It collects information about: @@ -244,7 +244,7 @@ Close spider extension .. module:: scrapy.extensions.closespider :synopsis: Close spider extension -.. class:: scrapy.extensions.closespider.CloseSpider +.. class:: CloseSpider Closes a spider automatically when some conditions are met, using a specific closing reason for each condition. @@ -317,7 +317,7 @@ StatsMailer extension .. module:: scrapy.extensions.statsmailer :synopsis: StatsMailer extension -.. class:: scrapy.extensions.statsmailer.StatsMailer +.. class:: StatsMailer This simple extension can be used to send a notification e-mail every time a domain has finished scraping, including the Scrapy stats collected. The email @@ -333,7 +333,7 @@ Debugging extensions Stack trace dump extension ~~~~~~~~~~~~~~~~~~~~~~~~~~ -.. class:: scrapy.extensions.debug.StackTraceDump +.. class:: StackTraceDump Dumps information about the running process when a `SIGQUIT`_ or `SIGUSR2`_ signal is received. The information dumped is the following: @@ -362,7 +362,7 @@ There are at least two ways to send Scrapy the `SIGQUIT`_ signal: Debugger extension ~~~~~~~~~~~~~~~~~~ -.. class:: scrapy.extensions.debug.Debugger +.. class:: Debugger Invokes a `Python debugger`_ inside a running Scrapy process when a `SIGUSR2`_ signal is received. After the debugger is exited, the Scrapy process continues From 2061f2a382b136382e18d6302fcd53053654e1ce Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Sat, 31 Aug 2019 02:10:18 -0300 Subject: [PATCH 061/387] [doc] cb_kwargs contract --- docs/topics/contracts.rst | 10 +++++++++- 1 file changed, 9 insertions(+), 1 deletion(-) diff --git a/docs/topics/contracts.rst b/docs/topics/contracts.rst index 957761b76..62f9a743b 100644 --- a/docs/topics/contracts.rst +++ b/docs/topics/contracts.rst @@ -35,12 +35,20 @@ This callback is tested using three built-in contracts: .. class:: UrlContract - This contract (``@url``) sets the sample url used when checking other + This contract (``@url``) sets the sample URL used when checking other contract conditions for this spider. This contract is mandatory. All callbacks lacking this contract are ignored when running the checks:: @url url +.. class:: CallbackKeywordArgumentsContract + + This contract (``@cb_kwargs``) sets the :attr:`cb_kwargs ` + attribute for the sample request. It must be a valid JSON dictionary. + :: + + @cb_kwargs {"arg1": "value1", "arg2": "value2", ...} + .. class:: ReturnsContract This contract (``@returns``) sets lower and upper bounds for the items and From b92b1146335f6b651ba8e6e56e579a6844ec3ebd Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Sat, 31 Aug 2019 02:44:09 -0300 Subject: [PATCH 062/387] [test] cb_kwargs contract --- tests/test_contracts.py | 78 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 78 insertions(+) diff --git a/tests/test_contracts.py b/tests/test_contracts.py index a728099c0..b2e358700 100644 --- a/tests/test_contracts.py +++ b/tests/test_contracts.py @@ -14,6 +14,7 @@ from scrapy.item import Item, Field from scrapy.contracts import ContractsManager, Contract from scrapy.contracts.default import ( UrlContract, + CallbackKeywordArgumentsContract, ReturnsContract, ScrapesContract, ) @@ -70,6 +71,37 @@ class TestSpider(Spider): """ return TestItem(url=response.url) + def returns_request_cb_kwargs(self, response, url): + """ method which returns request + @url https://example.org + @cb_kwargs {"url": "http://scrapy.org"} + @returns requests 1 + """ + return Request(url, callback=self.returns_item_cb_kwargs) + + def returns_item_cb_kwargs(self, response, name): + """ method which returns item + @url http://scrapy.org + @cb_kwargs {"name": "Scrapy"} + @returns items 1 1 + """ + return TestItem(name=name, url=response.url) + + def returns_item_cb_kwargs_error_unexpected_keyword(self, response): + """ method which returns item + @url http://scrapy.org + @cb_kwargs {"arg": "value"} + @returns items 1 1 + """ + return TestItem(url=response.url) + + def returns_item_cb_kwargs_error_missing_argument(self, response, arg): + """ method which returns item + @url http://scrapy.org + @returns items 1 1 + """ + return TestItem(url=response.url) + def returns_dict_item(self, response): """ method which returns item @url http://scrapy.org @@ -172,6 +204,7 @@ class InheritsTestSpider(TestSpider): class ContractsManagerTest(unittest.TestCase): contracts = [ UrlContract, + CallbackKeywordArgumentsContract, ReturnsContract, ScrapesContract, CustomFormContract, @@ -211,6 +244,51 @@ class ContractsManagerTest(unittest.TestCase): request = self.conman.from_method(spider.parse_no_url, self.results) self.assertEqual(request, None) + def test_cb_kwargs(self): + spider = TestSpider() + response = ResponseMock() + + # extract contracts correctly + contracts = self.conman.extract_contracts(spider.returns_request_cb_kwargs) + self.assertEqual(len(contracts), 3) + self.assertEqual(frozenset(type(x) for x in contracts), + frozenset([UrlContract, CallbackKeywordArgumentsContract, ReturnsContract])) + + contracts = self.conman.extract_contracts(spider.returns_item_cb_kwargs) + self.assertEqual(len(contracts), 3) + self.assertEqual(frozenset(type(x) for x in contracts), + frozenset([UrlContract, CallbackKeywordArgumentsContract, ReturnsContract])) + + contracts = self.conman.extract_contracts(spider.returns_item_cb_kwargs_error_unexpected_keyword) + self.assertEqual(len(contracts), 3) + self.assertEqual(frozenset(type(x) for x in contracts), + frozenset([UrlContract, CallbackKeywordArgumentsContract, ReturnsContract])) + + contracts = self.conman.extract_contracts(spider.returns_item_cb_kwargs_error_missing_argument) + self.assertEqual(len(contracts), 2) + self.assertEqual(frozenset(type(x) for x in contracts), + frozenset([UrlContract, ReturnsContract])) + + # returns_request + request = self.conman.from_method(spider.returns_request_cb_kwargs, self.results) + request.callback(response, **request.cb_kwargs) + self.should_succeed() + + # returns_item + request = self.conman.from_method(spider.returns_item_cb_kwargs, self.results) + request.callback(response, **request.cb_kwargs) + self.should_succeed() + + # returns_item (error, callback doesn't take keyword arguments) + request = self.conman.from_method(spider.returns_item_cb_kwargs_error_unexpected_keyword, self.results) + request.callback(response, **request.cb_kwargs) + self.should_error() + + # returns_item (error, contract doesn't provide keyword arguments) + request = self.conman.from_method(spider.returns_item_cb_kwargs_error_missing_argument, self.results) + request.callback(response, **request.cb_kwargs) + self.should_error() + def test_returns(self): spider = TestSpider() response = ResponseMock() From 9578f490991fbefc4ad643588c87862a57a9f032 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Mon, 9 Sep 2019 07:36:55 +0000 Subject: [PATCH 063/387] use protego as a default robots.txt parser --- docs/topics/downloader-middleware.rst | 17 ++++++++--------- requirements-py2.txt | 1 + requirements-py3.txt | 1 + scrapy/settings/default_settings.py | 2 +- tox.ini | 2 -- 5 files changed, 11 insertions(+), 12 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 0845ef6e4..be96425cf 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -513,7 +513,7 @@ defines the methods described below. .. method:: open_spider(spider) - This method gets called after a spider has been opened for crawling. It handles + This method gets called after a spider has been opened for crawling. It handles the :signal:`open_spider ` signal. :param spider: the spider which has been opened @@ -521,8 +521,8 @@ defines the methods described below. .. method:: close_spider(spider) - This method gets called after a spider has been closed. It handles - the :signal:`close_spider ` signal. + This method gets called after a spider has been closed. It handles + the :signal:`close_spider ` signal. :param spider: the spider which has been closed :type spider: :class:`~scrapy.spiders.Spider` object @@ -1020,10 +1020,10 @@ the request will be ignored by this middleware even if RobotFileParser ~~~~~~~~~~~~~~~ -`RobotFileParser `_ is -Python's inbuilt robots.txt_ parser. The parser is fully compliant with `Martijn Koster's +`RobotFileParser `_ is +Python's inbuilt robots.txt_ parser. The parser is fully compliant with `Martijn Koster's 1996 draft specification `_. It lacks -support for wildcard matching. Scrapy uses this parser by default. +support for wildcard matching. In order to use this parser, set: @@ -1074,13 +1074,12 @@ Protego parser `Protego `_ is a pure-Python robots.txt_ parser. The parser is fully compliant with `Google's Robots.txt Specification -`_ hence supports wildcard +`_ hence supports wildcard matching, and uses the length based rule similar to `Reppy `_. +Scrapy uses this parser by default. In order to use this parser: -* Install `Protego `_ by running ``pip install protego`` - * Set :setting:`ROBOTSTXT_PARSER` setting to ``scrapy.robotstxt.ProtegoRobotParser`` diff --git a/requirements-py2.txt b/requirements-py2.txt index 9e6944240..61176bdba 100644 --- a/requirements-py2.txt +++ b/requirements-py2.txt @@ -15,3 +15,4 @@ service_identity>=16.0.0 six>=1.10.0 Twisted>=16.0.0 zope.interface>=4.1.3 +protego diff --git a/requirements-py3.txt b/requirements-py3.txt index cd183a525..61e2e32d8 100644 --- a/requirements-py3.txt +++ b/requirements-py3.txt @@ -15,3 +15,4 @@ lxml>=3.5.0 service_identity>=16.0.0 six>=1.10.0 zope.interface>=4.1.3 +protego diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 05ab4b628..9c22999cb 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -246,7 +246,7 @@ RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408, 429] RETRY_PRIORITY_ADJUST = -1 ROBOTSTXT_OBEY = False -ROBOTSTXT_PARSER = 'scrapy.robotstxt.PythonRobotParser' +ROBOTSTXT_PARSER = 'scrapy.robotstxt.ProtegoRobotParser' ROBOTSTXT_USER_AGENT = None SCHEDULER = 'scrapy.core.scheduler.Scheduler' diff --git a/tox.ini b/tox.ini index cc845faf1..c3502c2ca 100644 --- a/tox.ini +++ b/tox.ini @@ -125,7 +125,6 @@ deps = {[testenv:py35]deps} reppy robotexclusionrulesparser - protego [testenv:py27-extra-deps] basepython = python2.7 @@ -133,4 +132,3 @@ deps = {[testenv]deps} reppy robotexclusionrulesparser - protego \ No newline at end of file From 7af8c76649caf772627a189bb2d88c2d8fd620a2 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Mon, 9 Sep 2019 08:10:09 +0000 Subject: [PATCH 064/387] add pinned versions --- tox.ini | 2 ++ 1 file changed, 2 insertions(+) diff --git a/tox.ini b/tox.ini index c3502c2ca..fdd227d02 100644 --- a/tox.ini +++ b/tox.ini @@ -33,6 +33,7 @@ deps = cssselect==0.9.1 lxml==3.5.0 parsel==1.5.0 + Protego=0.1.15 PyDispatcher==2.0.5 pyOpenSSL==16.2.0 queuelib==1.4.2 @@ -69,6 +70,7 @@ deps = cssselect==0.9.1 lxml==3.5.0 parsel==1.5.0 + Protego=0.1.15 PyDispatcher==2.0.5 pyOpenSSL==16.2.0 queuelib==1.4.2 From e418554c21cdd9da87b914472e32668cc41d7e87 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Mon, 9 Sep 2019 08:12:32 +0000 Subject: [PATCH 065/387] use proper equal --- tox.ini | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/tox.ini b/tox.ini index fdd227d02..ffe7360d3 100644 --- a/tox.ini +++ b/tox.ini @@ -33,7 +33,7 @@ deps = cssselect==0.9.1 lxml==3.5.0 parsel==1.5.0 - Protego=0.1.15 + Protego==0.1.15 PyDispatcher==2.0.5 pyOpenSSL==16.2.0 queuelib==1.4.2 @@ -70,7 +70,7 @@ deps = cssselect==0.9.1 lxml==3.5.0 parsel==1.5.0 - Protego=0.1.15 + Protego==0.1.15 PyDispatcher==2.0.5 pyOpenSSL==16.2.0 queuelib==1.4.2 From 38828d3fd45e7846ef3556d384e4c1fdc7a0c8ba Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Mon, 9 Sep 2019 17:04:13 +0300 Subject: [PATCH 066/387] Update docs/topics/downloader-middleware.rst Co-Authored-By: elacuesta --- docs/topics/downloader-middleware.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index be96425cf..192bfd19a 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -1021,7 +1021,7 @@ RobotFileParser ~~~~~~~~~~~~~~~ `RobotFileParser `_ is -Python's inbuilt robots.txt_ parser. The parser is fully compliant with `Martijn Koster's +Python's built-in robots.txt_ parser. The parser is fully compliant with `Martijn Koster's 1996 draft specification `_. It lacks support for wildcard matching. From 7b33fa58fa46ae7fb96cb40d82ecf98eafcbd49d Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Mon, 9 Sep 2019 17:04:27 +0300 Subject: [PATCH 067/387] Update requirements-py2.txt Co-Authored-By: elacuesta --- requirements-py2.txt | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/requirements-py2.txt b/requirements-py2.txt index 61176bdba..c865cbaef 100644 --- a/requirements-py2.txt +++ b/requirements-py2.txt @@ -15,4 +15,4 @@ service_identity>=16.0.0 six>=1.10.0 Twisted>=16.0.0 zope.interface>=4.1.3 -protego +protego>=0.1.15 From db202487f06a1710b32d991220dbd6656da8b2a0 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Mon, 9 Sep 2019 14:05:45 +0000 Subject: [PATCH 068/387] newer version of protego and move up to top --- requirements-py2.txt | 2 +- requirements-py3.txt | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/requirements-py2.txt b/requirements-py2.txt index c865cbaef..dde8d1c9c 100644 --- a/requirements-py2.txt +++ b/requirements-py2.txt @@ -1,6 +1,7 @@ parsel>=1.5.0 PyDispatcher>=2.0.5 w3lib>=1.17.0 +protego>=0.1.15 pyOpenSSL>=16.2.0 # Earlier versions fail with "AttributeError: module 'lib' has no attribute 'SSL_ST_INIT'" queuelib>=1.4.2 # Earlier versions fail with "AttributeError: '...QueueTest' object has no attribute 'qpath'" @@ -15,4 +16,3 @@ service_identity>=16.0.0 six>=1.10.0 Twisted>=16.0.0 zope.interface>=4.1.3 -protego>=0.1.15 diff --git a/requirements-py3.txt b/requirements-py3.txt index 61e2e32d8..2c98e6f6d 100644 --- a/requirements-py3.txt +++ b/requirements-py3.txt @@ -2,6 +2,7 @@ parsel>=1.5.0 PyDispatcher>=2.0.5 Twisted>=17.9.0 w3lib>=1.17.0 +protego>=0.1.15 pyOpenSSL>=16.2.0 # Earlier versions fail with "AttributeError: module 'lib' has no attribute 'SSL_ST_INIT'" queuelib>=1.4.2 # Earlier versions fail with "AttributeError: '...QueueTest' object has no attribute 'qpath'" @@ -15,4 +16,3 @@ lxml>=3.5.0 service_identity>=16.0.0 six>=1.10.0 zope.interface>=4.1.3 -protego From 6bd88711f2f5946bce6c72d46cd8a5ad3ce4ce86 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Tue, 10 Sep 2019 08:55:37 +0000 Subject: [PATCH 069/387] update documentation --- docs/topics/settings.rst | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 943ba13ee..75e0af63b 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -951,7 +951,7 @@ LOGSTATS_INTERVAL Default: ``60.0`` -The interval (in seconds) between each logging printout of the stats +The interval (in seconds) between each logging printout of the stats by :class:`~scrapy.extensions.logstats.LogStats`. .. setting:: MEMDEBUG_ENABLED @@ -1165,7 +1165,7 @@ If enabled, Scrapy will respect robots.txt policies. For more information see ROBOTSTXT_PARSER ---------------- -Default: ``'scrapy.robotstxt.PythonRobotParser'`` +Default: ``'scrapy.robotstxt.ProtegoRobotParser'`` The parser backend to use for parsing ``robots.txt`` files. For more information see :ref:`topics-dlmw-robots`. From c7f2bdfdbed0e82a385b05ea080768d38d6626cd Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Tue, 10 Sep 2019 08:58:52 +0000 Subject: [PATCH 070/387] add protego to install_requires --- setup.py | 1 + 1 file changed, 1 insertion(+) diff --git a/setup.py b/setup.py index 37892cfbf..850456503 100644 --- a/setup.py +++ b/setup.py @@ -77,6 +77,7 @@ setup( 'six>=1.10.0', 'w3lib>=1.17.0', 'zope.interface>=4.1.3', + 'protego>=0.1.15', ], extras_require=extras_require, ) From 171fa1cd106f5a4ba34e6d57bdd65c8688564b82 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Tue, 10 Sep 2019 09:59:36 +0000 Subject: [PATCH 071/387] documentation rework --- docs/topics/downloader-middleware.rst | 116 +++++++++++++++++--------- 1 file changed, 76 insertions(+), 40 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 192bfd19a..52be8ded2 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -1000,10 +1000,10 @@ RobotsTxtMiddleware Scrapy ships with support for the following robots.txt_ parsers: - * :ref:`RobotFileParser ` (default) + * :ref:`Protego ` (default) + * :ref:`RobotFileParser ` * :ref:`Reppy ` * :ref:`Robotexclusionrulesparser ` - * :ref:`Protego ` You can change the robots.txt_ parser with the :setting:`ROBOTSTXT_PARSER` setting. Or you can also :ref:`implement support for a new parser `. @@ -1015,50 +1015,78 @@ If :attr:`Request.meta ` has the request will be ignored by this middleware even if :setting:`ROBOTSTXT_OBEY` is enabled. +Parsers varies in several aspects: + +* Language of implementation + +* Supported specification + +* Support for wildcard matching + +* usage of length based rule: in particular for ``Allow`` and + ``Disallow`` directives, where the most specific rule based on the length of + the path trumps the less specific (shorter) rule + + +.. _protego-parser: + +Protego parser +~~~~~~~~~~~~~~ + +based on `Protego `_: + +* implemented in Python + +* is compliant with `Google's Robots.txt Specification + `_ + +* supports wildcard matching + +* uses the length based rule, + +Scrapy uses this parser by default. + .. _python-robotfileparser: RobotFileParser ~~~~~~~~~~~~~~~ -`RobotFileParser `_ is -Python's built-in robots.txt_ parser. The parser is fully compliant with `Martijn Koster's -1996 draft specification `_. It lacks -support for wildcard matching. +based on `RobotFileParser +`_: + +* is Python's built-in robots.txt_ parser. + +* is compliant with `Martijn Koster's 1996 draft specification + `_. + +* lacks support for wildcard matching. + +* doesn't use the length based rule, + +It is faster than Protego and backward-compatible with versions of Scrapy before 1.8.0 . In order to use this parser, set: * :setting:`ROBOTSTXT_PARSER` to ``scrapy.robotstxt.PythonRobotParser`` -.. _rerp-parser: - -Robotexclusionrulesparser -~~~~~~~~~~~~~~~~~~~~~~~~~ - -`Robotexclusionrulesparser `_ is fully compliant -with `Martijn Koster's 1996 draft specification `_, -with support for wildcard matching. - -In order to use this parser: - -* Install `Robotexclusionrulesparser `_ by running - ``pip install robotexclusionrulesparser`` - -* Set :setting:`ROBOTSTXT_PARSER` setting to - ``scrapy.robotstxt.RerpRobotParser`` - .. _reppy-parser: Reppy parser ~~~~~~~~~~~~ -`Reppy `_ is a Python wrapper around `Robots Exclusion -Protocol Parser for C++ `_. The parser is fully compliant -with `Martijn Koster's 1996 draft specification `_, -with support for wildcard matching. Unlike -`RobotFileParser `_ and -`Robotexclusionrulesparser `_, it uses the length based -rule, in particular for ``Allow`` and ``Disallow`` directives, where the most specific -rule based on the length of the path trumps the less specific (shorter) rule. +based on `Reppy `_: + +* is a Python wrapper around `Robots Exclusion Protocol Parser for C++ + `_. + +* is compliant with `Martijn Koster's 1996 draft specification + `_. + +* supports wildcard matching + +* uses the length based rule, + +Native implementation provides better speed than Protego. In order to use this parser: @@ -1067,21 +1095,29 @@ In order to use this parser: * Set :setting:`ROBOTSTXT_PARSER` setting to ``scrapy.robotstxt.ReppyRobotParser`` -.. _protego-parser: +.. _rerp-parser: -Protego parser -~~~~~~~~~~~~~~ +Robotexclusionrulesparser +~~~~~~~~~~~~~~~~~~~~~~~~~ -`Protego `_ is a pure-Python robots.txt_ parser. -The parser is fully compliant with `Google's Robots.txt Specification -`_ hence supports wildcard -matching, and uses the length based rule similar to `Reppy `_. -Scrapy uses this parser by default. +based on `Robotexclusionrulesparser `_: + +* implemented in Python + +* is compliant with `Martijn Koster's 1996 draft specification + `_. + +* supports wildcard matching + +* doesn't use the length based rule, In order to use this parser: +* Install `Robotexclusionrulesparser `_ by running + ``pip install robotexclusionrulesparser`` + * Set :setting:`ROBOTSTXT_PARSER` setting to - ``scrapy.robotstxt.ProtegoRobotParser`` + ``scrapy.robotstxt.RerpRobotParser`` .. _support-for-new-robots-parser: From 66145b4eaf813929c45c17624d3cb4645b2e3ba0 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Thu, 12 Sep 2019 18:51:00 +0300 Subject: [PATCH 072/387] Update docs/topics/downloader-middleware.rst Co-Authored-By: Mikhail Korobov --- docs/topics/downloader-middleware.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 52be8ded2..76ee77a35 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -1015,7 +1015,7 @@ If :attr:`Request.meta ` has the request will be ignored by this middleware even if :setting:`ROBOTSTXT_OBEY` is enabled. -Parsers varies in several aspects: +Parsers vary in several aspects: * Language of implementation From c5612f387bcf34dc47044b81f802804a719e858f Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Fri, 13 Sep 2019 14:21:09 -0300 Subject: [PATCH 073/387] Remove deprecated xlib module --- .coveragerc | 1 - conftest.py | 10 ---------- debian/scrapy.lintian-overrides | 1 - scrapy/xlib/__init__.py | 2 -- scrapy/xlib/pydispatch.py | 19 ------------------- scrapy/xlib/tx.py | 19 ------------------- tests/test_pydispatch_deprecated.py | 12 ------------ 7 files changed, 64 deletions(-) delete mode 100644 scrapy/xlib/__init__.py delete mode 100644 scrapy/xlib/pydispatch.py delete mode 100644 scrapy/xlib/tx.py delete mode 100644 tests/test_pydispatch_deprecated.py diff --git a/.coveragerc b/.coveragerc index 914d697a0..02acbff8e 100644 --- a/.coveragerc +++ b/.coveragerc @@ -3,4 +3,3 @@ branch = true include = scrapy/* omit = tests/* - scrapy/xlib/* diff --git a/conftest.py b/conftest.py index d8531d6cc..06d65ba1d 100644 --- a/conftest.py +++ b/conftest.py @@ -1,22 +1,12 @@ -import glob import six import pytest -from twisted import version as twisted_version - - -def _py_files(folder): - return glob.glob(folder + "/*.py") + glob.glob(folder + "/*/*.py") collect_ignore = [ # not a test, but looks like a test "scrapy/utils/testsite.py", - ] -if (twisted_version.major, twisted_version.minor, twisted_version.micro) >= (15, 5, 0): - collect_ignore += _py_files("scrapy/xlib/tx") - if six.PY3: for line in open('tests/py3-ignores.txt'): diff --git a/debian/scrapy.lintian-overrides b/debian/scrapy.lintian-overrides index 955e7def0..b5de7f67d 100644 --- a/debian/scrapy.lintian-overrides +++ b/debian/scrapy.lintian-overrides @@ -1,2 +1 @@ new-package-should-close-itp-bug -extra-license-file usr/share/pyshared/scrapy/xlib/pydispatch/license.txt diff --git a/scrapy/xlib/__init__.py b/scrapy/xlib/__init__.py deleted file mode 100644 index 11f022087..000000000 --- a/scrapy/xlib/__init__.py +++ /dev/null @@ -1,2 +0,0 @@ -"""This package contains some third party modules that are distributed along -with Scrapy""" diff --git a/scrapy/xlib/pydispatch.py b/scrapy/xlib/pydispatch.py deleted file mode 100644 index 5ffeaf579..000000000 --- a/scrapy/xlib/pydispatch.py +++ /dev/null @@ -1,19 +0,0 @@ -from __future__ import absolute_import - -import warnings -from scrapy.exceptions import ScrapyDeprecationWarning - -from pydispatch import ( - dispatcher, - errors, - robust, - robustapply, - saferef, -) - -warnings.warn("Importing from scrapy.xlib.pydispatch is deprecated and will" - " no longer be supported in future Scrapy versions." - " If you just want to connect signals use the from_crawler class method," - " otherwise import pydispatch directly if needed." - " See: https://github.com/scrapy/scrapy/issues/1762", - ScrapyDeprecationWarning, stacklevel=2) diff --git a/scrapy/xlib/tx.py b/scrapy/xlib/tx.py deleted file mode 100644 index 0d94307b7..000000000 --- a/scrapy/xlib/tx.py +++ /dev/null @@ -1,19 +0,0 @@ -from __future__ import absolute_import - -import warnings -from scrapy.exceptions import ScrapyDeprecationWarning - -from twisted.web import client -from twisted.internet import endpoints - -Agent = client.Agent # since < 11.1 -ProxyAgent = client.ProxyAgent # since 11.1 -ResponseDone = client.ResponseDone # since 11.1 -ResponseFailed = client.ResponseFailed # since 11.1 -HTTPConnectionPool = client.HTTPConnectionPool # since 12.1 -TCP4ClientEndpoint = endpoints.TCP4ClientEndpoint # since 10.1 - -warnings.warn("Importing from scrapy.xlib.tx is deprecated and will" - " no longer be supported in future Scrapy versions." - " Update your code to import from twisted proper.", - ScrapyDeprecationWarning, stacklevel=2) diff --git a/tests/test_pydispatch_deprecated.py b/tests/test_pydispatch_deprecated.py deleted file mode 100644 index 6d3237fe1..000000000 --- a/tests/test_pydispatch_deprecated.py +++ /dev/null @@ -1,12 +0,0 @@ -import unittest -import warnings -from six.moves import reload_module - - -class DeprecatedPydispatchTest(unittest.TestCase): - def test_import_xlib_pydispatch_show_warning(self): - with warnings.catch_warnings(record=True) as w: - from scrapy.xlib import pydispatch - reload_module(pydispatch) - self.assertIn('Importing from scrapy.xlib.pydispatch is deprecated', - str(w[0].message)) From 21ad8e20b9189eeca9bc3f2641eb6d27fe40ce4f Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Fri, 13 Sep 2019 16:02:49 -0300 Subject: [PATCH 074/387] Crawling rules: make link extractors optional --- docs/topics/spiders.rst | 4 +++- scrapy/spiders/crawl.py | 12 ++++++++---- tests/test_spider.py | 20 ++++++++++++++++++++ 3 files changed, 31 insertions(+), 5 deletions(-) diff --git a/docs/topics/spiders.rst b/docs/topics/spiders.rst index 869a61441..45eea3e60 100644 --- a/docs/topics/spiders.rst +++ b/docs/topics/spiders.rst @@ -374,12 +374,14 @@ CrawlSpider Crawling rules ~~~~~~~~~~~~~~ -.. class:: Rule(link_extractor, callback=None, cb_kwargs=None, follow=None, process_links=None, process_request=None) +.. autoclass:: Rule ``link_extractor`` is a :ref:`Link Extractor ` object which defines how links will be extracted from each crawled page. Each produced link will be used to generate a :class:`~scrapy.http.Request` object, which will contain the link's text in its ``meta`` dictionary (under the ``link_text`` key). + If omitted, a default link extractor created with no arguments will be used, + resulting in all links being extracted. ``callback`` is a callable or a string (in which case a method from the spider object with that name will be used) to be called for each link extracted with diff --git a/scrapy/spiders/crawl.py b/scrapy/spiders/crawl.py index 90a6eb806..03000ce54 100644 --- a/scrapy/spiders/crawl.py +++ b/scrapy/spiders/crawl.py @@ -12,9 +12,10 @@ import six from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.http import Request, HtmlResponse -from scrapy.utils.spider import iterate_spider_output -from scrapy.utils.python import get_func_args +from scrapy.linkextractors import LinkExtractor from scrapy.spiders import Spider +from scrapy.utils.python import get_func_args +from scrapy.utils.spider import iterate_spider_output def _identity(request, response): @@ -28,10 +29,13 @@ def _get_method(method, spider): return getattr(spider, method, None) +_default_link_extractor = LinkExtractor() + + class Rule(object): - def __init__(self, link_extractor, callback=None, cb_kwargs=None, follow=None, process_links=None, process_request=None): - self.link_extractor = link_extractor + def __init__(self, link_extractor=None, callback=None, cb_kwargs=None, follow=None, process_links=None, process_request=None): + self.link_extractor = link_extractor or _default_link_extractor self.callback = callback self.cb_kwargs = cb_kwargs or {} self.process_links = process_links diff --git a/tests/test_spider.py b/tests/test_spider.py index e81e6d5f9..2220b8ffc 100644 --- a/tests/test_spider.py +++ b/tests/test_spider.py @@ -177,6 +177,26 @@ class CrawlSpiderTest(SpiderTest): """ spider_class = CrawlSpider + def test_rule_without_link_extractor(self): + + response = HtmlResponse("http://example.org/somepage/index.html", body=self.test_body) + + class _CrawlSpider(self.spider_class): + name = "test" + allowed_domains = ['example.org'] + rules = ( + Rule(), + ) + + spider = _CrawlSpider() + output = list(spider._requests_to_follow(response)) + self.assertEqual(len(output), 3) + self.assertTrue(all(map(lambda r: isinstance(r, Request), output))) + self.assertEqual([r.url for r in output], + ['http://example.org/somepage/item/12.html', + 'http://example.org/about.html', + 'http://example.org/nofollow.html']) + def test_process_links(self): response = HtmlResponse("http://example.org/somepage/index.html", body=self.test_body) From 13735bcf34b0caf60a52e95a907ef390324fdddc Mon Sep 17 00:00:00 2001 From: OmarFarrag Date: Mon, 16 Sep 2019 14:04:06 +0200 Subject: [PATCH 075/387] Disallow media extensions unregistered with IANA (#3954) Co-Authored-By: s-sanjay --- scrapy/pipelines/files.py | 9 +++++++++ tests/test_pipeline_files.py | 6 ++++++ 2 files changed, 15 insertions(+) diff --git a/scrapy/pipelines/files.py b/scrapy/pipelines/files.py index ea06d2ae8..cc3d10b63 100644 --- a/scrapy/pipelines/files.py +++ b/scrapy/pipelines/files.py @@ -5,6 +5,7 @@ See documentation in topics/media-pipeline.rst """ import functools import hashlib +import mimetypes import os import os.path import time @@ -14,6 +15,7 @@ from six.moves.urllib.parse import urlparse from collections import defaultdict import six + try: from cStringIO import StringIO as BytesIO except ImportError: @@ -473,4 +475,11 @@ class FilesPipeline(MediaPipeline): def file_path(self, request, response=None, info=None): media_guid = hashlib.sha1(to_bytes(request.url)).hexdigest() media_ext = os.path.splitext(request.url)[1] + # Handles empty and wild extensions by trying to guess the + # mime type then extension or default to empty string otherwise + if media_ext not in mimetypes.types_map: + media_ext = '' + media_type = mimetypes.guess_type(request.url)[0] + if media_type: + media_ext = mimetypes.guess_extension(media_type) return 'full/%s%s' % (media_guid, media_ext) diff --git a/tests/test_pipeline_files.py b/tests/test_pipeline_files.py index 0c5aaaa44..cb8f8da18 100644 --- a/tests/test_pipeline_files.py +++ b/tests/test_pipeline_files.py @@ -57,6 +57,12 @@ class FilesPipelineTestCase(unittest.TestCase): response=Response("http://www.dorma.co.uk/images/product_details/2532"), info=object()), 'full/244e0dd7d96a3b7b01f54eded250c9e272577aa1') + self.assertEqual(file_path(Request("http://www.dfsonline.co.uk/get_prod_image.php?img=status_0907_mdm.jpg.bohaha")), + 'full/76c00cef2ef669ae65052661f68d451162829507') + self.assertEqual(file_path(Request("data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAR0AAACxCAMAAADOHZloAAACClBMVEX/\ + //+F0tzCwMK76ZKQ21AMqr7oAAC96JvD5aWM2kvZ78J0N7fmAAC46Y4Ap7y")), + 'full/178059cbeba2e34120a67f2dc1afc3ecc09b61cb.png') + def test_fs_store(self): assert isinstance(self.pipeline.store, FSFilesStore) From 0b52fa6ca9ec8916e4ebcaa7ee148e0b20c2068f Mon Sep 17 00:00:00 2001 From: watsta Date: Mon, 16 Sep 2019 14:12:04 +0200 Subject: [PATCH 076/387] LogFormatter: Add the ability to skip log messages (#3987) --- scrapy/core/engine.py | 7 ++-- scrapy/core/scraper.py | 6 ++-- scrapy/logformatter.py | 4 +++ tests/test_logformatter.py | 66 +++++++++++++++++++++++++++++++++++++- 4 files changed, 77 insertions(+), 6 deletions(-) diff --git a/scrapy/core/engine.py b/scrapy/core/engine.py index 37fe0a873..fa913e528 100644 --- a/scrapy/core/engine.py +++ b/scrapy/core/engine.py @@ -233,10 +233,11 @@ class ExecutionEngine(object): def _on_success(response): assert isinstance(response, (Response, Request)) if isinstance(response, Response): - response.request = request # tie request to response received + response.request = request # tie request to response received logkws = self.logformatter.crawled(request, response, spider) - logger.log(*logformatter_adapter(logkws), extra={'spider': spider}) - self.signals.send_catch_log(signal=signals.response_received, \ + if logkws is not None: + logger.log(*logformatter_adapter(logkws), extra={'spider': spider}) + self.signals.send_catch_log(signal=signals.response_received, response=response, request=request, spider=spider) return response diff --git a/scrapy/core/scraper.py b/scrapy/core/scraper.py index 66f5d0e05..1f389cf2e 100644 --- a/scrapy/core/scraper.py +++ b/scrapy/core/scraper.py @@ -225,7 +225,8 @@ class Scraper(object): ex = output.value if isinstance(ex, DropItem): logkws = self.logformatter.dropped(item, ex, response, spider) - logger.log(*logformatter_adapter(logkws), extra={'spider': spider}) + if logkws is not None: + logger.log(*logformatter_adapter(logkws), extra={'spider': spider}) return self.signals.send_catch_log_deferred( signal=signals.item_dropped, item=item, response=response, spider=spider, exception=output.value) @@ -238,7 +239,8 @@ class Scraper(object): spider=spider, failure=output) else: logkws = self.logformatter.scraped(output, response, spider) - logger.log(*logformatter_adapter(logkws), extra={'spider': spider}) + if logkws is not None: + logger.log(*logformatter_adapter(logkws), extra={'spider': spider}) return self.signals.send_catch_log_deferred( signal=signals.item_scraped, item=output, response=response, spider=spider) diff --git a/scrapy/logformatter.py b/scrapy/logformatter.py index b4d6787ff..f15940ed1 100644 --- a/scrapy/logformatter.py +++ b/scrapy/logformatter.py @@ -29,6 +29,10 @@ class LogFormatter(object): * ``args`` should be a tuple or dict with the formatting placeholders for ``msg``. The final log message is computed as ``msg % args``. + Users can define their own ``LogFormatter`` class if they want to customise how + each action is logged or if they want to omit it entirely. In order to omit + logging an action the method must return ``None``. + Here is an example on how to create a custom log formatter to lower the severity level of the log message when an item is dropped from the pipeline:: diff --git a/tests/test_logformatter.py b/tests/test_logformatter.py index 94e6c9fde..eb9c4a561 100644 --- a/tests/test_logformatter.py +++ b/tests/test_logformatter.py @@ -1,10 +1,18 @@ import unittest + +from testfixtures import LogCapture +from twisted.internet import defer +from twisted.trial.unittest import TestCase as TwistedTestCase import six -from scrapy.spiders import Spider +from scrapy.crawler import CrawlerRunner +from scrapy.exceptions import DropItem from scrapy.http import Request, Response from scrapy.item import Item, Field from scrapy.logformatter import LogFormatter +from scrapy.spiders import Spider +from tests.mockserver import MockServer +from tests.spiders import ItemSpider class CustomItem(Item): @@ -89,5 +97,61 @@ class LogformatterSubclassTest(LoggingContribTest): pass +class SkipMessagesLogFormatter(LogFormatter): + def crawled(self, *args, **kwargs): + return None + + def scraped(self, *args, **kwargs): + return None + + def dropped(self, *args, **kwargs): + return None + + +class DropSomeItemsPipeline(object): + drop = True + + def process_item(self, item, spider): + if self.drop: + self.drop = False + raise DropItem("Ignoring item") + else: + self.drop = True + +class ShowOrSkipMessagesTestCase(TwistedTestCase): + def setUp(self): + self.mockserver = MockServer() + self.mockserver.__enter__() + self.base_settings = { + 'LOG_LEVEL': 'DEBUG', + 'ITEM_PIPELINES': { + __name__ + '.DropSomeItemsPipeline': 300, + }, + } + + def tearDown(self): + self.mockserver.__exit__(None, None, None) + + @defer.inlineCallbacks + def test_show_messages(self): + crawler = CrawlerRunner(self.base_settings).create_crawler(ItemSpider) + with LogCapture() as lc: + yield crawler.crawl(mockserver=self.mockserver) + self.assertIn("Scraped from <200 http://127.0.0.1:", str(lc)) + self.assertIn("Crawled (200) Date: Mon, 16 Sep 2019 14:24:25 +0000 Subject: [PATCH 077/387] fix capitalization, remove commas --- docs/topics/downloader-middleware.rst | 18 +++++++++--------- 1 file changed, 9 insertions(+), 9 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 52be8ded2..f2f754573 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -1023,7 +1023,7 @@ Parsers varies in several aspects: * Support for wildcard matching -* usage of length based rule: in particular for ``Allow`` and +* Usage of length based rule: in particular for ``Allow`` and ``Disallow`` directives, where the most specific rule based on the length of the path trumps the less specific (shorter) rule @@ -1033,7 +1033,7 @@ Parsers varies in several aspects: Protego parser ~~~~~~~~~~~~~~ -based on `Protego `_: +Based on `Protego `_: * implemented in Python @@ -1042,7 +1042,7 @@ based on `Protego `_: * supports wildcard matching -* uses the length based rule, +* uses the length based rule Scrapy uses this parser by default. @@ -1051,7 +1051,7 @@ Scrapy uses this parser by default. RobotFileParser ~~~~~~~~~~~~~~~ -based on `RobotFileParser +Based on `RobotFileParser `_: * is Python's built-in robots.txt_ parser. @@ -1061,7 +1061,7 @@ based on `RobotFileParser * lacks support for wildcard matching. -* doesn't use the length based rule, +* doesn't use the length based rule It is faster than Protego and backward-compatible with versions of Scrapy before 1.8.0 . @@ -1074,7 +1074,7 @@ In order to use this parser, set: Reppy parser ~~~~~~~~~~~~ -based on `Reppy `_: +Based on `Reppy `_: * is a Python wrapper around `Robots Exclusion Protocol Parser for C++ `_. @@ -1084,7 +1084,7 @@ based on `Reppy `_: * supports wildcard matching -* uses the length based rule, +* uses the length based rule Native implementation provides better speed than Protego. @@ -1100,7 +1100,7 @@ In order to use this parser: Robotexclusionrulesparser ~~~~~~~~~~~~~~~~~~~~~~~~~ -based on `Robotexclusionrulesparser `_: +Based on `Robotexclusionrulesparser `_: * implemented in Python @@ -1109,7 +1109,7 @@ based on `Robotexclusionrulesparser `_: * supports wildcard matching -* doesn't use the length based rule, +* doesn't use the length based rule In order to use this parser: From f6872189b96595625e428331ddd4d1c620047642 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 29 Aug 2019 10:38:49 -0300 Subject: [PATCH 078/387] Add LogFormatter.error method --- scrapy/core/scraper.py | 6 +++--- scrapy/logformatter.py | 11 +++++++++++ 2 files changed, 14 insertions(+), 3 deletions(-) diff --git a/scrapy/core/scraper.py b/scrapy/core/scraper.py index 1f389cf2e..3273a1506 100644 --- a/scrapy/core/scraper.py +++ b/scrapy/core/scraper.py @@ -231,9 +231,9 @@ class Scraper(object): signal=signals.item_dropped, item=item, response=response, spider=spider, exception=output.value) else: - logger.error('Error processing %(item)s', {'item': item}, - exc_info=failure_to_exc_info(output), - extra={'spider': spider}) + logkws = self.logformatter.error(item, ex, response, spider) + logger.log(*logformatter_adapter(logkws), extra={'spider': spider}, + exc_info=failure_to_exc_info(output)) return self.signals.send_catch_log_deferred( signal=signals.item_error, item=item, response=response, spider=spider, failure=output) diff --git a/scrapy/logformatter.py b/scrapy/logformatter.py index f15940ed1..4437d1106 100644 --- a/scrapy/logformatter.py +++ b/scrapy/logformatter.py @@ -8,6 +8,7 @@ from scrapy.utils.request import referer_str SCRAPEDMSG = u"Scraped from %(src)s" + os.linesep + "%(item)s" DROPPEDMSG = u"Dropped: %(exception)s" + os.linesep + "%(item)s" CRAWLEDMSG = u"Crawled (%(status)s) %(request)s%(request_flags)s (referer: %(referer)s)%(response_flags)s" +ERRORMSG = u"'Error processing %(item)s'" class LogFormatter(object): @@ -92,6 +93,16 @@ class LogFormatter(object): } } + def error(self, item, exception, response, spider): + """Logs a message when an item causes an error while it is passing through the item pipeline.""" + return { + 'level': logging.ERROR, + 'msg': ERRORMSG, + 'args': { + 'item': item, + } + } + @classmethod def from_crawler(cls, crawler): return cls() From 27436cbbc9e7331d18d4b63f3c894e6621226efb Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 29 Aug 2019 13:51:42 -0300 Subject: [PATCH 079/387] [test] LogFormatter.error --- tests/test_logformatter.py | 30 ++++++++++++++++++++++-------- 1 file changed, 22 insertions(+), 8 deletions(-) diff --git a/tests/test_logformatter.py b/tests/test_logformatter.py index eb9c4a561..502bc4ccc 100644 --- a/tests/test_logformatter.py +++ b/tests/test_logformatter.py @@ -23,7 +23,7 @@ class CustomItem(Item): return "name: %s" % self['name'] -class LoggingContribTest(unittest.TestCase): +class LoggingFormatterTest(unittest.TestCase): def setUp(self): self.formatter = LogFormatter() @@ -61,6 +61,16 @@ class LoggingContribTest(unittest.TestCase): lines = logline.splitlines() assert all(isinstance(x, six.text_type) for x in lines) self.assertEqual(lines, [u"Dropped: \u2018", '{}']) + + def test_error(self): + # In practice, the complete traceback is shown by passing the + # 'exc_info' argument to the logging function + item = {'key': 'value'} + exception = Exception() + response = Response("http://www.example.com") + logkws = self.formatter.error(item, exception, response, self.spider) + logline = logkws['msg'] % logkws['args'] + self.assertEqual(logline, u"'Error processing {'key': 'value'}'") def test_scraped(self): item = CustomItem() @@ -75,26 +85,30 @@ class LoggingContribTest(unittest.TestCase): class LogFormatterSubclass(LogFormatter): def crawled(self, request, response, spider): - kwargs = super(LogFormatterSubclass, self).crawled( - request, response, spider) + kwargs = super(LogFormatterSubclass, self).crawled(request, response, spider) CRAWLEDMSG = ( - u"Crawled (%(status)s) %(request)s (referer: " - u"%(referer)s)%(flags)s" + u"Crawled (%(status)s) %(request)s (referer: %(referer)s) %(flags)s" ) + log_args = kwargs['args'] + log_args['flags'] = str(request.flags) return { 'level': kwargs['level'], 'msg': CRAWLEDMSG, - 'args': kwargs['args'] + 'args': log_args, } -class LogformatterSubclassTest(LoggingContribTest): +class LogformatterSubclassTest(unittest.TestCase): def setUp(self): self.formatter = LogFormatterSubclass() self.spider = Spider('default') def test_flags_in_request(self): - pass + req = Request("http://www.example.com", flags=['test','flag']) + res = Response("http://www.example.com") + logkws = self.formatter.crawled(req, res, self.spider) + logline = logkws['msg'] % logkws['args'] + self.assertEqual(logline, "Crawled (200) (referer: None) ['test', 'flag']") class SkipMessagesLogFormatter(LogFormatter): From b792dba5281c91dc6845e3e64771dde2157b20f7 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Tue, 17 Sep 2019 06:28:33 +0000 Subject: [PATCH 080/387] remove periods --- docs/topics/downloader-middleware.rst | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index b67723b86..de5f72b80 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -1054,12 +1054,12 @@ RobotFileParser Based on `RobotFileParser `_: -* is Python's built-in robots.txt_ parser. +* is Python's built-in robots.txt_ parser * is compliant with `Martijn Koster's 1996 draft specification - `_. + `_ -* lacks support for wildcard matching. +* lacks support for wildcard matching * doesn't use the length based rule @@ -1077,10 +1077,10 @@ Reppy parser Based on `Reppy `_: * is a Python wrapper around `Robots Exclusion Protocol Parser for C++ - `_. + `_ * is compliant with `Martijn Koster's 1996 draft specification - `_. + `_ * supports wildcard matching @@ -1105,7 +1105,7 @@ Based on `Robotexclusionrulesparser `_: * implemented in Python * is compliant with `Martijn Koster's 1996 draft specification - `_. + `_ * supports wildcard matching From d39ef77e6be5f1a24858d92018ea4ed7a38bd127 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Tue, 17 Sep 2019 06:34:33 +0000 Subject: [PATCH 081/387] add link to google description of lenght-based rule --- docs/topics/downloader-middleware.rst | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index de5f72b80..9ce4293d2 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -1023,10 +1023,10 @@ Parsers vary in several aspects: * Support for wildcard matching -* Usage of length based rule: in particular for ``Allow`` and - ``Disallow`` directives, where the most specific rule based on the length of - the path trumps the less specific (shorter) rule - +* Usage of `length based rule `_: + in particular for ``Allow`` and ``Disallow`` directives, where the most + specific rule based on the length of the path trumps the less specific + (shorter) rule .. _protego-parser: From 57e6f4c75087b7f2182d2ebc6b28e3afa18e12bd Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Tue, 17 Sep 2019 07:18:37 +0000 Subject: [PATCH 082/387] add link to performance comparison --- docs/topics/downloader-middleware.rst | 3 +++ 1 file changed, 3 insertions(+) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 9ce4293d2..ec302f2eb 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -1028,6 +1028,9 @@ Parsers vary in several aspects: specific rule based on the length of the path trumps the less specific (shorter) rule +Performance comparison of different parsers is available at `the following link +`_. + .. _protego-parser: Protego parser From d1d0bf8491da34d1a5a4bcf3d1241241346b62c8 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Tue, 17 Sep 2019 12:27:12 +0500 Subject: [PATCH 083/387] Update docs/topics/downloader-middleware.rst MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves --- docs/topics/downloader-middleware.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index ec302f2eb..c08e13a9a 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -1066,7 +1066,7 @@ Based on `RobotFileParser * doesn't use the length based rule -It is faster than Protego and backward-compatible with versions of Scrapy before 1.8.0 . +It is faster than Protego and backward-compatible with versions of Scrapy before 1.8.0. In order to use this parser, set: From 2438ac529a647c6c665d402b59164d509404d584 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita Date: Tue, 17 Sep 2019 12:27:22 +0500 Subject: [PATCH 084/387] Update docs/topics/downloader-middleware.rst MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves --- docs/topics/downloader-middleware.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index c08e13a9a..539832618 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -1089,7 +1089,7 @@ Based on `Reppy `_: * uses the length based rule -Native implementation provides better speed than Protego. +Native implementation, provides better speed than Protego. In order to use this parser: From 9b65f9aa5b03353dcbb7daff0b85d3c3599b08fa Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 19 Sep 2019 09:17:23 +0200 Subject: [PATCH 085/387] Fix the item exporter example (#4022) --- docs/topics/exporters.rst | 1 - 1 file changed, 1 deletion(-) diff --git a/docs/topics/exporters.rst b/docs/topics/exporters.rst index 11b3045ec..a698a6a4e 100644 --- a/docs/topics/exporters.rst +++ b/docs/topics/exporters.rst @@ -51,7 +51,6 @@ value of one of their fields:: def close_spider(self, spider): for exporter in self.year_to_exporter.values(): exporter.finish_exporting() - exporter.file.close() def _exporter_for_item(self, item): year = item['year'] From c26a9015ad2f1f50e0dfc1510e9551414da0d04a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 19 Sep 2019 11:08:06 +0200 Subject: [PATCH 086/387] Clarify the effects of dont_merge_cookies --- docs/topics/request-response.rst | 22 +++++++++++++--------- 1 file changed, 13 insertions(+), 9 deletions(-) diff --git a/docs/topics/request-response.rst b/docs/topics/request-response.rst index ad1b9af10..284d3479b 100644 --- a/docs/topics/request-response.rst +++ b/docs/topics/request-response.rst @@ -83,17 +83,21 @@ Request objects .. reqmeta:: dont_merge_cookies When some site returns cookies (in a response) those are stored in the - cookies for that domain and will be sent again in future requests. That's - the typical behaviour of any regular web browser. However, if, for some - reason, you want to avoid merging with existing cookies you can instruct - Scrapy to do so by setting the ``dont_merge_cookies`` key to True in the - :attr:`Request.meta`. + cookies for that domain and will be sent again in future requests. + That's the typical behaviour of any regular web browser. - Example of request without merging cookies:: + To create a request that does not send stored cookies and does not + store received cookies, set the ``dont_merge_cookies`` key to ``True`` + in :attr:`request.meta `. - request_with_cookies = Request(url="http://www.example.com", - cookies={'currency': 'USD', 'country': 'UY'}, - meta={'dont_merge_cookies': True}) + Example of a request that sends manually-defined cookies and ignores + cookie storage:: + + Request( + url="http://www.example.com", + cookies={'currency': 'USD', 'country': 'UY'}, + meta={'dont_merge_cookies': True}, + ) For more info see :ref:`cookies-mw`. :type cookies: dict or list From 447b3d9d8133da0448cba871314df01887b645d7 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Wed, 25 Sep 2019 11:13:37 +0200 Subject: [PATCH 087/387] =?UTF-8?q?Fix=20documentation=20typo:=20accesible?= =?UTF-8?q?=20=E2=86=92=20accessible=20(#4033)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- scrapy/utils/request.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/utils/request.py b/scrapy/utils/request.py index 50bc3cb1e..9c143b83a 100644 --- a/scrapy/utils/request.py +++ b/scrapy/utils/request.py @@ -30,7 +30,7 @@ def request_fingerprint(request, include_headers=None): and are equivalent (ie. they should return the same response). Another example are cookies used to store session ids. Suppose the - following page is only accesible to authenticated users: + following page is only accessible to authenticated users: http://www.example.com/members/offers.html From 1236e9e81ecd83de716cabc61374fbbacd82d0aa Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Tue, 3 Sep 2019 14:08:08 +0200 Subject: [PATCH 088/387] Provide complete API documentation coverage of scrapy.item --- docs/conf.py | 3 +++ docs/topics/items.rst | 6 ++++++ scrapy/item.py | 33 +++++++++++++++++++++++++++++---- tests/test_item.py | 30 +++++++++++++++++++++++++++++- 4 files changed, 67 insertions(+), 5 deletions(-) diff --git a/docs/conf.py b/docs/conf.py index fa257dead..34dd5bcb7 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -262,6 +262,9 @@ coverage_ignore_pyobjects = [ # details that are not documented. r'^scrapy\.extensions\.[a-z]\w*?\.[A-Z]\w*?\.', # methods r'^scrapy\.extensions\.[a-z]\w*?\.[a-z]', # helper functions + + # Never documented before, and deprecated now. + r'^scrapy\.item\.DictItem$', ] diff --git a/docs/topics/items.rst b/docs/topics/items.rst index 60fbc82f8..260f5882c 100644 --- a/docs/topics/items.rst +++ b/docs/topics/items.rst @@ -263,3 +263,9 @@ Field objects .. _dict: https://docs.python.org/2/library/stdtypes.html#dict +Other classes related to Item +============================= + +.. autoclass:: BaseItem + +.. autoclass:: ItemMeta diff --git a/scrapy/item.py b/scrapy/item.py index 9d4786788..73b8f54b0 100644 --- a/scrapy/item.py +++ b/scrapy/item.py @@ -4,13 +4,15 @@ Scrapy Item See documentation in docs/topics/item.rst """ -from abc import ABCMeta -from pprint import pformat -from copy import deepcopy import collections +from abc import ABCMeta +from copy import deepcopy +from pprint import pformat +from warnings import warn import six +from scrapy.utils.deprecate import ScrapyDeprecationWarning from scrapy.utils.trackref import object_ref @@ -21,7 +23,19 @@ else: class BaseItem(object_ref): - """Base class for all scraped items.""" + """Base class for all scraped items. + + In Scrapy, an object is considered an *item* if it is an instance of either + :class:`BaseItem` or :class:`dict`. For example, when the output of a + spider callback is evaluated, only instances of :class:`BaseItem` or + :class:`dict` are passed to :ref:`item pipelines `. + + If you need instances of a custom class to be considered items by Scrapy, + you must inherit from either :class:`BaseItem` or :class:`dict`. + + Unlike instances of :class:`dict`, instances of :class:`BaseItem` may be + :ref:`tracked ` to debug memory leaks. + """ pass @@ -30,6 +44,10 @@ class Field(dict): class ItemMeta(ABCMeta): + """Metaclass_ of :class:`Item` that handles field definitions. + + .. _metaclass: https://realpython.com/python-metaclasses + """ def __new__(mcs, class_name, bases, attrs): classcell = attrs.pop('__classcell__', None) @@ -56,6 +74,13 @@ class DictItem(MutableMapping, BaseItem): fields = {} + def __new__(cls, *args, **kwargs): + if issubclass(cls, DictItem) and not issubclass(cls, Item): + warn('scrapy.item.DictItem is deprecated, please use ' + 'scrapy.item.Item instead', + ScrapyDeprecationWarning, stacklevel=2) + return super(DictItem, cls).__new__(cls, *args, **kwargs) + def __init__(self, *args, **kwargs): self._values = {} if args or kwargs: # avoid creating dict for most common case diff --git a/tests/test_item.py b/tests/test_item.py index 010d3b141..947566686 100644 --- a/tests/test_item.py +++ b/tests/test_item.py @@ -1,9 +1,11 @@ import sys import unittest +from warnings import catch_warnings import six -from scrapy.item import ABCMeta, Item, ItemMeta, Field +from scrapy.exceptions import ScrapyDeprecationWarning +from scrapy.item import ABCMeta, DictItem, Field, Item, ItemMeta from tests import mock @@ -257,6 +259,17 @@ class ItemTest(unittest.TestCase): item['tags'].append('tag2') assert item['tags'] != copied_item['tags'] + def test_dictitem_deprecation_warning(self): + """Make sure the DictItem deprecation warning is not issued for + Item""" + with catch_warnings(record=True) as warnings: + item = Item() + self.assertEqual(len(warnings), 0) + class SubclassedItem(Item): + pass + subclassed_item = SubclassedItem() + self.assertEqual(len(warnings), 0) + class ItemMetaTest(unittest.TestCase): @@ -302,5 +315,20 @@ class ItemMetaClassCellRegression(unittest.TestCase): super(MyItem, self).__init__(*args, **kwargs) +class DictItemTest(unittest.TestCase): + + def test_deprecation_warning(self): + with catch_warnings(record=True) as warnings: + dict_item = DictItem() + self.assertEqual(len(warnings), 1) + self.assertEqual(warnings[0].category, ScrapyDeprecationWarning) + with catch_warnings(record=True) as warnings: + class SubclassedDictItem(DictItem): + pass + subclassed_dict_item = SubclassedDictItem() + self.assertEqual(len(warnings), 1) + self.assertEqual(warnings[0].category, ScrapyDeprecationWarning) + + if __name__ == "__main__": unittest.main() From 2c14692e603873cf5b5783ad62ebe47474d2d7a8 Mon Sep 17 00:00:00 2001 From: s-sanjay <7111850+s-sanjay@users.noreply.github.com> Date: Fri, 27 Sep 2019 00:56:43 -0700 Subject: [PATCH 089/387] remove .keys() to avoid creating a tmp list/keyview obj (#4031) Also add --verbose and --nolinks for code coverage --- scrapy/commands/parse.py | 15 ++++++++------- scrapy/commands/runspider.py | 5 ++--- tests/test_command_parse.py | 6 ++++++ 3 files changed, 16 insertions(+), 10 deletions(-) diff --git a/scrapy/commands/parse.py b/scrapy/commands/parse.py index d4f2234b0..ef8acd29c 100644 --- a/scrapy/commands/parse.py +++ b/scrapy/commands/parse.py @@ -60,10 +60,12 @@ class Command(ScrapyCommand): @property def max_level(self): - levels = list(self.items.keys()) + list(self.requests.keys()) - if not levels: - return 0 - return max(levels) + max_items, max_requests = 0, 0 + if self.items: + max_items = max(self.items) + if self.requests: + max_requests = max(self.requests) + return max(max_items, max_requests) def add_items(self, lvl, new_items): old_items = self.items.get(lvl, []) @@ -84,9 +86,8 @@ class Command(ScrapyCommand): def print_requests(self, lvl=None, colour=True): if lvl is None: - levels = list(self.requests.keys()) - if levels: - requests = self.requests[max(levels)] + if self.requests: + requests = self.requests[max(self.requests)] else: requests = [] else: diff --git a/scrapy/commands/runspider.py b/scrapy/commands/runspider.py index 376d3c84e..57d8471ca 100644 --- a/scrapy/commands/runspider.py +++ b/scrapy/commands/runspider.py @@ -60,14 +60,13 @@ class Command(ScrapyCommand): else: self.settings.set('FEED_URI', opts.output, priority='cmdline') feed_exporters = without_none_values(self.settings.getwithbase('FEED_EXPORTERS')) - valid_output_formats = feed_exporters.keys() if not opts.output_format: opts.output_format = os.path.splitext(opts.output)[1].replace(".", "") - if opts.output_format not in valid_output_formats: + if opts.output_format not in feed_exporters: raise UsageError("Unrecognized output format '%s', set one" " using the '-t' switch or as a file extension" " from the supported list %s" % (opts.output_format, - tuple(valid_output_formats))) + tuple(feed_exporters))) self.settings.set('FEED_FORMAT', opts.output_format, priority='cmdline') def run(self, args, opts): diff --git a/tests/test_command_parse.py b/tests/test_command_parse.py index 98e415ad3..62d5d76b4 100644 --- a/tests/test_command_parse.py +++ b/tests/test_command_parse.py @@ -108,6 +108,7 @@ ITEM_PIPELINES = {'%s.pipelines.MyPipeline': 1} _, _, stderr = yield self.execute(['--spider', self.spider_name, '-a', 'test_arg=1', '-c', 'parse', + '--verbose', self.url('/html')]) self.assertIn("DEBUG: It Works!", _textmode(stderr)) @@ -117,12 +118,14 @@ ITEM_PIPELINES = {'%s.pipelines.MyPipeline': 1} _, _, stderr = yield self.execute(['--spider', self.spider_name, '--meta', raw_json_string, '-c', 'parse_request_with_meta', + '--verbose', self.url('/html')]) self.assertIn("DEBUG: It Works!", _textmode(stderr)) _, _, stderr = yield self.execute(['--spider', self.spider_name, '-m', raw_json_string, '-c', 'parse_request_with_meta', + '--verbose', self.url('/html')]) self.assertIn("DEBUG: It Works!", _textmode(stderr)) @@ -132,6 +135,7 @@ ITEM_PIPELINES = {'%s.pipelines.MyPipeline': 1} _, _, stderr = yield self.execute(['--spider', self.spider_name, '--cbkwargs', raw_json_string, '-c', 'parse_request_with_cb_kwargs', + '--verbose', self.url('/html')]) self.assertIn("DEBUG: It Works!", _textmode(stderr)) @@ -139,6 +143,7 @@ ITEM_PIPELINES = {'%s.pipelines.MyPipeline': 1} def test_request_without_meta(self): _, _, stderr = yield self.execute(['--spider', self.spider_name, '-c', 'parse_request_without_meta', + '--nolinks', self.url('/html')]) self.assertIn("DEBUG: It Works!", _textmode(stderr)) @@ -148,6 +153,7 @@ ITEM_PIPELINES = {'%s.pipelines.MyPipeline': 1} _, _, stderr = yield self.execute(['--spider', self.spider_name, '--pipelines', '-c', 'parse', + '--verbose', self.url('/html')]) self.assertIn("INFO: It Works!", _textmode(stderr)) From 7f4f98fd38d3fdf6b45a9d0289df0cbb48bcd22b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 30 Sep 2019 18:22:28 +0200 Subject: [PATCH 090/387] Provide complete API documentation coverage of scrapy.linkextractors --- docs/conf.py | 4 +++ docs/topics/link-extractors.rst | 45 ++++++++++++------------------- scrapy/linkextractors/__init__.py | 11 ++++++++ scrapy/linkextractors/lxmlhtml.py | 8 ++++++ tests/test_linkextractors.py | 30 +++++++++++++++++++++ 5 files changed, 70 insertions(+), 28 deletions(-) diff --git a/docs/conf.py b/docs/conf.py index 34dd5bcb7..5ba6b8d43 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -265,6 +265,10 @@ coverage_ignore_pyobjects = [ # Never documented before, and deprecated now. r'^scrapy\.item\.DictItem$', + r'^scrapy\.linkextractors\.FilteringLinkExtractor$', + + # Implementation detail of LxmlLinkExtractor + r'^scrapy\.linkextractors\.lxmlhtml\.LxmlParserLinkExtractor', ] diff --git a/docs/topics/link-extractors.rst b/docs/topics/link-extractors.rst index 713a94e10..f9936a498 100644 --- a/docs/topics/link-extractors.rst +++ b/docs/topics/link-extractors.rst @@ -4,46 +4,33 @@ Link Extractors =============== -Link extractors are objects whose only purpose is to extract links from web -pages (:class:`scrapy.http.Response` objects) which will be eventually -followed. +A link extractor is an object that extracts links from responses. -There is ``scrapy.linkextractors.LinkExtractor`` available -in Scrapy, but you can create your own custom Link Extractors to suit your -needs by implementing a simple interface. - -The only public method that every link extractor has is ``extract_links``, -which receives a :class:`~scrapy.http.Response` object and returns a list -of :class:`scrapy.link.Link` objects. Link extractors are meant to be -instantiated once and their ``extract_links`` method called several times -with different responses to extract links to follow. - -Link extractors are used in the :class:`~scrapy.spiders.CrawlSpider` -class (available in Scrapy), through a set of rules, but you can also use it in -your spiders, even if you don't subclass from -:class:`~scrapy.spiders.CrawlSpider`, as its purpose is very simple: to -extract links. +The constructor of :class:`~scrapy.linkextractors.lxmlhtml.LxmlLinkExtractor` +takes settings that determine which links may be extracted. +:class:`LxmlLinkExtractor.extract_links +` returns a +list of matching :class:`scrapy.link.Link` objects from a +:class:`~scrapy.http.Response` object. +Link extractors are used in :class:`~scrapy.spiders.CrawlSpider` spiders +through a set of :class:`~scrapy.spiders.Rule` objects. You can also use link +extractors in regular spiders. .. _topics-link-extractors-ref: -Built-in link extractors reference -================================== +Link extractor reference +======================== .. module:: scrapy.linkextractors :synopsis: Link extractors classes -Link extractors classes bundled with Scrapy are provided in the -:mod:`scrapy.linkextractors` module. - -The default link extractor is ``LinkExtractor``, which is the same as -:class:`~.LxmlLinkExtractor`:: +The link extractor class is +:class:`scrapy.linkextractors.lxmlhtml.LxmlLinkExtractor`. For convenience it +can also be imported as ``scrapy.linkextractors.LinkExtractor``:: from scrapy.linkextractors import LinkExtractor -There used to be other link extractor classes in previous Scrapy versions, -but they are deprecated now. - LxmlLinkExtractor ----------------- @@ -152,4 +139,6 @@ LxmlLinkExtractor from elements or attributes which allow leading/trailing whitespaces). :type strip: boolean + .. automethod:: extract_links + .. _scrapy.linkextractors: https://github.com/scrapy/scrapy/blob/master/scrapy/linkextractors/__init__.py diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index ebf3cd7d8..ca80dc339 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -6,11 +6,13 @@ This package contains a collection of Link Extractors. For more info see docs/topics/link-extractors.rst """ import re +from warnings import warn from six.moves.urllib.parse import urlparse from parsel.csstranslator import HTMLTranslator from w3lib.url import canonicalize_url +from scrapy.utils.deprecate import ScrapyDeprecationWarning from scrapy.utils.misc import arg_to_iter from scrapy.utils.url import ( url_is_from_any_domain, url_has_any_extension, @@ -49,6 +51,15 @@ class FilteringLinkExtractor(object): _csstranslator = HTMLTranslator() + def __new__(cls, *args, **kwargs): + from scrapy.linkextractors.lxmlhtml import LxmlLinkExtractor + if (issubclass(cls, FilteringLinkExtractor) and + not issubclass(cls, LxmlLinkExtractor)): + warn('scrapy.linkextractors.FilteringLinkExtractor is deprecated, ' + 'please use scrapy.linkextractors.LinkExtractor instead', + ScrapyDeprecationWarning, stacklevel=2) + return super(FilteringLinkExtractor, cls).__new__(cls) + def __init__(self, link_extractor, allow, deny, allow_domains, deny_domains, restrict_xpaths, canonicalize, deny_extensions, restrict_css, restrict_text): diff --git a/scrapy/linkextractors/lxmlhtml.py b/scrapy/linkextractors/lxmlhtml.py index 8f6f93a44..41091ba23 100644 --- a/scrapy/linkextractors/lxmlhtml.py +++ b/scrapy/linkextractors/lxmlhtml.py @@ -117,6 +117,14 @@ class LxmlLinkExtractor(FilteringLinkExtractor): restrict_text=restrict_text) def extract_links(self, response): + """Returns a list of :class:`~scrapy.link.Link` objects from the + specified :class:`response `. + + Only links that match the settings passed to the link extractor + constructor are returned. + + Duplicate links are omitted. + """ base_url = get_base_url(response) if self.restrict_xpaths: docs = [subdoc diff --git a/tests/test_linkextractors.py b/tests/test_linkextractors.py index d96e259f6..ea6db28c0 100644 --- a/tests/test_linkextractors.py +++ b/tests/test_linkextractors.py @@ -1,10 +1,13 @@ import re import unittest +from warnings import catch_warnings import pytest +from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.http import HtmlResponse, XmlResponse from scrapy.link import Link +from scrapy.linkextractors import FilteringLinkExtractor from scrapy.linkextractors.lxmlhtml import LxmlLinkExtractor from tests import get_testdata @@ -506,3 +509,30 @@ class LxmlLinkExtractorTestCase(Base.LinkExtractorTestCase): @pytest.mark.xfail def test_restrict_xpaths_with_html_entities(self): super(LxmlLinkExtractorTestCase, self).test_restrict_xpaths_with_html_entities() + + def test_filteringlinkextractor_deprecation_warning(self): + """Make sure the FilteringLinkExtractor deprecation warning is not + issued for LxmlLinkExtractor""" + with catch_warnings(record=True) as warnings: + extractor = LxmlLinkExtractor() + self.assertEqual(len(warnings), 0) + class SubclassedItem(LxmlLinkExtractor): + pass + subclassed_extractor = SubclassedItem() + self.assertEqual(len(warnings), 0) + + +class FilteringLinkExtractorTest(unittest.TestCase): + + def test_deprecation_warning(self): + args = [None]*10 + with catch_warnings(record=True) as warnings: + extractor = FilteringLinkExtractor(*args) + self.assertEqual(len(warnings), 1) + self.assertEqual(warnings[0].category, ScrapyDeprecationWarning) + with catch_warnings(record=True) as warnings: + class SubclassedFilteringLinkExtractor(FilteringLinkExtractor): + pass + subclassed_extractor = SubclassedFilteringLinkExtractor(*args) + self.assertEqual(len(warnings), 1) + self.assertEqual(warnings[0].category, ScrapyDeprecationWarning) From 7632e375c268d31c570904452902799db180c415 Mon Sep 17 00:00:00 2001 From: Eugen Date: Tue, 1 Oct 2019 17:29:48 +0300 Subject: [PATCH 091/387] fix typo in ScrapyCommand.help docstring (#4046) --- scrapy/commands/__init__.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/commands/__init__.py b/scrapy/commands/__init__.py index 43b420821..0b24193c2 100644 --- a/scrapy/commands/__init__.py +++ b/scrapy/commands/__init__.py @@ -47,7 +47,7 @@ class ScrapyCommand(object): def help(self): """An extensive help for the command. It will be shown when using the - "help" command. It can contain newlines, since not post-formatting will + "help" command. It can contain newlines, since no post-formatting will be applied to its contents. """ return self.long_desc() From c232bbdc426d91b98703d5f8e892b0c431c5a43a Mon Sep 17 00:00:00 2001 From: Kristobal Junta Date: Tue, 1 Oct 2019 17:41:38 +0300 Subject: [PATCH 092/387] fix typo in docs/topics/spiders --- docs/topics/spiders.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/topics/spiders.rst b/docs/topics/spiders.rst index 45eea3e60..d60c93be6 100644 --- a/docs/topics/spiders.rst +++ b/docs/topics/spiders.rst @@ -692,7 +692,7 @@ SitemapSpider .. method:: sitemap_filter(entries) - This is a filter funtion that could be overridden to select sitemap entries + This is a filter function that could be overridden to select sitemap entries based on their attributes. For example:: From 07a31b13db96b2aef3074157808f7ad300780c3d Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Tue, 1 Oct 2019 17:55:57 -0300 Subject: [PATCH 093/387] Update LogFormatter tests --- tests/test_logformatter.py | 23 ++++++++++++++++++++--- 1 file changed, 20 insertions(+), 3 deletions(-) diff --git a/tests/test_logformatter.py b/tests/test_logformatter.py index 502bc4ccc..f12ffc11b 100644 --- a/tests/test_logformatter.py +++ b/tests/test_logformatter.py @@ -23,13 +23,13 @@ class CustomItem(Item): return "name: %s" % self['name'] -class LoggingFormatterTest(unittest.TestCase): +class LogFormatterTestCase(unittest.TestCase): def setUp(self): self.formatter = LogFormatter() self.spider = Spider('default') - def test_crawled(self): + def test_crawled_with_referer(self): req = Request("http://www.example.com") res = Response("http://www.example.com") logkws = self.formatter.crawled(req, res, self.spider) @@ -37,6 +37,7 @@ class LoggingFormatterTest(unittest.TestCase): self.assertEqual(logline, "Crawled (200) (referer: None)") + def test_crawled_without_referer(self): req = Request("http://www.example.com", headers={'referer': 'http://example.com'}) res = Response("http://www.example.com", flags=['cached']) logkws = self.formatter.crawled(req, res, self.spider) @@ -98,11 +99,27 @@ class LogFormatterSubclass(LogFormatter): } -class LogformatterSubclassTest(unittest.TestCase): +class LogformatterSubclassTest(LogFormatterTestCase): def setUp(self): self.formatter = LogFormatterSubclass() self.spider = Spider('default') + def test_crawled_with_referer(self): + req = Request("http://www.example.com") + res = Response("http://www.example.com") + logkws = self.formatter.crawled(req, res, self.spider) + logline = logkws['msg'] % logkws['args'] + self.assertEqual(logline, + "Crawled (200) (referer: None) []") + + def test_crawled_without_referer(self): + req = Request("http://www.example.com", headers={'referer': 'http://example.com'}, flags=['cached']) + res = Response("http://www.example.com") + logkws = self.formatter.crawled(req, res, self.spider) + logline = logkws['msg'] % logkws['args'] + self.assertEqual(logline, + "Crawled (200) (referer: http://example.com) ['cached']") + def test_flags_in_request(self): req = Request("http://www.example.com", flags=['test','flag']) res = Response("http://www.example.com") From e0fabab5cc740bfbef4c13d9d7611278bc7cc335 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Tue, 1 Oct 2019 11:54:15 -0300 Subject: [PATCH 094/387] Fix TypeError when using DummyStatsCollector --- scrapy/extensions/corestats.py | 11 +++++--- tests/test_stats.py | 49 ++++++++++++++++++++++++++++++++++ 2 files changed, 56 insertions(+), 4 deletions(-) diff --git a/scrapy/extensions/corestats.py b/scrapy/extensions/corestats.py index 8cc5e18ac..20adfbe4b 100644 --- a/scrapy/extensions/corestats.py +++ b/scrapy/extensions/corestats.py @@ -1,14 +1,16 @@ """ Extension for collecting core stats like items scraped and start/finish times """ -import datetime +from datetime import datetime from scrapy import signals + class CoreStats(object): def __init__(self, stats): self.stats = stats + self.start_time = None @classmethod def from_crawler(cls, crawler): @@ -21,11 +23,12 @@ class CoreStats(object): return o def spider_opened(self, spider): - self.stats.set_value('start_time', datetime.datetime.utcnow(), spider=spider) + self.start_time = datetime.utcnow() + self.stats.set_value('start_time', self.start_time, spider=spider) def spider_closed(self, spider, reason): - finish_time = datetime.datetime.utcnow() - elapsed_time = finish_time - self.stats.get_value('start_time') + finish_time = datetime.utcnow() + elapsed_time = finish_time - self.start_time elapsed_time_seconds = elapsed_time.total_seconds() self.stats.set_value('elapsed_time_seconds', elapsed_time_seconds, spider=spider) self.stats.set_value('finish_time', finish_time, spider=spider) diff --git a/tests/test_stats.py b/tests/test_stats.py index 9f950ebc9..2033dbe07 100644 --- a/tests/test_stats.py +++ b/tests/test_stats.py @@ -1,10 +1,59 @@ +from datetime import datetime import unittest +try: + from unittest import mock +except ImportError: + import mock + +from scrapy.extensions.corestats import CoreStats from scrapy.spiders import Spider from scrapy.statscollectors import StatsCollector, DummyStatsCollector from scrapy.utils.test import get_crawler +class CoreStatsExtensionTest(unittest.TestCase): + + def setUp(self): + self.crawler = get_crawler(Spider) + self.spider = self.crawler._create_spider('foo') + + @mock.patch('scrapy.extensions.corestats.datetime') + def test_core_stats_default_stats_collector(self, mock_datetime): + fixed_datetime = datetime(2019, 12, 1, 11, 38) + mock_datetime.utcnow = mock.Mock(return_value=fixed_datetime) + self.crawler.stats = StatsCollector(self.crawler) + ext = CoreStats.from_crawler(self.crawler) + ext.spider_opened(self.spider) + ext.item_scraped({}, self.spider) + ext.response_received(self.spider) + ext.item_dropped({}, self.spider, ZeroDivisionError()) + ext.spider_closed(self.spider, 'finished') + self.assertEqual( + ext.stats._stats, + { + 'start_time': fixed_datetime, + 'finish_time': fixed_datetime, + 'item_scraped_count': 1, + 'response_received_count': 1, + 'item_dropped_count': 1, + 'item_dropped_reasons_count/ZeroDivisionError': 1, + 'finish_reason': 'finished', + 'elapsed_time_seconds': 0.0, + } + ) + + def test_core_stats_dummy_stats_collector(self): + self.crawler.stats = DummyStatsCollector(self.crawler) + ext = CoreStats.from_crawler(self.crawler) + ext.spider_opened(self.spider) + ext.item_scraped({}, self.spider) + ext.response_received(self.spider) + ext.item_dropped({}, self.spider, ZeroDivisionError()) + ext.spider_closed(self.spider, 'finished') + self.assertEqual(ext.stats._stats, {}) + + class StatsCollectorTest(unittest.TestCase): def setUp(self): From 6ad5a89cb0aeac18a72ed41819c07e9c22990f10 Mon Sep 17 00:00:00 2001 From: elacuesta Date: Wed, 2 Oct 2019 07:18:36 -0300 Subject: [PATCH 095/387] [Doc] Use autoclass in topics/request-response.rst (#4055) --- docs/topics/request-response.rst | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/topics/request-response.rst b/docs/topics/request-response.rst index 284d3479b..727c67482 100644 --- a/docs/topics/request-response.rst +++ b/docs/topics/request-response.rst @@ -24,7 +24,7 @@ below in :ref:`topics-request-response-ref-request-subclasses` and Request objects =============== -.. class:: Request(url[, callback, method='GET', headers, body, cookies, meta, encoding='utf-8', priority=0, dont_filter=False, errback, flags, cb_kwargs]) +.. autoclass:: Request A :class:`Request` object represents an HTTP request, which is usually generated in the Spider and executed by the Downloader, and thus generating @@ -400,7 +400,7 @@ fields with form data from :class:`Response` objects. .. class:: FormRequest(url, [formdata, ...]) - The :class:`FormRequest` class adds a new argument to the constructor. The + The :class:`FormRequest` class adds a new keyword parameter to the constructor. The remaining arguments are the same as for the :class:`Request` class and are not documented here. @@ -547,7 +547,7 @@ dealing with JSON requests. .. class:: JsonRequest(url, [... data, dumps_kwargs]) - The :class:`JsonRequest` class adds two new argument to the constructor. The + The :class:`JsonRequest` class adds two new keyword parameters to the constructor. The remaining arguments are the same as for the :class:`Request` class and are not documented here. @@ -581,7 +581,7 @@ Sending a JSON POST request with a JSON payload:: Response objects ================ -.. class:: Response(url, [status=200, headers=None, body=b'', flags=None, request=None]) +.. autoclass:: Response A :class:`Response` object represents an HTTP response, which is usually downloaded (by the Downloader) and fed to the Spiders for processing. From a4aa5b8926153560c493fa84ed1d1d6c730d6f5b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Wed, 2 Oct 2019 14:51:06 +0200 Subject: [PATCH 096/387] Fix internal links in the tutorial and release notes --- docs/intro/tutorial.rst | 6 +++--- docs/news.rst | 44 ++++++++++++++++++++++++---------------- docs/topics/items.rst | 4 ++++ docs/topics/settings.rst | 2 ++ 4 files changed, 35 insertions(+), 21 deletions(-) diff --git a/docs/intro/tutorial.rst b/docs/intro/tutorial.rst index a190ce407..0629b9e19 100644 --- a/docs/intro/tutorial.rst +++ b/docs/intro/tutorial.rst @@ -78,9 +78,9 @@ Our first Spider Spiders are classes that you define and that Scrapy uses to scrape information from a website (or a group of websites). They must subclass -:class:`scrapy.Spider` and define the initial requests to make, optionally how -to follow links in the pages, and how to parse the downloaded page content to -extract data. +:class:`~scrapy.spiders.Spider` and define the initial requests to make, +optionally how to follow links in the pages, and how to parse the downloaded +page content to extract data. This is the code for our first Spider. Save it in a file named ``quotes_spider.py`` under the ``tutorial/spiders`` directory in your project:: diff --git a/docs/news.rst b/docs/news.rst index aac750601..c107e907c 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -71,7 +71,7 @@ New features ~~~~~~~~~~~~ * A new scheduler priority queue, - :class:`scrapy.pqueues.DownloaderAwarePriorityQueue`, may be + ``scrapy.pqueues.DownloaderAwarePriorityQueue``, may be :ref:`enabled ` for a significant scheduling improvement on crawls targetting multiple web domains, at the cost of no :setting:`CONCURRENT_REQUESTS_PER_IP` support (:issue:`3520`) @@ -150,9 +150,9 @@ Bug fixes :setting:`AWS_REGION_NAME`, :setting:`AWS_USE_SSL`, :setting:`AWS_VERIFY` (:issue:`3625`) -* Fixed a memory leak in :class:`~scrapy.pipelines.media.MediaPipeline` - affecting, for example, non-200 responses and exceptions from custom - middlewares (:issue:`3813`) +* Fixed a memory leak in ``scrapy.pipelines.media.MediaPipeline`` affecting, + for example, non-200 responses and exceptions from custom middlewares + (:issue:`3813`) * Requests with private callbacks are now correctly unserialized from disk (:issue:`3790`) @@ -301,7 +301,7 @@ Deprecations * The ``queuelib.PriorityQueue`` value for the :setting:`SCHEDULER_PRIORITY_QUEUE` setting is deprecated. Use - :class:`scrapy.pqueues.ScrapyPriorityQueue` instead. + ``scrapy.pqueues.ScrapyPriorityQueue`` instead. * ``process_request`` callbacks passed to :class:`~scrapy.spiders.Rule` that do not accept two arguments are deprecated. @@ -551,7 +551,7 @@ Scrapy 1.5.2 (2019-01-22) *The fix is backward incompatible*, it enables telnet user-password authentication by default with a random generated password. If you can't - upgrade right away, please consider setting :setting:`TELNET_CONSOLE_PORT` + upgrade right away, please consider setting :setting:`TELNETCONSOLE_PORT` out of its default value. See :ref:`telnet console ` documentation for more info @@ -758,7 +758,9 @@ Enjoy! (Or read on for the rest of changes in this release.) Deprecations and Backward Incompatible Changes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -- Default to ``canonicalize=False`` in :class:`scrapy.linkextractors.LinkExtractor` +- Default to ``canonicalize=False`` in + :class:`scrapy.linkextractors.LinkExtractor + ` (:issue:`2537`, fixes :issue:`1941` and :issue:`1982`): **warning, this is technically backward-incompatible** - Enable memusage extension by default (:issue:`2539`, fixes :issue:`2187`); @@ -794,10 +796,13 @@ New Features - New ``data:`` URI download handler (:issue:`2334`, fixes :issue:`2156`) - Log cache directory when HTTP Cache is used (:issue:`2611`, fixes :issue:`2604`) - Warn users when project contains duplicate spider names (fixes :issue:`2181`) -- :class:`CaselessDict` now accepts ``Mapping`` instances and not only dicts (:issue:`2646`) -- :ref:`Media downloads `, with :class:`FilesPipelines` - or :class:`ImagesPipelines`, can now optionally handle HTTP redirects - using the new :setting:`MEDIA_ALLOW_REDIRECTS` setting (:issue:`2616`, fixes :issue:`2004`) +- ``scrapy.utils.datatypes.CaselessDict`` now accepts ``Mapping`` instances and + not only dicts (:issue:`2646`) +- :ref:`Media downloads `, with + :class:`~scrapy.pipelines.files.FilesPipeline` or + :class:`~scrapy.pipelines.images.ImagesPipeline`, can now optionally handle + HTTP redirects using the new :setting:`MEDIA_ALLOW_REDIRECTS` setting + (:issue:`2616`, fixes :issue:`2004`) - Accept non-complete responses from websites using a new :setting:`DOWNLOAD_FAIL_ON_DATALOSS` setting (:issue:`2590`, fixes :issue:`2586`) - Optional pretty-printing of JSON and XML items via @@ -817,8 +822,8 @@ Bug fixes - LinkExtractor now strips leading and trailing whitespaces from attributes (:issue:`2547`, fixes :issue:`1614`) -- Properly handle whitespaces in action attribute in :class:`FormRequest` - (:issue:`2548`) +- Properly handle whitespaces in action attribute in + :class:`~scrapy.http.FormRequest` (:issue:`2548`) - Buffer CONNECT response bytes from proxy until all HTTP headers are received (:issue:`2495`, fixes :issue:`2491`) - FTP downloader now works on Python 3, provided you use Twisted>=17.1 @@ -851,7 +856,8 @@ Cleanups & Refactoring fixes :issue:`2560`) - Add omitted ``self`` arguments in default project middleware template (:issue:`2595`) - Remove redundant ``slot.add_request()`` call in ExecutionEngine (:issue:`2617`) -- Catch more specific ``os.error`` exception in :class:`FSFilesStore` (:issue:`2644`) +- Catch more specific ``os.error`` exception in + ``scrapy.pipelines.files.FSFilesStore`` (:issue:`2644`) - Change "localhost" test server certificate (:issue:`2720`) - Remove unused ``MEMUSAGE_REPORT`` setting (:issue:`2576`) @@ -868,7 +874,8 @@ Documentation (:issue:`2477`, fixes :issue:`2475`) - FAQ: rewrite note on Python 3 support on Windows (:issue:`2690`) - Rearrange selector sections (:issue:`2705`) -- Remove ``__nonzero__`` from :class:`SelectorList` docs (:issue:`2683`) +- Remove ``__nonzero__`` from :class:`~scrapy.selector.SelectorList` + docs (:issue:`2683`) - Mention how to disable request filtering in documentation of :setting:`DUPEFILTER_CLASS` setting (:issue:`2714`) - Add sphinx_rtd_theme to docs setup readme (:issue:`2668`) @@ -2327,7 +2334,7 @@ Scrapy 0.18.0 (released 2013-08-09) - Moved persistent (on disk) queues to a separate project (queuelib_) which scrapy now depends on - Add scrapy commands using external libraries (:issue:`260`) - Added ``--pdb`` option to ``scrapy`` command line tool -- Added :meth:`XPathSelector.remove_namespaces` which allows to remove all namespaces from XML documents for convenience (to work with namespace-less XPaths). Documented in :ref:`topics-selectors`. +- Added :meth:`XPathSelector.remove_namespaces ` which allows to remove all namespaces from XML documents for convenience (to work with namespace-less XPaths). Documented in :ref:`topics-selectors`. - Several improvements to spider contracts - New default middleware named MetaRefreshMiddldeware that handles meta-refresh html tag redirections, - MetaRefreshMiddldeware and RedirectMiddleware have different priorities to address #62 @@ -2448,7 +2455,7 @@ Scrapy changes: - added options ``-o`` and ``-t`` to the :command:`runspider` command - documented :doc:`topics/autothrottle` and added to extensions installed by default. You still need to enable it with :setting:`AUTOTHROTTLE_ENABLED` - major Stats Collection refactoring: removed separation of global/per-spider stats, removed stats-related signals (``stats_spider_opened``, etc). Stats are much simpler now, backward compatibility is kept on the Stats Collector API and signals. -- added :meth:`~scrapy.contrib.spidermiddleware.SpiderMiddleware.process_start_requests` method to spider middlewares +- added :meth:`~scrapy.spidermiddlewares.SpiderMiddleware.process_start_requests` method to spider middlewares - dropped Signals singleton. Signals should now be accesed through the Crawler.signals attribute. See the signals documentation for more info. - dropped Signals singleton. Signals should now be accesed through the Crawler.signals attribute. See the signals documentation for more info. - dropped Stats Collector singleton. Stats can now be accessed through the Crawler.stats attribute. See the stats collection documentation for more info. @@ -2609,7 +2616,8 @@ The numbers like #NNN reference tickets in the old issue tracker (Trac) which is New features and improvements ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -- Passed item is now sent in the ``item`` argument of the :signal:`item_passed` (#273) +- Passed item is now sent in the ``item`` argument of the :signal:`item_passed + ` (#273) - Added verbose option to ``scrapy version`` command, useful for bug reports (#298) - HTTP cache now stored by default in the project data dir (#279) - Added project data storage directory (#276, #277) diff --git a/docs/topics/items.rst b/docs/topics/items.rst index 260f5882c..fbc888f47 100644 --- a/docs/topics/items.rst +++ b/docs/topics/items.rst @@ -240,6 +240,10 @@ Item objects Items replicate the standard `dict API`_, including its constructor. The only additional attribute provided by Items is: + .. automethod:: copy + + .. automethod:: deepcopy + .. attribute:: fields A dictionary containing *all declared fields* for this Item, not only diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 75e0af63b..350c3c4b9 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -796,6 +796,7 @@ Default: ``True`` Whether or not to use passive mode when initiating FTP transfers. +.. reqmeta:: ftp_password .. setting:: FTP_PASSWORD FTP_PASSWORD @@ -814,6 +815,7 @@ in ``Request`` meta. .. _RFC 1635: https://tools.ietf.org/html/rfc1635 +.. reqmeta:: ftp_user .. setting:: FTP_USER FTP_USER From 39c9a3cc1c6a59e03abc6c2d4080ddc211289a8d Mon Sep 17 00:00:00 2001 From: John Bampton Date: Sat, 5 Oct 2019 10:09:14 +1000 Subject: [PATCH 097/387] Fix case of GitHub. --- .github/ISSUE_TEMPLATE/bug_report.md | 2 +- .github/ISSUE_TEMPLATE/feature_request.md | 2 +- docs/news.rst | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/.github/ISSUE_TEMPLATE/bug_report.md b/.github/ISSUE_TEMPLATE/bug_report.md index 66821171f..8ca10109b 100644 --- a/.github/ISSUE_TEMPLATE/bug_report.md +++ b/.github/ISSUE_TEMPLATE/bug_report.md @@ -8,7 +8,7 @@ about: Report a problem to help us improve Thanks for taking an interest in Scrapy! If you have a question that starts with "How to...", please see the Scrapy Community page: https://scrapy.org/community/. -The Github issue tracker's purpose is to deal with bug reports and feature requests for the project itself. +The GitHub issue tracker's purpose is to deal with bug reports and feature requests for the project itself. Keep in mind that by filing an issue, you are expected to comply with Scrapy's Code of Conduct, including treating everyone with respect: https://github.com/scrapy/scrapy/blob/master/CODE_OF_CONDUCT.md diff --git a/.github/ISSUE_TEMPLATE/feature_request.md b/.github/ISSUE_TEMPLATE/feature_request.md index df5127b4c..e05273fe2 100644 --- a/.github/ISSUE_TEMPLATE/feature_request.md +++ b/.github/ISSUE_TEMPLATE/feature_request.md @@ -8,7 +8,7 @@ about: Suggest an idea for an enhancement or new feature Thanks for taking an interest in Scrapy! If you have a question that starts with "How to...", please see the Scrapy Community page: https://scrapy.org/community/. -The Github issue tracker's purpose is to deal with bug reports and feature requests for the project itself. +The GitHub issue tracker's purpose is to deal with bug reports and feature requests for the project itself. Keep in mind that by filing an issue, you are expected to comply with Scrapy's Code of Conduct, including treating everyone with respect: https://github.com/scrapy/scrapy/blob/master/CODE_OF_CONDUCT.md diff --git a/docs/news.rst b/docs/news.rst index aac750601..2bcfe4d1c 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -2590,7 +2590,7 @@ Code rearranged and removed - `w3lib`_ (several functions from ``scrapy.utils.{http,markup,multipart,response,url}``, done in :rev:`2584`) - `scrapely`_ (was ``scrapy.contrib.ibl``, done in :rev:`2586`) - Removed unused function: ``scrapy.utils.request.request_info()`` (:rev:`2577`) -- Removed googledir project from ``examples/googledir``. There's now a new example project called ``dirbot`` available on github: https://github.com/scrapy/dirbot +- Removed googledir project from ``examples/googledir``. There's now a new example project called ``dirbot`` available on GitHub: https://github.com/scrapy/dirbot - Removed support for default field values in Scrapy items (:rev:`2616`) - Removed experimental crawlspider v2 (:rev:`2632`) - Removed scheduler middleware to simplify architecture. Duplicates filter is now done in the scheduler itself, using the same dupe fltering class as before (``DUPEFILTER_CLASS`` setting) (:rev:`2640`) From f52148143be614779ccfdb34cac995b16f8264ff Mon Sep 17 00:00:00 2001 From: akhter wahab Date: Mon, 7 Oct 2019 23:28:33 +0500 Subject: [PATCH 098/387] Add dmg, iso & apk to ignored other extensions --- scrapy/linkextractors/__init__.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index ebf3cd7d8..6049c312c 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -35,7 +35,7 @@ IGNORED_EXTENSIONS = [ 'odp', # other - 'css', 'pdf', 'exe', 'bin', 'rss', 'zip', 'rar', + 'css', 'pdf', 'exe', 'bin', 'rss', 'zip', 'rar', 'dmg', 'iso', 'apk' ] From 5f168cd459cfc026de1e4d0b43c45ea740fe1dd5 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Wed, 2 Oct 2019 14:08:08 -0300 Subject: [PATCH 099/387] Response.follow_all --- docs/topics/request-response.rst | 4 ++ scrapy/http/response/__init__.py | 63 +++++++++++++++++----- scrapy/http/response/text.py | 61 ++++++++++++++++++--- tests/test_http_response.py | 91 ++++++++++++++++++++++++++++++++ 4 files changed, 199 insertions(+), 20 deletions(-) diff --git a/docs/topics/request-response.rst b/docs/topics/request-response.rst index 727c67482..9fe3c7518 100644 --- a/docs/topics/request-response.rst +++ b/docs/topics/request-response.rst @@ -701,6 +701,8 @@ Response objects .. automethod:: Response.follow + .. automethod:: Response.follow_all + .. _urlparse.urljoin: https://docs.python.org/2/library/urlparse.html#urlparse.urljoin @@ -790,6 +792,8 @@ TextResponse objects .. automethod:: TextResponse.follow + .. automethod:: TextResponse.follow_all + .. method:: TextResponse.body_as_unicode() The same as :attr:`text`, but available as a method. This method is diff --git a/scrapy/http/response/__init__.py b/scrapy/http/response/__init__.py index b0a526b72..96359705e 100644 --- a/scrapy/http/response/__init__.py +++ b/scrapy/http/response/__init__.py @@ -113,8 +113,8 @@ class Response(object_ref): It accepts the same arguments as ``Request.__init__`` method, but ``url`` can be a relative URL or a ``scrapy.link.Link`` object, not only an absolute URL. - - :class:`~.TextResponse` provides a :meth:`~.TextResponse.follow` + + :class:`~.TextResponse` provides a :meth:`~.TextResponse.follow` method which supports selectors in addition to absolute/relative URLs and Link objects. """ @@ -123,14 +123,51 @@ class Response(object_ref): elif url is None: raise ValueError("url can't be None") url = self.urljoin(url) - return Request(url, callback, - method=method, - headers=headers, - body=body, - cookies=cookies, - meta=meta, - encoding=encoding, - priority=priority, - dont_filter=dont_filter, - errback=errback, - cb_kwargs=cb_kwargs) + return Request( + url=url, + callback=callback, + method=method, + headers=headers, + body=body, + cookies=cookies, + meta=meta, + encoding=encoding, + priority=priority, + dont_filter=dont_filter, + errback=errback, + cb_kwargs=cb_kwargs, + ) + + def follow_all(self, urls, callback=None, method='GET', headers=None, body=None, + cookies=None, meta=None, encoding='utf-8', priority=0, + dont_filter=False, errback=None, cb_kwargs=None): + # type: (...) -> Generator[Request, None, None] + """ + Return an iterable of :class:`~.Request` instance to follow all links + in ``urls``. It accepts the same arguments as ``Request.__init__`` method, + but elements of ``urls`` can be relative URLs or ``scrapy.link.Link`` objects + not only absolute URLs. + + :class:`~.TextResponse` provides a :meth:`~.TextResponse.follow_all` + method which supports selectors in addition to absolute/relative URLs + and Link objects. + """ + if not hasattr(urls, '__iter__'): + raise TypeError("'urls' argument must be an iterable") + return ( + self.follow( + url=url, + callback=callback, + method=method, + headers=headers, + body=body, + cookies=cookies, + meta=meta, + encoding=encoding, + priority=priority, + dont_filter=dont_filter, + errback=errback, + cb_kwargs=cb_kwargs, + ) + for url in urls + ) diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 339913d4e..81e3f9b28 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -13,7 +13,6 @@ from w3lib.encoding import html_to_unicode, resolve_encoding, \ html_body_declared_encoding, http_content_type_encoding from w3lib.html import strip_html5_whitespace -from scrapy.http.request import Request from scrapy.http.response import Response from scrapy.utils.response import get_base_url from scrapy.utils.python import memoizemethod_noargs, to_native_str @@ -44,7 +43,7 @@ class TextResponse(Response): if isinstance(body, six.text_type): if self._encoding is None: raise TypeError('Cannot convert unicode body - %s has no encoding' % - type(self).__name__) + type(self).__name__) self._body = body.encode(self._encoding) else: super(TextResponse, self)._set_body(body) @@ -90,8 +89,8 @@ class TextResponse(Response): if self._cached_benc is None: content_type = to_native_str(self.headers.get(b'Content-Type', b'')) benc, ubody = html_to_unicode(content_type, self.body, - auto_detect_fun=self._auto_detect_fun, - default_encoding=self._DEFAULT_ENCODING) + auto_detect_fun=self._auto_detect_fun, + default_encoding=self._DEFAULT_ENCODING) self._cached_benc = benc self._cached_ubody = ubody return self._cached_benc @@ -129,7 +128,7 @@ class TextResponse(Response): Return a :class:`~.Request` instance to follow a link ``url``. It accepts the same arguments as ``Request.__init__`` method, but ``url`` can be not only an absolute URL, but also - + * a relative URL; * a scrapy.link.Link object (e.g. a link extractor result); * an attribute Selector (not SelectorList) - e.g. @@ -137,7 +136,7 @@ class TextResponse(Response): ``response.xpath('//img/@src')[0]``. * a Selector for ```` or ```` element, e.g. ``response.css('a.my_link')[0]``. - + See :ref:`response-follow-example` for usage examples. """ if isinstance(url, parsel.Selector): @@ -145,7 +144,9 @@ class TextResponse(Response): elif isinstance(url, parsel.SelectorList): raise ValueError("SelectorList is not supported") encoding = self.encoding if encoding is None else encoding - return super(TextResponse, self).follow(url, callback, + return super(TextResponse, self).follow( + url=url, + callback=callback, method=method, headers=headers, body=body, @@ -158,6 +159,52 @@ class TextResponse(Response): cb_kwargs=cb_kwargs, ) + def follow_all(self, urls=None, callback=None, method='GET', headers=None, body=None, + cookies=None, meta=None, encoding=None, priority=0, + dont_filter=False, errback=None, cb_kwargs=None, + css=None, xpath=None): + # type: (...) -> Generator[Request, None, None] + """ + A generator that produces :class:`~.Request` instances to follow all + links in ``urls``. It accepts the same arguments as the :class:`~.Request` + initializer, except that each ``urls`` element does not need to be an absolute + URL, it can be any of the following: + + * a relative URL; + * a scrapy.link.Link object (e.g. a link extractor result); + * an attribute Selector (not SelectorList) - e.g. + ``response.css('a::attr(href)')[0]`` or + ``response.xpath('//img/@src')[0]``. + * a Selector for ```` or ```` element, e.g. + ``response.css('a.my_link')[0]``. + + In addition, ``css`` and ``xpath`` arguments are accepted to perform the link extraction + within the ``follow_all`` method (only one of ``urls``, ``css`` and ``xpath`` are accepted). + """ + if len(list(filter(None, (urls, css, xpath)))) > 1: + raise ValueError('Please supply only one of the following arguments: {urls, css, xpath}') + if css: + urls = self.css(css) + elif xpath: + urls = self.xpath(xpath) + return ( + self.follow( + url=url, + callback=callback, + method=method, + headers=headers, + body=body, + cookies=cookies, + meta=meta, + encoding=encoding, + priority=priority, + dont_filter=dont_filter, + errback=errback, + cb_kwargs=cb_kwargs, + ) + for url in urls + ) + def _url_from_selector(sel): # type: (parsel.Selector) -> str diff --git a/tests/test_http_response.py b/tests/test_http_response.py index cd5c3486e..134856c78 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -143,6 +143,8 @@ class BaseResponseTest(unittest.TestCase): r.css('body') r.xpath('//body') + # Response.follow + def test_follow_url_absolute(self): self._assert_followed_url('http://foo.example.com', 'http://foo.example.com') @@ -166,6 +168,64 @@ class BaseResponseTest(unittest.TestCase): def test_follow_whitespace_link(self): self._assert_followed_url(Link('http://example.com/foo '), 'http://example.com/foo%20') + + # Response.follow_all + + def test_follow_all_absolute(self): + url_list = ['http://example.org', 'http://www.example.org', + 'http://example.com', 'http://www.example.com'] + self._assert_followed_all_urls(url_list, url_list) + + def test_follow_all_relative(self): + relative = ['foo', 'bar', 'foo/bar', 'bar/foo'] + absolute = [ + 'http://example.com/foo', + 'http://example.com/bar', + 'http://example.com/foo/bar', + 'http://example.com/bar/foo', + ] + self._assert_followed_all_urls(relative, absolute) + + def test_follow_all_links(self): + absolute = [ + 'http://example.com/foo', + 'http://example.com/bar', + 'http://example.com/foo/bar', + 'http://example.com/bar/foo', + ] + links = map(Link, absolute) + self._assert_followed_all_urls(links, absolute) + + def test_follow_all_invalid(self): + r = self.response_class("http://example.com") + with self.assertRaises(TypeError): + list(r.follow_all(None)) + with self.assertRaises(TypeError): + list(r.follow_all(12345)) + with self.assertRaises(ValueError): + list(r.follow_all([None])) + + def test_follow_all_whitespace(self): + relative = ['foo ', 'bar ', 'foo/bar ', 'bar/foo '] + absolute = [ + 'http://example.com/foo%20', + 'http://example.com/bar%20', + 'http://example.com/foo/bar%20', + 'http://example.com/bar/foo%20', + ] + self._assert_followed_all_urls(relative, absolute) + + def test_follow_all_whitespace_links(self): + absolute = [ + 'http://example.com/foo ', + 'http://example.com/bar ', + 'http://example.com/foo/bar ', + 'http://example.com/bar/foo ', + ] + links = map(Link, absolute) + expected = [u.replace(' ', '%20') for u in absolute] + self._assert_followed_all_urls(links, expected) + def _assert_followed_url(self, follow_obj, target_url, response=None): if response is None: response = self._links_response() @@ -173,6 +233,14 @@ class BaseResponseTest(unittest.TestCase): self.assertEqual(req.url, target_url) return req + def _assert_followed_all_urls(self, follow_obj, target_urls, response=None): + if response is None: + response = self._links_response() + followed = response.follow_all(follow_obj) + for req, target in zip(followed, target_urls): + self.assertEqual(req.url, target) + yield req + def _links_response(self): body = get_testdata('link_extractor', 'sgml_linkextractor.html') resp = self.response_class('http://example.com/index', body=body) @@ -483,6 +551,29 @@ class TextResponseTest(BaseResponseTest): ) self.assertEqual(req.encoding, 'cp1251') + def test_follow_all_css(self): + expected = [ + 'http://example.com/sample3.html', + 'http://example.com/innertag.html', + ] + response = self._links_response() + extracted = [r.url for r in response.follow_all(css='a[href*="example.com"]')] + self.assertEqual(expected, extracted) + + def test_follow_all_xpath(self): + expected = [ + 'http://example.com/sample3.html', + 'http://example.com/innertag.html', + ] + response = self._links_response() + extracted = response.follow_all(xpath='//a[contains(@href, "example.com")]') + self.assertEqual(expected, [r.url for r in extracted]) + + def test_follow_all_exception(self): + response = self._links_response() + with self.assertRaises(ValueError): + response.follow_all(css='a[href*="example.com"]', xpath='//a[contains(@href, "example.com")]') + class HtmlResponseTest(TextResponseTest): From 877ef4269e5683bf87832e816cf20a3cac352440 Mon Sep 17 00:00:00 2001 From: akhter wahab Date: Wed, 9 Oct 2019 16:03:44 +0500 Subject: [PATCH 100/387] Add .webm to ignored video extensions --- scrapy/linkextractors/__init__.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index 6049c312c..5f6df9c73 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -28,7 +28,7 @@ IGNORED_EXTENSIONS = [ # video '3gp', 'asf', 'asx', 'avi', 'mov', 'mp4', 'mpg', 'qt', 'rm', 'swf', 'wmv', - 'm4a', 'm4v', 'flv', + 'm4a', 'm4v', 'flv', 'webm', # office suites 'xls', 'xlsx', 'ppt', 'pptx', 'pps', 'doc', 'docx', 'odt', 'ods', 'odg', From a25a2d5ee4755046d736efa0a03898d9d18671f9 Mon Sep 17 00:00:00 2001 From: akhter wahab Date: Wed, 9 Oct 2019 16:05:39 +0500 Subject: [PATCH 101/387] Add .tar.xz to ignored other extensions --- scrapy/linkextractors/__init__.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index 5f6df9c73..0405462d6 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -35,7 +35,7 @@ IGNORED_EXTENSIONS = [ 'odp', # other - 'css', 'pdf', 'exe', 'bin', 'rss', 'zip', 'rar', 'dmg', 'iso', 'apk' + 'css', 'pdf', 'exe', 'bin', 'rss', 'zip', 'rar', 'dmg', 'iso', 'apk', 'tar.xz' ] From 532770df5226757498a941746cf94ae47043e101 Mon Sep 17 00:00:00 2001 From: akhter wahab Date: Wed, 9 Oct 2019 22:53:14 +0500 Subject: [PATCH 102/387] instead of .tar.xz adding .xz in others extensions --- scrapy/linkextractors/__init__.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index 0405462d6..3c75e683d 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -35,7 +35,7 @@ IGNORED_EXTENSIONS = [ 'odp', # other - 'css', 'pdf', 'exe', 'bin', 'rss', 'zip', 'rar', 'dmg', 'iso', 'apk', 'tar.xz' + 'css', 'pdf', 'exe', 'bin', 'rss', 'zip', 'rar', 'dmg', 'iso', 'apk', 'xz' ] From e1fa1fd8ad62cb8958d21d3a33648ed4de6af757 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Thu, 10 Oct 2019 00:36:38 -0300 Subject: [PATCH 103/387] TextResponse.follow_all: skip invalid links --- scrapy/http/response/text.py | 26 +++++++--- .../sgml_linkextractor_no_href.html | 25 ++++++++++ tests/test_http_response.py | 47 ++++++++++++++++--- 3 files changed, 85 insertions(+), 13 deletions(-) create mode 100644 tests/sample_data/link_extractor/sgml_linkextractor_no_href.html diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 81e3f9b28..ccacce550 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -180,13 +180,27 @@ class TextResponse(Response): In addition, ``css`` and ``xpath`` arguments are accepted to perform the link extraction within the ``follow_all`` method (only one of ``urls``, ``css`` and ``xpath`` are accepted). + + Note that when using the ``css`` or ``xpath`` parameters, this method will not produce + requests for selectors from which links cannot be obtained (for instance, anchor tags + without ``href`` attribute) """ - if len(list(filter(None, (urls, css, xpath)))) > 1: - raise ValueError('Please supply only one of the following arguments: {urls, css, xpath}') - if css: - urls = self.css(css) - elif xpath: - urls = self.xpath(xpath) + arg_count = len(list(filter(None, (urls, css, xpath)))) + if arg_count != 1: + raise ValueError('Please supply exactly one of the following arguments: {urls, css, xpath}') + if not urls: + urls = [] + if css: + selector_method = getattr(self, 'css') + expression = css + elif xpath: + selector_method = getattr(self, 'xpath') + expression = xpath + for selector in selector_method(expression): + try: + urls.append(_url_from_selector(selector)) + except ValueError: + pass return ( self.follow( url=url, diff --git a/tests/sample_data/link_extractor/sgml_linkextractor_no_href.html b/tests/sample_data/link_extractor/sgml_linkextractor_no_href.html new file mode 100644 index 000000000..0b01cede8 --- /dev/null +++ b/tests/sample_data/link_extractor/sgml_linkextractor_no_href.html @@ -0,0 +1,25 @@ + + + + Sample page with anchor tags containing no href attribute, to test the TextResponse.follow_all method + + + + + + + \ No newline at end of file diff --git a/tests/test_http_response.py b/tests/test_http_response.py index 134856c78..6d3c5cb9d 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -198,12 +198,20 @@ class BaseResponseTest(unittest.TestCase): def test_follow_all_invalid(self): r = self.response_class("http://example.com") - with self.assertRaises(TypeError): - list(r.follow_all(None)) - with self.assertRaises(TypeError): - list(r.follow_all(12345)) - with self.assertRaises(ValueError): - list(r.follow_all([None])) + if self.response_class == Response: + with self.assertRaises(TypeError): + list(r.follow_all(urls=None)) + with self.assertRaises(TypeError): + list(r.follow_all(urls=12345)) + with self.assertRaises(ValueError): + list(r.follow_all(urls=[None])) + else: + with self.assertRaises(ValueError): + list(r.follow_all(urls=None)) + with self.assertRaises(TypeError): + list(r.follow_all(urls=12345)) + with self.assertRaises(ValueError): + list(r.follow_all(urls=[None])) def test_follow_all_whitespace(self): relative = ['foo ', 'bar ', 'foo/bar ', 'bar/foo '] @@ -246,6 +254,11 @@ class BaseResponseTest(unittest.TestCase): resp = self.response_class('http://example.com/index', body=body) return resp + def _links_response_no_href(self): + body = get_testdata('link_extractor', 'sgml_linkextractor_no_href.html') + resp = self.response_class('http://example.com/index', body=body) + return resp + class TextResponseTest(BaseResponseTest): @@ -560,6 +573,16 @@ class TextResponseTest(BaseResponseTest): extracted = [r.url for r in response.follow_all(css='a[href*="example.com"]')] self.assertEqual(expected, extracted) + def test_follow_all_css_skip_invalid(self): + expected = [ + 'http://example.com/page/1/', + 'http://example.com/page/3/', + 'http://example.com/page/4/', + ] + response = self._links_response_no_href() + extracted = [r.url for r in response.follow_all(css='.pagination a')] + self.assertEqual(expected, extracted) + def test_follow_all_xpath(self): expected = [ 'http://example.com/sample3.html', @@ -569,7 +592,17 @@ class TextResponseTest(BaseResponseTest): extracted = response.follow_all(xpath='//a[contains(@href, "example.com")]') self.assertEqual(expected, [r.url for r in extracted]) - def test_follow_all_exception(self): + def test_follow_all_xpath_skip_invalid(self): + expected = [ + 'http://example.com/page/1/', + 'http://example.com/page/3/', + 'http://example.com/page/4/', + ] + response = self._links_response_no_href() + extracted = [r.url for r in response.follow_all(xpath='//div[@id="pagination"]/a')] + self.assertEqual(expected, extracted) + + def test_follow_all_too_many_arguments(self): response = self._links_response() with self.assertRaises(ValueError): response.follow_all(css='a[href*="example.com"]', xpath='//a[contains(@href, "example.com")]') From f0db1b4b6601e71caafe91f7970c9bfaa3d55395 Mon Sep 17 00:00:00 2001 From: "Matsievskiy S.V" Date: Fri, 11 Oct 2019 00:55:18 +0300 Subject: [PATCH 104/387] update zsh completion --- extras/scrapy_zsh_completion | 223 ++++++++++++++++++++++++++++++++--- 1 file changed, 204 insertions(+), 19 deletions(-) diff --git a/extras/scrapy_zsh_completion b/extras/scrapy_zsh_completion index 564991aa8..86c52c36c 100644 --- a/extras/scrapy_zsh_completion +++ b/extras/scrapy_zsh_completion @@ -1,25 +1,210 @@ #compdef scrapy - -# zsh completion for the Scrapy command-line tool - _scrapy() { - local curcontext="$curcontext" cmd spiders + local context state state_descr line typeset -A opt_args - cmd=$words[2] - - case "$cmd" in - crawl|edit|check) - spiders=$(scrapy list 2>/dev/null) || spiders="" - if [[ -n "$spiders" ]]; then - compadd `echo $spiders` - fi - ;; - *) - if [[ CURRENT -eq 2 ]]; then - _arguments '*: :(check crawl edit fetch genspider list parse runspider settings shell startproject version view)' - fi - ;; + _arguments \ + "(- 1 *)--help[Help]" \ + "1: :->command" \ + "*:: :->args" + + case $state in + command) + _scrapy_cmds + ;; + args) + case $words[1] in + bench) + _scrapy_glb_opts + ;; + fetch) + local options=( + '--headers[print response HTTP headers instead of body]' + '--no-redirect[do not handle HTTP 3xx status codes and print response as-is]' + '--spider[use this spider]:spider:_scrapy_spiders' + '1::URL:_httpie_urls' + ) + _scrapy_glb_opts $options + ;; + genspider) + local options=( + {-l,--list}'[List available templates]' + {-e,--edit}'[Edit spider after creating it]' + '--force[If the spider already exists, overwrite it with the template]' + {-d,--dump=}'[Dump template to standard output]:template:(basic crawl csvfeed xmlfeed)' + {-t,--template=}'[Uses a custom template]:template:(basic crawl csvfeed xmlfeed)' + '1:name:(NAME)' + '2:domain:_httpie_urls' + ) + _scrapy_glb_opts $options + ;; + runspider) + local options=( + {-o,--output}'[dump scraped items into FILE (use - for stdout)]:file:_files' + {-t,--output-format}'[format to use for dumping items with -o]:format:(FORMAT)' + '*-a[set spider argument (may be repeated)]:value pair:(NAME=VALUE)' + '1:spider file:_files -g \*.py' + ) + _scrapy_glb_opts $options + ;; + settings) + local options=( + '--get=[print raw setting value]:option:(SETTING)' + '--getbool=[print setting value, interpreted as a boolean]:option:(SETTING)' + '--getint=[print setting value, interpreted as an integer]:option:(SETTING)' + '--getfloat=[print setting value, interpreted as a float]:option:(SETTING)' + '--getlist=[print setting value, interpreted as a list]:option:(SETTING)' + ) + _scrapy_glb_opts $options + ;; + shell) + local options=( + '-c[evaluate the code in the shell, print the result and exit]:code:(CODE)' + '--no-redirect[do not handle HTTP 3xx status codes and print response as-is]' + '--spider[use this spider]:spider:_scrapy_spiders' + '::file:_files -g \*.http' + '::URL:_httpie_urls' + ) + _scrapy_glb_opts $options + ;; + startproject) + local options=( + '1:name:(NAME)' + '2:dir:_dir_list' + ) + _scrapy_glb_opts $options + ;; + version) + local options=( + {-v,--verbose}'[also display twisted/python/platform info (useful for bug reports)]' + ) + _scrapy_glb_opts $options + ;; + view) + local options=( + '--no-redirect[do not handle HTTP 3xx status codes and print response as-is]' + '--spider[use this spider]:spider:_scrapy_spiders' + '1:URL:_httpie_urls' + ) + _scrapy_glb_opts $options + ;; + check) + local options=( + '(- 1 *)'{-l,--list}'[only list contracts, without checking them]' + {-v,--verbose}'[print contract tests for all spiders]' + '1:spider:_scrapy_spiders' + ) + _scrapy_glb_opts $options + ;; + crawl) + local options=( + {-o,--output}'[dump scraped items into FILE (use - for stdout)]:file:_files' + {-t,--output-format}'[format to use for dumping items with -o]:format:(FORMAT)' + '*-a[set spider argument (may be repeated)]:value pair:(NAME=VALUE)' + '1:spider:_scrapy_spiders' + ) + _scrapy_glb_opts $options + ;; + edit) + local options=( + '1:spider:_scrapy_spiders' + ) + _scrapy_glb_opts $options + ;; + list) + _scrapy_glb_opts + ;; + parse) + local options=( + '*-a[set spider argument (may be repeated)]:value pair:(NAME=VALUE)' + '--spider[use this spider without looking for one]:spider:_scrapy_spiders' + '--pipelines[process items through pipelines]' + "--nolinks[don't show links to follow (extracted requests)]" + "--noitems[don't show scraped items]" + '--nocolour[avoid using pygments to colorize the output]' + {-r,--rules}'[use CrawlSpider rules to discover the callback]' + {-c,--callback=}'[use this callback for parsing, instead looking for a callback]:callback:(CALLBACK)' + {-m,--meta=}'[inject extra meta into the Request, it must be a valid raw json string]:meta:(META)' + '--cbkwargs=[inject extra callback kwargs into the Request, it must be a valid raw json string]:arguments:(CBKWARGS)' + {-d,--depth=}'[maximum depth for parsing requests (default: 1)]:depth:(DEPTH)' + {-v,--verbose}'[print each depth level one by one]' + '1:URL:_httpie_urls' + ) + _scrapy_glb_opts $options + ;; + esac + ;; esac } -_scrapy \ No newline at end of file +_scrapy_cmds() { + local -a commands project_commands + commands=( + 'bench:Run quick benchmark test' + 'fetch:Fetch a URL using the Scrapy downloader' + 'genspider:Generate new spider using pre-defined templates' + 'runspider:Run a self-contained spider (without creating a project)' + 'settings:Get settings values' + 'shell:Interactive scraping console' + 'startproject:Create new project' + 'version:Print Scrapy version' + 'view:Open URL in browser, as seen by Scrapy' + ) + project_commands=( + 'check:Check spider contracts' + 'crawl:Run a spider' + 'edit:Edit spider' + 'list:List available spiders' + 'parse:Parse URL (using its spider) and print the results' + ) + if [[ $(scrapy -h | grep -s "no active project") == "" ]]; then + commands=(${commands[@]} ${project_commands[@]}) + fi + _describe -t common-commands 'common commands' commands +} + +_scrapy_glb_opts() { + local -a options + options=( + '(- *)'{-h,--help}'[show this help message and exit]' + '(--nolog)--logfile=[log file. if omitted stderr will be used]:file:_files' + '--pidfile=[write process ID to FILE]:file:_files' + '--profile=[write python cProfile stats to FILE]:file:_files' + '(--nolog)'{-L,--loglevel=}'[log level (default: INFO)]:log level:(DEBUG INFO WARN ERROR)' + '(-L --loglevel --logfile)--nolog[disable logging completely]' + '--pdb[enable pdb on failure]' + '*'{-s,--set=}'[set/override setting (may be repeated)]:value pair:(NAME=VALUE)' + ) + options=(${options[@]} "$@") + _arguments $options +} + +_httpie_urls() { + + local ret=1 + + if ! [[ -prefix [-+.a-z0-9]#:// ]]; then + local expl + compset -S '[^:/]*' && compstate[to_end]='' + _wanted url-schemas expl 'URL schema' compadd -S '' http:// https:// && ret=0 + else + _urls && ret=0 + fi + + return $ret + +} + +_scrapy_spiders() { + + local ret=1 + + if [[ $(scrapy -h | grep -s "no active project") == "" ]]; then + compadd -S '' $(scrapy list) && ret=0 + else + compadd -S '' SPIDER && ret=0 + fi + + return $ret +} + +_scrapy $@ From 12f1e468e9f071e8f5c00921d0fbada460e41c57 Mon Sep 17 00:00:00 2001 From: Purva Udai Date: Sun, 13 Oct 2019 15:55:27 +0530 Subject: [PATCH 105/387] Issue #3731 --- scrapy/extensions/feedexport.py | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/scrapy/extensions/feedexport.py b/scrapy/extensions/feedexport.py index ce2846eba..981efee55 100644 --- a/scrapy/extensions/feedexport.py +++ b/scrapy/extensions/feedexport.py @@ -199,9 +199,9 @@ class FeedExporter(object): def __init__(self, settings): self.settings = settings - self.urifmt = settings['FEED_URI'] - if not self.urifmt: + if not settings['FEED_URI']: raise NotConfigured + self.urifmt=str(settings['FEED_URI']) self.format = settings['FEED_FORMAT'].lower() self.export_encoding = settings['FEED_EXPORT_ENCODING'] self.storages = self._load_components('FEED_STORAGES') From b970851299bf561af7ff0bb21caca9b6c3f296af Mon Sep 17 00:00:00 2001 From: elacuesta Date: Mon, 14 Oct 2019 13:35:06 -0300 Subject: [PATCH 106/387] Update scrapy/http/response/__init__.py MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves --- scrapy/http/response/__init__.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/http/response/__init__.py b/scrapy/http/response/__init__.py index 96359705e..cdaababac 100644 --- a/scrapy/http/response/__init__.py +++ b/scrapy/http/response/__init__.py @@ -143,7 +143,7 @@ class Response(object_ref): dont_filter=False, errback=None, cb_kwargs=None): # type: (...) -> Generator[Request, None, None] """ - Return an iterable of :class:`~.Request` instance to follow all links + Return an iterable of :class:`~.Request` instances to follow all links in ``urls``. It accepts the same arguments as ``Request.__init__`` method, but elements of ``urls`` can be relative URLs or ``scrapy.link.Link`` objects not only absolute URLs. From ba840c5a6bea95333b3b78fb767ac6f092d02177 Mon Sep 17 00:00:00 2001 From: elacuesta Date: Mon, 14 Oct 2019 13:35:24 -0300 Subject: [PATCH 107/387] Update scrapy/http/response/__init__.py MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves --- scrapy/http/response/__init__.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/http/response/__init__.py b/scrapy/http/response/__init__.py index cdaababac..92fa01621 100644 --- a/scrapy/http/response/__init__.py +++ b/scrapy/http/response/__init__.py @@ -145,7 +145,7 @@ class Response(object_ref): """ Return an iterable of :class:`~.Request` instances to follow all links in ``urls``. It accepts the same arguments as ``Request.__init__`` method, - but elements of ``urls`` can be relative URLs or ``scrapy.link.Link`` objects + but elements of ``urls`` can be relative URLs or :class:`~scrapy.link.Link` objects, not only absolute URLs. :class:`~.TextResponse` provides a :meth:`~.TextResponse.follow_all` From 498d33aac37d5d7a0955e9b409fc3d5ff8627dc0 Mon Sep 17 00:00:00 2001 From: elacuesta Date: Mon, 14 Oct 2019 13:35:54 -0300 Subject: [PATCH 108/387] Update scrapy/http/response/text.py MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves --- scrapy/http/response/text.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index ccacce550..bb5a4eb9d 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -183,7 +183,7 @@ class TextResponse(Response): Note that when using the ``css`` or ``xpath`` parameters, this method will not produce requests for selectors from which links cannot be obtained (for instance, anchor tags - without ``href`` attribute) + without an ``href`` attribute) """ arg_count = len(list(filter(None, (urls, css, xpath)))) if arg_count != 1: From c7c54f5453eaa16f2b6d6441be68b0a272606ec9 Mon Sep 17 00:00:00 2001 From: elacuesta Date: Mon, 14 Oct 2019 13:47:44 -0300 Subject: [PATCH 109/387] Update scrapy/http/response/text.py MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves --- scrapy/http/response/text.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index bb5a4eb9d..ecb582b4a 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -179,7 +179,7 @@ class TextResponse(Response): ``response.css('a.my_link')[0]``. In addition, ``css`` and ``xpath`` arguments are accepted to perform the link extraction - within the ``follow_all`` method (only one of ``urls``, ``css`` and ``xpath`` are accepted). + within the ``follow_all`` method (only one of ``urls``, ``css`` and ``xpath`` is accepted). Note that when using the ``css`` or ``xpath`` parameters, this method will not produce requests for selectors from which links cannot be obtained (for instance, anchor tags From 9d5398e7f2834940cbf9c2efa312bf5d96ffba98 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Mon, 14 Oct 2019 13:57:46 -0300 Subject: [PATCH 110/387] TextResponse.follow_all: improve docs --- scrapy/http/response/text.py | 28 +++++++++++++++------------- 1 file changed, 15 insertions(+), 13 deletions(-) diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index ecb582b4a..b2907baa4 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -129,13 +129,14 @@ class TextResponse(Response): It accepts the same arguments as ``Request.__init__`` method, but ``url`` can be not only an absolute URL, but also - * a relative URL; - * a scrapy.link.Link object (e.g. a link extractor result); - * an attribute Selector (not SelectorList) - e.g. + * a relative URL + * a :class:`~scrapy.link.Link` object, e.g. the result of + :ref:`topics-link-extractors` + * a :class:`~scrapy.selector.Selector` object for a ```` or ```` element, e.g. + ``response.css('a.my_link')[0]`` + * an attribute :class:`~scrapy.selector.Selector` (not SelectorList), e.g. ``response.css('a::attr(href)')[0]`` or - ``response.xpath('//img/@src')[0]``. - * a Selector for ```` or ```` element, e.g. - ``response.css('a.my_link')[0]``. + ``response.xpath('//img/@src')[0]`` See :ref:`response-follow-example` for usage examples. """ @@ -170,13 +171,14 @@ class TextResponse(Response): initializer, except that each ``urls`` element does not need to be an absolute URL, it can be any of the following: - * a relative URL; - * a scrapy.link.Link object (e.g. a link extractor result); - * an attribute Selector (not SelectorList) - e.g. + * a relative URL + * a :class:`~scrapy.link.Link` object, e.g. the result of + :ref:`topics-link-extractors` + * a :class:`~scrapy.selector.Selector` object for a ```` or ```` element, e.g. + ``response.css('a.my_link')[0]`` + * an attribute :class:`~scrapy.selector.Selector` (not SelectorList), e.g. ``response.css('a::attr(href)')[0]`` or - ``response.xpath('//img/@src')[0]``. - * a Selector for ```` or ```` element, e.g. - ``response.css('a.my_link')[0]``. + ``response.xpath('//img/@src')[0]`` In addition, ``css`` and ``xpath`` arguments are accepted to perform the link extraction within the ``follow_all`` method (only one of ``urls``, ``css`` and ``xpath`` is accepted). @@ -187,7 +189,7 @@ class TextResponse(Response): """ arg_count = len(list(filter(None, (urls, css, xpath)))) if arg_count != 1: - raise ValueError('Please supply exactly one of the following arguments: {urls, css, xpath}') + raise ValueError('Please supply exactly one of the following arguments: urls, css, xpath') if not urls: urls = [] if css: From 7b1e69dec4cb1f1c66e9df6f754cdd280cff7e71 Mon Sep 17 00:00:00 2001 From: Baron Hou Date: Tue, 15 Oct 2019 20:51:15 +0800 Subject: [PATCH 111/387] =?UTF-8?q?reponse=20=E2=86=92=20response=20(#4079?= =?UTF-8?q?)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/topics/developer-tools.rst | 2 +- docs/topics/downloader-middleware.rst | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/topics/developer-tools.rst b/docs/topics/developer-tools.rst index dcf8af365..bf14643be 100644 --- a/docs/topics/developer-tools.rst +++ b/docs/topics/developer-tools.rst @@ -203,7 +203,7 @@ where our quotes are coming from: First click on the request with the name ``scroll``. On the right you can now inspect the request. In ``Headers`` you'll find details about the request headers, such as the URL, the method, the IP-address, -and so on. We'll ignore the other tabs and click directly on ``Reponse``. +and so on. We'll ignore the other tabs and click directly on ``Response``. What you should see in the ``Preview`` pane is the rendered HTML-code, that is exactly what we saw when we called ``view(response)`` in the diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 539832618..2892b9b79 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -534,7 +534,7 @@ defines the methods described below. :param spider: the spider which generated the request :type spider: :class:`~scrapy.spiders.Spider` object - :param request: the request to find cached reponse for + :param request: the request to find cached response for :type request: :class:`~scrapy.http.Request` object .. method:: store_response(spider, request, response) From d72ed46fe01f95e77554e58588675d5a1ff47e0a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Tue, 15 Oct 2019 16:03:42 +0200 Subject: [PATCH 112/387] Improve how extra Item API members are introduced in the documentation --- docs/topics/items.rst | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/topics/items.rst b/docs/topics/items.rst index fbc888f47..d70e7428b 100644 --- a/docs/topics/items.rst +++ b/docs/topics/items.rst @@ -237,8 +237,8 @@ Item objects Return a new Item optionally initialized from the given argument. - Items replicate the standard `dict API`_, including its constructor. The - only additional attribute provided by Items is: + Items replicate the standard `dict API`_, including its constructor, and + also provide the following additional API members: .. automethod:: copy From 2a4d4a466aa6cede4829b62c7287332a627877ed Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Tue, 15 Oct 2019 11:52:12 -0300 Subject: [PATCH 113/387] TextResponse.follow_all: Simplify implementation --- scrapy/http/response/text.py | 12 +++++------- 1 file changed, 5 insertions(+), 7 deletions(-) diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index b2907baa4..25e115bf9 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -191,14 +191,12 @@ class TextResponse(Response): if arg_count != 1: raise ValueError('Please supply exactly one of the following arguments: urls, css, xpath') if not urls: - urls = [] if css: - selector_method = getattr(self, 'css') - expression = css - elif xpath: - selector_method = getattr(self, 'xpath') - expression = xpath - for selector in selector_method(expression): + selector_list = self.css(css) + if xpath: + selector_list = self.xpath(xpath) + urls = [] + for selector in selector_list: try: urls.append(_url_from_selector(selector)) except ValueError: From 2c6f7fee6456168f4870293c442258a2470b1b48 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Tue, 15 Oct 2019 13:48:14 -0300 Subject: [PATCH 114/387] TextResponse.follow_all: invoke Response.follow_all --- scrapy/http/response/text.py | 29 +++++++++++++---------------- 1 file changed, 13 insertions(+), 16 deletions(-) diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 25e115bf9..f782f6217 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -201,22 +201,19 @@ class TextResponse(Response): urls.append(_url_from_selector(selector)) except ValueError: pass - return ( - self.follow( - url=url, - callback=callback, - method=method, - headers=headers, - body=body, - cookies=cookies, - meta=meta, - encoding=encoding, - priority=priority, - dont_filter=dont_filter, - errback=errback, - cb_kwargs=cb_kwargs, - ) - for url in urls + return super(TextResponse, self).follow_all( + urls=urls, + callback=callback, + method=method, + headers=headers, + body=body, + cookies=cookies, + meta=meta, + encoding=encoding, + priority=priority, + dont_filter=dont_filter, + errback=errback, + cb_kwargs=cb_kwargs, ) From c9614a5bdd1c8a3f50eeace148c59f0b52e293a2 Mon Sep 17 00:00:00 2001 From: Bulat Date: Wed, 16 Oct 2019 12:07:19 +0300 Subject: [PATCH 115/387] Fixed BOT_NAME documentation --- docs/topics/settings.rst | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 75e0af63b..dbe3e9e44 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -229,8 +229,7 @@ BOT_NAME Default: ``'scrapybot'`` The name of the bot implemented by this Scrapy project (also known as the -project name). This will be used to construct the User-Agent by default, and -also for logging. +project name) and also for logging. It's automatically populated with your project name when you create your project with the :command:`startproject` command. From 84be6a941e3c4e80802dd52e43eb8e7064892cdb Mon Sep 17 00:00:00 2001 From: Bulat Date: Wed, 16 Oct 2019 14:04:07 +0300 Subject: [PATCH 116/387] Refactor sentence. --- docs/topics/settings.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index dbe3e9e44..d41381dd1 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -229,7 +229,7 @@ BOT_NAME Default: ``'scrapybot'`` The name of the bot implemented by this Scrapy project (also known as the -project name) and also for logging. +project name). This name will be used for the logging too. It's automatically populated with your project name when you create your project with the :command:`startproject` command. From 865d58fd1b5676dcca44b64970b5cd48fc1ab715 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Jos=C3=A9=20Alberto=20Orejuela=20Garc=C3=ADa?= Date: Thu, 3 Oct 2019 18:06:19 +0200 Subject: [PATCH 117/387] Make punctuation consistent --- CODE_OF_CONDUCT.md | 2 +- README.rst | 18 +++++++++--------- 2 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CODE_OF_CONDUCT.md b/CODE_OF_CONDUCT.md index d477168eb..d1cd3e517 100644 --- a/CODE_OF_CONDUCT.md +++ b/CODE_OF_CONDUCT.md @@ -68,7 +68,7 @@ members of the project's leadership. ## Attribution This Code of Conduct is adapted from the [Contributor Covenant][homepage], version 1.4, -available at [http://contributor-covenant.org/version/1/4][version] +available at [http://contributor-covenant.org/version/1/4][version]. [homepage]: http://contributor-covenant.org [version]: http://contributor-covenant.org/version/1/4/ diff --git a/README.rst b/README.rst index bd82bff06..87eaac2af 100644 --- a/README.rst +++ b/README.rst @@ -34,8 +34,8 @@ Scrapy is a fast high-level web crawling and web scraping framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing. -For more information including a list of features check the Scrapy homepage at: -https://scrapy.org +Check the Scrapy homepage at https://scrapy.org for more information, +including a list of features. Requirements ============ @@ -50,8 +50,8 @@ The quick way:: pip install scrapy -For more details see the install section in the documentation: -https://docs.scrapy.org/en/latest/intro/install.html +See the install section in the documentation at +https://docs.scrapy.org/en/latest/intro/install.html for more details. Documentation ============= @@ -62,17 +62,17 @@ directory. Releases ======== -You can find release notes at https://docs.scrapy.org/en/latest/news.html +You can check https://docs.scrapy.org/en/latest/news.html for release notes. Community (blog, twitter, mail list, IRC) ========================================= -See https://scrapy.org/community/ +See https://scrapy.org/community/ for details. Contributing ============ -See https://docs.scrapy.org/en/master/contributing.html +See https://docs.scrapy.org/en/master/contributing.html for details. Code of Conduct --------------- @@ -86,9 +86,9 @@ Please report unacceptable behavior to opensource@scrapinghub.com. Companies using Scrapy ====================== -See https://scrapy.org/companies/ +See https://scrapy.org/companies/ for a list. Commercial Support ================== -See https://scrapy.org/support/ +See https://scrapy.org/support/ for details. From 6df6b6dd6a5afd9172c03f6b9e8ce8e180867339 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Mon, 21 Oct 2019 03:56:45 -0300 Subject: [PATCH 118/387] Initializer -> __init__ --- scrapy/http/response/text.py | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index f782f6217..74017b5aa 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -167,9 +167,9 @@ class TextResponse(Response): # type: (...) -> Generator[Request, None, None] """ A generator that produces :class:`~.Request` instances to follow all - links in ``urls``. It accepts the same arguments as the :class:`~.Request` - initializer, except that each ``urls`` element does not need to be an absolute - URL, it can be any of the following: + links in ``urls``. It accepts the same arguments as the :class:`~.Request`'s + ``__init__`` method, except that each ``urls`` element does not need to be + an absolute URL, it can be any of the following: * a relative URL * a :class:`~scrapy.link.Link` object, e.g. the result of From 0fbd1ff4a91c68e552a3f919824c60437fc9a141 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 21 Oct 2019 14:06:45 +0200 Subject: [PATCH 119/387] =?UTF-8?q?constructor=20=E2=86=92=20=5F=5Finit=5F?= =?UTF-8?q?=5F=20method?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/topics/link-extractors.rst | 6 +++--- scrapy/linkextractors/lxmlhtml.py | 4 ++-- 2 files changed, 5 insertions(+), 5 deletions(-) diff --git a/docs/topics/link-extractors.rst b/docs/topics/link-extractors.rst index f9936a498..2119cb8f8 100644 --- a/docs/topics/link-extractors.rst +++ b/docs/topics/link-extractors.rst @@ -6,9 +6,9 @@ Link Extractors A link extractor is an object that extracts links from responses. -The constructor of :class:`~scrapy.linkextractors.lxmlhtml.LxmlLinkExtractor` -takes settings that determine which links may be extracted. -:class:`LxmlLinkExtractor.extract_links +The ``__init__`` method of +:class:`~scrapy.linkextractors.lxmlhtml.LxmlLinkExtractor` takes settings that +determine which links may be extracted. :class:`LxmlLinkExtractor.extract_links ` returns a list of matching :class:`scrapy.link.Link` objects from a :class:`~scrapy.http.Response` object. diff --git a/scrapy/linkextractors/lxmlhtml.py b/scrapy/linkextractors/lxmlhtml.py index 41091ba23..37003720f 100644 --- a/scrapy/linkextractors/lxmlhtml.py +++ b/scrapy/linkextractors/lxmlhtml.py @@ -120,8 +120,8 @@ class LxmlLinkExtractor(FilteringLinkExtractor): """Returns a list of :class:`~scrapy.link.Link` objects from the specified :class:`response `. - Only links that match the settings passed to the link extractor - constructor are returned. + Only links that match the settings passed to the ``__init__`` method of + the link extractor are returned. Duplicate links are omitted. """ From 68a7d05ed86331ae70778c83a131e6dcc8b4ffc8 Mon Sep 17 00:00:00 2001 From: Ammar Najjar Date: Mon, 21 Oct 2019 15:42:24 +0200 Subject: [PATCH 120/387] docs: use __init__ method instead of constructor Issue #4086 --- docs/conf.py | 2 +- docs/news.rst | 10 ++++---- docs/topics/email.rst | 4 ++-- docs/topics/exporters.rst | 40 +++++++++++++++---------------- docs/topics/extensions.rst | 2 +- docs/topics/items.rst | 8 +++---- docs/topics/loaders.rst | 26 ++++++++++---------- docs/topics/request-response.rst | 14 +++++------ scrapy/exporters.py | 2 +- scrapy/extensions/feedexport.py | 2 +- scrapy/utils/datatypes.py | 2 +- scrapy/utils/misc.py | 6 ++--- scrapy/utils/python.py | 2 +- sep/sep-009.rst | 12 +++++----- tests/test_http_request.py | 6 ++--- tests/test_http_response.py | 2 +- tests/test_loader.py | 12 +++++----- tests/test_spider.py | 4 ++-- tests/test_utils_misc/__init__.py | 10 ++++---- 19 files changed, 83 insertions(+), 83 deletions(-) diff --git a/docs/conf.py b/docs/conf.py index 34dd5bcb7..6ab5959d5 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -237,7 +237,7 @@ coverage_ignore_pyobjects = [ r'\bContractsManager\b$', # For default contracts we only want to document their general purpose in - # their constructor, the methods they reimplement to achieve that purpose + # their __init__ method, the methods they reimplement to achieve that purpose # should be irrelevant to developers using those contracts. r'\w+Contract\.(adjust_request_args|(pre|post)_process)$', diff --git a/docs/news.rst b/docs/news.rst index 59317f5eb..c1f548060 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -84,12 +84,12 @@ New features convenient way to build JSON requests (:issue:`3504`, :issue:`3505`) * A ``process_request`` callback passed to the :class:`~scrapy.spiders.Rule` - constructor now receives the :class:`~scrapy.http.Response` object that + ``__init__`` method now receives the :class:`~scrapy.http.Response` object that originated the request as its second argument (:issue:`3682`) * A new ``restrict_text`` parameter for the :attr:`LinkExtractor ` - constructor allows filtering links by linking text (:issue:`3622`, + ``__init__`` method allows filtering links by linking text (:issue:`3622`, :issue:`3635`) * A new :setting:`FEED_STORAGE_S3_ACL` setting allows defining a custom ACL @@ -255,7 +255,7 @@ The following deprecated APIs have been removed (:issue:`3578`): * From :class:`~scrapy.selector.Selector`: - * ``_root`` (both the constructor argument and the object property, use + * ``_root`` (both the ``__init__`` method argument and the object property, use ``root``) * ``extract_unquoted`` (use ``getall``) @@ -2479,7 +2479,7 @@ Scrapy changes: - removed ``ENCODING_ALIASES`` setting, as encoding auto-detection has been moved to the `w3lib`_ library - promoted :ref:`topics-djangoitem` to main contrib - LogFormatter method now return dicts(instead of strings) to support lazy formatting (:issue:`164`, :commit:`dcef7b0`) -- downloader handlers (:setting:`DOWNLOAD_HANDLERS` setting) now receive settings as the first argument of the constructor +- downloader handlers (:setting:`DOWNLOAD_HANDLERS` setting) now receive settings as the first argument of the ``__init__`` method - replaced memory usage acounting with (more portable) `resource`_ module, removed ``scrapy.utils.memory`` module - removed signal: ``scrapy.mail.mail_sent`` - removed ``TRACK_REFS`` setting, now :ref:`trackrefs ` is always enabled @@ -2693,7 +2693,7 @@ API changes - ``Request.copy()`` and ``Request.replace()`` now also copies their ``callback`` and ``errback`` attributes (#231) - Removed ``UrlFilterMiddleware`` from ``scrapy.contrib`` (already disabled by default) - Offsite middelware doesn't filter out any request coming from a spider that doesn't have a allowed_domains attribute (#225) -- Removed Spider Manager ``load()`` method. Now spiders are loaded in the constructor itself. +- Removed Spider Manager ``load()`` method. Now spiders are loaded in the ``__init__`` method itself. - Changes to Scrapy Manager (now called "Crawler"): - ``scrapy.core.manager.ScrapyManager`` class renamed to ``scrapy.crawler.Crawler`` - ``scrapy.core.manager.scrapymanager`` singleton moved to ``scrapy.project.crawler`` diff --git a/docs/topics/email.rst b/docs/topics/email.rst index 949cdc638..12eedf2cd 100644 --- a/docs/topics/email.rst +++ b/docs/topics/email.rst @@ -21,7 +21,7 @@ Quick example ============= There are two ways to instantiate the mail sender. You can instantiate it using -the standard constructor:: +the standard ``__init__`` method:: from scrapy.mail import MailSender mailer = MailSender() @@ -111,7 +111,7 @@ uses `Twisted non-blocking IO`_, like the rest of the framework. Mail settings ============= -These settings define the default constructor values of the :class:`MailSender` +These settings define the default ``__init__`` method values of the :class:`MailSender` class, and can be used to configure e-mail notifications in your project without writing any code (for those extensions and code that uses :class:`MailSender`). diff --git a/docs/topics/exporters.rst b/docs/topics/exporters.rst index a698a6a4e..d7ab7a5cc 100644 --- a/docs/topics/exporters.rst +++ b/docs/topics/exporters.rst @@ -87,8 +87,8 @@ described next. 1. Declaring a serializer in the field -------------------------------------- -If you use :class:`~.Item` you can declare a serializer in the -:ref:`field metadata `. The serializer must be +If you use :class:`~.Item` you can declare a serializer in the +:ref:`field metadata `. The serializer must be a callable which receives a value and returns its serialized form. Example:: @@ -144,7 +144,7 @@ BaseItemExporter defining what fields to export, whether to export empty fields, or which encoding to use. - These features can be configured through the constructor arguments which + These features can be configured through the `__init__` method arguments which populate their respective instance attributes: :attr:`fields_to_export`, :attr:`export_empty_fields`, :attr:`encoding`, :attr:`indent`. @@ -246,8 +246,8 @@ XmlItemExporter :param item_element: The name of each item element in the exported XML. :type item_element: str - The additional keyword arguments of this constructor are passed to the - :class:`BaseItemExporter` constructor. + The additional keyword arguments of this `__init__` method are passed to the + :class:`BaseItemExporter` `__init__` method A typical output of this exporter would be:: @@ -306,9 +306,9 @@ CsvItemExporter multi-valued fields, if found. :type include_headers_line: str - The additional keyword arguments of this constructor are passed to the - :class:`BaseItemExporter` constructor, and the leftover arguments to the - `csv.writer`_ constructor, so you can use any ``csv.writer`` constructor + The additional keyword arguments of this `__init__` method are passed to the + :class:`BaseItemExporter` `__init__` method, and the leftover arguments to the + `csv.writer`_ `__init__` method, so you can use any ``csv.writer`` `__init__` method argument to customize this exporter. A typical output of this exporter would be:: @@ -334,8 +334,8 @@ PickleItemExporter For more information, refer to the `pickle module documentation`_. - The additional keyword arguments of this constructor are passed to the - :class:`BaseItemExporter` constructor. + The additional keyword arguments of this `__init__` method are passed to the + :class:`BaseItemExporter` `__init__` method. Pickle isn't a human readable format, so no output examples are provided. @@ -351,8 +351,8 @@ PprintItemExporter :param file: the file-like object to use for exporting the data. Its ``write`` method should accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) - The additional keyword arguments of this constructor are passed to the - :class:`BaseItemExporter` constructor. + The additional keyword arguments of this `__init__` method are passed to the + :class:`BaseItemExporter` `__init__` method A typical output of this exporter would be:: @@ -367,10 +367,10 @@ JsonItemExporter .. class:: JsonItemExporter(file, \**kwargs) Exports Items in JSON format to the specified file-like object, writing all - objects as a list of objects. The additional constructor arguments are - passed to the :class:`BaseItemExporter` constructor, and the leftover - arguments to the `JSONEncoder`_ constructor, so you can use any - `JSONEncoder`_ constructor argument to customize this exporter. + objects as a list of objects. The additional `__init__` method arguments are + passed to the :class:`BaseItemExporter` `__init__` method, and the leftover + arguments to the `JSONEncoder`_ `__init__` method, so you can use any + `JSONEncoder`_ `__init__` method argument to customize this exporter. :param file: the file-like object to use for exporting the data. Its ``write`` method should accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) @@ -398,10 +398,10 @@ JsonLinesItemExporter .. class:: JsonLinesItemExporter(file, \**kwargs) Exports Items in JSON format to the specified file-like object, writing one - JSON-encoded item per line. The additional constructor arguments are passed - to the :class:`BaseItemExporter` constructor, and the leftover arguments to - the `JSONEncoder`_ constructor, so you can use any `JSONEncoder`_ - constructor argument to customize this exporter. + JSON-encoded item per line. The additional `__init__` method arguments are passed + to the :class:`BaseItemExporter` `__init__` method and the leftover arguments to + the `JSONEncoder`_ `__init__` method, so you can use any `JSONEncoder`_ + `__init__` method argument to customize this exporter. :param file: the file-like object to use for exporting the data. Its ``write`` method should accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) diff --git a/docs/topics/extensions.rst b/docs/topics/extensions.rst index 72c2290b5..0a7455ec9 100644 --- a/docs/topics/extensions.rst +++ b/docs/topics/extensions.rst @@ -28,7 +28,7 @@ Loading & activating extensions Extensions are loaded and activated at startup by instantiating a single instance of the extension class. Therefore, all the extension initialization -code must be performed in the class constructor (``__init__`` method). +code must be performed in the class ``__init__`` method. To make an extension available, add it to the :setting:`EXTENSIONS` setting in your Scrapy settings. In :setting:`EXTENSIONS`, each extension is represented diff --git a/docs/topics/items.rst b/docs/topics/items.rst index d70e7428b..8ea2635cc 100644 --- a/docs/topics/items.rst +++ b/docs/topics/items.rst @@ -16,12 +16,12 @@ especially in a larger project with many spiders. To define common output data format Scrapy provides the :class:`Item` class. :class:`Item` objects are simple containers used to collect the scraped data. They provide a `dictionary-like`_ API with a convenient syntax for declaring -their available fields. +their available fields. -Various Scrapy components use extra information provided by Items: +Various Scrapy components use extra information provided by Items: exporters look at declared fields to figure out columns to export, serialization can be customized using Item fields metadata, :mod:`trackref` -tracks Item instances to help find memory leaks +tracks Item instances to help find memory leaks (see :ref:`topics-leaks-trackrefs`), etc. .. _dictionary-like: https://docs.python.org/2/library/stdtypes.html#dict @@ -237,7 +237,7 @@ Item objects Return a new Item optionally initialized from the given argument. - Items replicate the standard `dict API`_, including its constructor, and + Items replicate the standard `dict API`_, including its `__init__` method and also provide the following additional API members: .. automethod:: copy diff --git a/docs/topics/loaders.rst b/docs/topics/loaders.rst index 1c2f1da4d..db1175a13 100644 --- a/docs/topics/loaders.rst +++ b/docs/topics/loaders.rst @@ -26,7 +26,7 @@ Using Item Loaders to populate items To use an Item Loader, you must first instantiate it. You can either instantiate it with a dict-like object (e.g. Item or dict) or without one, in -which case an Item is automatically instantiated in the Item Loader constructor +which case an Item is automatically instantiated in the Item Loader ``__init__`` method using the Item class specified in the :attr:`ItemLoader.default_item_class` attribute. @@ -265,7 +265,7 @@ There are several ways to modify Item Loader context values: loader.context['unit'] = 'cm' 2. On Item Loader instantiation (the keyword arguments of Item Loader - constructor are stored in the Item Loader context):: + ``__init__`` methodare stored in the Item Loader context):: loader = ItemLoader(product, unit='cm') @@ -494,7 +494,7 @@ ItemLoader objects .. attribute:: default_item_class An Item class (or factory), used to instantiate items when not given in - the constructor. + the `__init__` method .. attribute:: default_input_processor @@ -509,15 +509,15 @@ ItemLoader objects .. attribute:: default_selector_class The class used to construct the :attr:`selector` of this - :class:`ItemLoader`, if only a response is given in the constructor. - If a selector is given in the constructor this attribute is ignored. + :class:`ItemLoader`, if only a response is given in the `__init__` method + If a selector is given in the `__init__` method this attribute is ignored. This attribute is sometimes overridden in subclasses. .. attribute:: selector The :class:`~scrapy.selector.Selector` object to extract data from. - It's either the selector given in the constructor or one created from - the response given in the constructor using the + It's either the selector given in the `__init__` methodor one created from + the response given in the `__init__` methodusing the :attr:`default_selector_class`. This attribute is meant to be read-only. @@ -642,7 +642,7 @@ Here is a list of all built-in processors: .. class:: Identity The simplest processor, which doesn't do anything. It returns the original - values unchanged. It doesn't receive any constructor arguments, nor does it + values unchanged. It doesn't receive any `__init__` method arguments, nor does it accept Loader contexts. Example:: @@ -656,7 +656,7 @@ Here is a list of all built-in processors: Returns the first non-null/non-empty value from the values received, so it's typically used as an output processor to single-valued fields. - It doesn't receive any constructor arguments, nor does it accept Loader contexts. + It doesn't receive any `__init__` methodarguments, nor does it accept Loader contexts. Example:: @@ -667,7 +667,7 @@ Here is a list of all built-in processors: .. class:: Join(separator=u' ') - Returns the values joined with the separator given in the constructor, which + Returns the values joined with the separator given in the `__init__` method which defaults to ``u' '``. It doesn't accept Loader contexts. When using the default separator, this processor is equivalent to the @@ -705,7 +705,7 @@ Here is a list of all built-in processors: those which do, this processor will pass the currently active :ref:`Loader context ` through that parameter. - The keyword arguments passed in the constructor are used as the default + The keyword arguments passed in the `__init__` methodare used as the default Loader context values passed to each function call. However, the final Loader context values passed to functions are overridden with the currently active Loader context accessible through the :meth:`ItemLoader.context` @@ -749,12 +749,12 @@ Here is a list of all built-in processors: ['HELLO, 'THIS', 'IS', 'SCRAPY'] As with the Compose processor, functions can receive Loader contexts, and - constructor keyword arguments are used as default context values. See + `__init__` method keyword arguments are used as default context values. See :class:`Compose` processor for more info. .. class:: SelectJmes(json_path) - Queries the value using the json path provided to the constructor and returns the output. + Queries the value using the json path provided to the `__init__` methodand returns the output. Requires jmespath (https://github.com/jmespath/jmespath.py) to run. This processor takes only one input at a time. diff --git a/docs/topics/request-response.rst b/docs/topics/request-response.rst index 727c67482..d253064f2 100644 --- a/docs/topics/request-response.rst +++ b/docs/topics/request-response.rst @@ -137,7 +137,7 @@ Request objects A string containing the URL of this request. Keep in mind that this attribute contains the escaped URL, so it can differ from the URL passed in - the constructor. + the `__init__` method This attribute is read-only. To change the URL of a Request use :meth:`replace`. @@ -400,7 +400,7 @@ fields with form data from :class:`Response` objects. .. class:: FormRequest(url, [formdata, ...]) - The :class:`FormRequest` class adds a new keyword parameter to the constructor. The + The :class:`FormRequest` class adds a new keyword parameter to the `__init__` method The remaining arguments are the same as for the :class:`Request` class and are not documented here. @@ -473,7 +473,7 @@ fields with form data from :class:`Response` objects. :type dont_click: boolean The other parameters of this class method are passed directly to the - :class:`FormRequest` constructor. + :class:`FormRequest` `__init__` method .. versionadded:: 0.10.3 The ``formname`` parameter. @@ -547,7 +547,7 @@ dealing with JSON requests. .. class:: JsonRequest(url, [... data, dumps_kwargs]) - The :class:`JsonRequest` class adds two new keyword parameters to the constructor. The + The :class:`JsonRequest` class adds two new keyword parameters to the `__init__` method The remaining arguments are the same as for the :class:`Request` class and are not documented here. @@ -556,7 +556,7 @@ dealing with JSON requests. :param data: is any JSON serializable object that needs to be JSON encoded and assigned to body. if :attr:`Request.body` argument is provided this parameter will be ignored. - if :attr:`Request.body` argument is not provided and data argument is provided :attr:`Request.method` will be + if :attr:`Request.body` argument is not provided and data argument is provided :attr:`Request.method` will be set to ``'POST'`` automatically. :type data: JSON serializable object @@ -721,7 +721,7 @@ TextResponse objects :class:`Response` class, which is meant to be used only for binary data, such as images, sounds or any media file. - :class:`TextResponse` objects support a new constructor argument, in + :class:`TextResponse` objects support a new `__init__` method argument, in addition to the base :class:`Response` objects. The remaining functionality is the same as for the :class:`Response` class and is not documented here. @@ -755,7 +755,7 @@ TextResponse objects A string with the encoding of this response. The encoding is resolved by trying the following mechanisms, in order: - 1. the encoding passed in the constructor ``encoding`` argument + 1. the encoding passed in the `__init__` method`` ``encoding`` argument 2. the encoding declared in the Content-Type HTTP header. If this encoding is not valid (ie. unknown), it is ignored and the next diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 6fc87ed18..2fdc86b63 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -31,7 +31,7 @@ class BaseItemExporter(object): def _configure(self, options, dont_fail=False): """Configure the exporter by poping options from the ``options`` dict. If dont_fail is set, it won't raise an exception on unexpected options - (useful for using with keyword arguments in subclasses constructors) + (useful for using with keyword arguments in subclasses __init__ methods) """ self.encoding = options.pop('encoding', None) self.fields_to_export = options.pop('fields_to_export', None) diff --git a/scrapy/extensions/feedexport.py b/scrapy/extensions/feedexport.py index ce2846eba..655c10482 100644 --- a/scrapy/extensions/feedexport.py +++ b/scrapy/extensions/feedexport.py @@ -106,7 +106,7 @@ class S3FeedStorage(BlockingFeedStorage): warnings.warn( "Initialising `scrapy.extensions.feedexport.S3FeedStorage` " "without AWS keys is deprecated. Please supply credentials or " - "use the `from_crawler()` constructor.", + "use the `from_crawler()` __init__ method.", category=ScrapyDeprecationWarning, stacklevel=2 ) diff --git a/scrapy/utils/datatypes.py b/scrapy/utils/datatypes.py index df2b99c28..0dfee1b27 100644 --- a/scrapy/utils/datatypes.py +++ b/scrapy/utils/datatypes.py @@ -246,7 +246,7 @@ class CaselessDict(dict): class MergeDict(object): """ A simple class for creating new "virtual" dictionaries that actually look - up values in more than one dictionary, passed in the constructor. + up values in more than one dictionary, passed in the __init__ method. If a key appears in more than one of the given dictionaries, only the first occurrence will be used. diff --git a/scrapy/utils/misc.py b/scrapy/utils/misc.py index f638adb25..c892bd438 100644 --- a/scrapy/utils/misc.py +++ b/scrapy/utils/misc.py @@ -123,14 +123,14 @@ def rel_has_nofollow(rel): def create_instance(objcls, settings, crawler, *args, **kwargs): """Construct a class instance using its ``from_crawler`` or - ``from_settings`` constructors, if available. + ``from_settings`` __init__ method, if available. At least one of ``settings`` and ``crawler`` needs to be different from ``None``. If ``settings `` is ``None``, ``crawler.settings`` will be used. - If ``crawler`` is ``None``, only the ``from_settings`` constructor will be + If ``crawler`` is ``None``, only the ``from_settings`` __init__ method will be tried. - ``*args`` and ``**kwargs`` are forwarded to the constructors. + ``*args`` and ``**kwargs`` are forwarded to the __init__ methods. Raises ``ValueError`` if both ``settings`` and ``crawler`` are ``None``. """ diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index c6140f885..0d645543a 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -303,7 +303,7 @@ class WeakKeyCache(object): def stringify_dict(dct_or_tuples, encoding='utf-8', keys_only=True): """Return a (new) dict with unicode keys (and values when "keys_only" is False) of the given dict converted to strings. ``dct_or_tuples`` can be a - dict or a list of tuples, like any dict constructor supports. + dict or a list of tuples, like any dict __init__ method supports. """ d = {} for k, v in six.iteritems(dict(dct_or_tuples)): diff --git a/sep/sep-009.rst b/sep/sep-009.rst index 232a536a8..929cedc61 100644 --- a/sep/sep-009.rst +++ b/sep/sep-009.rst @@ -38,7 +38,7 @@ singletons members of that object, as explained below: ``scrapy.core.manager.ExecutionManager``) - instantiated with a ``Settings`` object - - **crawler.settings**: ``scrapy.conf.Settings`` instance (passed in the constructor) + - **crawler.settings**: ``scrapy.conf.Settings`` instance (passed in the ``__init__`` method) - **crawler.extensions**: ``scrapy.extension.ExtensionManager`` instance - **crawler.engine**: ``scrapy.core.engine.ExecutionEngine`` instance - ``crawler.engine.scheduler`` @@ -55,7 +55,7 @@ singletons members of that object, as explained below: ``STATS_CLASS`` setting) - **crawler.log**: Logger class with methods replacing the current ``scrapy.log`` functions. Logging would be started (if enabled) on - ``Crawler`` constructor, so no log starting functions are required. + ``Crawler`` __init__ method, so no log starting functions are required. - ``crawler.log.msg`` - **crawler.signals**: signal handling @@ -69,12 +69,12 @@ Required code changes after singletons removal ============================================== All components (extensions, middlewares, etc) will receive this ``Crawler`` -object in their constructors, and this will be the only mechanism for accessing +object in their ``__init__`` methods, and this will be the only mechanism for accessing any other components (as opposed to importing each singleton from their respective module). This will also serve to stabilize the core API, something which we haven't documented so far (partly because of this). -So, for a typical middleware constructor code, instead of this: +So, for a typical middleware ``__init__`` method code, instead of this: :: @@ -125,13 +125,13 @@ Open issues to resolve - Should we pass ``Settings`` object to ``ScrapyCommand.add_options()``? - How should spiders access settings? - - Option 1. Pass ``Crawler`` object to spider constructors too + - Option 1. Pass ``Crawler`` object to spider ``__init__`` methods too - pro: one way to access all components (settings and signals being the most relevant to spiders) - con?: spider code can access (and control) any crawler component - since we don't want to support spiders messing with the crawler (write an extension or spider middleware if you need that) - - Option 2. Pass ``Settings`` object to spider constructors, which would + - Option 2. Pass ``Settings`` object to spider ``__init__`` methods, which would then be accessed through ``self.settings``, like logging which is accessed through ``self.log`` diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 16d7a1cb8..96a4fb141 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -25,7 +25,7 @@ class RequestTest(unittest.TestCase): default_meta = {} def test_init(self): - # Request requires url in the constructor + # Request requires url in the __init__ method self.assertRaises(Exception, self.request_class) # url argument must be basestring @@ -500,7 +500,7 @@ class FormRequestTest(RequestTest): formdata=(('foo', 'bar'), ('foo', 'baz'))) self.assertEqual(urlparse(req.url).hostname, 'www.example.com') self.assertEqual(urlparse(req.url).query, 'foo=bar&foo=baz') - + def test_from_response_override_duplicate_form_key(self): response = _buildresponse( """
@@ -657,7 +657,7 @@ class FormRequestTest(RequestTest): req = self.request_class.from_response(response, dont_click=True) fs = _qs(req) self.assertEqual(fs, {b'i1': [b'i1v'], b'i2': [b'i2v']}) - + def test_from_response_clickdata_does_not_ignore_image(self): response = _buildresponse( """ diff --git a/tests/test_http_response.py b/tests/test_http_response.py index cd5c3486e..dfc8562f3 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -534,7 +534,7 @@ class XmlResponseTest(TextResponseTest): r2 = self.response_class("http://www.example.com", body=body) self._assert_response_values(r2, 'iso-8859-1', body) - # make sure replace() preserves the explicit encoding passed in the constructor + # make sure replace() preserves the explicit encoding passed in the __init__ method body = b"""""" r3 = self.response_class("http://www.example.com", body=body, encoding='utf-8') body2 = b"New body" diff --git a/tests/test_loader.py b/tests/test_loader.py index 2725b001a..bcc1e6421 100644 --- a/tests/test_loader.py +++ b/tests/test_loader.py @@ -548,11 +548,11 @@ class SelectortemLoaderTest(unittest.TestCase): """) - def test_constructor(self): + def test_init_method(self): l = TestItemLoader() self.assertEqual(l.selector, None) - def test_constructor_errors(self): + def test_init_method_errors(self): l = TestItemLoader() self.assertRaises(RuntimeError, l.add_xpath, 'url', '//a/@href') self.assertRaises(RuntimeError, l.replace_xpath, 'url', '//a/@href') @@ -561,7 +561,7 @@ class SelectortemLoaderTest(unittest.TestCase): self.assertRaises(RuntimeError, l.replace_css, 'name', '#name::text') self.assertRaises(RuntimeError, l.get_css, '#name::text') - def test_constructor_with_selector(self): + def test_init_method_with_selector(self): sel = Selector(text=u"
marta
") l = TestItemLoader(selector=sel) self.assertIs(l.selector, sel) @@ -569,7 +569,7 @@ class SelectortemLoaderTest(unittest.TestCase): l.add_xpath('name', '//div/text()') self.assertEqual(l.get_output_value('name'), [u'Marta']) - def test_constructor_with_selector_css(self): + def test_init_method_with_selector_css(self): sel = Selector(text=u"
marta
") l = TestItemLoader(selector=sel) self.assertIs(l.selector, sel) @@ -577,14 +577,14 @@ class SelectortemLoaderTest(unittest.TestCase): l.add_css('name', 'div::text') self.assertEqual(l.get_output_value('name'), [u'Marta']) - def test_constructor_with_response(self): + def test_init_method_with_response(self): l = TestItemLoader(response=self.response) self.assertTrue(l.selector) l.add_xpath('name', '//div/text()') self.assertEqual(l.get_output_value('name'), [u'Marta']) - def test_constructor_with_response_css(self): + def test_init_method_with_response_css(self): l = TestItemLoader(response=self.response) self.assertTrue(l.selector) diff --git a/tests/test_spider.py b/tests/test_spider.py index 2220b8ffc..64ff40b61 100644 --- a/tests/test_spider.py +++ b/tests/test_spider.py @@ -42,12 +42,12 @@ class SpiderTest(unittest.TestCase): self.assertEqual(list(start_requests), []) def test_spider_args(self): - """Constructor arguments are assigned to spider attributes""" + """__init__ method arguments are assigned to spider attributes""" spider = self.spider_class('example.com', foo='bar') self.assertEqual(spider.foo, 'bar') def test_spider_without_name(self): - """Constructor arguments are assigned to spider attributes""" + """__init__ method arguments are assigned to spider attributes""" self.assertRaises(ValueError, self.spider_class) self.assertRaises(ValueError, self.spider_class, somearg='foo') diff --git a/tests/test_utils_misc/__init__.py b/tests/test_utils_misc/__init__.py index e109d5343..457a2aa78 100644 --- a/tests/test_utils_misc/__init__.py +++ b/tests/test_utils_misc/__init__.py @@ -109,11 +109,11 @@ class UtilsMiscTestCase(unittest.TestCase): else: mock.assert_called_once_with(*args, **kwargs) - # Check usage of correct constructor using four mocks: - # 1. with no alternative constructors - # 2. with from_settings() constructor - # 3. with from_crawler() constructor - # 4. with from_settings() and from_crawler() constructor + # Check usage of correct __init__ method using four mocks: + # 1. with no alternative __init__ methods + # 2. with from_settings() __init__ method + # 3. with from_crawler() __init__ method + # 4. with from_settings() and from_crawler() __init__ method spec_sets = ([], ['from_settings'], ['from_crawler'], ['from_settings', 'from_crawler']) for specs in spec_sets: From 3d4317bfe4697955de1c2450ce4cb0d12471a9a3 Mon Sep 17 00:00:00 2001 From: Roy Date: Mon, 21 Oct 2019 18:32:30 +0100 Subject: [PATCH 121/387] [tox.ini] Added python 3.8 fields https://github.com/scrapy/scrapy/issues/4085 --- tox.ini | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/tox.ini b/tox.ini index ffe7360d3..14afec23f 100644 --- a/tox.ini +++ b/tox.ini @@ -92,6 +92,10 @@ deps = {[testenv:py35]deps} basepython = python3.7 deps = {[testenv:py35]deps} +[testenv:py38] +basepython = python3.8 +deps = {[testenv:py35]deps} + [testenv:pypy3] basepython = pypy3 deps = {[testenv:py35]deps} @@ -128,6 +132,13 @@ deps = reppy robotexclusionrulesparser +[testenv:py38-extra-deps] +basepython = python3.8 +deps = + {[testenv:py35]deps} + reppy + robotexclusionrulesparser + [testenv:py27-extra-deps] basepython = python2.7 deps = From 4e939ca75d14daf5bc658fdbe5a97af6f4c3c498 Mon Sep 17 00:00:00 2001 From: Roy Date: Mon, 21 Oct 2019 18:33:18 +0100 Subject: [PATCH 122/387] [setup.py] Added python 3.8 fields https://github.com/scrapy/scrapy/issues/4085 --- setup.py | 1 + 1 file changed, 1 insertion(+) diff --git a/setup.py b/setup.py index 850456503..2f5fca4c9 100644 --- a/setup.py +++ b/setup.py @@ -56,6 +56,7 @@ setup( 'Programming Language :: Python :: 3.5', 'Programming Language :: Python :: 3.6', 'Programming Language :: Python :: 3.7', + 'Programming Language :: Python :: 3.8', 'Programming Language :: Python :: Implementation :: CPython', 'Programming Language :: Python :: Implementation :: PyPy', 'Topic :: Internet :: WWW/HTTP', From c12a075164e95e98a136499faf7bf8eefdce0300 Mon Sep 17 00:00:00 2001 From: Roy Date: Mon, 21 Oct 2019 18:34:15 +0100 Subject: [PATCH 123/387] [.travis.yml] Added python 3.8 fields https://github.com/scrapy/scrapy/issues/4085 --- .travis.yml | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/.travis.yml b/.travis.yml index 0190a7f4d..2ba504972 100644 --- a/.travis.yml +++ b/.travis.yml @@ -27,6 +27,10 @@ matrix: python: 3.7 - env: TOXENV=py37-extra-deps python: 3.7 + - env: TOXENV=py38 + python: 3.8 + - env: TOXENV=py37-extra-deps + python: 3.8 - env: TOXENV=docs python: 3.6 install: From da8cd9448de28bdeac7343bcd367e95237423789 Mon Sep 17 00:00:00 2001 From: Ammar Najjar Date: Mon, 21 Oct 2019 19:48:13 +0200 Subject: [PATCH 124/387] docs: always surround __init__ with `` in docs Issue #4086 --- docs/topics/exporters.rst | 36 ++++++++++++++++---------------- docs/topics/items.rst | 2 +- docs/topics/loaders.rst | 22 +++++++++---------- docs/topics/request-response.rst | 12 +++++------ scrapy/exporters.py | 2 +- scrapy/extensions/feedexport.py | 2 +- scrapy/utils/datatypes.py | 2 +- scrapy/utils/misc.py | 6 +++--- scrapy/utils/python.py | 2 +- sep/sep-009.rst | 2 +- tests/test_spider.py | 4 ++-- 11 files changed, 46 insertions(+), 46 deletions(-) diff --git a/docs/topics/exporters.rst b/docs/topics/exporters.rst index d7ab7a5cc..1b8a69ca3 100644 --- a/docs/topics/exporters.rst +++ b/docs/topics/exporters.rst @@ -144,7 +144,7 @@ BaseItemExporter defining what fields to export, whether to export empty fields, or which encoding to use. - These features can be configured through the `__init__` method arguments which + These features can be configured through the ``__init__`` method arguments which populate their respective instance attributes: :attr:`fields_to_export`, :attr:`export_empty_fields`, :attr:`encoding`, :attr:`indent`. @@ -246,8 +246,8 @@ XmlItemExporter :param item_element: The name of each item element in the exported XML. :type item_element: str - The additional keyword arguments of this `__init__` method are passed to the - :class:`BaseItemExporter` `__init__` method + The additional keyword arguments of this ``__init__`` method are passed to the + :class:`BaseItemExporter` ``__init__`` method A typical output of this exporter would be:: @@ -306,9 +306,9 @@ CsvItemExporter multi-valued fields, if found. :type include_headers_line: str - The additional keyword arguments of this `__init__` method are passed to the - :class:`BaseItemExporter` `__init__` method, and the leftover arguments to the - `csv.writer`_ `__init__` method, so you can use any ``csv.writer`` `__init__` method + The additional keyword arguments of this ``__init__`` method are passed to the + :class:`BaseItemExporter` ``__init__`` method, and the leftover arguments to the + `csv.writer`_ ``__init__`` method, so you can use any ``csv.writer`` ``__init__`` method argument to customize this exporter. A typical output of this exporter would be:: @@ -334,8 +334,8 @@ PickleItemExporter For more information, refer to the `pickle module documentation`_. - The additional keyword arguments of this `__init__` method are passed to the - :class:`BaseItemExporter` `__init__` method. + The additional keyword arguments of this ``__init__`` method are passed to the + :class:`BaseItemExporter` ``__init__`` method. Pickle isn't a human readable format, so no output examples are provided. @@ -351,8 +351,8 @@ PprintItemExporter :param file: the file-like object to use for exporting the data. Its ``write`` method should accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) - The additional keyword arguments of this `__init__` method are passed to the - :class:`BaseItemExporter` `__init__` method + The additional keyword arguments of this ``__init__`` method are passed to the + :class:`BaseItemExporter` ``__init__`` method A typical output of this exporter would be:: @@ -367,10 +367,10 @@ JsonItemExporter .. class:: JsonItemExporter(file, \**kwargs) Exports Items in JSON format to the specified file-like object, writing all - objects as a list of objects. The additional `__init__` method arguments are - passed to the :class:`BaseItemExporter` `__init__` method, and the leftover - arguments to the `JSONEncoder`_ `__init__` method, so you can use any - `JSONEncoder`_ `__init__` method argument to customize this exporter. + objects as a list of objects. The additional ``__init__`` method arguments are + passed to the :class:`BaseItemExporter` ``__init__`` method, and the leftover + arguments to the `JSONEncoder`_ ``__init__`` method, so you can use any + `JSONEncoder`_ ``__init__`` method argument to customize this exporter. :param file: the file-like object to use for exporting the data. Its ``write`` method should accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) @@ -398,10 +398,10 @@ JsonLinesItemExporter .. class:: JsonLinesItemExporter(file, \**kwargs) Exports Items in JSON format to the specified file-like object, writing one - JSON-encoded item per line. The additional `__init__` method arguments are passed - to the :class:`BaseItemExporter` `__init__` method and the leftover arguments to - the `JSONEncoder`_ `__init__` method, so you can use any `JSONEncoder`_ - `__init__` method argument to customize this exporter. + JSON-encoded item per line. The additional ``__init__`` method arguments are passed + to the :class:`BaseItemExporter` ``__init__`` method and the leftover arguments to + the `JSONEncoder`_ ``__init__`` method, so you can use any `JSONEncoder`_ + ``__init__`` method argument to customize this exporter. :param file: the file-like object to use for exporting the data. Its ``write`` method should accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) diff --git a/docs/topics/items.rst b/docs/topics/items.rst index 8ea2635cc..370409026 100644 --- a/docs/topics/items.rst +++ b/docs/topics/items.rst @@ -237,7 +237,7 @@ Item objects Return a new Item optionally initialized from the given argument. - Items replicate the standard `dict API`_, including its `__init__` method and + Items replicate the standard `dict API`_, including its ``__init__`` method and also provide the following additional API members: .. automethod:: copy diff --git a/docs/topics/loaders.rst b/docs/topics/loaders.rst index db1175a13..72610f645 100644 --- a/docs/topics/loaders.rst +++ b/docs/topics/loaders.rst @@ -494,7 +494,7 @@ ItemLoader objects .. attribute:: default_item_class An Item class (or factory), used to instantiate items when not given in - the `__init__` method + the ``__init__`` method .. attribute:: default_input_processor @@ -509,15 +509,15 @@ ItemLoader objects .. attribute:: default_selector_class The class used to construct the :attr:`selector` of this - :class:`ItemLoader`, if only a response is given in the `__init__` method - If a selector is given in the `__init__` method this attribute is ignored. + :class:`ItemLoader`, if only a response is given in the ``__init__`` method + If a selector is given in the ``__init__`` method this attribute is ignored. This attribute is sometimes overridden in subclasses. .. attribute:: selector The :class:`~scrapy.selector.Selector` object to extract data from. - It's either the selector given in the `__init__` methodor one created from - the response given in the `__init__` methodusing the + It's either the selector given in the ``__init__`` methodor one created from + the response given in the ``__init__`` methodusing the :attr:`default_selector_class`. This attribute is meant to be read-only. @@ -642,7 +642,7 @@ Here is a list of all built-in processors: .. class:: Identity The simplest processor, which doesn't do anything. It returns the original - values unchanged. It doesn't receive any `__init__` method arguments, nor does it + values unchanged. It doesn't receive any ``__init__`` method arguments, nor does it accept Loader contexts. Example:: @@ -656,7 +656,7 @@ Here is a list of all built-in processors: Returns the first non-null/non-empty value from the values received, so it's typically used as an output processor to single-valued fields. - It doesn't receive any `__init__` methodarguments, nor does it accept Loader contexts. + It doesn't receive any ``__init__`` methodarguments, nor does it accept Loader contexts. Example:: @@ -667,7 +667,7 @@ Here is a list of all built-in processors: .. class:: Join(separator=u' ') - Returns the values joined with the separator given in the `__init__` method which + Returns the values joined with the separator given in the ``__init__`` method which defaults to ``u' '``. It doesn't accept Loader contexts. When using the default separator, this processor is equivalent to the @@ -705,7 +705,7 @@ Here is a list of all built-in processors: those which do, this processor will pass the currently active :ref:`Loader context ` through that parameter. - The keyword arguments passed in the `__init__` methodare used as the default + The keyword arguments passed in the ``__init__`` methodare used as the default Loader context values passed to each function call. However, the final Loader context values passed to functions are overridden with the currently active Loader context accessible through the :meth:`ItemLoader.context` @@ -749,12 +749,12 @@ Here is a list of all built-in processors: ['HELLO, 'THIS', 'IS', 'SCRAPY'] As with the Compose processor, functions can receive Loader contexts, and - `__init__` method keyword arguments are used as default context values. See + ``__init__`` method keyword arguments are used as default context values. See :class:`Compose` processor for more info. .. class:: SelectJmes(json_path) - Queries the value using the json path provided to the `__init__` methodand returns the output. + Queries the value using the json path provided to the ``__init__`` methodand returns the output. Requires jmespath (https://github.com/jmespath/jmespath.py) to run. This processor takes only one input at a time. diff --git a/docs/topics/request-response.rst b/docs/topics/request-response.rst index d253064f2..bf6a02a1d 100644 --- a/docs/topics/request-response.rst +++ b/docs/topics/request-response.rst @@ -137,7 +137,7 @@ Request objects A string containing the URL of this request. Keep in mind that this attribute contains the escaped URL, so it can differ from the URL passed in - the `__init__` method + the ``__init__`` method This attribute is read-only. To change the URL of a Request use :meth:`replace`. @@ -400,7 +400,7 @@ fields with form data from :class:`Response` objects. .. class:: FormRequest(url, [formdata, ...]) - The :class:`FormRequest` class adds a new keyword parameter to the `__init__` method The + The :class:`FormRequest` class adds a new keyword parameter to the ``__init__`` method The remaining arguments are the same as for the :class:`Request` class and are not documented here. @@ -473,7 +473,7 @@ fields with form data from :class:`Response` objects. :type dont_click: boolean The other parameters of this class method are passed directly to the - :class:`FormRequest` `__init__` method + :class:`FormRequest` ``__init__`` method .. versionadded:: 0.10.3 The ``formname`` parameter. @@ -547,7 +547,7 @@ dealing with JSON requests. .. class:: JsonRequest(url, [... data, dumps_kwargs]) - The :class:`JsonRequest` class adds two new keyword parameters to the `__init__` method The + The :class:`JsonRequest` class adds two new keyword parameters to the ``__init__`` method The remaining arguments are the same as for the :class:`Request` class and are not documented here. @@ -721,7 +721,7 @@ TextResponse objects :class:`Response` class, which is meant to be used only for binary data, such as images, sounds or any media file. - :class:`TextResponse` objects support a new `__init__` method argument, in + :class:`TextResponse` objects support a new ``__init__`` method argument, in addition to the base :class:`Response` objects. The remaining functionality is the same as for the :class:`Response` class and is not documented here. @@ -755,7 +755,7 @@ TextResponse objects A string with the encoding of this response. The encoding is resolved by trying the following mechanisms, in order: - 1. the encoding passed in the `__init__` method`` ``encoding`` argument + 1. the encoding passed in the ``__init__`` method`` ``encoding`` argument 2. the encoding declared in the Content-Type HTTP header. If this encoding is not valid (ie. unknown), it is ignored and the next diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 2fdc86b63..8ed8d55f1 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -31,7 +31,7 @@ class BaseItemExporter(object): def _configure(self, options, dont_fail=False): """Configure the exporter by poping options from the ``options`` dict. If dont_fail is set, it won't raise an exception on unexpected options - (useful for using with keyword arguments in subclasses __init__ methods) + (useful for using with keyword arguments in subclasses ``__init__`` methods) """ self.encoding = options.pop('encoding', None) self.fields_to_export = options.pop('fields_to_export', None) diff --git a/scrapy/extensions/feedexport.py b/scrapy/extensions/feedexport.py index 655c10482..bceb648a0 100644 --- a/scrapy/extensions/feedexport.py +++ b/scrapy/extensions/feedexport.py @@ -106,7 +106,7 @@ class S3FeedStorage(BlockingFeedStorage): warnings.warn( "Initialising `scrapy.extensions.feedexport.S3FeedStorage` " "without AWS keys is deprecated. Please supply credentials or " - "use the `from_crawler()` __init__ method.", + "use the `from_crawler()` ``__init__`` method.", category=ScrapyDeprecationWarning, stacklevel=2 ) diff --git a/scrapy/utils/datatypes.py b/scrapy/utils/datatypes.py index 0dfee1b27..70c2aebc8 100644 --- a/scrapy/utils/datatypes.py +++ b/scrapy/utils/datatypes.py @@ -246,7 +246,7 @@ class CaselessDict(dict): class MergeDict(object): """ A simple class for creating new "virtual" dictionaries that actually look - up values in more than one dictionary, passed in the __init__ method. + up values in more than one dictionary, passed in the ``__init__`` method. If a key appears in more than one of the given dictionaries, only the first occurrence will be used. diff --git a/scrapy/utils/misc.py b/scrapy/utils/misc.py index c892bd438..b3ba2ccec 100644 --- a/scrapy/utils/misc.py +++ b/scrapy/utils/misc.py @@ -123,14 +123,14 @@ def rel_has_nofollow(rel): def create_instance(objcls, settings, crawler, *args, **kwargs): """Construct a class instance using its ``from_crawler`` or - ``from_settings`` __init__ method, if available. + ``from_settings`` ``__init__`` method, if available. At least one of ``settings`` and ``crawler`` needs to be different from ``None``. If ``settings `` is ``None``, ``crawler.settings`` will be used. - If ``crawler`` is ``None``, only the ``from_settings`` __init__ method will be + If ``crawler`` is ``None``, only the ``from_settings`` ``__init__`` method will be tried. - ``*args`` and ``**kwargs`` are forwarded to the __init__ methods. + ``*args`` and ``**kwargs`` are forwarded to the ``__init__`` methods. Raises ``ValueError`` if both ``settings`` and ``crawler`` are ``None``. """ diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index 0d645543a..ea5193f12 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -303,7 +303,7 @@ class WeakKeyCache(object): def stringify_dict(dct_or_tuples, encoding='utf-8', keys_only=True): """Return a (new) dict with unicode keys (and values when "keys_only" is False) of the given dict converted to strings. ``dct_or_tuples`` can be a - dict or a list of tuples, like any dict __init__ method supports. + dict or a list of tuples, like any dict ``__init__`` method supports. """ d = {} for k, v in six.iteritems(dict(dct_or_tuples)): diff --git a/sep/sep-009.rst b/sep/sep-009.rst index 929cedc61..e46479a74 100644 --- a/sep/sep-009.rst +++ b/sep/sep-009.rst @@ -55,7 +55,7 @@ singletons members of that object, as explained below: ``STATS_CLASS`` setting) - **crawler.log**: Logger class with methods replacing the current ``scrapy.log`` functions. Logging would be started (if enabled) on - ``Crawler`` __init__ method, so no log starting functions are required. + ``Crawler`` ``__init__`` method, so no log starting functions are required. - ``crawler.log.msg`` - **crawler.signals**: signal handling diff --git a/tests/test_spider.py b/tests/test_spider.py index 64ff40b61..6f6cdb8ff 100644 --- a/tests/test_spider.py +++ b/tests/test_spider.py @@ -42,12 +42,12 @@ class SpiderTest(unittest.TestCase): self.assertEqual(list(start_requests), []) def test_spider_args(self): - """__init__ method arguments are assigned to spider attributes""" + """``__init__`` method arguments are assigned to spider attributes""" spider = self.spider_class('example.com', foo='bar') self.assertEqual(spider.foo, 'bar') def test_spider_without_name(self): - """__init__ method arguments are assigned to spider attributes""" + """``__init__`` method arguments are assigned to spider attributes""" self.assertRaises(ValueError, self.spider_class) self.assertRaises(ValueError, self.spider_class, somearg='foo') From 85ac5c5c5778a116f762cb6200492aa6a2439e83 Mon Sep 17 00:00:00 2001 From: Roy Healy Date: Mon, 21 Oct 2019 19:06:43 +0100 Subject: [PATCH 125/387] Update .travis.yml Co-Authored-By: Mikhail Korobov --- .travis.yml | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/.travis.yml b/.travis.yml index 2ba504972..044fa9e95 100644 --- a/.travis.yml +++ b/.travis.yml @@ -29,7 +29,7 @@ matrix: python: 3.7 - env: TOXENV=py38 python: 3.8 - - env: TOXENV=py37-extra-deps + - env: TOXENV=py38-extra-deps python: 3.8 - env: TOXENV=docs python: 3.6 From 2ee38e8ddbac6841c8d2bb067ddc0ddba813151f Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 22 Oct 2019 14:43:47 +0530 Subject: [PATCH 126/387] Added Pathlib.Path test --- tests/test_feedexport.py | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index f32ac2a4b..842e8d8c0 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -29,6 +29,8 @@ from scrapy.utils.test import assert_aws_environ, get_s3_content_and_delete, get from scrapy.utils.python import to_native_str from scrapy.utils.project import get_project_settings +from Pathlib import Path + class FileFeedStorageTest(unittest.TestCase): @@ -843,3 +845,17 @@ class FeedExportTest(unittest.TestCase): yield self.exported_data({}, settings) self.assertTrue(FromCrawlerCsvItemExporter.init_with_crawler) self.assertTrue(FromCrawlerFileFeedStorage.init_with_crawler) + + @defer.inlineCallbacks + def test_pathlib_uri(self): + tmpdir = tempfile.mkdtemp() + feed_uri = Path(tmpdir) / 'res' + settings = { + 'FEED_FORMAT': 'csv', + 'FEED_STORE_EMPTY': True, + 'FEED_URI': feed_uri, + } + + data = yield self.exported_no_data(settings) + self.assertEqual(data, b'') + shutil.rmtree(tmpdir, ignore_errors=True) \ No newline at end of file From ad96d6ef594c29e8648c3c3b310c42e8fcb210d3 Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 22 Oct 2019 14:53:59 +0530 Subject: [PATCH 127/387] Added Pathlib.Path test correctly --- tests/test_feedexport.py | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 842e8d8c0..664cfd6de 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -856,6 +856,6 @@ class FeedExportTest(unittest.TestCase): 'FEED_URI': feed_uri, } - data = yield self.exported_no_data(settings) - self.assertEqual(data, b'') - shutil.rmtree(tmpdir, ignore_errors=True) \ No newline at end of file + data = yield self.exported_no_data(settings) + self.assertEqual(data, b'') + shutil.rmtree(tmpdir, ignore_errors=True) \ No newline at end of file From 4226791481bb440b9542dda063d1b5ac920a8411 Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 22 Oct 2019 15:07:13 +0530 Subject: [PATCH 128/387] Added Pathlib.Path test --- requirements-py2.txt | 1 + requirements-py3.txt | 1 + 2 files changed, 2 insertions(+) diff --git a/requirements-py2.txt b/requirements-py2.txt index dde8d1c9c..42e057417 100644 --- a/requirements-py2.txt +++ b/requirements-py2.txt @@ -16,3 +16,4 @@ service_identity>=16.0.0 six>=1.10.0 Twisted>=16.0.0 zope.interface>=4.1.3 +pathlib2>=2.0 diff --git a/requirements-py3.txt b/requirements-py3.txt index 2c98e6f6d..77296b91b 100644 --- a/requirements-py3.txt +++ b/requirements-py3.txt @@ -16,3 +16,4 @@ lxml>=3.5.0 service_identity>=16.0.0 six>=1.10.0 zope.interface>=4.1.3 +pathlib2>=2.0 From 0b7d8a51b4dabeb4097f5fc2de396904f308c56d Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 22 Oct 2019 15:12:53 +0530 Subject: [PATCH 129/387] Added Pathlib.Path test --- requirements-py2.txt | 2 +- requirements-py3.txt | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/requirements-py2.txt b/requirements-py2.txt index 42e057417..2a0bb49d3 100644 --- a/requirements-py2.txt +++ b/requirements-py2.txt @@ -16,4 +16,4 @@ service_identity>=16.0.0 six>=1.10.0 Twisted>=16.0.0 zope.interface>=4.1.3 -pathlib2>=2.0 +pathlib==1.0.1 diff --git a/requirements-py3.txt b/requirements-py3.txt index 77296b91b..c57cff5da 100644 --- a/requirements-py3.txt +++ b/requirements-py3.txt @@ -16,4 +16,4 @@ lxml>=3.5.0 service_identity>=16.0.0 six>=1.10.0 zope.interface>=4.1.3 -pathlib2>=2.0 +pathlib==1.0.1 From a776554282aadc651f9e142aea9bc436f9f587a7 Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 22 Oct 2019 15:31:55 +0530 Subject: [PATCH 130/387] Added Pathlib.Path test --- tests/test_feedexport.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 664cfd6de..fcada3a45 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -29,7 +29,7 @@ from scrapy.utils.test import assert_aws_environ, get_s3_content_and_delete, get from scrapy.utils.python import to_native_str from scrapy.utils.project import get_project_settings -from Pathlib import Path +from pathlib import Path class FileFeedStorageTest(unittest.TestCase): From 07822935eca497f3fe31de536afd76a2c31cc5a9 Mon Sep 17 00:00:00 2001 From: illgitthat Date: Tue, 22 Oct 2019 06:05:34 -0400 Subject: [PATCH 131/387] Updating link for miniconda (#4089) --- docs/intro/install.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/intro/install.rst b/docs/intro/install.rst index 2bf98dbdc..51b41b4d7 100644 --- a/docs/intro/install.rst +++ b/docs/intro/install.rst @@ -290,5 +290,5 @@ For details, see `Issue #2473 `_. .. _zsh: https://www.zsh.org/ .. _Scrapinghub: https://scrapinghub.com .. _Anaconda: https://docs.anaconda.com/anaconda/ -.. _Miniconda: https://conda.io/docs/user-guide/install/index.html +.. _Miniconda: https://docs.conda.io/projects/conda/en/latest/user-guide/install/index.html .. _conda-forge: https://conda-forge.org/ From cd4c211f4b5c9d5716d0e49472a978c11d7351f8 Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 22 Oct 2019 15:38:06 +0530 Subject: [PATCH 132/387] Added Pathlib.Path test --- tests/test_feedexport.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index fcada3a45..f497bb32e 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -29,7 +29,7 @@ from scrapy.utils.test import assert_aws_environ, get_s3_content_and_delete, get from scrapy.utils.python import to_native_str from scrapy.utils.project import get_project_settings -from pathlib import Path +from pathlib2 import Path class FileFeedStorageTest(unittest.TestCase): From 5d75ed4cba3943ac88b4e9cd1d5e87259b85f11f Mon Sep 17 00:00:00 2001 From: WinterComes Date: Tue, 22 Oct 2019 13:19:07 +0300 Subject: [PATCH 133/387] Remove an old note about contracts (#4093) --- docs/topics/contracts.rst | 4 ---- 1 file changed, 4 deletions(-) diff --git a/docs/topics/contracts.rst b/docs/topics/contracts.rst index 62f9a743b..371ae62d5 100644 --- a/docs/topics/contracts.rst +++ b/docs/topics/contracts.rst @@ -6,10 +6,6 @@ Spiders Contracts .. versionadded:: 0.15 -.. note:: This is a new feature (introduced in Scrapy 0.15) and may be subject - to minor functionality/API updates. Check the :ref:`release notes ` to - be notified of updates. - Testing spiders can get particularly annoying and while nothing prevents you from writing unit tests the task gets cumbersome quickly. Scrapy offers an integrated way of testing your spiders by the means of contracts. From cd0964643879b5513807c8955c7e887c77624970 Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 22 Oct 2019 16:19:41 +0530 Subject: [PATCH 134/387] Added Pathlib.Path test --- requirements-py2.txt | 2 +- requirements-py3.txt | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/requirements-py2.txt b/requirements-py2.txt index 2a0bb49d3..42e057417 100644 --- a/requirements-py2.txt +++ b/requirements-py2.txt @@ -16,4 +16,4 @@ service_identity>=16.0.0 six>=1.10.0 Twisted>=16.0.0 zope.interface>=4.1.3 -pathlib==1.0.1 +pathlib2>=2.0 diff --git a/requirements-py3.txt b/requirements-py3.txt index c57cff5da..77296b91b 100644 --- a/requirements-py3.txt +++ b/requirements-py3.txt @@ -16,4 +16,4 @@ lxml>=3.5.0 service_identity>=16.0.0 six>=1.10.0 zope.interface>=4.1.3 -pathlib==1.0.1 +pathlib2>=2.0 From 7031e3a12422ae977448e7b338457f6c09af9b0e Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 22 Oct 2019 16:31:14 +0530 Subject: [PATCH 135/387] Added Pathlib.Path test --- tests/test_feedexport.py | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index f497bb32e..c4bfdb4fe 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -850,10 +850,11 @@ class FeedExportTest(unittest.TestCase): def test_pathlib_uri(self): tmpdir = tempfile.mkdtemp() feed_uri = Path(tmpdir) / 'res' + res_uri = urljoin('file:', pathname2url(feed_uri)) settings = { 'FEED_FORMAT': 'csv', 'FEED_STORE_EMPTY': True, - 'FEED_URI': feed_uri, + 'FEED_URI': res_uri, } data = yield self.exported_no_data(settings) From 85f56a92f0c753dfa55012e647f425a6a3d23076 Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 22 Oct 2019 16:43:17 +0530 Subject: [PATCH 136/387] Added Pathlib.Path test --- tests/test_feedexport.py | 1 + 1 file changed, 1 insertion(+) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index c4bfdb4fe..17526716b 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -850,6 +850,7 @@ class FeedExportTest(unittest.TestCase): def test_pathlib_uri(self): tmpdir = tempfile.mkdtemp() feed_uri = Path(tmpdir) / 'res' + feed_uri=str(feeduri) res_uri = urljoin('file:', pathname2url(feed_uri)) settings = { 'FEED_FORMAT': 'csv', From 4184bac0687d55a442f599ebcfc913231a7c98e2 Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 22 Oct 2019 16:57:14 +0530 Subject: [PATCH 137/387] Added Pathlib.Path test --- tests/test_feedexport.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 17526716b..8c0e5cd3d 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -850,7 +850,7 @@ class FeedExportTest(unittest.TestCase): def test_pathlib_uri(self): tmpdir = tempfile.mkdtemp() feed_uri = Path(tmpdir) / 'res' - feed_uri=str(feeduri) + feed_uri=str(feed_uri) res_uri = urljoin('file:', pathname2url(feed_uri)) settings = { 'FEED_FORMAT': 'csv', From d21e1034f046c7e8d778e698614c47ebd6875c87 Mon Sep 17 00:00:00 2001 From: Ammar Najjar Date: Tue, 22 Oct 2019 13:24:57 +0200 Subject: [PATCH 138/387] docs: correct point,comma and plural replacements Issue #4086 --- docs/topics/exporters.rst | 6 +++--- docs/topics/items.rst | 2 +- docs/topics/loaders.rst | 18 +++++++++--------- docs/topics/request-response.rst | 10 +++++----- scrapy/utils/misc.py | 2 +- 5 files changed, 19 insertions(+), 19 deletions(-) diff --git a/docs/topics/exporters.rst b/docs/topics/exporters.rst index 1b8a69ca3..b8d898022 100644 --- a/docs/topics/exporters.rst +++ b/docs/topics/exporters.rst @@ -247,7 +247,7 @@ XmlItemExporter :type item_element: str The additional keyword arguments of this ``__init__`` method are passed to the - :class:`BaseItemExporter` ``__init__`` method + :class:`BaseItemExporter` ``__init__`` method. A typical output of this exporter would be:: @@ -352,7 +352,7 @@ PprintItemExporter accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) The additional keyword arguments of this ``__init__`` method are passed to the - :class:`BaseItemExporter` ``__init__`` method + :class:`BaseItemExporter` ``__init__`` method. A typical output of this exporter would be:: @@ -399,7 +399,7 @@ JsonLinesItemExporter Exports Items in JSON format to the specified file-like object, writing one JSON-encoded item per line. The additional ``__init__`` method arguments are passed - to the :class:`BaseItemExporter` ``__init__`` method and the leftover arguments to + to the :class:`BaseItemExporter` ``__init__`` method, and the leftover arguments to the `JSONEncoder`_ ``__init__`` method, so you can use any `JSONEncoder`_ ``__init__`` method argument to customize this exporter. diff --git a/docs/topics/items.rst b/docs/topics/items.rst index 370409026..cdf60208e 100644 --- a/docs/topics/items.rst +++ b/docs/topics/items.rst @@ -237,7 +237,7 @@ Item objects Return a new Item optionally initialized from the given argument. - Items replicate the standard `dict API`_, including its ``__init__`` method and + Items replicate the standard `dict API`_, including its ``__init__`` method, and also provide the following additional API members: .. automethod:: copy diff --git a/docs/topics/loaders.rst b/docs/topics/loaders.rst index 72610f645..a4465f88d 100644 --- a/docs/topics/loaders.rst +++ b/docs/topics/loaders.rst @@ -265,7 +265,7 @@ There are several ways to modify Item Loader context values: loader.context['unit'] = 'cm' 2. On Item Loader instantiation (the keyword arguments of Item Loader - ``__init__`` methodare stored in the Item Loader context):: + ``__init__`` method are stored in the Item Loader context):: loader = ItemLoader(product, unit='cm') @@ -494,7 +494,7 @@ ItemLoader objects .. attribute:: default_item_class An Item class (or factory), used to instantiate items when not given in - the ``__init__`` method + the ``__init__`` method. .. attribute:: default_input_processor @@ -509,15 +509,15 @@ ItemLoader objects .. attribute:: default_selector_class The class used to construct the :attr:`selector` of this - :class:`ItemLoader`, if only a response is given in the ``__init__`` method + :class:`ItemLoader`, if only a response is given in the ``__init__`` method. If a selector is given in the ``__init__`` method this attribute is ignored. This attribute is sometimes overridden in subclasses. .. attribute:: selector The :class:`~scrapy.selector.Selector` object to extract data from. - It's either the selector given in the ``__init__`` methodor one created from - the response given in the ``__init__`` methodusing the + It's either the selector given in the ``__init__`` method or one created from + the response given in the ``__init__`` method using the :attr:`default_selector_class`. This attribute is meant to be read-only. @@ -656,7 +656,7 @@ Here is a list of all built-in processors: Returns the first non-null/non-empty value from the values received, so it's typically used as an output processor to single-valued fields. - It doesn't receive any ``__init__`` methodarguments, nor does it accept Loader contexts. + It doesn't receive any ``__init__`` method arguments, nor does it accept Loader contexts. Example:: @@ -667,7 +667,7 @@ Here is a list of all built-in processors: .. class:: Join(separator=u' ') - Returns the values joined with the separator given in the ``__init__`` method which + Returns the values joined with the separator given in the ``__init__`` method, which defaults to ``u' '``. It doesn't accept Loader contexts. When using the default separator, this processor is equivalent to the @@ -705,7 +705,7 @@ Here is a list of all built-in processors: those which do, this processor will pass the currently active :ref:`Loader context ` through that parameter. - The keyword arguments passed in the ``__init__`` methodare used as the default + The keyword arguments passed in the ``__init__`` method are used as the default Loader context values passed to each function call. However, the final Loader context values passed to functions are overridden with the currently active Loader context accessible through the :meth:`ItemLoader.context` @@ -754,7 +754,7 @@ Here is a list of all built-in processors: .. class:: SelectJmes(json_path) - Queries the value using the json path provided to the ``__init__`` methodand returns the output. + Queries the value using the json path provided to the ``__init__`` method and returns the output. Requires jmespath (https://github.com/jmespath/jmespath.py) to run. This processor takes only one input at a time. diff --git a/docs/topics/request-response.rst b/docs/topics/request-response.rst index bf6a02a1d..123c2dde1 100644 --- a/docs/topics/request-response.rst +++ b/docs/topics/request-response.rst @@ -137,7 +137,7 @@ Request objects A string containing the URL of this request. Keep in mind that this attribute contains the escaped URL, so it can differ from the URL passed in - the ``__init__`` method + the ``__init__`` method. This attribute is read-only. To change the URL of a Request use :meth:`replace`. @@ -400,7 +400,7 @@ fields with form data from :class:`Response` objects. .. class:: FormRequest(url, [formdata, ...]) - The :class:`FormRequest` class adds a new keyword parameter to the ``__init__`` method The + The :class:`FormRequest` class adds a new keyword parameter to the ``__init__`` method. The remaining arguments are the same as for the :class:`Request` class and are not documented here. @@ -473,7 +473,7 @@ fields with form data from :class:`Response` objects. :type dont_click: boolean The other parameters of this class method are passed directly to the - :class:`FormRequest` ``__init__`` method + :class:`FormRequest` ``__init__`` method. .. versionadded:: 0.10.3 The ``formname`` parameter. @@ -547,7 +547,7 @@ dealing with JSON requests. .. class:: JsonRequest(url, [... data, dumps_kwargs]) - The :class:`JsonRequest` class adds two new keyword parameters to the ``__init__`` method The + The :class:`JsonRequest` class adds two new keyword parameters to the ``__init__`` method. The remaining arguments are the same as for the :class:`Request` class and are not documented here. @@ -755,7 +755,7 @@ TextResponse objects A string with the encoding of this response. The encoding is resolved by trying the following mechanisms, in order: - 1. the encoding passed in the ``__init__`` method`` ``encoding`` argument + 1. the encoding passed in the ``__init__`` method ``encoding`` argument 2. the encoding declared in the Content-Type HTTP header. If this encoding is not valid (ie. unknown), it is ignored and the next diff --git a/scrapy/utils/misc.py b/scrapy/utils/misc.py index b3ba2ccec..8060553ad 100644 --- a/scrapy/utils/misc.py +++ b/scrapy/utils/misc.py @@ -123,7 +123,7 @@ def rel_has_nofollow(rel): def create_instance(objcls, settings, crawler, *args, **kwargs): """Construct a class instance using its ``from_crawler`` or - ``from_settings`` ``__init__`` method, if available. + ``from_settings`` ``__init__`` methods, if available. At least one of ``settings`` and ``crawler`` needs to be different from ``None``. If ``settings `` is ``None``, ``crawler.settings`` will be used. From 1d5c270ce8caf954ce83c8db262e2a35707e0c5e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Tue, 22 Oct 2019 15:12:52 +0200 Subject: [PATCH 139/387] Fix dangling file descriptor in FeedExporter when FEED_STORE_EMPTY is False (#4023) --- scrapy/extensions/feedexport.py | 4 +++- tests/test_feedexport.py | 2 +- 2 files changed, 4 insertions(+), 2 deletions(-) diff --git a/scrapy/extensions/feedexport.py b/scrapy/extensions/feedexport.py index ce2846eba..6fb6397b1 100644 --- a/scrapy/extensions/feedexport.py +++ b/scrapy/extensions/feedexport.py @@ -242,7 +242,9 @@ class FeedExporter(object): def close_spider(self, spider): slot = self.slot if not slot.itemcount and not self.store_empty: - return + # We need to call slot.storage.store nonetheless to get the file + # properly closed. + return defer.maybeDeferred(slot.storage.store, slot.file) if self._exporting: slot.exporter.finish_exporting() self._exporting = False diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index f32ac2a4b..e1436fbe5 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -417,7 +417,7 @@ class FeedExportTest(unittest.TestCase): content = f.read() finally: - shutil.rmtree(tmpdir, ignore_errors=True) + shutil.rmtree(tmpdir) defer.returnValue(content) From 7a84a4bdba45bab56cee0daed9d9c3a04d8d61c0 Mon Sep 17 00:00:00 2001 From: Ammar Najjar Date: Tue, 22 Oct 2019 15:31:34 +0200 Subject: [PATCH 140/387] docs: use "constructor" for from_settings() & rom_crawler() factory methods Issue #4086 --- scrapy/utils/misc.py | 6 +++--- tests/test_utils_misc/__init__.py | 10 +++++----- 2 files changed, 8 insertions(+), 8 deletions(-) diff --git a/scrapy/utils/misc.py b/scrapy/utils/misc.py index 8060553ad..f638adb25 100644 --- a/scrapy/utils/misc.py +++ b/scrapy/utils/misc.py @@ -123,14 +123,14 @@ def rel_has_nofollow(rel): def create_instance(objcls, settings, crawler, *args, **kwargs): """Construct a class instance using its ``from_crawler`` or - ``from_settings`` ``__init__`` methods, if available. + ``from_settings`` constructors, if available. At least one of ``settings`` and ``crawler`` needs to be different from ``None``. If ``settings `` is ``None``, ``crawler.settings`` will be used. - If ``crawler`` is ``None``, only the ``from_settings`` ``__init__`` method will be + If ``crawler`` is ``None``, only the ``from_settings`` constructor will be tried. - ``*args`` and ``**kwargs`` are forwarded to the ``__init__`` methods. + ``*args`` and ``**kwargs`` are forwarded to the constructors. Raises ``ValueError`` if both ``settings`` and ``crawler`` are ``None``. """ diff --git a/tests/test_utils_misc/__init__.py b/tests/test_utils_misc/__init__.py index 457a2aa78..e109d5343 100644 --- a/tests/test_utils_misc/__init__.py +++ b/tests/test_utils_misc/__init__.py @@ -109,11 +109,11 @@ class UtilsMiscTestCase(unittest.TestCase): else: mock.assert_called_once_with(*args, **kwargs) - # Check usage of correct __init__ method using four mocks: - # 1. with no alternative __init__ methods - # 2. with from_settings() __init__ method - # 3. with from_crawler() __init__ method - # 4. with from_settings() and from_crawler() __init__ method + # Check usage of correct constructor using four mocks: + # 1. with no alternative constructors + # 2. with from_settings() constructor + # 3. with from_crawler() constructor + # 4. with from_settings() and from_crawler() constructor spec_sets = ([], ['from_settings'], ['from_crawler'], ['from_settings', 'from_crawler']) for specs in spec_sets: From f701f5b0db10faef08e4ed9a21b98fd72f9cfc9a Mon Sep 17 00:00:00 2001 From: Victor Torres Date: Tue, 22 Oct 2019 10:48:02 -0300 Subject: [PATCH 141/387] fix #2552 by improving request schema check on its initialization --- scrapy/http/request/__init__.py | 2 +- tests/test_http_request.py | 2 ++ 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/scrapy/http/request/__init__.py b/scrapy/http/request/__init__.py index d09eaf849..76a428199 100644 --- a/scrapy/http/request/__init__.py +++ b/scrapy/http/request/__init__.py @@ -66,7 +66,7 @@ class Request(object_ref): s = safe_url_string(url, self.encoding) self._url = escape_ajax(s) - if ':' not in self._url: + if ('://' not in self._url) and (not self._url.startswith('data:')): raise ValueError('Missing scheme in request url: %s' % self._url) url = property(_get_url, obsolete_setter(_set_url, 'url')) diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 16d7a1cb8..64f1184c3 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -52,6 +52,8 @@ class RequestTest(unittest.TestCase): def test_url_no_scheme(self): self.assertRaises(ValueError, self.request_class, 'foo') + self.assertRaises(ValueError, self.request_class, '/foo/') + self.assertRaises(ValueError, self.request_class, '/foo:bar') def test_headers(self): # Different ways of setting headers attribute From 7fba8434f3997197f98d2b3e73099ad85429e377 Mon Sep 17 00:00:00 2001 From: Ammar Najjar Date: Tue, 22 Oct 2019 15:55:52 +0200 Subject: [PATCH 142/387] use instantiation for "Crawler" Issue #4086 Co-Authored-By: Mikhail Korobov --- sep/sep-009.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/sep/sep-009.rst b/sep/sep-009.rst index e46479a74..da87fa9aa 100644 --- a/sep/sep-009.rst +++ b/sep/sep-009.rst @@ -55,7 +55,7 @@ singletons members of that object, as explained below: ``STATS_CLASS`` setting) - **crawler.log**: Logger class with methods replacing the current ``scrapy.log`` functions. Logging would be started (if enabled) on - ``Crawler`` ``__init__`` method, so no log starting functions are required. + ``Crawler`` instantiation, so no log starting functions are required. - ``crawler.log.msg`` - **crawler.signals**: signal handling From bf5c1a3dec02165b81dffb98fc560c43fb065f38 Mon Sep 17 00:00:00 2001 From: Ammar Najjar Date: Tue, 22 Oct 2019 15:56:46 +0200 Subject: [PATCH 143/387] docs: use "constructor" for "from_crawler" Issue #4086 --- scrapy/extensions/feedexport.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/extensions/feedexport.py b/scrapy/extensions/feedexport.py index bceb648a0..ce2846eba 100644 --- a/scrapy/extensions/feedexport.py +++ b/scrapy/extensions/feedexport.py @@ -106,7 +106,7 @@ class S3FeedStorage(BlockingFeedStorage): warnings.warn( "Initialising `scrapy.extensions.feedexport.S3FeedStorage` " "without AWS keys is deprecated. Please supply credentials or " - "use the `from_crawler()` ``__init__`` method.", + "use the `from_crawler()` constructor.", category=ScrapyDeprecationWarning, stacklevel=2 ) From c623a16a223df31669d9f204b8af7d89062ee56c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Tue, 22 Oct 2019 17:52:34 +0200 Subject: [PATCH 144/387] Remove unused method from scrapy.pqueues._SlotPriorityQueues --- scrapy/pqueues.py | 3 --- 1 file changed, 3 deletions(-) diff --git a/scrapy/pqueues.py b/scrapy/pqueues.py index 6ecd1b51a..717ed4d27 100644 --- a/scrapy/pqueues.py +++ b/scrapy/pqueues.py @@ -86,9 +86,6 @@ class _SlotPriorityQueues(object): def __len__(self): return sum(len(x) for x in self.pqueues.values()) if self.pqueues else 0 - def __contains__(self, slot): - return slot in self.pqueues - class ScrapyPriorityQueue(PriorityQueue): """ From 84fe4011b0063dae1a8efbcb563e772d3b9fce09 Mon Sep 17 00:00:00 2001 From: Kevin Lloyd Bernal Date: Wed, 23 Oct 2019 20:39:53 +0800 Subject: [PATCH 145/387] update docs of scrapy.loader.ItemLoader.item --- docs/topics/loaders.rst | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/topics/loaders.rst b/docs/topics/loaders.rst index 1c2f1da4d..4bd564014 100644 --- a/docs/topics/loaders.rst +++ b/docs/topics/loaders.rst @@ -485,6 +485,8 @@ ItemLoader objects .. attribute:: item The :class:`~scrapy.item.Item` object being parsed by this Item Loader. + This is mostly used as a property so when attempting to override this + value, you may want to check out :attr:`default_item_class` first. .. attribute:: context From 179dc916ff3e3518d8874f3fcd93948e11d5a259 Mon Sep 17 00:00:00 2001 From: Roy Healy Date: Sun, 27 Oct 2019 16:48:20 +0000 Subject: [PATCH 146/387] Update setup.py Updating setup.py to remove python 3.8 support for now --- setup.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/setup.py b/setup.py index 2f5fca4c9..4127d3191 100644 --- a/setup.py +++ b/setup.py @@ -56,7 +56,7 @@ setup( 'Programming Language :: Python :: 3.5', 'Programming Language :: Python :: 3.6', 'Programming Language :: Python :: 3.7', - 'Programming Language :: Python :: 3.8', + # 'Programming Language :: Python :: 3.8', not supported yet 'Programming Language :: Python :: Implementation :: CPython', 'Programming Language :: Python :: Implementation :: PyPy', 'Topic :: Internet :: WWW/HTTP', From 4068797558f6d8d58f81e12930f487fc0f59f0a5 Mon Sep 17 00:00:00 2001 From: Roy Healy Date: Sun, 27 Oct 2019 17:02:17 +0000 Subject: [PATCH 147/387] Update test_downloadermiddleware_httpcache.py Adding xfail denoting that leveldb is not supported in 3.8 --- tests/test_downloadermiddleware_httpcache.py | 1 + 1 file changed, 1 insertion(+) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index 22946b98c..34bf5776a 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -154,6 +154,7 @@ class FilesystemStorageGzipTest(FilesystemStorageTest): new_settings.setdefault('HTTPCACHE_GZIP', True) return super(FilesystemStorageTest, self)._get_settings(**new_settings) +@pytest.mark.xfail(reason='leveldb not supported in python 3.8') class LeveldbStorageTest(DefaultStorageTest): pytest.importorskip('leveldb') From deacd34c8d2400b07e911dea9173b840cec7dece Mon Sep 17 00:00:00 2001 From: Roy Date: Sun, 27 Oct 2019 17:39:47 +0000 Subject: [PATCH 148/387] [test_downloadermiddleware_httpcache] Attempting to add xfail for leveldb related tests https://github.com/scrapy/scrapy/issues/4085 --- tests/test_downloadermiddleware_httpcache.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index 34bf5776a..e350da72c 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -104,6 +104,7 @@ class _BaseTest(unittest.TestCase): class DefaultStorageTest(_BaseTest): + @pytest.mark.xfail(reason='leveldb not supported in python 3.8') def test_storage(self): with self._storage() as storage: request2 = self.request.copy() @@ -154,7 +155,6 @@ class FilesystemStorageGzipTest(FilesystemStorageTest): new_settings.setdefault('HTTPCACHE_GZIP', True) return super(FilesystemStorageTest, self)._get_settings(**new_settings) -@pytest.mark.xfail(reason='leveldb not supported in python 3.8') class LeveldbStorageTest(DefaultStorageTest): pytest.importorskip('leveldb') From 11942c436c78e998e3d58f0aad7cba490d74cc67 Mon Sep 17 00:00:00 2001 From: Roy Date: Sun, 27 Oct 2019 18:02:13 +0000 Subject: [PATCH 149/387] [test_downloadermiddleware_httpcache] Trying hack to handle systemerror whjen importing leveldb https://github.com/scrapy/scrapy/issues/4085 --- tests/test_downloadermiddleware_httpcache.py | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index e350da72c..eec0feafc 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -6,6 +6,7 @@ import unittest import email.utils from contextlib import contextmanager import pytest +import sys from scrapy.http import Response, HtmlResponse, Request from scrapy.spiders import Spider @@ -157,7 +158,10 @@ class FilesystemStorageGzipTest(FilesystemStorageTest): class LeveldbStorageTest(DefaultStorageTest): - pytest.importorskip('leveldb') + try: + pytest.importorskip('leveldb') + except SystemError: + pass storage_class = 'scrapy.extensions.httpcache.LeveldbCacheStorage' From b3df0a84150f15ff00cdfaffc586f91b653baecd Mon Sep 17 00:00:00 2001 From: Roy Date: Sun, 27 Oct 2019 18:28:47 +0000 Subject: [PATCH 150/387] [test_downloadermiddleware_httpcache] Adding xfails to impacted tests following hack fix https://github.com/scrapy/scrapy/issues/4085 --- tests/test_downloadermiddleware_httpcache.py | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index eec0feafc..605223088 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -6,7 +6,6 @@ import unittest import email.utils from contextlib import contextmanager import pytest -import sys from scrapy.http import Response, HtmlResponse, Request from scrapy.spiders import Spider @@ -90,6 +89,7 @@ class _BaseTest(unittest.TestCase): assert any(h in request2.headers for h in (b'If-None-Match', b'If-Modified-Since')) self.assertEqual(request1.body, request2.body) + @pytest.mark.xfail(reason='leveldb not supported in python 3.8') def test_dont_cache(self): with self._middleware() as mw: self.request.meta['dont_cache'] = True @@ -119,6 +119,7 @@ class DefaultStorageTest(_BaseTest): time.sleep(2) # wait for cache to expire assert storage.retrieve_response(self.spider, request2) is None + @pytest.mark.xfail(reason='leveldb not supported in python 3.8') def test_storage_never_expire(self): with self._storage(HTTPCACHE_EXPIRATION_SECS=0) as storage: assert storage.retrieve_response(self.spider, self.request) is None @@ -161,6 +162,8 @@ class LeveldbStorageTest(DefaultStorageTest): try: pytest.importorskip('leveldb') except SystemError: + # Happens in python 3.8 + # This will cause xfail in DefaultStorageTest to trigger pass storage_class = 'scrapy.extensions.httpcache.LeveldbCacheStorage' From 70b2854590c1aa10324e6a4b50b6d74a5079c8e9 Mon Sep 17 00:00:00 2001 From: Roy Date: Sun, 27 Oct 2019 18:51:26 +0000 Subject: [PATCH 151/387] [test_downloadermiddleware_httpcache] Making xfails more informative https://github.com/scrapy/scrapy/issues/4085 --- tests/test_downloadermiddleware_httpcache.py | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index 605223088..faee37446 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -6,6 +6,7 @@ import unittest import email.utils from contextlib import contextmanager import pytest +import sys from scrapy.http import Response, HtmlResponse, Request from scrapy.spiders import Spider @@ -89,7 +90,7 @@ class _BaseTest(unittest.TestCase): assert any(h in request2.headers for h in (b'If-None-Match', b'If-Modified-Since')) self.assertEqual(request1.body, request2.body) - @pytest.mark.xfail(reason='leveldb not supported in python 3.8') + @pytest.mark.xfail(sys.version_info >= (3, 8), raises=RuntimeError, reason='leveldb not supported in python 3.8') def test_dont_cache(self): with self._middleware() as mw: self.request.meta['dont_cache'] = True @@ -105,7 +106,7 @@ class _BaseTest(unittest.TestCase): class DefaultStorageTest(_BaseTest): - @pytest.mark.xfail(reason='leveldb not supported in python 3.8') + @pytest.mark.xfail(sys.version_info >= (3, 8), raises=RuntimeError, reason='leveldb not supported in python 3.8') def test_storage(self): with self._storage() as storage: request2 = self.request.copy() @@ -119,7 +120,7 @@ class DefaultStorageTest(_BaseTest): time.sleep(2) # wait for cache to expire assert storage.retrieve_response(self.spider, request2) is None - @pytest.mark.xfail(reason='leveldb not supported in python 3.8') + @pytest.mark.xfail(sys.version_info >= (3, 8), raises=RuntimeError, reason='leveldb not supported in python 3.8') def test_storage_never_expire(self): with self._storage(HTTPCACHE_EXPIRATION_SECS=0) as storage: assert storage.retrieve_response(self.spider, self.request) is None From 20ea912513053aa2b2122ba5bb0387189bbc5467 Mon Sep 17 00:00:00 2001 From: Roy Date: Sun, 27 Oct 2019 18:52:01 +0000 Subject: [PATCH 152/387] [test_downloadermiddleware_httpcache] Making xfails more informative https://github.com/scrapy/scrapy/issues/4085 --- tests/test_downloadermiddleware_httpcache.py | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index faee37446..b832fc38c 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -90,7 +90,7 @@ class _BaseTest(unittest.TestCase): assert any(h in request2.headers for h in (b'If-None-Match', b'If-Modified-Since')) self.assertEqual(request1.body, request2.body) - @pytest.mark.xfail(sys.version_info >= (3, 8), raises=RuntimeError, reason='leveldb not supported in python 3.8') + @pytest.mark.xfail(sys.version_info >= (3, 8), raises=SystemError, reason='leveldb not supported in python 3.8') def test_dont_cache(self): with self._middleware() as mw: self.request.meta['dont_cache'] = True @@ -106,7 +106,7 @@ class _BaseTest(unittest.TestCase): class DefaultStorageTest(_BaseTest): - @pytest.mark.xfail(sys.version_info >= (3, 8), raises=RuntimeError, reason='leveldb not supported in python 3.8') + @pytest.mark.xfail(sys.version_info >= (3, 8), raises=SystemError, reason='leveldb not supported in python 3.8') def test_storage(self): with self._storage() as storage: request2 = self.request.copy() @@ -120,7 +120,7 @@ class DefaultStorageTest(_BaseTest): time.sleep(2) # wait for cache to expire assert storage.retrieve_response(self.spider, request2) is None - @pytest.mark.xfail(sys.version_info >= (3, 8), raises=RuntimeError, reason='leveldb not supported in python 3.8') + @pytest.mark.xfail(sys.version_info >= (3, 8), raises=SystemError, reason='leveldb not supported in python 3.8') def test_storage_never_expire(self): with self._storage(HTTPCACHE_EXPIRATION_SECS=0) as storage: assert storage.retrieve_response(self.spider, self.request) is None From bb91f9c78c9c8c892495f6e6252cbeebbe12725b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 21 Oct 2019 12:08:35 +0200 Subject: [PATCH 153/387] Cover Scrapy 1.7.4 in the release notes --- docs/news.rst | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/docs/news.rst b/docs/news.rst index 59317f5eb..8dfe8693c 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -6,6 +6,18 @@ Release notes .. note:: Scrapy 1.x will be the last series supporting Python 2. Scrapy 2.0, planned for Q4 2019 or Q1 2020, will support **Python 3 only**. +Scrapy 1.7.4 (2019-10-21) +------------------------- + +Revert the fix for :issue:`3804` (:issue:`3819`), which has a few undesired +side effects (:issue:`3897`, :issue:`3976`). + +As a result, when an item loader is initialized with an item, +:meth:`ItemLoader.load_item() ` once again +makes later calls to :meth:`ItemLoader.get_output_value() +` or :meth:`ItemLoader.load_item() +` return empty data. + Scrapy 1.7.3 (2019-08-01) ------------------------- From 7731814cc25c57fe31db9ba749450cd5a27eed39 Mon Sep 17 00:00:00 2001 From: elacuesta Date: Mon, 28 Oct 2019 06:53:53 -0300 Subject: [PATCH 154/387] ItemLoader: improve handling of initial item (#4036) --- docs/topics/loaders.rst | 10 +- scrapy/loader/__init__.py | 31 ++-- tests/test_loader.py | 321 +++++++++++++++++++++++++++++--------- 3 files changed, 273 insertions(+), 89 deletions(-) diff --git a/docs/topics/loaders.rst b/docs/topics/loaders.rst index 1c2f1da4d..0318e37aa 100644 --- a/docs/topics/loaders.rst +++ b/docs/topics/loaders.rst @@ -35,6 +35,12 @@ Then, you start collecting values into the Item Loader, typically using the same item field; the Item Loader will know how to "join" those values later using a proper processing function. +.. note:: Collected data is internally stored as lists, + allowing to add several values to the same field. + If an ``item`` argument is passed when creating a loader, + each of the item's values will be stored as-is if it's already + an iterable, or wrapped with a list if it's a single value. + Here is a typical Item Loader usage in a :ref:`Spider `, using the :ref:`Product item ` declared in the :ref:`Items chapter `:: @@ -128,9 +134,9 @@ So what happens is: It's worth noticing that processors are just callable objects, which are called with the data to be parsed, and return a parsed value. So you can use any function as input or output processor. The only requirement is that they must -accept one (and only one) positional argument, which will be an iterator. +accept one (and only one) positional argument, which will be an iterable. -.. note:: Both input and output processors must receive an iterator as their +.. note:: Both input and output processors must receive an iterable as their first argument. The output of those functions can be anything. The result of input processors will be appended to an internal list (in the Loader) containing the collected values (for that field). The result of the output diff --git a/scrapy/loader/__init__.py b/scrapy/loader/__init__.py index 6665eba16..60fd6d222 100644 --- a/scrapy/loader/__init__.py +++ b/scrapy/loader/__init__.py @@ -1,19 +1,19 @@ -"""Item Loader +""" +Item Loader See documentation in docs/topics/loaders.rst - """ from collections import defaultdict + import six from scrapy.item import Item +from scrapy.loader.common import wrap_loader_context +from scrapy.loader.processors import Identity from scrapy.selector import Selector from scrapy.utils.misc import arg_to_iter, extract_regex from scrapy.utils.python import flatten -from .common import wrap_loader_context -from .processors import Identity - class ItemLoader(object): @@ -33,10 +33,9 @@ class ItemLoader(object): self.parent = parent self._local_item = context['item'] = item self._local_values = defaultdict(list) - # Preprocess values if item built from dict - # Values need to be added to item._values if added them from dict (not with add_values) + # values from initial item for field_name, value in item.items(): - self._values[field_name] = self._process_input_value(field_name, value) + self._values[field_name] += arg_to_iter(value) @property def _values(self): @@ -132,8 +131,8 @@ class ItemLoader(object): try: return proc(self._values[field_name]) except Exception as e: - raise ValueError("Error with output processor: field=%r value=%r error='%s: %s'" % \ - (field_name, self._values[field_name], type(e).__name__, str(e))) + raise ValueError("Error with output processor: field=%r value=%r error='%s: %s'" % + (field_name, self._values[field_name], type(e).__name__, str(e))) def get_collected_values(self, field_name): return self._values[field_name] @@ -141,15 +140,15 @@ class ItemLoader(object): def get_input_processor(self, field_name): proc = getattr(self, '%s_in' % field_name, None) if not proc: - proc = self._get_item_field_attr(field_name, 'input_processor', \ - self.default_input_processor) + proc = self._get_item_field_attr(field_name, 'input_processor', + self.default_input_processor) return proc def get_output_processor(self, field_name): proc = getattr(self, '%s_out' % field_name, None) if not proc: - proc = self._get_item_field_attr(field_name, 'output_processor', \ - self.default_output_processor) + proc = self._get_item_field_attr(field_name, 'output_processor', + self.default_output_processor) return proc def _process_input_value(self, field_name, value): @@ -174,8 +173,8 @@ class ItemLoader(object): def _check_selector_method(self): if self.selector is None: raise RuntimeError("To use XPath or CSS selectors, " - "%s must be instantiated with a selector " - "or a response" % self.__class__.__name__) + "%s must be instantiated with a selector " + "or a response" % self.__class__.__name__) def add_xpath(self, field_name, xpath, *processors, **kw): values = self._get_xpathvalues(xpath, **kw) diff --git a/tests/test_loader.py b/tests/test_loader.py index 2725b001a..4a4264a2a 100644 --- a/tests/test_loader.py +++ b/tests/test_loader.py @@ -1,13 +1,15 @@ -import unittest -import six from functools import partial +import unittest + +import six -from scrapy.loader import ItemLoader -from scrapy.loader.processors import Join, Identity, TakeFirst, \ - Compose, MapCompose, SelectJmes -from scrapy.item import Item, Field -from scrapy.selector import Selector from scrapy.http import HtmlResponse +from scrapy.item import Item, Field +from scrapy.loader import ItemLoader +from scrapy.loader.processors import (Compose, Identity, Join, + MapCompose, SelectJmes, TakeFirst) +from scrapy.selector import Selector + # test items class NameItem(Item): @@ -61,7 +63,7 @@ class BasicItemLoaderTest(unittest.TestCase): il.add_value('name', u'marta') item = il.load_item() assert item is i - self.assertEqual(item['summary'], u'lala') + self.assertEqual(item['summary'], [u'lala']) self.assertEqual(item['name'], [u'marta']) def test_load_item_using_custom_loader(self): @@ -419,43 +421,6 @@ class BasicItemLoaderTest(unittest.TestCase): self.assertEqual(item['url'], u'rabbit.hole') self.assertEqual(item['summary'], u'rabbithole') - def test_create_item_from_dict(self): - class TestItem(Item): - title = Field() - - class TestItemLoader(ItemLoader): - default_item_class = TestItem - - input_item = {'title': 'Test item title 1'} - il = TestItemLoader(item=input_item) - # Getting output value mustn't remove value from item - self.assertEqual(il.load_item(), { - 'title': 'Test item title 1', - }) - self.assertEqual(il.get_output_value('title'), 'Test item title 1') - self.assertEqual(il.load_item(), { - 'title': 'Test item title 1', - }) - - input_item = {'title': 'Test item title 2'} - il = TestItemLoader(item=input_item) - # Values from dict must be added to item _values - self.assertEqual(il._values.get('title'), 'Test item title 2') - - input_item = {'title': [u'Test item title 3', u'Test item 4']} - il = TestItemLoader(item=input_item) - # Same rules must work for lists - self.assertEqual(il._values.get('title'), - [u'Test item title 3', u'Test item 4']) - self.assertEqual(il.load_item(), { - 'title': [u'Test item title 3', u'Test item 4'], - }) - self.assertEqual(il.get_output_value('title'), - [u'Test item title 3', u'Test item 4']) - self.assertEqual(il.load_item(), { - 'title': [u'Test item title 3', u'Test item 4'], - }) - def test_error_input_processor(self): class TestItem(Item): name = Field() @@ -493,6 +458,220 @@ class BasicItemLoaderTest(unittest.TestCase): [u'marta', u'other'], Compose(float)) +class InitializationTestMixin(object): + + item_class = None + + def test_keep_single_value(self): + """Loaded item should contain values from the initial item""" + input_item = self.item_class(name='foo') + il = ItemLoader(item=input_item) + loaded_item = il.load_item() + self.assertIsInstance(loaded_item, self.item_class) + self.assertEqual(dict(loaded_item), {'name': ['foo']}) + + def test_keep_list(self): + """Loaded item should contain values from the initial item""" + input_item = self.item_class(name=['foo', 'bar']) + il = ItemLoader(item=input_item) + loaded_item = il.load_item() + self.assertIsInstance(loaded_item, self.item_class) + self.assertEqual(dict(loaded_item), {'name': ['foo', 'bar']}) + + def test_add_value_singlevalue_singlevalue(self): + """Values added after initialization should be appended""" + input_item = self.item_class(name='foo') + il = ItemLoader(item=input_item) + il.add_value('name', 'bar') + loaded_item = il.load_item() + self.assertIsInstance(loaded_item, self.item_class) + self.assertEqual(dict(loaded_item), {'name': ['foo', 'bar']}) + + def test_add_value_singlevalue_list(self): + """Values added after initialization should be appended""" + input_item = self.item_class(name='foo') + il = ItemLoader(item=input_item) + il.add_value('name', ['item', 'loader']) + loaded_item = il.load_item() + self.assertIsInstance(loaded_item, self.item_class) + self.assertEqual(dict(loaded_item), {'name': ['foo', 'item', 'loader']}) + + def test_add_value_list_singlevalue(self): + """Values added after initialization should be appended""" + input_item = self.item_class(name=['foo', 'bar']) + il = ItemLoader(item=input_item) + il.add_value('name', 'qwerty') + loaded_item = il.load_item() + self.assertIsInstance(loaded_item, self.item_class) + self.assertEqual(dict(loaded_item), {'name': ['foo', 'bar', 'qwerty']}) + + def test_add_value_list_list(self): + """Values added after initialization should be appended""" + input_item = self.item_class(name=['foo', 'bar']) + il = ItemLoader(item=input_item) + il.add_value('name', ['item', 'loader']) + loaded_item = il.load_item() + self.assertIsInstance(loaded_item, self.item_class) + self.assertEqual(dict(loaded_item), {'name': ['foo', 'bar', 'item', 'loader']}) + + def test_get_output_value_singlevalue(self): + """Getting output value must not remove value from item""" + input_item = self.item_class(name='foo') + il = ItemLoader(item=input_item) + self.assertEqual(il.get_output_value('name'), ['foo']) + loaded_item = il.load_item() + self.assertIsInstance(loaded_item, self.item_class) + self.assertEqual(loaded_item, dict({'name': ['foo']})) + + def test_get_output_value_list(self): + """Getting output value must not remove value from item""" + input_item = self.item_class(name=['foo', 'bar']) + il = ItemLoader(item=input_item) + self.assertEqual(il.get_output_value('name'), ['foo', 'bar']) + loaded_item = il.load_item() + self.assertIsInstance(loaded_item, self.item_class) + self.assertEqual(loaded_item, dict({'name': ['foo', 'bar']})) + + def test_values_single(self): + """Values from initial item must be added to loader._values""" + input_item = self.item_class(name='foo') + il = ItemLoader(item=input_item) + self.assertEqual(il._values.get('name'), ['foo']) + + def test_values_list(self): + """Values from initial item must be added to loader._values""" + input_item = self.item_class(name=['foo', 'bar']) + il = ItemLoader(item=input_item) + self.assertEqual(il._values.get('name'), ['foo', 'bar']) + + +class InitializationFromDictTest(InitializationTestMixin, unittest.TestCase): + item_class = dict + + +class InitializationFromItemTest(InitializationTestMixin, unittest.TestCase): + item_class = NameItem + + +class BaseNoInputReprocessingLoader(ItemLoader): + title_in = MapCompose(str.upper) + title_out = TakeFirst() + + +class NoInputReprocessingDictLoader(BaseNoInputReprocessingLoader): + default_item_class = dict + + +class NoInputReprocessingFromDictTest(unittest.TestCase): + """ + Loaders initialized from loaded items must not reprocess fields (dict instances) + """ + def test_avoid_reprocessing_with_initial_values_single(self): + il = NoInputReprocessingDictLoader(item=dict(title='foo')) + il_loaded = il.load_item() + self.assertEqual(il_loaded, dict(title='foo')) + self.assertEqual(NoInputReprocessingDictLoader(item=il_loaded).load_item(), dict(title='foo')) + + def test_avoid_reprocessing_with_initial_values_list(self): + il = NoInputReprocessingDictLoader(item=dict(title=['foo', 'bar'])) + il_loaded = il.load_item() + self.assertEqual(il_loaded, dict(title='foo')) + self.assertEqual(NoInputReprocessingDictLoader(item=il_loaded).load_item(), dict(title='foo')) + + def test_avoid_reprocessing_without_initial_values_single(self): + il = NoInputReprocessingDictLoader() + il.add_value('title', 'foo') + il_loaded = il.load_item() + self.assertEqual(il_loaded, dict(title='FOO')) + self.assertEqual(NoInputReprocessingDictLoader(item=il_loaded).load_item(), dict(title='FOO')) + + def test_avoid_reprocessing_without_initial_values_list(self): + il = NoInputReprocessingDictLoader() + il.add_value('title', ['foo', 'bar']) + il_loaded = il.load_item() + self.assertEqual(il_loaded, dict(title='FOO')) + self.assertEqual(NoInputReprocessingDictLoader(item=il_loaded).load_item(), dict(title='FOO')) + + +class NoInputReprocessingItem(Item): + title = Field() + + +class NoInputReprocessingItemLoader(BaseNoInputReprocessingLoader): + default_item_class = NoInputReprocessingItem + + +class NoInputReprocessingFromItemTest(unittest.TestCase): + """ + Loaders initialized from loaded items must not reprocess fields (BaseItem instances) + """ + def test_avoid_reprocessing_with_initial_values_single(self): + il = NoInputReprocessingItemLoader(item=NoInputReprocessingItem(title='foo')) + il_loaded = il.load_item() + self.assertEqual(il_loaded, {'title': 'foo'}) + self.assertEqual(NoInputReprocessingItemLoader(item=il_loaded).load_item(), {'title': 'foo'}) + + def test_avoid_reprocessing_with_initial_values_list(self): + il = NoInputReprocessingItemLoader(item=NoInputReprocessingItem(title=['foo', 'bar'])) + il_loaded = il.load_item() + self.assertEqual(il_loaded, {'title': 'foo'}) + self.assertEqual(NoInputReprocessingItemLoader(item=il_loaded).load_item(), {'title': 'foo'}) + + def test_avoid_reprocessing_without_initial_values_single(self): + il = NoInputReprocessingItemLoader() + il.add_value('title', 'FOO') + il_loaded = il.load_item() + self.assertEqual(il_loaded, {'title': 'FOO'}) + self.assertEqual(NoInputReprocessingItemLoader(item=il_loaded).load_item(), {'title': 'FOO'}) + + def test_avoid_reprocessing_without_initial_values_list(self): + il = NoInputReprocessingItemLoader() + il.add_value('title', ['foo', 'bar']) + il_loaded = il.load_item() + self.assertEqual(il_loaded, {'title': 'FOO'}) + self.assertEqual(NoInputReprocessingItemLoader(item=il_loaded).load_item(), {'title': 'FOO'}) + + +class TestOutputProcessorDict(unittest.TestCase): + def test_output_processor(self): + + class TempDict(dict): + def __init__(self, *args, **kwargs): + super(TempDict, self).__init__(self, *args, **kwargs) + self.setdefault('temp', 0.3) + + class TempLoader(ItemLoader): + default_item_class = TempDict + default_input_processor = Identity() + default_output_processor = Compose(TakeFirst()) + + loader = TempLoader() + item = loader.load_item() + self.assertIsInstance(item, TempDict) + self.assertEqual(dict(item), {'temp': 0.3}) + + +class TestOutputProcessorItem(unittest.TestCase): + def test_output_processor(self): + + class TempItem(Item): + temp = Field() + + def __init__(self, *args, **kwargs): + super(TempItem, self).__init__(self, *args, **kwargs) + self.setdefault('temp', 0.3) + + class TempLoader(ItemLoader): + default_item_class = TempItem + default_input_processor = Identity() + default_output_processor = Compose(TakeFirst()) + + loader = TempLoader() + item = loader.load_item() + self.assertIsInstance(item, TempItem) + self.assertEqual(dict(item), {'temp': 0.3}) + + class ProcessorsTest(unittest.TestCase): def test_take_first(self): @@ -523,7 +702,8 @@ class ProcessorsTest(unittest.TestCase): self.assertRaises(ValueError, proc, 'hello') def test_mapcompose(self): - filter_world = lambda x: None if x == 'world' else x + def filter_world(x): + return None if x == 'world' else x proc = MapCompose(filter_world, six.text_type.upper) self.assertEqual(proc([u'hello', u'world', u'this', u'is', u'scrapy']), [u'HELLO', u'THIS', u'IS', u'SCRAPY']) @@ -535,7 +715,6 @@ class ProcessorsTest(unittest.TestCase): self.assertRaises(ValueError, proc, 'hello') - class SelectortemLoaderTest(unittest.TestCase): response = HtmlResponse(url="", encoding='utf-8', body=b""" @@ -672,7 +851,7 @@ class SelectortemLoaderTest(unittest.TestCase): self.assertEqual(l.get_css(['p::text', 'div::text']), [u'paragraph', 'marta']) self.assertEqual(l.get_css(['a::attr(href)', 'img::attr(src)']), - [u'http://www.scrapy.org', u'/images/logo.png']) + [u'http://www.scrapy.org', u'/images/logo.png']) def test_replace_css_multi_fields(self): l = TestItemLoader(response=self.response) @@ -720,7 +899,7 @@ class SubselectorLoaderTest(unittest.TestCase): self.assertEqual(l.get_output_value('name'), [u'marta']) self.assertEqual(l.get_output_value('name_div'), [u'
marta
']) - self.assertEqual(l.get_output_value('name_value'), [u'marta']) + self.assertEqual(l.get_output_value('name_value'), [u'marta']) self.assertEqual(l.get_output_value('name'), nl.get_output_value('name')) self.assertEqual(l.get_output_value('name_div'), nl.get_output_value('name_div')) @@ -735,7 +914,7 @@ class SubselectorLoaderTest(unittest.TestCase): self.assertEqual(l.get_output_value('name'), [u'marta']) self.assertEqual(l.get_output_value('name_div'), [u'
marta
']) - self.assertEqual(l.get_output_value('name_value'), [u'marta']) + self.assertEqual(l.get_output_value('name_value'), [u'marta']) self.assertEqual(l.get_output_value('name'), nl.get_output_value('name')) self.assertEqual(l.get_output_value('name_div'), nl.get_output_value('name_div')) @@ -791,28 +970,28 @@ class SubselectorLoaderTest(unittest.TestCase): class SelectJmesTestCase(unittest.TestCase): - test_list_equals = { - 'simple': ('foo.bar', {"foo": {"bar": "baz"}}, "baz"), - 'invalid': ('foo.bar.baz', {"foo": {"bar": "baz"}}, None), - 'top_level': ('foo', {"foo": {"bar": "baz"}}, {"bar": "baz"}), - 'double_vs_single_quote_string': ('foo.bar', {"foo": {"bar": "baz"}}, "baz"), - 'dict': ( - 'foo.bar[*].name', - {"foo": {"bar": [{"name": "one"}, {"name": "two"}]}}, - ['one', 'two'] - ), - 'list': ('[1]', [1, 2], 2) - } + test_list_equals = { + 'simple': ('foo.bar', {"foo": {"bar": "baz"}}, "baz"), + 'invalid': ('foo.bar.baz', {"foo": {"bar": "baz"}}, None), + 'top_level': ('foo', {"foo": {"bar": "baz"}}, {"bar": "baz"}), + 'double_vs_single_quote_string': ('foo.bar', {"foo": {"bar": "baz"}}, "baz"), + 'dict': ( + 'foo.bar[*].name', + {"foo": {"bar": [{"name": "one"}, {"name": "two"}]}}, + ['one', 'two'] + ), + 'list': ('[1]', [1, 2], 2) + } - def test_output(self): - for l in self.test_list_equals: - expr, test_list, expected = self.test_list_equals[l] - test = SelectJmes(expr)(test_list) - self.assertEqual( - test, - expected, - msg='test "{}" got {} expected {}'.format(l, test, expected) - ) + def test_output(self): + for l in self.test_list_equals: + expr, test_list, expected = self.test_list_equals[l] + test = SelectJmes(expr)(test_list) + self.assertEqual( + test, + expected, + msg='test "{}" got {} expected {}'.format(l, test, expected) + ) if __name__ == "__main__": From b5a00262ec48534b89750037060318326b4e349c Mon Sep 17 00:00:00 2001 From: Roy Healy Date: Mon, 28 Oct 2019 09:59:33 +0000 Subject: [PATCH 155/387] Update .travis.yml Reverting change to 3.8 extra dependency environment. --- .travis.yml | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/.travis.yml b/.travis.yml index 044fa9e95..2ba504972 100644 --- a/.travis.yml +++ b/.travis.yml @@ -29,7 +29,7 @@ matrix: python: 3.7 - env: TOXENV=py38 python: 3.8 - - env: TOXENV=py38-extra-deps + - env: TOXENV=py37-extra-deps python: 3.8 - env: TOXENV=docs python: 3.6 From 3d0df419c4eeb0ccf7934c8f06a38707eae8d722 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 28 Oct 2019 11:24:47 +0100 Subject: [PATCH 156/387] Mark the LevelDB storage backend as deprecated --- docs/topics/downloader-middleware.rst | 22 ---------------------- scrapy/extensions/httpcache.py | 17 ++++++++++++----- 2 files changed, 12 insertions(+), 27 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 2892b9b79..8048e1c86 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -348,7 +348,6 @@ HttpCacheMiddleware * :ref:`httpcache-storage-fs` * :ref:`httpcache-storage-dbm` - * :ref:`httpcache-storage-leveldb` You can change the HTTP cache storage backend with the :setting:`HTTPCACHE_STORAGE` setting. Or you can also :ref:`implement your own storage backend. ` @@ -478,27 +477,6 @@ DBM storage backend By default, it uses the anydbm_ module, but you can change it with the :setting:`HTTPCACHE_DBM_MODULE` setting. -.. _httpcache-storage-leveldb: - -LevelDB storage backend -~~~~~~~~~~~~~~~~~~~~~~~ - -.. class:: LeveldbCacheStorage - - .. versionadded:: 0.23 - - A LevelDB_ storage backend is also available for the HTTP cache middleware. - - This backend is not recommended for development because only one process - can access LevelDB databases at the same time, so you can't run a crawl and - open the scrapy shell in parallel for the same spider. - - In order to use this storage backend, install the `LevelDB python - bindings`_ (e.g. ``pip install leveldb``). - - .. _LevelDB: https://github.com/google/leveldb - .. _leveldb python bindings: https://pypi.python.org/pypi/leveldb - .. _httpcache-storage-custom: Writing your own storage backend diff --git a/scrapy/extensions/httpcache.py b/scrapy/extensions/httpcache.py index c6094643d..7c650a91e 100644 --- a/scrapy/extensions/httpcache.py +++ b/scrapy/extensions/httpcache.py @@ -1,19 +1,24 @@ from __future__ import print_function -import os + import gzip import logging -from six.moves import cPickle as pickle +import os +from email.utils import mktime_tz, parsedate_tz from importlib import import_module from time import time +from warnings import warn from weakref import WeakKeyDictionary -from email.utils import mktime_tz, parsedate_tz + +from six.moves import cPickle as pickle from w3lib.http import headers_raw_to_dict, headers_dict_to_raw + +from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.http import Headers, Response from scrapy.responsetypes import responsetypes -from scrapy.utils.request import request_fingerprint -from scrapy.utils.project import data_path from scrapy.utils.httpobj import urlparse_cached +from scrapy.utils.project import data_path from scrapy.utils.python import to_bytes, to_unicode, garbage_collect +from scrapy.utils.request import request_fingerprint logger = logging.getLogger(__name__) @@ -345,6 +350,8 @@ class FilesystemCacheStorage(object): class LeveldbCacheStorage(object): def __init__(self, settings): + warn("The LevelDB storage backend is deprecated.", + ScrapyDeprecationWarning, stacklevel=2) import leveldb self._leveldb = leveldb self.cachedir = data_path(settings['HTTPCACHE_DIR'], createdir=True) From 16bb3ac20dae8b7c5fbccf4ab85b3a0393e7c55d Mon Sep 17 00:00:00 2001 From: Roy Date: Mon, 28 Oct 2019 11:24:09 +0000 Subject: [PATCH 157/387] [test_downloadermiddleware_httpcache] Using skipif approach https://github.com/scrapy/scrapy/issues/4085 --- tests/test_downloadermiddleware_httpcache.py | 9 ++++----- 1 file changed, 4 insertions(+), 5 deletions(-) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index b832fc38c..ba5027307 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -90,7 +90,6 @@ class _BaseTest(unittest.TestCase): assert any(h in request2.headers for h in (b'If-None-Match', b'If-Modified-Since')) self.assertEqual(request1.body, request2.body) - @pytest.mark.xfail(sys.version_info >= (3, 8), raises=SystemError, reason='leveldb not supported in python 3.8') def test_dont_cache(self): with self._middleware() as mw: self.request.meta['dont_cache'] = True @@ -106,7 +105,6 @@ class _BaseTest(unittest.TestCase): class DefaultStorageTest(_BaseTest): - @pytest.mark.xfail(sys.version_info >= (3, 8), raises=SystemError, reason='leveldb not supported in python 3.8') def test_storage(self): with self._storage() as storage: request2 = self.request.copy() @@ -120,7 +118,6 @@ class DefaultStorageTest(_BaseTest): time.sleep(2) # wait for cache to expire assert storage.retrieve_response(self.spider, request2) is None - @pytest.mark.xfail(sys.version_info >= (3, 8), raises=SystemError, reason='leveldb not supported in python 3.8') def test_storage_never_expire(self): with self._storage(HTTPCACHE_EXPIRATION_SECS=0) as storage: assert storage.retrieve_response(self.spider, self.request) is None @@ -164,8 +161,10 @@ class LeveldbStorageTest(DefaultStorageTest): pytest.importorskip('leveldb') except SystemError: # Happens in python 3.8 - # This will cause xfail in DefaultStorageTest to trigger - pass + pytest.mark.skipif( + sys.version_info >= (3, 8), + reason='leveldb not supported in python 3.8', + ) storage_class = 'scrapy.extensions.httpcache.LeveldbCacheStorage' From 9b47dc6a703310d13c9470e50d4b14f81ee893c6 Mon Sep 17 00:00:00 2001 From: Roy Date: Mon, 28 Oct 2019 11:24:52 +0000 Subject: [PATCH 158/387] [travis, setup] Adding official python 3.8 support https://github.com/scrapy/scrapy/issues/4085 --- .travis.yml | 4 +--- setup.py | 2 +- 2 files changed, 2 insertions(+), 4 deletions(-) diff --git a/.travis.yml b/.travis.yml index 2ba504972..4c2498053 100644 --- a/.travis.yml +++ b/.travis.yml @@ -25,11 +25,9 @@ matrix: python: 3.6 - env: TOXENV=py37 python: 3.7 - - env: TOXENV=py37-extra-deps - python: 3.7 - env: TOXENV=py38 python: 3.8 - - env: TOXENV=py37-extra-deps + - env: TOXENV=py38-extra-deps python: 3.8 - env: TOXENV=docs python: 3.6 diff --git a/setup.py b/setup.py index 4127d3191..2f5fca4c9 100644 --- a/setup.py +++ b/setup.py @@ -56,7 +56,7 @@ setup( 'Programming Language :: Python :: 3.5', 'Programming Language :: Python :: 3.6', 'Programming Language :: Python :: 3.7', - # 'Programming Language :: Python :: 3.8', not supported yet + 'Programming Language :: Python :: 3.8', 'Programming Language :: Python :: Implementation :: CPython', 'Programming Language :: Python :: Implementation :: PyPy', 'Topic :: Internet :: WWW/HTTP', From 4432136ffff4d8af42f7a485c17ab7fbbb228078 Mon Sep 17 00:00:00 2001 From: Roy Date: Mon, 28 Oct 2019 12:22:21 +0000 Subject: [PATCH 159/387] [test_downloadermiddleware_httpcache] Fixing pytest skip behaviour https://github.com/scrapy/scrapy/issues/4085 --- tests/test_downloadermiddleware_httpcache.py | 6 +----- 1 file changed, 1 insertion(+), 5 deletions(-) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index ba5027307..32085b095 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -6,7 +6,6 @@ import unittest import email.utils from contextlib import contextmanager import pytest -import sys from scrapy.http import Response, HtmlResponse, Request from scrapy.spiders import Spider @@ -161,10 +160,7 @@ class LeveldbStorageTest(DefaultStorageTest): pytest.importorskip('leveldb') except SystemError: # Happens in python 3.8 - pytest.mark.skipif( - sys.version_info >= (3, 8), - reason='leveldb not supported in python 3.8', - ) + pytest.skip("'SystemError: bad call flags' occurs on Python 3.8") storage_class = 'scrapy.extensions.httpcache.LeveldbCacheStorage' From c51fb959e2985faf6f21fe7f03d2fb8160de064f Mon Sep 17 00:00:00 2001 From: Roy Date: Mon, 28 Oct 2019 12:36:54 +0000 Subject: [PATCH 160/387] [test_downloadermiddleware_httpcache] Fixing pytest skip behaviour https://github.com/scrapy/scrapy/issues/4085 --- tests/test_downloadermiddleware_httpcache.py | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index 32085b095..047526569 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -6,6 +6,7 @@ import unittest import email.utils from contextlib import contextmanager import pytest +import sys from scrapy.http import Response, HtmlResponse, Request from scrapy.spiders import Spider @@ -160,7 +161,7 @@ class LeveldbStorageTest(DefaultStorageTest): pytest.importorskip('leveldb') except SystemError: # Happens in python 3.8 - pytest.skip("'SystemError: bad call flags' occurs on Python 3.8") + pytestmark = pytest.skip("'SystemError: bad call flags' occurs on Python 3.8") storage_class = 'scrapy.extensions.httpcache.LeveldbCacheStorage' From 74909030a55b59e3b858fc736b5b1f685d9596a6 Mon Sep 17 00:00:00 2001 From: Roy Date: Mon, 28 Oct 2019 12:52:36 +0000 Subject: [PATCH 161/387] [tox.ini] Removing obsolete py37 extra deps enviornment https://github.com/scrapy/scrapy/issues/4085 --- tox.ini | 7 ------- 1 file changed, 7 deletions(-) diff --git a/tox.ini b/tox.ini index 14afec23f..fe925951b 100644 --- a/tox.ini +++ b/tox.ini @@ -125,13 +125,6 @@ deps = {[docs]deps} commands = sphinx-build -W -b linkcheck . {envtmpdir}/linkcheck -[testenv:py37-extra-deps] -basepython = python3.7 -deps = - {[testenv:py35]deps} - reppy - robotexclusionrulesparser - [testenv:py38-extra-deps] basepython = python3.8 deps = From b73d217de5647a68c7b8dfda747cd3d0685c226d Mon Sep 17 00:00:00 2001 From: Roy Date: Mon, 28 Oct 2019 12:55:54 +0000 Subject: [PATCH 162/387] [test_downloadermiddleware_httpcache.py] Fixing pytest mark behaviour https://github.com/scrapy/scrapy/issues/4085 --- tests/test_downloadermiddleware_httpcache.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index 047526569..f5917d0f0 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -161,7 +161,7 @@ class LeveldbStorageTest(DefaultStorageTest): pytest.importorskip('leveldb') except SystemError: # Happens in python 3.8 - pytestmark = pytest.skip("'SystemError: bad call flags' occurs on Python 3.8") + pytestmark = pytest.mark.skip("'SystemError: bad call flags' occurs on Python 3.8") storage_class = 'scrapy.extensions.httpcache.LeveldbCacheStorage' From 93e3dc1b826e44d1a5a24fbb39c090ce426aa862 Mon Sep 17 00:00:00 2001 From: Roy Date: Mon, 28 Oct 2019 16:12:03 +0000 Subject: [PATCH 163/387] [test_downloadermiddleware_httpcache.py] Cleaning text https://github.com/scrapy/scrapy/issues/4085 --- tests/test_downloadermiddleware_httpcache.py | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index f5917d0f0..972d400a4 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -155,13 +155,13 @@ class FilesystemStorageGzipTest(FilesystemStorageTest): new_settings.setdefault('HTTPCACHE_GZIP', True) return super(FilesystemStorageTest, self)._get_settings(**new_settings) + class LeveldbStorageTest(DefaultStorageTest): try: pytest.importorskip('leveldb') except SystemError: - # Happens in python 3.8 - pytestmark = pytest.mark.skip("'SystemError: bad call flags' occurs on Python 3.8") + pytestmark = pytest.mark.skip("Test module skipped - 'SystemError: bad call flags' occurs when >= Python 3.8") storage_class = 'scrapy.extensions.httpcache.LeveldbCacheStorage' From 94f060fcc84853f28f3f91b6dde1d61c8e19251e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Tue, 29 Oct 2019 12:53:46 +0100 Subject: [PATCH 164/387] Cover Scrapy 1.8.0 in the release notes (#3952) --- docs/news.rst | 226 +++++++++++++++++++++++++++++++++++++++- docs/topics/logging.rst | 5 +- scrapy/logformatter.py | 2 +- 3 files changed, 229 insertions(+), 4 deletions(-) diff --git a/docs/news.rst b/docs/news.rst index 8dfe8693c..669844045 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -6,6 +6,209 @@ Release notes .. note:: Scrapy 1.x will be the last series supporting Python 2. Scrapy 2.0, planned for Q4 2019 or Q1 2020, will support **Python 3 only**. +.. _release-1.8.0: + +Scrapy 1.8.0 (2019-10-28) +------------------------- + +Highlights: + +* Dropped Python 3.4 support and updated minimum requirements; made Python 3.8 + support official +* New :meth:`Request.from_curl ` class method +* New :setting:`ROBOTSTXT_PARSER` and :setting:`ROBOTSTXT_USER_AGENT` settings +* New :setting:`DOWNLOADER_CLIENT_TLS_CIPHERS` and + :setting:`DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING` settings + +Backward-incompatible changes +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +* Python 3.4 is no longer supported, and some of the minimum requirements of + Scrapy have also changed: + + * cssselect_ 0.9.1 + * cryptography_ 2.0 + * lxml_ 3.5.0 + * pyOpenSSL_ 16.2.0 + * queuelib_ 1.4.2 + * service_identity_ 16.0.0 + * six_ 1.10.0 + * Twisted_ 17.9.0 (16.0.0 with Python 2) + * zope.interface_ 4.1.3 + + (:issue:`3892`) + +* ``JSONRequest`` is now called :class:`~scrapy.http.JsonRequest` for + consistency with similar classes (:issue:`3929`, :issue:`3982`) + +* If you are using a custom context factory + (:setting:`DOWNLOADER_CLIENTCONTEXTFACTORY`), its ``__init__`` method must + accept two new parameters: ``tls_verbose_logging`` and ``tls_ciphers`` + (:issue:`2111`, :issue:`3392`, :issue:`3442`, :issue:`3450`) + +* :class:`~scrapy.loader.ItemLoader` now turns the values of its input item + into lists:: + + >>> item = MyItem() + >>> item['field'] = 'value1' + >>> loader = ItemLoader(item=item) + >>> item['field'] + ['value1'] + + This is needed to allow adding values to existing fields + (``loader.add_value('field', 'value2')``). + + (:issue:`3804`, :issue:`3819`, :issue:`3897`, :issue:`3976`, :issue:`3998`, + :issue:`4036`) + +See also :ref:`1.8-deprecation-removals` below. + + +New features +~~~~~~~~~~~~ + +* A new :meth:`Request.from_curl ` class + method allows :ref:`creating a request from a cURL command + ` (:issue:`2985`, :issue:`3862`) + +* A new :setting:`ROBOTSTXT_PARSER` setting allows choosing which robots.txt_ + parser to use. It includes built-in support for + :ref:`RobotFileParser `, + :ref:`Protego ` (default), :ref:`Reppy `, and + :ref:`Robotexclusionrulesparser `, and allows you to + :ref:`implement support for additional parsers + ` (:issue:`754`, :issue:`2669`, + :issue:`3796`, :issue:`3935`, :issue:`3969`, :issue:`4006`) + +* A new :setting:`ROBOTSTXT_USER_AGENT` setting allows defining a separate + user agent string to use for robots.txt_ parsing (:issue:`3931`, + :issue:`3966`) + +* :class:`~scrapy.spiders.Rule` no longer requires a :class:`LinkExtractor + ` parameter + (:issue:`781`, :issue:`4016`) + +* Use the new :setting:`DOWNLOADER_CLIENT_TLS_CIPHERS` setting to customize + the TLS/SSL ciphers used by the default HTTP/1.1 downloader (:issue:`3392`, + :issue:`3442`) + +* Set the new :setting:`DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING` setting to + ``True`` to enable debug-level messages about TLS connection parameters + after establishing HTTPS connections (:issue:`2111`, :issue:`3450`) + +* Callbacks that receive keyword arguments + (see :attr:`Request.cb_kwargs `) can now be + tested using the new :class:`@cb_kwargs + ` + :ref:`spider contract ` (:issue:`3985`, :issue:`3988`) + +* When a :class:`@scrapes ` spider + contract fails, all missing fields are now reported (:issue:`766`, + :issue:`3939`) + +* :ref:`Custom log formats ` can now drop messages by + having the corresponding methods of the configured :setting:`LOG_FORMATTER` + return ``None`` (:issue:`3984`, :issue:`3987`) + +* A much improved completion definition is now available for Zsh_ + (:issue:`4069`) + + +Bug fixes +~~~~~~~~~ + +* :meth:`ItemLoader.load_item() ` no + longer makes later calls to :meth:`ItemLoader.get_output_value() + ` or + :meth:`ItemLoader.load_item() ` return + empty data (:issue:`3804`, :issue:`3819`, :issue:`3897`, :issue:`3976`, + :issue:`3998`, :issue:`4036`) + +* Fixed :class:`~scrapy.statscollectors.DummyStatsCollector` raising a + :exc:`TypeError` exception (:issue:`4007`, :issue:`4052`) + +* :meth:`FilesPipeline.file_path + ` and + :meth:`ImagesPipeline.file_path + ` no longer choose + file extensions that are not `registered with IANA`_ (:issue:`1287`, + :issue:`3953`, :issue:`3954`) + +* When using botocore_ to persist files in S3, all botocore-supported headers + are properly mapped now (:issue:`3904`, :issue:`3905`) + +* FTP passwords in :setting:`FEED_URI` containing percent-escaped characters + are now properly decoded (:issue:`3941`) + +* A memory-handling and error-handling issue in + :func:`scrapy.utils.ssl.get_temp_key_info` has been fixed (:issue:`3920`) + + +Documentation +~~~~~~~~~~~~~ + +* The documentation now covers how to define and configure a :ref:`custom log + format ` (:issue:`3616`, :issue:`3660`) + +* API documentation added for :class:`~scrapy.exporters.MarshalItemExporter` + and :class:`~scrapy.exporters.PythonItemExporter` (:issue:`3973`) + +* API documentation added for :class:`~scrapy.item.BaseItem` and + :class:`~scrapy.item.ItemMeta` (:issue:`3999`) + +* Minor documentation fixes (:issue:`2998`, :issue:`3398`, :issue:`3597`, + :issue:`3894`, :issue:`3934`, :issue:`3978`, :issue:`3993`, :issue:`4022`, + :issue:`4028`, :issue:`4033`, :issue:`4046`, :issue:`4050`, :issue:`4055`, + :issue:`4056`, :issue:`4061`, :issue:`4072`, :issue:`4071`, :issue:`4079`, + :issue:`4081`, :issue:`4089`, :issue:`4093`) + + +.. _1.8-deprecation-removals: + +Deprecation removals +~~~~~~~~~~~~~~~~~~~~ + +* ``scrapy.xlib`` has been removed (:issue:`4015`) + + +Deprecations +~~~~~~~~~~~~ + +* The LevelDB_ storage backend + (``scrapy.extensions.httpcache.LeveldbCacheStorage``) of + :class:`~scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware` is + deprecated (:issue:`4085`, :issue:`4092`) + +* Use of the undocumented ``SCRAPY_PICKLED_SETTINGS_TO_OVERRIDE`` environment + variable is deprecated (:issue:`3910`) + +* ``scrapy.item.DictItem`` is deprecated, use :class:`~scrapy.item.Item` + instead (:issue:`3999`) + + +Other changes +~~~~~~~~~~~~~ + +* Minimum versions of optional Scrapy requirements that are covered by + continuous integration tests have been updated: + + * botocore_ 1.3.23 + * Pillow_ 3.4.2 + + Lower versions of these optional requirements may work, but it is not + guaranteed (:issue:`3892`) + +* GitHub templates for bug reports and feature requests (:issue:`3126`, + :issue:`3471`, :issue:`3749`, :issue:`3754`) + +* Continuous integration fixes (:issue:`3923`) + +* Code cleanup (:issue:`3391`, :issue:`3907`, :issue:`3946`, :issue:`3950`, + :issue:`4023`, :issue:`4031`) + + +.. _release-1.7.4: + Scrapy 1.7.4 (2019-10-21) ------------------------- @@ -18,22 +221,31 @@ makes later calls to :meth:`ItemLoader.get_output_value() ` or :meth:`ItemLoader.load_item() ` return empty data. + +.. _release-1.7.3: + Scrapy 1.7.3 (2019-08-01) ------------------------- Enforce lxml 4.3.5 or lower for Python 3.4 (:issue:`3912`, :issue:`3918`). + +.. _release-1.7.2: + Scrapy 1.7.2 (2019-07-23) ------------------------- Fix Python 2 support (:issue:`3889`, :issue:`3893`, :issue:`3896`). +.. _release-1.7.1: + Scrapy 1.7.1 (2019-07-18) ------------------------- Re-packaging of Scrapy 1.7.0, which was missing some changes in PyPI. + .. _release-1.7.0: Scrapy 1.7.0 (2019-07-18) @@ -568,7 +780,7 @@ Scrapy 1.5.2 (2019-01-22) See :ref:`telnet console ` documentation for more info -* Backport CI build failure under GCE environemnt due to boto import error. +* Backport CI build failure under GCE environment due to boto import error. .. _release-1.5.1: @@ -2830,23 +3042,35 @@ First release of Scrapy. .. _AJAX crawleable urls: https://developers.google.com/webmasters/ajax-crawling/docs/getting-started?csw=1 +.. _botocore: https://github.com/boto/botocore .. _chunked transfer encoding: https://en.wikipedia.org/wiki/Chunked_transfer_encoding .. _ClientForm: http://wwwsearch.sourceforge.net/old/ClientForm/ .. _Creating a pull request: https://help.github.com/en/articles/creating-a-pull-request +.. _cryptography: https://cryptography.io/en/latest/ .. _cssselect: https://github.com/scrapy/cssselect/ .. _docstrings: https://docs.python.org/glossary.html#term-docstring .. _KeyboardInterrupt: https://docs.python.org/library/exceptions.html#KeyboardInterrupt +.. _LevelDB: https://github.com/google/leveldb .. _lxml: http://lxml.de/ .. _marshal: https://docs.python.org/2/library/marshal.html .. _parsel.csstranslator.GenericTranslator: https://parsel.readthedocs.io/en/latest/parsel.html#parsel.csstranslator.GenericTranslator .. _parsel.csstranslator.HTMLTranslator: https://parsel.readthedocs.io/en/latest/parsel.html#parsel.csstranslator.HTMLTranslator .. _parsel.csstranslator.XPathExpr: https://parsel.readthedocs.io/en/latest/parsel.html#parsel.csstranslator.XPathExpr .. _PEP 257: https://www.python.org/dev/peps/pep-0257/ +.. _Pillow: https://python-pillow.org/ +.. _pyOpenSSL: https://www.pyopenssl.org/en/stable/ .. _queuelib: https://github.com/scrapy/queuelib +.. _registered with IANA: https://www.iana.org/assignments/media-types/media-types.xhtml .. _resource: https://docs.python.org/2/library/resource.html +.. _robots.txt: http://www.robotstxt.org/ .. _scrapely: https://github.com/scrapy/scrapely +.. _service_identity: https://service-identity.readthedocs.io/en/stable/ +.. _six: https://six.readthedocs.io/ .. _tox: https://pypi.python.org/pypi/tox +.. _Twisted: https://twistedmatrix.com/trac/ .. _Twisted - hello, asynchronous programming: http://jessenoller.com/blog/2009/02/11/twisted-hello-asynchronous-programming/ .. _w3lib: https://github.com/scrapy/w3lib .. _w3lib.encoding: https://github.com/scrapy/w3lib/blob/master/w3lib/encoding.py .. _What is cacheable: https://www.w3.org/Protocols/rfc2616/rfc2616-sec14.html#sec14.9.1 +.. _zope.interface: https://zopeinterface.readthedocs.io/en/latest/ +.. _Zsh: https://www.zsh.org/ diff --git a/docs/topics/logging.rst b/docs/topics/logging.rst index 87ea43c7d..2db0ffddd 100644 --- a/docs/topics/logging.rst +++ b/docs/topics/logging.rst @@ -198,8 +198,9 @@ to override some of the Scrapy settings regarding logging. Custom Log Formats ------------------ -A custom log format can be set for different actions by extending :class:`~scrapy.logformatter.LogFormatter` class -and making :setting:`LOG_FORMATTER` point to your new class. +A custom log format can be set for different actions by extending +:class:`~scrapy.logformatter.LogFormatter` class and making +:setting:`LOG_FORMATTER` point to your new class. .. autoclass:: scrapy.logformatter.LogFormatter :members: diff --git a/scrapy/logformatter.py b/scrapy/logformatter.py index f15940ed1..3c61ed7e0 100644 --- a/scrapy/logformatter.py +++ b/scrapy/logformatter.py @@ -29,7 +29,7 @@ class LogFormatter(object): * ``args`` should be a tuple or dict with the formatting placeholders for ``msg``. The final log message is computed as ``msg % args``. - Users can define their own ``LogFormatter`` class if they want to customise how + Users can define their own ``LogFormatter`` class if they want to customize how each action is logged or if they want to omit it entirely. In order to omit logging an action the method must return ``None``. From be2e910dd06ba4904e7b10eb5a7e3251e8dab099 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Tue, 29 Oct 2019 12:57:02 +0100 Subject: [PATCH 165/387] =?UTF-8?q?Bump=20version:=201.7.0=20=E2=86=92=201?= =?UTF-8?q?.8.0?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .bumpversion.cfg | 2 +- scrapy/VERSION | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/.bumpversion.cfg b/.bumpversion.cfg index 70affe63f..c9f1abea5 100644 --- a/.bumpversion.cfg +++ b/.bumpversion.cfg @@ -1,5 +1,5 @@ [bumpversion] -current_version = 1.7.0 +current_version = 1.8.0 commit = True tag = True tag_name = {new_version} diff --git a/scrapy/VERSION b/scrapy/VERSION index bd8bf882d..27f9cd322 100644 --- a/scrapy/VERSION +++ b/scrapy/VERSION @@ -1 +1 @@ -1.7.0 +1.8.0 From 66cbceeb0a9104fc0fa238898e38d0d9ce9cbcf6 Mon Sep 17 00:00:00 2001 From: Amardeep Bhowmick Date: Wed, 30 Oct 2019 13:39:12 +0530 Subject: [PATCH 166/387] Fix redirection error when the Location header value starts with 3 slashes (#4042) --- scrapy/downloadermiddlewares/redirect.py | 7 +++++-- tests/test_downloadermiddleware_redirect.py | 16 ++++++++++++++++ 2 files changed, 21 insertions(+), 2 deletions(-) diff --git a/scrapy/downloadermiddlewares/redirect.py b/scrapy/downloadermiddlewares/redirect.py index 49468a2e4..b73f864dd 100644 --- a/scrapy/downloadermiddlewares/redirect.py +++ b/scrapy/downloadermiddlewares/redirect.py @@ -1,5 +1,5 @@ import logging -from six.moves.urllib.parse import urljoin +from six.moves.urllib.parse import urljoin, urlparse from w3lib.url import safe_url_string @@ -70,7 +70,10 @@ class RedirectMiddleware(BaseRedirectMiddleware): if 'Location' not in response.headers or response.status not in allowed_status: return response - location = safe_url_string(response.headers['location']) + location = safe_url_string(response.headers['Location']) + if response.headers['Location'].startswith(b'//'): + request_scheme = urlparse(request.url).scheme + location = request_scheme + '://' + location.lstrip('/') redirected_url = urljoin(request.url, location) diff --git a/tests/test_downloadermiddleware_redirect.py b/tests/test_downloadermiddleware_redirect.py index 0e841489d..e7faf14a7 100644 --- a/tests/test_downloadermiddleware_redirect.py +++ b/tests/test_downloadermiddleware_redirect.py @@ -106,6 +106,22 @@ class RedirectMiddlewareTest(unittest.TestCase): del rsp.headers['Location'] assert self.mw.process_response(req, rsp, self.spider) is rsp + def test_redirect_302_relative(self): + url = 'http://www.example.com/302' + url2 = '///i8n.example2.com/302' + url3 = 'http://i8n.example2.com/302' + req = Request(url, method='HEAD') + rsp = Response(url, headers={'Location': url2}, status=302) + + req2 = self.mw.process_response(req, rsp, self.spider) + assert isinstance(req2, Request) + self.assertEqual(req2.url, url3) + self.assertEqual(req2.method, 'HEAD') + + # response without Location header but with status code is 3XX should be ignored + del rsp.headers['Location'] + assert self.mw.process_response(req, rsp, self.spider) is rsp + def test_max_redirect_times(self): self.mw.max_redirect_times = 1 From 6d6da78eda3cc0bba1bfdf70194fdf655fac8aeb Mon Sep 17 00:00:00 2001 From: Benjamin Ooghe-Tabanou Date: Wed, 30 Oct 2019 09:13:36 +0100 Subject: [PATCH 167/387] Add a keep_fragments parameter to the request_fingerprint function (#4104) --- scrapy/utils/request.py | 16 +++++++++++----- tests/test_utils_request.py | 9 ++++++++- 2 files changed, 19 insertions(+), 6 deletions(-) diff --git a/scrapy/utils/request.py b/scrapy/utils/request.py index 9c143b83a..fb5af66a2 100644 --- a/scrapy/utils/request.py +++ b/scrapy/utils/request.py @@ -16,7 +16,7 @@ from scrapy.utils.httpobj import urlparse_cached _fingerprint_cache = weakref.WeakKeyDictionary() -def request_fingerprint(request, include_headers=None): +def request_fingerprint(request, include_headers=None, keep_fragments=False): """ Return the request fingerprint. @@ -42,15 +42,21 @@ def request_fingerprint(request, include_headers=None): the fingeprint. If you want to include specific headers use the include_headers argument, which is a list of Request headers to include. + Also, servers usually ignore fragments in urls when handling requests, + so they are also ignored by default when calculating the fingerprint. + If you want to include them, set the keep_fragments argument to True + (for instance when handling requests with a headless browser). + """ if include_headers: include_headers = tuple(to_bytes(h.lower()) for h in sorted(include_headers)) cache = _fingerprint_cache.setdefault(request, {}) - if include_headers not in cache: + cache_key = (include_headers, keep_fragments) + if cache_key not in cache: fp = hashlib.sha1() fp.update(to_bytes(request.method)) - fp.update(to_bytes(canonicalize_url(request.url))) + fp.update(to_bytes(canonicalize_url(request.url, keep_fragments=keep_fragments))) fp.update(request.body or b'') if include_headers: for hdr in include_headers: @@ -58,8 +64,8 @@ def request_fingerprint(request, include_headers=None): fp.update(hdr) for v in request.headers.getlist(hdr): fp.update(v) - cache[include_headers] = fp.hexdigest() - return cache[include_headers] + cache[cache_key] = fp.hexdigest() + return cache[cache_key] def request_authenticate(request, username, password): diff --git a/tests/test_utils_request.py b/tests/test_utils_request.py index e8a4eb3ea..625a32048 100644 --- a/tests/test_utils_request.py +++ b/tests/test_utils_request.py @@ -17,7 +17,7 @@ class UtilsRequestTest(unittest.TestCase): self.assertNotEqual(request_fingerprint(r1), request_fingerprint(r2)) # make sure caching is working - self.assertEqual(request_fingerprint(r1), _fingerprint_cache[r1][None]) + self.assertEqual(request_fingerprint(r1), _fingerprint_cache[r1][(None, False)]) r1 = Request("http://www.example.com/members/offers.html") r2 = Request("http://www.example.com/members/offers.html") @@ -42,6 +42,13 @@ class UtilsRequestTest(unittest.TestCase): self.assertEqual(request_fingerprint(r3, include_headers=['accept-language', 'sessionid']), request_fingerprint(r3, include_headers=['SESSIONID', 'Accept-Language'])) + r1 = Request("http://www.example.com/test.html") + r2 = Request("http://www.example.com/test.html#fragment") + self.assertEqual(request_fingerprint(r1), request_fingerprint(r2)) + self.assertEqual(request_fingerprint(r1), request_fingerprint(r1, keep_fragments=True)) + self.assertNotEqual(request_fingerprint(r2), request_fingerprint(r2, keep_fragments=True)) + self.assertNotEqual(request_fingerprint(r1), request_fingerprint(r2, keep_fragments=True)) + r1 = Request("http://www.example.com") r2 = Request("http://www.example.com", method='POST') r3 = Request("http://www.example.com", method='POST', body=b'request body') From 229e722a03aced0fb62ec8c631b036c3e13e2188 Mon Sep 17 00:00:00 2001 From: Andrey Rahmatullin Date: Thu, 31 Oct 2019 14:46:02 +0500 Subject: [PATCH 168/387] Initial Python 2 removal (#4091) --- .travis.yml | 10 +----- README.rst | 2 +- docs/faq.rst | 4 +-- docs/intro/install.rst | 16 +++------- docs/topics/feed-exports.rst | 7 ++--- docs/topics/leaks.rst | 2 +- docs/topics/media-pipeline.rst | 2 +- docs/topics/request-response.rst | 4 +-- docs/topics/settings.rst | 5 +-- docs/topics/spiders.rst | 2 -- requirements-py2.txt | 18 ----------- scrapy/__init__.py | 4 +-- setup.py | 7 ++--- tests/requirements-py2.txt | 15 --------- tox.ini | 53 ++------------------------------ 15 files changed, 24 insertions(+), 127 deletions(-) delete mode 100644 requirements-py2.txt delete mode 100644 tests/requirements-py2.txt diff --git a/.travis.yml b/.travis.yml index 4c2498053..2352ef124 100644 --- a/.travis.yml +++ b/.travis.yml @@ -7,14 +7,6 @@ branches: - /^\d\.\d+\.\d+(rc\d+|\.dev\d+)?$/ matrix: include: - - env: TOXENV=py27 - python: 2.7 - - env: TOXENV=py27-pinned - python: 2.7 - - env: TOXENV=py27-extra-deps - python: 2.7 - - env: TOXENV=pypy - python: 2.7 - env: TOXENV=pypy3 python: 3.5 - env: TOXENV=py35 @@ -70,4 +62,4 @@ deploy: on: tags: true repo: scrapy/scrapy - condition: "$TOXENV == py27 && $TRAVIS_TAG =~ ^[0-9]+[.][0-9]+[.][0-9]+(rc[0-9]+|[.]dev[0-9]+)?$" + condition: "$TOXENV == py37 && $TRAVIS_TAG =~ ^[0-9]+[.][0-9]+[.][0-9]+(rc[0-9]+|[.]dev[0-9]+)?$" diff --git a/README.rst b/README.rst index bd82bff06..fb4ca8e4f 100644 --- a/README.rst +++ b/README.rst @@ -40,7 +40,7 @@ https://scrapy.org Requirements ============ -* Python 2.7 or Python 3.5+ +* Python 3.5+ * Works on Linux, Windows, Mac OSX, BSD Install diff --git a/docs/faq.rst b/docs/faq.rst index 9733471bf..080d81981 100644 --- a/docs/faq.rst +++ b/docs/faq.rst @@ -69,11 +69,11 @@ Here's an example spider using BeautifulSoup API, with ``lxml`` as the HTML pars What Python versions does Scrapy support? ----------------------------------------- -Scrapy is supported under Python 2.7 and Python 3.5+ +Scrapy is supported under Python 3.5+ under CPython (default Python implementation) and PyPy (starting with PyPy 5.9). -Python 2.6 support was dropped starting at Scrapy 0.20. Python 3 support was added in Scrapy 1.1. PyPy support was added in Scrapy 1.4, PyPy3 support was added in Scrapy 1.5. +Python 2 support was dropped in Scrapy 2.0. .. note:: For Python 3 support on Windows, it is recommended to use diff --git a/docs/intro/install.rst b/docs/intro/install.rst index 51b41b4d7..e924b5303 100644 --- a/docs/intro/install.rst +++ b/docs/intro/install.rst @@ -7,7 +7,7 @@ Installation guide Installing Scrapy ================= -Scrapy runs on Python 2.7 and Python 3.5 or above +Scrapy runs on Python 3.5 or above under CPython (default Python implementation) and PyPy (starting with PyPy 5.9). If you're using `Anaconda`_ or `Miniconda`_, you can install the package from @@ -102,10 +102,8 @@ just like any other Python package. (See :ref:`platform-specific guides ` below for non-Python dependencies that you may need to install beforehand). -Python virtualenvs can be created to use Python 2 by default, or Python 3 by default. - -* If you want to install scrapy with Python 3, install scrapy within a Python 3 virtualenv. -* And if you want to install scrapy with Python 2, install scrapy within a Python 2 virtualenv. +Python virtualenvs can be created to use Python 2 by default, or Python 3 by default. As Scrapy +only supports Python 3, make sure you created a Python 3 virtualenv. .. _virtualenv: https://virtualenv.pypa.io .. _virtualenv installation instructions: https://virtualenv.pypa.io/en/stable/installation/ @@ -149,16 +147,12 @@ typically too old and slow to catch up with latest Scrapy. To install scrapy on Ubuntu (or Ubuntu-based) systems, you need to install these dependencies:: - sudo apt-get install python-dev python-pip libxml2-dev libxslt1-dev zlib1g-dev libffi-dev libssl-dev + sudo apt-get install python3 python3-dev python3-pip libxml2-dev libxslt1-dev zlib1g-dev libffi-dev libssl-dev -- ``python-dev``, ``zlib1g-dev``, ``libxml2-dev`` and ``libxslt1-dev`` +- ``python3-dev``, ``zlib1g-dev``, ``libxml2-dev`` and ``libxslt1-dev`` are required for ``lxml`` - ``libssl-dev`` and ``libffi-dev`` are required for ``cryptography`` -If you want to install scrapy on Python 3, you’ll also need Python 3 development headers:: - - sudo apt-get install python3 python3-dev - Inside a :ref:`virtualenv `, you can install Scrapy with ``pip`` after that:: diff --git a/docs/topics/feed-exports.rst b/docs/topics/feed-exports.rst index af541db78..7481b1a99 100644 --- a/docs/topics/feed-exports.rst +++ b/docs/topics/feed-exports.rst @@ -99,12 +99,12 @@ The storages backends supported out of the box are: * :ref:`topics-feed-storage-fs` * :ref:`topics-feed-storage-ftp` - * :ref:`topics-feed-storage-s3` (requires botocore_ or boto_) + * :ref:`topics-feed-storage-s3` (requires botocore_) * :ref:`topics-feed-storage-stdout` Some storage backends may be unavailable if the required external libraries are not available. For example, the S3 backend is only available if the botocore_ -or boto_ library is installed (Scrapy supports boto_ only on Python 2). +library is installed. .. _topics-feed-uri-params: @@ -182,7 +182,7 @@ The feeds are stored on `Amazon S3`_. * ``s3://mybucket/path/to/export.csv`` * ``s3://aws_key:aws_secret@mybucket/path/to/export.csv`` - * Required external libraries: `botocore`_ (Python 2 and Python 3) or `boto`_ (Python 2 only) + * Required external libraries: `botocore`_ The AWS credentials can be passed as user/password in the URI, or they can be passed through the following settings: @@ -399,6 +399,5 @@ format in :setting:`FEED_EXPORTERS`. E.g., to disable the built-in CSV exporter .. _URI: https://en.wikipedia.org/wiki/Uniform_Resource_Identifier .. _Amazon S3: https://aws.amazon.com/s3/ -.. _boto: https://github.com/boto/boto .. _botocore: https://github.com/boto/botocore .. _Canned ACL: https://docs.aws.amazon.com/AmazonS3/latest/dev/acl-overview.html#canned-acl diff --git a/docs/topics/leaks.rst b/docs/topics/leaks.rst index 8278e9849..793636f59 100644 --- a/docs/topics/leaks.rst +++ b/docs/topics/leaks.rst @@ -260,7 +260,7 @@ knowledge about Python internals. For more info about Guppy, refer to the Debugging memory leaks with muppy ================================= -If you're using Python 3, you can use muppy from `Pympler`_. +You can use muppy from `Pympler`_. .. _Pympler: https://pypi.org/project/Pympler/ diff --git a/docs/topics/media-pipeline.rst b/docs/topics/media-pipeline.rst index 0ce431ff5..431cc6027 100644 --- a/docs/topics/media-pipeline.rst +++ b/docs/topics/media-pipeline.rst @@ -171,7 +171,7 @@ policy:: For more information, see `canned ACLs`_ in the Amazon S3 Developer Guide. -Because Scrapy uses ``boto`` / ``botocore`` internally you can also use other S3-like storages. Storages like +Because Scrapy uses ``botocore`` internally you can also use other S3-like storages. Storages like self-hosted `Minio`_ or `s3.scality`_. All you need to do is set endpoint option in you Scrapy settings:: AWS_ENDPOINT_URL = 'http://minio.example.com:9000' diff --git a/docs/topics/request-response.rst b/docs/topics/request-response.rst index 727c67482..5a76e189e 100644 --- a/docs/topics/request-response.rst +++ b/docs/topics/request-response.rst @@ -596,8 +596,8 @@ Response objects (for single valued headers) or lists (for multi-valued headers). :type headers: dict - :param body: the response body. To access the decoded text as str (unicode - in Python 2) you can use ``response.text`` from an encoding-aware + :param body: the response body. To access the decoded text as str you can use + ``response.text`` from an encoding-aware :ref:`Response subclass `, such as :class:`TextResponse`. :type body: bytes diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 56375664f..a1d15a760 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -188,7 +188,6 @@ AWS_ENDPOINT_URL Default: ``None`` Endpoint URL used for S3-like storage, for example Minio or s3.scality. -Only supported with ``botocore`` library. .. setting:: AWS_USE_SSL @@ -199,7 +198,6 @@ Default: ``None`` Use this option if you want to disable SSL connection for communication with S3 or S3-like storage. By default SSL will be used. -Only supported with ``botocore`` library. .. setting:: AWS_VERIFY @@ -209,7 +207,7 @@ AWS_VERIFY Default: ``None`` Verify SSL connection between Scrapy and S3 or S3-like storage. By default -SSL verification will occur. Only supported with ``botocore`` library. +SSL verification will occur. .. setting:: AWS_REGION_NAME @@ -219,7 +217,6 @@ AWS_REGION_NAME Default: ``None`` The name of the region associated with the AWS client. -Only supported with ``botocore`` library. .. setting:: BOT_NAME diff --git a/docs/topics/spiders.rst b/docs/topics/spiders.rst index d60c93be6..d65a43afd 100644 --- a/docs/topics/spiders.rst +++ b/docs/topics/spiders.rst @@ -72,8 +72,6 @@ scrapy.Spider spider that crawls ``mywebsite.com`` would often be called ``mywebsite``. - .. note:: In Python 2 this must be ASCII only. - .. attribute:: allowed_domains An optional list of strings containing domains that this spider is diff --git a/requirements-py2.txt b/requirements-py2.txt deleted file mode 100644 index dde8d1c9c..000000000 --- a/requirements-py2.txt +++ /dev/null @@ -1,18 +0,0 @@ -parsel>=1.5.0 -PyDispatcher>=2.0.5 -w3lib>=1.17.0 -protego>=0.1.15 - -pyOpenSSL>=16.2.0 # Earlier versions fail with "AttributeError: module 'lib' has no attribute 'SSL_ST_INIT'" -queuelib>=1.4.2 # Earlier versions fail with "AttributeError: '...QueueTest' object has no attribute 'qpath'" -cryptography>=2.0 # Earlier versions would fail to install - -# Reference versions taken from -# https://packages.ubuntu.com/xenial/python/ -# https://packages.ubuntu.com/xenial/zope/ -cssselect>=0.9.1 -lxml>=3.5.0 -service_identity>=16.0.0 -six>=1.10.0 -Twisted>=16.0.0 -zope.interface>=4.1.3 diff --git a/scrapy/__init__.py b/scrapy/__init__.py index 03ec6c667..230e5cee3 100644 --- a/scrapy/__init__.py +++ b/scrapy/__init__.py @@ -14,8 +14,8 @@ del pkgutil # Check minimum required Python version import sys -if sys.version_info < (2, 7): - print("Scrapy %s requires Python 2.7" % __version__) +if sys.version_info < (3, 5): + print("Scrapy %s requires Python 3.5" % __version__) sys.exit(1) # Ignore noisy twisted deprecation warnings diff --git a/setup.py b/setup.py index 2f5fca4c9..8f5f14f0d 100644 --- a/setup.py +++ b/setup.py @@ -50,8 +50,6 @@ setup( 'License :: OSI Approved :: BSD License', 'Operating System :: OS Independent', 'Programming Language :: Python', - 'Programming Language :: Python :: 2', - 'Programming Language :: Python :: 2.7', 'Programming Language :: Python :: 3', 'Programming Language :: Python :: 3.5', 'Programming Language :: Python :: 3.6', @@ -63,10 +61,9 @@ setup( 'Topic :: Software Development :: Libraries :: Application Frameworks', 'Topic :: Software Development :: Libraries :: Python Modules', ], - python_requires='>=2.7, !=3.0.*, !=3.1.*, !=3.2.*, !=3.3.*, !=3.4.*', + python_requires='>=3.5', install_requires=[ - 'Twisted>=16.0.0;python_version=="2.7"', - 'Twisted>=17.9.0;python_version>="3.5"', + 'Twisted>=17.9.0', 'cryptography>=2.0', 'cssselect>=0.9.1', 'lxml>=3.5.0', diff --git a/tests/requirements-py2.txt b/tests/requirements-py2.txt deleted file mode 100644 index f621eb4eb..000000000 --- a/tests/requirements-py2.txt +++ /dev/null @@ -1,15 +0,0 @@ -# Tests requirements -brotlipy -jmespath -mitmproxy==0.10.1 -mock -netlib==0.10.1 -pytest -pytest-cov -pytest-twisted -pytest-xdist -testfixtures - -# optional for shell wrapper tests -bpython -ipython<6.0 diff --git a/tox.ini b/tox.ini index fe925951b..821144381 100644 --- a/tox.ini +++ b/tox.ini @@ -4,18 +4,16 @@ # and then run "tox" from this directory. [tox] -envlist = py27 +envlist = py35 [testenv] deps = -ctests/constraints.txt - -rrequirements-py2.txt + -rrequirements-py3.txt + -rtests/requirements-py3.txt # Extras botocore>=1.3.23 - google-cloud-storage - leveldb Pillow>=3.4.2 - -rtests/requirements-py2.txt passenv = S3_TEST_FILE_URI AWS_ACCESS_KEY_ID @@ -25,42 +23,8 @@ passenv = commands = py.test --cov=scrapy --cov-report= {posargs:scrapy tests} -[testenv:py27-pinned] -basepython = python2.7 -deps = - -ctests/constraints.txt - cryptography==2.0 - cssselect==0.9.1 - lxml==3.5.0 - parsel==1.5.0 - Protego==0.1.15 - PyDispatcher==2.0.5 - pyOpenSSL==16.2.0 - queuelib==1.4.2 - service_identity==16.0.0 - six==1.10.0 - Twisted==16.0.0 - w3lib==1.17.0 - zope.interface==4.1.3 - -rtests/requirements-py2.txt - # Extras - botocore==1.3.23 - Pillow==3.4.2 - -[testenv:pypy] -basepython = pypy -commands = - py.test {posargs:scrapy tests} - [testenv:py35] basepython = python3.5 -deps = - -ctests/constraints.txt - -rrequirements-py3.txt - -rtests/requirements-py3.txt - # Extras - botocore>=1.3.23 - Pillow>=3.4.2 [testenv:py35-pinned] basepython = python3.5 @@ -86,19 +50,15 @@ deps = [testenv:py36] basepython = python3.6 -deps = {[testenv:py35]deps} [testenv:py37] basepython = python3.7 -deps = {[testenv:py35]deps} [testenv:py38] basepython = python3.8 -deps = {[testenv:py35]deps} [testenv:pypy3] basepython = pypy3 -deps = {[testenv:py35]deps} commands = py.test {posargs:scrapy tests} @@ -127,13 +87,6 @@ commands = [testenv:py38-extra-deps] basepython = python3.8 -deps = - {[testenv:py35]deps} - reppy - robotexclusionrulesparser - -[testenv:py27-extra-deps] -basepython = python2.7 deps = {[testenv]deps} reppy From 15c55d0c1d4a4384fd9725d47c305f52918e3831 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 31 Oct 2019 10:47:29 +0100 Subject: [PATCH 169/387] Remove LevelDB support (#4112) --- scrapy/extensions/httpcache.py | 71 -------------------- tests/requirements-py3.txt | 1 - tests/test_downloadermiddleware_httpcache.py | 9 --- 3 files changed, 81 deletions(-) diff --git a/scrapy/extensions/httpcache.py b/scrapy/extensions/httpcache.py index 7c650a91e..f3fabf710 100644 --- a/scrapy/extensions/httpcache.py +++ b/scrapy/extensions/httpcache.py @@ -347,77 +347,6 @@ class FilesystemCacheStorage(object): return pickle.load(f) -class LeveldbCacheStorage(object): - - def __init__(self, settings): - warn("The LevelDB storage backend is deprecated.", - ScrapyDeprecationWarning, stacklevel=2) - import leveldb - self._leveldb = leveldb - self.cachedir = data_path(settings['HTTPCACHE_DIR'], createdir=True) - self.expiration_secs = settings.getint('HTTPCACHE_EXPIRATION_SECS') - self.db = None - - def open_spider(self, spider): - dbpath = os.path.join(self.cachedir, '%s.leveldb' % spider.name) - self.db = self._leveldb.LevelDB(dbpath) - - logger.debug("Using LevelDB cache storage in %(cachepath)s" % {'cachepath': dbpath}, extra={'spider': spider}) - - def close_spider(self, spider): - # Do compactation each time to save space and also recreate files to - # avoid them being removed in storages with timestamp-based autoremoval. - self.db.CompactRange() - del self.db - garbage_collect() - - def retrieve_response(self, spider, request): - data = self._read_data(spider, request) - if data is None: - return # not cached - url = data['url'] - status = data['status'] - headers = Headers(data['headers']) - body = data['body'] - respcls = responsetypes.from_args(headers=headers, url=url) - response = respcls(url=url, headers=headers, status=status, body=body) - return response - - def store_response(self, spider, request, response): - key = self._request_key(request) - data = { - 'status': response.status, - 'url': response.url, - 'headers': dict(response.headers), - 'body': response.body, - } - batch = self._leveldb.WriteBatch() - batch.Put(key + b'_data', pickle.dumps(data, protocol=2)) - batch.Put(key + b'_time', to_bytes(str(time()))) - self.db.Write(batch) - - def _read_data(self, spider, request): - key = self._request_key(request) - try: - ts = self.db.Get(key + b'_time') - except KeyError: - return # not found or invalid entry - - if 0 < self.expiration_secs < time() - float(ts): - return # expired - - try: - data = self.db.Get(key + b'_data') - except KeyError: - return # invalid entry - else: - return pickle.loads(data) - - def _request_key(self, request): - return to_bytes(request_fingerprint(request)) - - - def parse_cachecontrol(header): """Parse Cache-Control header diff --git a/tests/requirements-py3.txt b/tests/requirements-py3.txt index cb67bc40e..dd5b23cc3 100644 --- a/tests/requirements-py3.txt +++ b/tests/requirements-py3.txt @@ -1,6 +1,5 @@ # Tests requirements jmespath -leveldb; sys_platform != "win32" pytest pytest-cov pytest-twisted diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index 972d400a4..950664ffe 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -156,15 +156,6 @@ class FilesystemStorageGzipTest(FilesystemStorageTest): return super(FilesystemStorageTest, self)._get_settings(**new_settings) -class LeveldbStorageTest(DefaultStorageTest): - - try: - pytest.importorskip('leveldb') - except SystemError: - pytestmark = pytest.mark.skip("Test module skipped - 'SystemError: bad call flags' occurs when >= Python 3.8") - storage_class = 'scrapy.extensions.httpcache.LeveldbCacheStorage' - - class DummyPolicyTest(_BaseTest): policy_class = 'scrapy.extensions.httpcache.DummyPolicy' From b44bd6f8250579dc9ddc25d5e0f10d8b3790f5ea Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Mon, 22 Jul 2019 20:51:03 +0500 Subject: [PATCH 170/387] Remove Python 2-only tests. --- conftest.py | 10 +- tests/py3-ignores.txt | 3 - tests/test_downloader_handlers.py | 22 +-- tests/test_linkextractors_deprecated.py | 233 ------------------------ tests/test_proxy_connect.py | 120 ------------ tests/test_utils_python.py | 27 --- tests/test_webclient.py | 20 -- 7 files changed, 14 insertions(+), 421 deletions(-) delete mode 100644 tests/test_linkextractors_deprecated.py delete mode 100644 tests/test_proxy_connect.py diff --git a/conftest.py b/conftest.py index 06d65ba1d..ede091e9f 100644 --- a/conftest.py +++ b/conftest.py @@ -1,4 +1,3 @@ -import six import pytest @@ -8,11 +7,10 @@ collect_ignore = [ ] -if six.PY3: - for line in open('tests/py3-ignores.txt'): - file_path = line.strip() - if file_path and file_path[0] != '#': - collect_ignore.append(file_path) +for line in open('tests/py3-ignores.txt'): + file_path = line.strip() + if file_path and file_path[0] != '#': + collect_ignore.append(file_path) @pytest.fixture() diff --git a/tests/py3-ignores.txt b/tests/py3-ignores.txt index 313e74ec9..45cf6fb92 100644 --- a/tests/py3-ignores.txt +++ b/tests/py3-ignores.txt @@ -1,6 +1,3 @@ -tests/test_linkextractors_deprecated.py -tests/test_proxy_connect.py - scrapy/linkextractors/sgml.py scrapy/linkextractors/regex.py scrapy/linkextractors/htmlparser.py diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 109469503..4d3c4d4aa 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -633,18 +633,16 @@ class Http11MockServerTestCase(unittest.TestCase): # download_maxsize < 100, hence the CancelledError self.assertIsInstance(failure.value, defer.CancelledError) - if six.PY2: - request.headers.setdefault(b'Accept-Encoding', b'gzip,deflate') - request = request.replace(url=self.mockserver.url('/xpayload')) - yield crawler.crawl(seed=request) - # download_maxsize = 50 is enough for the gzipped response - failure = crawler.spider.meta.get('failure') - self.assertTrue(failure == None) - reason = crawler.spider.meta['close_reason'] - self.assertTrue(reason, 'finished') - else: - # See issue https://twistedmatrix.com/trac/ticket/8175 - raise unittest.SkipTest("xpayload only enabled for PY2") + # See issue https://twistedmatrix.com/trac/ticket/8175 + raise unittest.SkipTest("xpayload only enabled for PY2") + request.headers.setdefault(b'Accept-Encoding', b'gzip,deflate') + request = request.replace(url=self.mockserver.url('/xpayload')) + yield crawler.crawl(seed=request) + # download_maxsize = 50 is enough for the gzipped response + failure = crawler.spider.meta.get('failure') + self.assertTrue(failure == None) + reason = crawler.spider.meta['close_reason'] + self.assertTrue(reason, 'finished') class UriResource(resource.Resource): diff --git a/tests/test_linkextractors_deprecated.py b/tests/test_linkextractors_deprecated.py deleted file mode 100644 index 1366971be..000000000 --- a/tests/test_linkextractors_deprecated.py +++ /dev/null @@ -1,233 +0,0 @@ -# -*- coding: utf-8 -*- -import unittest -from scrapy.linkextractors.regex import RegexLinkExtractor -from scrapy.http import HtmlResponse -from scrapy.link import Link -from scrapy.linkextractors.htmlparser import HtmlParserLinkExtractor -from scrapy.linkextractors.sgml import SgmlLinkExtractor, BaseSgmlLinkExtractor -from tests import get_testdata - -from tests.test_linkextractors import Base - - -class BaseSgmlLinkExtractorTestCase(unittest.TestCase): - # XXX: should we move some of these tests to base link extractor tests? - - def test_basic(self): - html = """Page title<title> - <body><p><a href="item/12.html">Item 12</a></p> - <p><a href="/about.html">About us</a></p> - <img src="/logo.png" alt="Company logo (not a link)" /> - <p><a href="../othercat.html">Other category</a></p> - <p><a href="/">>></a></p> - <p><a href="/" /></p> - </body></html>""" - response = HtmlResponse("http://example.org/somepage/index.html", body=html) - - lx = BaseSgmlLinkExtractor() # default: tag=a, attr=href - self.assertEqual(lx.extract_links(response), - [Link(url='http://example.org/somepage/item/12.html', text='Item 12'), - Link(url='http://example.org/about.html', text='About us'), - Link(url='http://example.org/othercat.html', text='Other category'), - Link(url='http://example.org/', text='>>'), - Link(url='http://example.org/', text='')]) - - def test_base_url(self): - html = """<html><head><title>Page title<title><base href="http://otherdomain.com/base/" /> - <body><p><a href="item/12.html">Item 12</a></p> - </body></html>""" - response = HtmlResponse("http://example.org/somepage/index.html", body=html) - - lx = BaseSgmlLinkExtractor() # default: tag=a, attr=href - self.assertEqual(lx.extract_links(response), - [Link(url='http://otherdomain.com/base/item/12.html', text='Item 12')]) - - # base url is an absolute path and relative to host - html = """<html><head><title>Page title<title><base href="/" /> - <body><p><a href="item/12.html">Item 12</a></p></body></html>""" - response = HtmlResponse("https://example.org/somepage/index.html", body=html) - self.assertEqual(lx.extract_links(response), - [Link(url='https://example.org/item/12.html', text='Item 12')]) - - # base url has no scheme - html = """<html><head><title>Page title<title><base href="//noschemedomain.com/path/to/" /> - <body><p><a href="item/12.html">Item 12</a></p></body></html>""" - response = HtmlResponse("https://example.org/somepage/index.html", body=html) - self.assertEqual(lx.extract_links(response), - [Link(url='https://noschemedomain.com/path/to/item/12.html', text='Item 12')]) - - def test_link_text_wrong_encoding(self): - html = """<body><p><a href="item/12.html">Wrong: \xed</a></p></body></html>""" - response = HtmlResponse("http://www.example.com", body=html, encoding='utf-8') - lx = BaseSgmlLinkExtractor() - self.assertEqual(lx.extract_links(response), [ - Link(url='http://www.example.com/item/12.html', text=u'Wrong: \ufffd'), - ]) - - def test_extraction_encoding(self): - body = get_testdata('link_extractor', 'linkextractor_noenc.html') - response_utf8 = HtmlResponse(url='http://example.com/utf8', body=body, headers={'Content-Type': ['text/html; charset=utf-8']}) - response_noenc = HtmlResponse(url='http://example.com/noenc', body=body) - body = get_testdata('link_extractor', 'linkextractor_latin1.html') - response_latin1 = HtmlResponse(url='http://example.com/latin1', body=body) - - lx = BaseSgmlLinkExtractor() - self.assertEqual(lx.extract_links(response_utf8), [ - Link(url='http://example.com/sample_%C3%B1.html', text=''), - Link(url='http://example.com/sample_%E2%82%AC.html', text='sample \xe2\x82\xac text'.decode('utf-8')), - ]) - - self.assertEqual(lx.extract_links(response_noenc), [ - Link(url='http://example.com/sample_%C3%B1.html', text=''), - Link(url='http://example.com/sample_%E2%82%AC.html', text='sample \xe2\x82\xac text'.decode('utf-8')), - ]) - - # document encoding does not affect URL path component, only query part - # >>> u'sample_ñ.html'.encode('utf8') - # b'sample_\xc3\xb1.html' - # >>> u"sample_á.html".encode('utf8') - # b'sample_\xc3\xa1.html' - # >>> u"sample_ö.html".encode('utf8') - # b'sample_\xc3\xb6.html' - # >>> u"£32".encode('latin1') - # b'\xa332' - # >>> u"µ".encode('latin1') - # b'\xb5' - self.assertEqual(lx.extract_links(response_latin1), [ - Link(url='http://example.com/sample_%C3%B1.html', text=''), - Link(url='http://example.com/sample_%C3%A1.html', text='sample \xe1 text'.decode('latin1')), - Link(url='http://example.com/sample_%C3%B6.html?price=%A332&%B5=unit', text=''), - ]) - - def test_matches(self): - url1 = 'http://lotsofstuff.com/stuff1/index' - url2 = 'http://evenmorestuff.com/uglystuff/index' - - lx = BaseSgmlLinkExtractor() - self.assertEqual(lx.matches(url1), True) - self.assertEqual(lx.matches(url2), True) - - -class HtmlParserLinkExtractorTestCase(unittest.TestCase): - - def setUp(self): - body = get_testdata('link_extractor', 'sgml_linkextractor.html') - self.response = HtmlResponse(url='http://example.com/index', body=body) - - def test_extraction(self): - # Default arguments - lx = HtmlParserLinkExtractor() - self.assertEqual(lx.extract_links(self.response), [ - Link(url='http://example.com/sample2.html', text=u'sample 2'), - Link(url='http://example.com/sample3.html', text=u'sample 3 text'), - Link(url='http://example.com/sample3.html', text=u'sample 3 repetition'), - Link(url='http://example.com/sample3.html#foo', text=u'sample 3 repetition with fragment'), - Link(url='http://www.google.com/something', text=u''), - Link(url='http://example.com/innertag.html', text=u'inner tag'), - Link(url='http://example.com/page%204.html', text=u'href with whitespaces'), - ]) - - def test_link_wrong_href(self): - html = """ - <a href="http://example.org/item1.html">Item 1</a> - <a href="http://[example.org/item2.html">Item 2</a> - <a href="http://example.org/item3.html">Item 3</a> - """ - response = HtmlResponse("http://example.org/index.html", body=html) - lx = HtmlParserLinkExtractor() - self.assertEqual([link for link in lx.extract_links(response)], [ - Link(url='http://example.org/item1.html', text=u'Item 1', nofollow=False), - Link(url='http://example.org/item3.html', text=u'Item 3', nofollow=False), - ]) - - -class SgmlLinkExtractorTestCase(Base.LinkExtractorTestCase): - extractor_cls = SgmlLinkExtractor - escapes_whitespace = True - - def test_deny_extensions(self): - html = """<a href="page.html">asd</a> and <a href="photo.jpg">""" - response = HtmlResponse("http://example.org/", body=html) - lx = SgmlLinkExtractor(deny_extensions="jpg") - self.assertEqual(lx.extract_links(response), [ - Link(url='http://example.org/page.html', text=u'asd'), - ]) - - def test_attrs_sgml(self): - html = """<html><area href="sample1.html"></area> - <a ref="sample2.html">sample text 2</a></html>""" - response = HtmlResponse("http://example.com/index.html", body=html) - lx = SgmlLinkExtractor(attrs="href") - self.assertEqual(lx.extract_links(response), [ - Link(url='http://example.com/sample1.html', text=u''), - ]) - - def test_link_nofollow(self): - html = """ - <a href="page.html?action=print" rel="nofollow">Printer-friendly page</a> - <a href="about.html">About us</a> - <a href="http://google.com/something" rel="external nofollow">Something</a> - """ - response = HtmlResponse("http://example.org/page.html", body=html) - lx = SgmlLinkExtractor() - self.assertEqual([link for link in lx.extract_links(response)], [ - Link(url='http://example.org/page.html?action=print', text=u'Printer-friendly page', nofollow=True), - Link(url='http://example.org/about.html', text=u'About us', nofollow=False), - Link(url='http://google.com/something', text=u'Something', nofollow=True), - ]) - - -class RegexLinkExtractorTestCase(unittest.TestCase): - # XXX: RegexLinkExtractor is not deprecated yet, but it must be rewritten - # not to depend on SgmlLinkExractor. Its speed is also much worse - # than it should be. - - def setUp(self): - body = get_testdata('link_extractor', 'sgml_linkextractor.html') - self.response = HtmlResponse(url='http://example.com/index', body=body) - - def test_extraction(self): - # Default arguments - lx = RegexLinkExtractor() - self.assertEqual(lx.extract_links(self.response), - [Link(url='http://example.com/sample2.html', text=u'sample 2'), - Link(url='http://example.com/sample3.html', text=u'sample 3 text'), - Link(url='http://example.com/sample3.html#foo', text=u'sample 3 repetition with fragment'), - Link(url='http://www.google.com/something', text=u''), - Link(url='http://example.com/innertag.html', text=u'inner tag'),]) - - def test_link_wrong_href(self): - html = """ - <a href="http://example.org/item1.html">Item 1</a> - <a href="http://[example.org/item2.html">Item 2</a> - <a href="http://example.org/item3.html">Item 3</a> - """ - response = HtmlResponse("http://example.org/index.html", body=html) - lx = RegexLinkExtractor() - self.assertEqual([link for link in lx.extract_links(response)], [ - Link(url='http://example.org/item1.html', text=u'Item 1', nofollow=False), - Link(url='http://example.org/item3.html', text=u'Item 3', nofollow=False), - ]) - - def test_html_base_href(self): - html = """ - <html> - <head> - <base href="http://b.com/"> - </head> - <body> - <a href="test.html"></a> - </body> - </html> - """ - response = HtmlResponse("http://a.com/", body=html) - lx = RegexLinkExtractor() - self.assertEqual([link for link in lx.extract_links(response)], [ - Link(url='http://b.com/test.html', text=u'', nofollow=False), - ]) - - @unittest.expectedFailure - def test_extraction(self): - # RegexLinkExtractor doesn't parse URLs with leading/trailing - # whitespaces correctly. - super(RegexLinkExtractorTestCase, self).test_extraction() diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py deleted file mode 100644 index ae1236bcb..000000000 --- a/tests/test_proxy_connect.py +++ /dev/null @@ -1,120 +0,0 @@ -import json -import os -import time - -from six.moves.urllib.parse import urlsplit, urlunsplit -from threading import Thread -from libmproxy import controller, proxy -from netlib import http_auth -from testfixtures import LogCapture - -from twisted.internet import defer -from twisted.trial.unittest import TestCase -from scrapy.utils.test import get_crawler -from scrapy.http import Request -from tests.spiders import SimpleSpider, SingleRequestSpider -from tests.mockserver import MockServer - - -class HTTPSProxy(controller.Master, Thread): - - def __init__(self): - password_manager = http_auth.PassManSingleUser('scrapy', 'scrapy') - authenticator = http_auth.BasicProxyAuth(password_manager, "mitmproxy") - cert_path = os.path.join(os.path.abspath(os.path.dirname(__file__)), - 'keys', 'mitmproxy-ca.pem') - server = proxy.ProxyServer(proxy.ProxyConfig( - authenticator = authenticator, - cacert = cert_path), - 0) - self.server = server - Thread.__init__(self) - controller.Master.__init__(self, server) - - def http_address(self): - return 'http://scrapy:scrapy@%s:%d' % self.server.socket.getsockname() - - -def _wrong_credentials(proxy_url): - bad_auth_proxy = list(urlsplit(proxy_url)) - bad_auth_proxy[1] = bad_auth_proxy[1].replace('scrapy:scrapy@', 'wrong:wronger@') - return urlunsplit(bad_auth_proxy) - -class ProxyConnectTestCase(TestCase): - - def setUp(self): - self.mockserver = MockServer() - self.mockserver.__enter__() - self._oldenv = os.environ.copy() - - self._proxy = HTTPSProxy() - self._proxy.start() - - # Wait for the proxy to start. - time.sleep(1.0) - os.environ['https_proxy'] = self._proxy.http_address() - os.environ['http_proxy'] = self._proxy.http_address() - - def tearDown(self): - self.mockserver.__exit__(None, None, None) - self._proxy.shutdown() - os.environ = self._oldenv - - @defer.inlineCallbacks - def test_https_connect_tunnel(self): - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(200, l) - - @defer.inlineCallbacks - def test_https_noconnect(self): - proxy = os.environ['https_proxy'] - os.environ['https_proxy'] = proxy + '?noconnect' - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(200, l) - - @defer.inlineCallbacks - def test_https_connect_tunnel_error(self): - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl("https://localhost:99999/status?n=200") - self._assert_got_tunnel_error(l) - - @defer.inlineCallbacks - def test_https_tunnel_auth_error(self): - os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - # The proxy returns a 407 error code but it does not reach the client; - # he just sees a TunnelError. - self._assert_got_tunnel_error(l) - - @defer.inlineCallbacks - def test_https_tunnel_without_leak_proxy_authorization_header(self): - request = Request(self.mockserver.url("/echo", is_secure=True)) - crawler = get_crawler(SingleRequestSpider) - with LogCapture() as l: - yield crawler.crawl(seed=request) - self._assert_got_response_code(200, l) - echo = json.loads(crawler.spider.meta['responses'][0].body) - self.assertTrue('Proxy-Authorization' not in echo['headers']) - - @defer.inlineCallbacks - def test_https_noconnect_auth_error(self): - os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) + '?noconnect' - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(407, l) - - def _assert_got_response_code(self, code, log): - print(log) - self.assertEqual(str(log).count('Crawled (%d)' % code), 1) - - def _assert_got_tunnel_error(self, log): - print(log) - self.assertIn('TunnelError', str(log)) diff --git a/tests/test_utils_python.py b/tests/test_utils_python.py index 3e1148354..6cb32cbdd 100644 --- a/tests/test_utils_python.py +++ b/tests/test_utils_python.py @@ -163,33 +163,6 @@ class UtilsPythonTestCase(unittest.TestCase): gc.collect() self.assertFalse(len(wk._weakdict)) - @unittest.skipUnless(six.PY2, "deprecated function") - def test_stringify_dict(self): - d = {'a': 123, u'b': b'c', u'd': u'e', object(): u'e'} - d2 = stringify_dict(d, keys_only=False) - self.assertEqual(d, d2) - self.assertIsNot(d, d2) # shouldn't modify in place - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.keys())) - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.values())) - - @unittest.skipUnless(six.PY2, "deprecated function") - def test_stringify_dict_tuples(self): - tuples = [('a', 123), (u'b', 'c'), (u'd', u'e'), (object(), u'e')] - d = dict(tuples) - d2 = stringify_dict(tuples, keys_only=False) - self.assertEqual(d, d2) - self.assertIsNot(d, d2) # shouldn't modify in place - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.keys()), d2.keys()) - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.values())) - - @unittest.skipUnless(six.PY2, "deprecated function") - def test_stringify_dict_keys_only(self): - d = {'a': 123, u'b': 'c', u'd': u'e', object(): u'e'} - d2 = stringify_dict(d) - self.assertEqual(d, d2) - self.assertIsNot(d, d2) # shouldn't modify in place - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.keys())) - def test_get_func_args(self): def f1(a, b, c): pass diff --git a/tests/test_webclient.py b/tests/test_webclient.py index a81946490..7b015ff8d 100644 --- a/tests/test_webclient.py +++ b/tests/test_webclient.py @@ -78,26 +78,6 @@ class ParseUrlTestCase(unittest.TestCase): to_bytes(x) if not isinstance(x, int) else x for x in test) self.assertEqual(client._parse(url), test, url) - def test_externalUnicodeInterference(self): - """ - L{client._parse} should return C{str} for the scheme, host, and path - elements of its return tuple, even when passed an URL which has - previously been passed to L{urlparse} as a C{unicode} string. - """ - if not six.PY2: - raise unittest.SkipTest( - "Applies only to Py2, as urls can be ONLY unicode on Py3") - badInput = u'http://example.com/path' - goodInput = badInput.encode('ascii') - self._parse(badInput) # cache badInput in urlparse_cached - scheme, netloc, host, port, path = self._parse(goodInput) - self.assertTrue(isinstance(scheme, str)) - self.assertTrue(isinstance(netloc, str)) - self.assertTrue(isinstance(host, str)) - self.assertTrue(isinstance(path, str)) - self.assertTrue(isinstance(port, int)) - - class ScrapyHTTPPageGetterTests(unittest.TestCase): From b0d6f4917d782f2396a702e90c36ffbc8bb42e7e Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 3 Sep 2019 15:17:03 +0500 Subject: [PATCH 171/387] Restore tests/test_proxy_connect.py and update it to modern mitmproxy. --- tests/requirements-py3.txt | 1 + tests/test_proxy_connect.py | 134 ++++++++++++++++++++++++++++++++++++ 2 files changed, 135 insertions(+) create mode 100644 tests/test_proxy_connect.py diff --git a/tests/requirements-py3.txt b/tests/requirements-py3.txt index dd5b23cc3..f27e45a54 100644 --- a/tests/requirements-py3.txt +++ b/tests/requirements-py3.txt @@ -1,5 +1,6 @@ # Tests requirements jmespath +mitmproxy pytest pytest-cov pytest-twisted diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py new file mode 100644 index 000000000..8142d9a41 --- /dev/null +++ b/tests/test_proxy_connect.py @@ -0,0 +1,134 @@ +import json +import os +import re +from subprocess import Popen, PIPE +import sys +import time + +from six.moves.urllib.parse import urlsplit, urlunsplit +from testfixtures import LogCapture + +from twisted.internet import defer +from twisted.trial.unittest import TestCase + +from scrapy.utils.test import get_crawler +from scrapy.http import Request +from tests.spiders import SimpleSpider, SingleRequestSpider +from tests.mockserver import MockServer + + +class MitmProxy: + auth_user = 'scrapy' + auth_pass = 'scrapy' + + def start(self): + from scrapy.utils.test import get_testenv + script = """ +import sys +from mitmproxy.tools.main import mitmdump +sys.argv[0] = "mitmdump" +sys.exit(mitmdump()) + """ + cert_path = os.path.join(os.path.abspath(os.path.dirname(__file__)), + 'keys', 'mitmproxy-ca.pem') + self.proc = Popen([sys.executable, + '-c', script, + '--listen-host', '127.0.0.1', + '--listen-port', '0', + '--proxyauth', '%s:%s' % (self.auth_user, self.auth_pass), + '--certs', cert_path, + '--ssl-insecure', + ], + stdout=PIPE, env=get_testenv()) + line = self.proc.stdout.readline().decode('utf-8') + host_port = re.search(r'listening at http://([^:]+:\d+)', line).group(1) + address = 'http://%s:%s@%s' % (self.auth_user, self.auth_pass, host_port) + return address + + def stop(self): + self.proc.kill() + self.proc.wait() + time.sleep(0.2) + + +def _wrong_credentials(proxy_url): + bad_auth_proxy = list(urlsplit(proxy_url)) + bad_auth_proxy[1] = bad_auth_proxy[1].replace('scrapy:scrapy@', 'wrong:wronger@') + return urlunsplit(bad_auth_proxy) + + +class ProxyConnectTestCase(TestCase): + + def setUp(self): + self.mockserver = MockServer() + self.mockserver.__enter__() + self._oldenv = os.environ.copy() + + self._proxy = MitmProxy() + proxy_url = self._proxy.start() + os.environ['https_proxy'] = proxy_url + os.environ['http_proxy'] = proxy_url + + def tearDown(self): + self.mockserver.__exit__(None, None, None) + self._proxy.stop() + os.environ = self._oldenv + + @defer.inlineCallbacks + def test_https_connect_tunnel(self): + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) + self._assert_got_response_code(200, l) + + @defer.inlineCallbacks + def test_https_noconnect(self): + proxy = os.environ['https_proxy'] + os.environ['https_proxy'] = proxy + '?noconnect' + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) + self._assert_got_response_code(200, l) + + @defer.inlineCallbacks + def test_https_connect_tunnel_error(self): + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl("https://localhost:99999/status?n=200") + self._assert_got_tunnel_error(l) + + @defer.inlineCallbacks + def test_https_tunnel_auth_error(self): + os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) + # The proxy returns a 407 error code but it does not reach the client; + # he just sees a TunnelError. + self._assert_got_tunnel_error(l) + + @defer.inlineCallbacks + def test_https_tunnel_without_leak_proxy_authorization_header(self): + request = Request(self.mockserver.url("/echo", is_secure=True)) + crawler = get_crawler(SingleRequestSpider) + with LogCapture() as l: + yield crawler.crawl(seed=request) + self._assert_got_response_code(200, l) + echo = json.loads(crawler.spider.meta['responses'][0].body) + self.assertTrue('Proxy-Authorization' not in echo['headers']) + + @defer.inlineCallbacks + def test_https_noconnect_auth_error(self): + os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) + '?noconnect' + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) + self._assert_got_response_code(407, l) + + def _assert_got_response_code(self, code, log): + print(log) + self.assertEqual(str(log).count('Crawled (%d)' % code), 1) + + def _assert_got_tunnel_error(self, log): + print(log) + self.assertIn('TunnelError', str(log)) From 439e37fc7b2c192b4591d29dde74c9f589763cc8 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 3 Sep 2019 15:23:24 +0500 Subject: [PATCH 172/387] Mark failing proxy tests. --- tests/test_proxy_connect.py | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index 8142d9a41..5e9470e39 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -5,6 +5,7 @@ from subprocess import Popen, PIPE import sys import time +import pytest from six.moves.urllib.parse import urlsplit, urlunsplit from testfixtures import LogCapture @@ -81,6 +82,7 @@ class ProxyConnectTestCase(TestCase): yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) self._assert_got_response_code(200, l) + @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') @defer.inlineCallbacks def test_https_noconnect(self): proxy = os.environ['https_proxy'] @@ -90,6 +92,7 @@ class ProxyConnectTestCase(TestCase): yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) self._assert_got_response_code(200, l) + @pytest.mark.xfail(reason='Python 3 fails this earlier') @defer.inlineCallbacks def test_https_connect_tunnel_error(self): crawler = get_crawler(SimpleSpider) @@ -117,6 +120,7 @@ class ProxyConnectTestCase(TestCase): echo = json.loads(crawler.spider.meta['responses'][0].body) self.assertTrue('Proxy-Authorization' not in echo['headers']) + @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') @defer.inlineCallbacks def test_https_noconnect_auth_error(self): os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) + '?noconnect' From 186f9d88acd9043cf5d8f21358c758c42e433105 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 3 Sep 2019 15:24:06 +0500 Subject: [PATCH 173/387] Fix the skip message for test_download_gzip_response. --- tests/test_downloader_handlers.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 4d3c4d4aa..9f74577a5 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -634,7 +634,7 @@ class Http11MockServerTestCase(unittest.TestCase): self.assertIsInstance(failure.value, defer.CancelledError) # See issue https://twistedmatrix.com/trac/ticket/8175 - raise unittest.SkipTest("xpayload only enabled for PY2") + raise unittest.SkipTest("xpayload fails on PY3") request.headers.setdefault(b'Accept-Encoding', b'gzip,deflate') request = request.replace(url=self.mockserver.url('/xpayload')) yield crawler.crawl(seed=request) From bbd9f4be90e734df5c290aeb399e8e512f834224 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Mon, 22 Jul 2019 22:27:29 +0500 Subject: [PATCH 174/387] Remove six.PY2 and six.PY3 conditionals. --- docs/topics/downloader-middleware.rst | 6 +++--- scrapy/_monkeypatches.py | 10 ---------- scrapy/commands/fetch.py | 5 ++--- scrapy/crawler.py | 11 ----------- scrapy/exporters.py | 2 +- scrapy/extensions/feedexport.py | 3 +-- scrapy/http/request/form.py | 3 +-- scrapy/http/response/text.py | 3 --- scrapy/item.py | 8 +------- scrapy/link.py | 15 ++------------- scrapy/mail.py | 9 ++------- scrapy/settings/__init__.py | 8 +------- scrapy/settings/default_settings.py | 4 +--- scrapy/utils/boto.py | 10 +--------- scrapy/utils/conf.py | 5 +---- scrapy/utils/datatypes.py | 20 +++++++------------- scrapy/utils/gz.py | 13 +++---------- scrapy/utils/iterators.py | 6 +----- scrapy/utils/python.py | 10 ++-------- tests/__init__.py | 15 --------------- tests/test_http_request.py | 11 +++-------- tests/test_http_response.py | 3 +-- tests/test_item.py | 8 ++------ tests/test_link.py | 13 ++----------- tests/test_middleware.py | 21 ++++++--------------- tests/test_request_cb_kwargs.py | 8 +------- tests/test_settings/__init__.py | 8 +------- tests/test_utils_datatypes.py | 7 +------ tests/test_utils_python.py | 7 +++---- 29 files changed, 50 insertions(+), 202 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 8048e1c86..366b95510 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -474,7 +474,7 @@ DBM storage backend A DBM_ storage backend is also available for the HTTP cache middleware. - By default, it uses the anydbm_ module, but you can change it with the + By default, it uses the dbm_ module, but you can change it with the :setting:`HTTPCACHE_DBM_MODULE` setting. .. _httpcache-storage-custom: @@ -626,7 +626,7 @@ HTTPCACHE_DBM_MODULE .. versionadded:: 0.13 -Default: ``'anydbm'`` +Default: ``'dbm'`` The database module to use in the :ref:`DBM storage backend <httpcache-storage-dbm>`. This setting is specific to the DBM backend. @@ -1202,4 +1202,4 @@ The default encoding for proxy authentication on :class:`HttpProxyMiddleware`. .. _DBM: https://en.wikipedia.org/wiki/Dbm -.. _anydbm: https://docs.python.org/2/library/anydbm.html +.. _dbm: https://docs.python.org/3/library/dbm.html diff --git a/scrapy/_monkeypatches.py b/scrapy/_monkeypatches.py index b68099cad..1f8067b35 100644 --- a/scrapy/_monkeypatches.py +++ b/scrapy/_monkeypatches.py @@ -1,16 +1,6 @@ -import six from six.moves import copyreg -if six.PY2: - from urlparse import urlparse - - # workaround for https://bugs.python.org/issue9374 - Python < 2.7.4 - if urlparse('s3://bucket/key?key=value').query != 'key=value': - from urlparse import uses_query - uses_query.append('s3') - - # Undo what Twisted's perspective broker adds to pickle register # to prevent bugs like Twisted#7989 while serializing requests import twisted.persisted.styles # NOQA diff --git a/scrapy/commands/fetch.py b/scrapy/commands/fetch.py index 7d4840529..d45133e0e 100644 --- a/scrapy/commands/fetch.py +++ b/scrapy/commands/fetch.py @@ -1,5 +1,5 @@ from __future__ import print_function -import sys, six +import sys from w3lib.url import is_url from scrapy.commands import ScrapyCommand @@ -45,8 +45,7 @@ class Command(ScrapyCommand): self._print_bytes(response.body) def _print_bytes(self, bytes_): - bytes_writer = sys.stdout if six.PY2 else sys.stdout.buffer - bytes_writer.write(bytes_ + b'\n') + sys.stdout.buffer.write(bytes_ + b'\n') def run(self, args, opts): if len(args) != 1 or not is_url(args[0]): diff --git a/scrapy/crawler.py b/scrapy/crawler.py index ded3c082b..19b998e0d 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -88,20 +88,9 @@ class Crawler(object): yield self.engine.open_spider(self.spider, start_requests) yield defer.maybeDeferred(self.engine.start) except Exception: - # In Python 2 reraising an exception after yield discards - # the original traceback (see https://bugs.python.org/issue7563), - # so sys.exc_info() workaround is used. - # This workaround also works in Python 3, but it is not needed, - # and it is slower, so in Python 3 we use native `raise`. - if six.PY2: - exc_info = sys.exc_info() - self.crawling = False if self.engine is not None: yield self.engine.close() - - if six.PY2: - six.reraise(*exc_info) raise def _create_spider(self, *args, **kwargs): diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 6fc87ed18..5d1f1ad8f 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -216,7 +216,7 @@ class CsvItemExporter(BaseItemExporter): write_through=True, encoding=self.encoding, newline='' # Windows needs this https://github.com/scrapy/scrapy/issues/3034 - ) if six.PY3 else file + ) self.csv_writer = csv.writer(self.stream, **kwargs) self._headers_not_written = True self._join_multivalued = join_multivalued diff --git a/scrapy/extensions/feedexport.py b/scrapy/extensions/feedexport.py index 6fb6397b1..07ffd3476 100644 --- a/scrapy/extensions/feedexport.py +++ b/scrapy/extensions/feedexport.py @@ -10,7 +10,6 @@ import logging import posixpath from tempfile import NamedTemporaryFile from datetime import datetime -import six from six.moves.urllib.parse import urlparse, unquote from ftplib import FTP @@ -65,7 +64,7 @@ class StdoutFeedStorage(object): def __init__(self, uri, _stdout=None): if not _stdout: - _stdout = sys.stdout if six.PY2 else sys.stdout.buffer + _stdout = sys.stdout.buffer self._stdout = _stdout def open(self, spider): diff --git a/scrapy/http/request/form.py b/scrapy/http/request/form.py index 3ce8fc48e..b6feede07 100644 --- a/scrapy/http/request/form.py +++ b/scrapy/http/request/form.py @@ -104,8 +104,7 @@ def _get_form(response, formname, formid, formnumber, formxpath): el = el.getparent() if el is None: break - encoded = formxpath if six.PY3 else formxpath.encode('unicode_escape') - raise ValueError('No <form> element found with %s' % encoded) + raise ValueError('No <form> element found with %s' % formxpath) # If we get here, it means that either formname was None # or invalid diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 339913d4e..a8010877c 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -32,9 +32,6 @@ class TextResponse(Response): def _set_url(self, url): if isinstance(url, six.text_type): - if six.PY2 and self.encoding is None: - raise TypeError("Cannot convert unicode url - %s " - "has no encoding" % type(self).__name__) self._url = to_native_str(url, self.encoding) else: super(TextResponse, self)._set_url(url) diff --git a/scrapy/item.py b/scrapy/item.py index 73b8f54b0..32f9b2ebb 100644 --- a/scrapy/item.py +++ b/scrapy/item.py @@ -4,8 +4,8 @@ Scrapy Item See documentation in docs/topics/item.rst """ -import collections from abc import ABCMeta +from collections.abc import MutableMapping from copy import deepcopy from pprint import pformat from warnings import warn @@ -16,12 +16,6 @@ from scrapy.utils.deprecate import ScrapyDeprecationWarning from scrapy.utils.trackref import object_ref -if six.PY2: - MutableMapping = collections.MutableMapping -else: - MutableMapping = collections.abc.MutableMapping - - class BaseItem(object_ref): """Base class for all scraped items. diff --git a/scrapy/link.py b/scrapy/link.py index 2c8301680..a175b8afd 100644 --- a/scrapy/link.py +++ b/scrapy/link.py @@ -4,12 +4,6 @@ This module defines the Link object used in Link extractors. For actual link extractors implementation see scrapy.linkextractors, or its documentation in: docs/topics/link-extractors.rst """ -import warnings -import six - -from scrapy.utils.python import to_bytes - - class Link(object): """Link objects represent an extracted link by the LinkExtractor.""" @@ -17,13 +11,8 @@ class Link(object): def __init__(self, url, text='', fragment='', nofollow=False): if not isinstance(url, str): - if six.PY2: - warnings.warn("Link urls must be str objects. " - "Assuming utf-8 encoding (which could be wrong)") - url = to_bytes(url, encoding='utf8') - else: - got = url.__class__.__name__ - raise TypeError("Link urls must be str objects, got %s" % got) + got = url.__class__.__name__ + raise TypeError("Link urls must be str objects, got %s" % got) self.url = url self.text = text self.fragment = fragment diff --git a/scrapy/mail.py b/scrapy/mail.py index 5b944e1c4..746468e25 100644 --- a/scrapy/mail.py +++ b/scrapy/mail.py @@ -9,18 +9,13 @@ try: from cStringIO import StringIO as BytesIO except ImportError: from io import BytesIO -import six from email.utils import COMMASPACE, formatdate from six.moves.email_mime_multipart import MIMEMultipart from six.moves.email_mime_text import MIMEText from six.moves.email_mime_base import MIMEBase -if six.PY2: - from email.MIMENonMultipart import MIMENonMultipart - from email import Encoders -else: - from email.mime.nonmultipart import MIMENonMultipart - from email import encoders as Encoders +from email.mime.nonmultipart import MIMENonMultipart +from email import encoders as Encoders from twisted.internet import defer, reactor, ssl diff --git a/scrapy/settings/__init__.py b/scrapy/settings/__init__.py index f28c7940d..c871e86e0 100644 --- a/scrapy/settings/__init__.py +++ b/scrapy/settings/__init__.py @@ -1,19 +1,13 @@ import six import json import copy -import collections +from collections.abc import MutableMapping from importlib import import_module from pprint import pformat from scrapy.settings import default_settings -if six.PY2: - MutableMapping = collections.MutableMapping -else: - MutableMapping = collections.abc.MutableMapping - - SETTINGS_PRIORITIES = { 'default': 0, 'command': 10, diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 9c22999cb..5c9678c01 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -17,8 +17,6 @@ import sys from importlib import import_module from os.path import join, abspath, dirname -import six - AJAXCRAWL_ENABLED = False AUTOTHROTTLE_ENABLED = False @@ -179,7 +177,7 @@ HTTPCACHE_ALWAYS_STORE = False HTTPCACHE_IGNORE_HTTP_CODES = [] HTTPCACHE_IGNORE_SCHEMES = ['file'] HTTPCACHE_IGNORE_RESPONSE_CACHE_CONTROLS = [] -HTTPCACHE_DBM_MODULE = 'anydbm' if six.PY2 else 'dbm' +HTTPCACHE_DBM_MODULE = 'dbm' HTTPCACHE_POLICY = 'scrapy.extensions.httpcache.DummyPolicy' HTTPCACHE_GZIP = False diff --git a/scrapy/utils/boto.py b/scrapy/utils/boto.py index 421ab2f7e..c8fc911bb 100644 --- a/scrapy/utils/boto.py +++ b/scrapy/utils/boto.py @@ -1,7 +1,6 @@ """Boto/botocore helpers""" from __future__ import absolute_import -import six from scrapy.exceptions import NotConfigured @@ -11,11 +10,4 @@ def is_botocore(): import botocore return True except ImportError: - if six.PY2: - try: - import boto - return False - except ImportError: - raise NotConfigured('missing botocore or boto library') - else: - raise NotConfigured('missing botocore library') + raise NotConfigured('missing botocore library') diff --git a/scrapy/utils/conf.py b/scrapy/utils/conf.py index fb7ca3310..561bb72fc 100644 --- a/scrapy/utils/conf.py +++ b/scrapy/utils/conf.py @@ -1,13 +1,10 @@ +from configparser import ConfigParser import os import sys import numbers from operator import itemgetter import six -if six.PY2: - from ConfigParser import SafeConfigParser as ConfigParser -else: - from configparser import ConfigParser from scrapy.settings import BaseSettings from scrapy.utils.deprecate import update_classpath diff --git a/scrapy/utils/datatypes.py b/scrapy/utils/datatypes.py index df2b99c28..6e9de47f3 100644 --- a/scrapy/utils/datatypes.py +++ b/scrapy/utils/datatypes.py @@ -7,6 +7,7 @@ This module must not depend on any module outside the Standard Library. import copy import collections +from collections.abc import Mapping import warnings import six @@ -14,12 +15,6 @@ import six from scrapy.exceptions import ScrapyDeprecationWarning -if six.PY2: - Mapping = collections.Mapping -else: - Mapping = collections.abc.Mapping - - class MultiValueDictKeyError(KeyError): def __init__(self, *args, **kwargs): warnings.warn( @@ -252,13 +247,12 @@ class MergeDict(object): first occurrence will be used. """ def __init__(self, *dicts): - if not six.PY2: - warnings.warn( - "scrapy.utils.datatypes.MergeDict is deprecated in favor " - "of collections.ChainMap (introduced in Python 3.3)", - category=ScrapyDeprecationWarning, - stacklevel=2, - ) + warnings.warn( + "scrapy.utils.datatypes.MergeDict is deprecated in favor " + "of collections.ChainMap (introduced in Python 3.3)", + category=ScrapyDeprecationWarning, + stacklevel=2, + ) self.dicts = dicts def __getitem__(self, key): diff --git a/scrapy/utils/gz.py b/scrapy/utils/gz.py index b3fb16b1e..9984492f0 100644 --- a/scrapy/utils/gz.py +++ b/scrapy/utils/gz.py @@ -6,7 +6,6 @@ except ImportError: from io import BytesIO from gzip import GzipFile -import six import re from scrapy.utils.decorators import deprecated @@ -17,14 +16,8 @@ from scrapy.utils.decorators import deprecated # (regression or bug-fix compared to Python 3.4) # - read1(), which fetches data before raising EOFError on next call # works here but is only available from Python>=3.3 -# - scrapy does not support Python 3.2 -# - Python 2.7 GzipFile works fine with standard read() + extrabuf -if six.PY2: - def read1(gzf, size=-1): - return gzf.read(size) -else: - def read1(gzf, size=-1): - return gzf.read1(size) +def read1(gzf, size=-1): + return gzf.read1(size) def gunzip(data): @@ -37,7 +30,7 @@ def gunzip(data): chunk = b'.' while chunk: try: - chunk = read1(f, 8196) + chunk = f.read1(8196) output_list.append(chunk) except (IOError, EOFError, struct.error): # complete only if there is some data, otherwise re-raise diff --git a/scrapy/utils/iterators.py b/scrapy/utils/iterators.py index a12e14005..dbc1e0d20 100644 --- a/scrapy/utils/iterators.py +++ b/scrapy/utils/iterators.py @@ -102,11 +102,7 @@ def csviter(obj, delimiter=None, headers=None, encoding=None, quotechar=None): def row_to_unicode(row_): return [to_unicode(field, encoding) for field in row_] - # Python 3 csv reader input object needs to return strings - if six.PY3: - lines = StringIO(_body_or_str(obj, unicode=True)) - else: - lines = BytesIO(_body_or_str(obj, unicode=False)) + lines = StringIO(_body_or_str(obj, unicode=True)) kwargs = {} if delimiter: kwargs["delimiter"] = delimiter diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index c6140f885..5009aab81 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -113,10 +113,7 @@ def to_bytes(text, encoding=None, errors='strict'): def to_native_str(text, encoding=None, errors='strict'): """ Return str representation of ``text`` (bytes in Python 2.x and unicode in Python 3.x). """ - if six.PY2: - return to_bytes(text, encoding, errors) - else: - return to_unicode(text, encoding, errors) + return to_unicode(text, encoding, errors) def re_rsearch(pattern, text, chunk_size=1024): @@ -189,7 +186,7 @@ def _getargspec_py23(func): """_getargspec_py23(function) -> named tuple ArgSpec(args, varargs, keywords, defaults) - Identical to inspect.getargspec() in python2, but uses + Was identical to inspect.getargspec() in python2, but uses inspect.getfullargspec() for python3 behind the scenes to avoid DeprecationWarning. @@ -199,9 +196,6 @@ def _getargspec_py23(func): >>> _getargspec_py23(f) ArgSpec(args=['a', 'b'], varargs='ar', keywords='kw', defaults=(2,)) """ - if six.PY2: - return inspect.getargspec(func) - return inspect.ArgSpec(*inspect.getfullargspec(func)[:4]) diff --git a/tests/__init__.py b/tests/__init__.py index 9c9e35c35..a54367f8c 100644 --- a/tests/__init__.py +++ b/tests/__init__.py @@ -35,18 +35,3 @@ def get_testdata(*paths): path = os.path.join(tests_datadir, *paths) with open(path, 'rb') as f: return f.read() - - -# FIXME: delete after dropping py2 support -# Monkey patch the unittest module to prevent the -# DeprecationWarning about assertRaisesRegexp -> assertRaisesRegex -import six -if six.PY2: - import unittest - import twisted.trial.unittest - if not getattr(unittest.TestCase, 'assertRegex', None): - unittest.TestCase.assertRegex = unittest.TestCase.assertRegexpMatches - if not getattr(unittest.TestCase, 'assertRaisesRegex', None): - unittest.TestCase.assertRaisesRegex = unittest.TestCase.assertRaisesRegexp - if not getattr(twisted.trial.unittest.TestCase, 'assertRaisesRegex', None): - twisted.trial.unittest.TestCase.assertRaisesRegex = twisted.trial.unittest.TestCase.assertRaisesRegexp diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 16d7a1cb8..828902b99 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -3,13 +3,12 @@ import cgi import unittest import re import json +from urllib.parse import unquote_to_bytes import warnings import six from six.moves import xmlrpc_client as xmlrpclib from six.moves.urllib.parse import urlparse, parse_qs, unquote -if six.PY3: - from urllib.parse import unquote_to_bytes from scrapy.http import Request, FormRequest, XmlRpcRequest, JsonRequest, Headers, HtmlResponse from scrapy.utils.python import to_bytes, to_native_str @@ -1064,8 +1063,7 @@ class FormRequestTest(RequestTest): self.assertEqual(fs, {}) xpath = u"//form[@name='\u03b1']" - encoded = xpath if six.PY3 else xpath.encode('unicode_escape') - self.assertRaisesRegex(ValueError, re.escape(encoded), + self.assertRaisesRegex(ValueError, re.escape(xpath), self.request_class.from_response, response, formxpath=xpath) @@ -1208,10 +1206,7 @@ def _qs(req, encoding='utf-8', to_unicode=False): qs = req.body else: qs = req.url.partition('?')[2] - if six.PY2: - uqs = unquote(to_native_str(qs, encoding)) - elif six.PY3: - uqs = unquote_to_bytes(qs) + uqs = unquote_to_bytes(qs) if to_unicode: uqs = uqs.decode(encoding) return parse_qs(uqs, True) diff --git a/tests/test_http_response.py b/tests/test_http_response.py index cd5c3486e..d6e77d6b8 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -21,8 +21,7 @@ class BaseResponseTest(unittest.TestCase): # Response requires url in the consturctor self.assertRaises(Exception, self.response_class) self.assertTrue(isinstance(self.response_class('http://example.com/'), self.response_class)) - if not six.PY2: - self.assertRaises(TypeError, self.response_class, b"http://example.com") + self.assertRaises(TypeError, self.response_class, b"http://example.com") # body can be str or None self.assertTrue(isinstance(self.response_class('http://example.com/', body=b''), self.response_class)) self.assertTrue(isinstance(self.response_class('http://example.com/', body=b'body'), self.response_class)) diff --git a/tests/test_item.py b/tests/test_item.py index 947566686..0ad278701 100644 --- a/tests/test_item.py +++ b/tests/test_item.py @@ -62,12 +62,8 @@ class ItemTest(unittest.TestCase): i['number'] = 123 itemrepr = repr(i) - if six.PY2: - self.assertEqual(itemrepr, - "{'name': u'John Doe', 'number': 123}") - else: - self.assertEqual(itemrepr, - "{'name': 'John Doe', 'number': 123}") + self.assertEqual(itemrepr, + "{'name': 'John Doe', 'number': 123}") i2 = eval(itemrepr) self.assertEqual(i2['name'], 'John Doe') diff --git a/tests/test_link.py b/tests/test_link.py index 955430b37..5e2ce5eeb 100644 --- a/tests/test_link.py +++ b/tests/test_link.py @@ -1,6 +1,4 @@ import unittest -import warnings -import six from scrapy.link import Link @@ -46,12 +44,5 @@ class LinkTest(unittest.TestCase): self._assert_same_links(l1, l2) def test_non_str_url_py2(self): - if six.PY2: - with warnings.catch_warnings(record=True) as w: - link = Link(u"http://www.example.com/\xa3") - self.assertIsInstance(link.url, str) - self.assertEqual(link.url, b'http://www.example.com/\xc2\xa3') - assert len(w) == 1, "warning not issued" - else: - with self.assertRaises(TypeError): - Link(b"http://www.example.com/\xc2\xa3") + with self.assertRaises(TypeError): + Link(b"http://www.example.com/\xc2\xa3") diff --git a/tests/test_middleware.py b/tests/test_middleware.py index aea0be825..af9b43d61 100644 --- a/tests/test_middleware.py +++ b/tests/test_middleware.py @@ -3,7 +3,6 @@ from twisted.trial import unittest from scrapy.settings import Settings from scrapy.exceptions import NotConfigured from scrapy.middleware import MiddlewareManager -import six class M1(object): @@ -66,20 +65,12 @@ class MiddlewareManagerTest(unittest.TestCase): def test_methods(self): mwman = TestMiddlewareManager(M1(), M2(), M3()) - if six.PY2: - self.assertEqual([x.im_class for x in mwman.methods['open_spider']], - [M1, M2]) - self.assertEqual([x.im_class for x in mwman.methods['close_spider']], - [M2, M1]) - self.assertEqual([x.im_class for x in mwman.methods['process']], - [M1, M3]) - else: - self.assertEqual([x.__self__.__class__ for x in mwman.methods['open_spider']], - [M1, M2]) - self.assertEqual([x.__self__.__class__ for x in mwman.methods['close_spider']], - [M2, M1]) - self.assertEqual([x.__self__.__class__ for x in mwman.methods['process']], - [M1, M3]) + self.assertEqual([x.__self__.__class__ for x in mwman.methods['open_spider']], + [M1, M2]) + self.assertEqual([x.__self__.__class__ for x in mwman.methods['close_spider']], + [M2, M1]) + self.assertEqual([x.__self__.__class__ for x in mwman.methods['process']], + [M1, M3]) def test_enabled(self): m1, m2, m3 = M1(), M2(), M3() diff --git a/tests/test_request_cb_kwargs.py b/tests/test_request_cb_kwargs.py index c9943faa8..a5cdc0de0 100644 --- a/tests/test_request_cb_kwargs.py +++ b/tests/test_request_cb_kwargs.py @@ -1,7 +1,6 @@ from testfixtures import LogCapture from twisted.internet import defer from twisted.trial.unittest import TestCase -import six from scrapy.http import Request from scrapy.crawler import CrawlerRunner @@ -161,9 +160,4 @@ class CallbackKeywordArgumentsTestCase(TestCase): self.assertEqual(exceptions['takes_less'].exc_info[0], TypeError) self.assertEqual(str(exceptions['takes_less'].exc_info[1]), "parse_takes_less() got an unexpected keyword argument 'number'") self.assertEqual(exceptions['takes_more'].exc_info[0], TypeError) - # py2 and py3 messages are different - exc_message = str(exceptions['takes_more'].exc_info[1]) - if six.PY2: - self.assertEqual(exc_message, "parse_takes_more() takes exactly 5 arguments (4 given)") - elif six.PY3: - self.assertEqual(exc_message, "parse_takes_more() missing 1 required positional argument: 'other'") + self.assertEqual(str(exceptions['takes_more'].exc_info[1]), "parse_takes_more() missing 1 required positional argument: 'other'") diff --git a/tests/test_settings/__init__.py b/tests/test_settings/__init__.py index 1dbacbea3..08286ff02 100644 --- a/tests/test_settings/__init__.py +++ b/tests/test_settings/__init__.py @@ -60,9 +60,6 @@ class SettingsAttributeTest(unittest.TestCase): class BaseSettingsTest(unittest.TestCase): - if six.PY3: - assertItemsEqual = unittest.TestCase.assertCountEqual - def setUp(self): self.settings = BaseSettings() @@ -152,7 +149,7 @@ class BaseSettingsTest(unittest.TestCase): self.settings.setmodule( 'tests.test_settings.default_settings', 10) - self.assertItemsEqual(six.iterkeys(self.settings.attributes), + self.assertCountEqual(six.iterkeys(self.settings.attributes), six.iterkeys(ctrl_attributes)) for key in six.iterkeys(ctrl_attributes): @@ -343,9 +340,6 @@ class BaseSettingsTest(unittest.TestCase): class SettingsTest(unittest.TestCase): - if six.PY3: - assertItemsEqual = unittest.TestCase.assertCountEqual - def setUp(self): self.settings = Settings() diff --git a/tests/test_utils_datatypes.py b/tests/test_utils_datatypes.py index 535095b8d..47877f555 100644 --- a/tests/test_utils_datatypes.py +++ b/tests/test_utils_datatypes.py @@ -1,12 +1,7 @@ +from collections.abc import Mapping, MutableMapping import copy import unittest -import six -if six.PY2: - from collections import Mapping, MutableMapping -else: - from collections.abc import Mapping, MutableMapping - from scrapy.utils.datatypes import CaselessDict, SequenceExclude diff --git a/tests/test_utils_python.py b/tests/test_utils_python.py index 3e1148354..096aa50b7 100644 --- a/tests/test_utils_python.py +++ b/tests/test_utils_python.py @@ -231,12 +231,11 @@ class UtilsPythonTestCase(unittest.TestCase): self.assertEqual(get_func_args(" ".join), []) self.assertEqual(get_func_args(operator.itemgetter(2)), []) else: - stripself = not six.PY2 # PyPy3 exposes them as methods self.assertEqual( - get_func_args(six.text_type.split, stripself), ['sep', 'maxsplit']) - self.assertEqual(get_func_args(" ".join, stripself), ['list']) + get_func_args(six.text_type.split, True), ['sep', 'maxsplit']) + self.assertEqual(get_func_args(" ".join, True), ['list']) self.assertEqual( - get_func_args(operator.itemgetter(2), stripself), ['obj']) + get_func_args(operator.itemgetter(2), True), ['obj']) def test_without_none_values(self): From de7789e52df5d8dc8595570d36a60ae92b4fcb50 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 20 Aug 2019 20:56:26 +0500 Subject: [PATCH 175/387] Remove unneeded and unused code from XmlItemExporter. --- scrapy/exporters.py | 22 ++++------------------ 1 file changed, 4 insertions(+), 18 deletions(-) diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 5d1f1ad8f..40567f53b 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -143,11 +143,11 @@ class XmlItemExporter(BaseItemExporter): def _beautify_newline(self, new_item=False): if self.indent is not None and (self.indent > 0 or new_item): - self._xg_characters('\n') + self.xg.characters('\n') def _beautify_indent(self, depth=1): if self.indent: - self._xg_characters(' ' * self.indent * depth) + self.xg.characters(' ' * self.indent * depth) def start_exporting(self): self.xg.startDocument() @@ -182,26 +182,12 @@ class XmlItemExporter(BaseItemExporter): self._export_xml_field('value', value, depth=depth+1) self._beautify_indent(depth=depth) elif isinstance(serialized_value, six.text_type): - self._xg_characters(serialized_value) + self.xg.characters(serialized_value) else: - self._xg_characters(str(serialized_value)) + self.xg.characters(str(serialized_value)) self.xg.endElement(name) self._beautify_newline() - # Workaround for https://bugs.python.org/issue17606 - # Before Python 2.7.4 xml.sax.saxutils required bytes; - # since 2.7.4 it requires unicode. The bug is likely to be - # fixed in 2.7.6, but 2.7.6 will still support unicode, - # and Python 3.x will require unicode, so ">= 2.7.4" should be fine. - if sys.version_info[:3] >= (2, 7, 4): - def _xg_characters(self, serialized_value): - if not isinstance(serialized_value, six.text_type): - serialized_value = serialized_value.decode(self.encoding) - return self.xg.characters(serialized_value) - else: # pragma: no cover - def _xg_characters(self, serialized_value): - return self.xg.characters(serialized_value) - class CsvItemExporter(BaseItemExporter): From c2898fdcf91cbbfa9d801196924960d3a3764b93 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 20 Aug 2019 21:06:52 +0500 Subject: [PATCH 176/387] Deprecate scrapy.utils.gz.read1. --- scrapy/utils/gz.py | 1 + 1 file changed, 1 insertion(+) diff --git a/scrapy/utils/gz.py b/scrapy/utils/gz.py index 9984492f0..dc8316d8c 100644 --- a/scrapy/utils/gz.py +++ b/scrapy/utils/gz.py @@ -16,6 +16,7 @@ from scrapy.utils.decorators import deprecated # (regression or bug-fix compared to Python 3.4) # - read1(), which fetches data before raising EOFError on next call # works here but is only available from Python>=3.3 +@deprecated('GzipFile.read1') def read1(gzf, size=-1): return gzf.read1(size) From cea2f5e244f30591bae667abf2b2689e566ab343 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 20 Aug 2019 21:09:49 +0500 Subject: [PATCH 177/387] Remove cStringIO imports. --- scrapy/downloadermiddlewares/decompression.py | 6 +----- scrapy/mail.py | 6 +----- scrapy/pipelines/files.py | 7 +------ scrapy/pipelines/images.py | 6 +----- scrapy/utils/gz.py | 9 ++------- scrapy/utils/iterators.py | 6 +----- tests/test_cmdline/__init__.py | 5 +---- tests/test_pipeline_media.py | 6 ++---- 8 files changed, 10 insertions(+), 41 deletions(-) diff --git a/scrapy/downloadermiddlewares/decompression.py b/scrapy/downloadermiddlewares/decompression.py index 49313cc04..e2d73f347 100644 --- a/scrapy/downloadermiddlewares/decompression.py +++ b/scrapy/downloadermiddlewares/decompression.py @@ -4,6 +4,7 @@ and extract the potentially compressed responses that may arrive. import bz2 import gzip +from io import BytesIO import zipfile import tarfile import logging @@ -11,11 +12,6 @@ from tempfile import mktemp import six -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO - from scrapy.responsetypes import responsetypes logger = logging.getLogger(__name__) diff --git a/scrapy/mail.py b/scrapy/mail.py index 746468e25..d24de2212 100644 --- a/scrapy/mail.py +++ b/scrapy/mail.py @@ -3,13 +3,9 @@ Mail sending helpers See documentation in docs/topics/email.rst """ +from io import BytesIO import logging -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO - from email.utils import COMMASPACE, formatdate from six.moves.email_mime_multipart import MIMEMultipart from six.moves.email_mime_text import MIMEText diff --git a/scrapy/pipelines/files.py b/scrapy/pipelines/files.py index cc3d10b63..8d74c5011 100644 --- a/scrapy/pipelines/files.py +++ b/scrapy/pipelines/files.py @@ -5,6 +5,7 @@ See documentation in topics/media-pipeline.rst """ import functools import hashlib +from io import BytesIO import mimetypes import os import os.path @@ -15,12 +16,6 @@ from six.moves.urllib.parse import urlparse from collections import defaultdict import six - -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO - from twisted.internet import defer, threads from scrapy.pipelines.media import MediaPipeline diff --git a/scrapy/pipelines/images.py b/scrapy/pipelines/images.py index fa4d12ad1..e77cef4ff 100644 --- a/scrapy/pipelines/images.py +++ b/scrapy/pipelines/images.py @@ -5,13 +5,9 @@ See documentation in topics/media-pipeline.rst """ import functools import hashlib +from io import BytesIO import six -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO - from PIL import Image from scrapy.utils.misc import md5sum diff --git a/scrapy/utils/gz.py b/scrapy/utils/gz.py index dc8316d8c..f41e62fe3 100644 --- a/scrapy/utils/gz.py +++ b/scrapy/utils/gz.py @@ -1,12 +1,7 @@ -import struct - -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO from gzip import GzipFile - +from io import BytesIO import re +import struct from scrapy.utils.decorators import deprecated diff --git a/scrapy/utils/iterators.py b/scrapy/utils/iterators.py index dbc1e0d20..9693ba768 100644 --- a/scrapy/utils/iterators.py +++ b/scrapy/utils/iterators.py @@ -1,11 +1,7 @@ import re import csv -import logging -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO from io import StringIO +import logging import six from scrapy.http import TextResponse, Response diff --git a/tests/test_cmdline/__init__.py b/tests/test_cmdline/__init__.py index 68dfb1cca..56cfe642a 100644 --- a/tests/test_cmdline/__init__.py +++ b/tests/test_cmdline/__init__.py @@ -1,3 +1,4 @@ +from io import StringIO import json import os import pstats @@ -7,10 +8,6 @@ from subprocess import Popen, PIPE import sys import tempfile import unittest -try: - from cStringIO import StringIO -except ImportError: - from io import StringIO from scrapy.utils.test import get_testenv diff --git a/tests/test_pipeline_media.py b/tests/test_pipeline_media.py index 28e39cefa..ad2618ec9 100644 --- a/tests/test_pipeline_media.py +++ b/tests/test_pipeline_media.py @@ -144,10 +144,8 @@ class BaseMediaPipelineTestCase(unittest.TestCase): # The Failure should encapsulate a FileException ... self.assertEqual(failure.value, file_exc) - # ... and if we're running on Python 3 ... - if sys.version_info.major >= 3: - # ... it should have the returnValue exception set as its context - self.assertEqual(failure.value.__context__, def_gen_return_exc) + # ... and it should have the returnValue exception set as its context + self.assertEqual(failure.value.__context__, def_gen_return_exc) # Let's calculate the request fingerprint and fake some runtime data... fp = request_fingerprint(request) From 5b70b051a6cec699c3cb14c3ac9cf530671a8487 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 20 Aug 2019 21:13:01 +0500 Subject: [PATCH 178/387] Some text function messages cleanup, deprecate to_native_str. --- scrapy/http/response/__init__.py | 2 +- scrapy/utils/python.py | 8 ++++---- 2 files changed, 5 insertions(+), 5 deletions(-) diff --git a/scrapy/http/response/__init__.py b/scrapy/http/response/__init__.py index b0a526b72..a81404afb 100644 --- a/scrapy/http/response/__init__.py +++ b/scrapy/http/response/__init__.py @@ -88,7 +88,7 @@ class Response(object_ref): @property def text(self): """For subclasses of TextResponse, this will return the body - as text (unicode object in Python 2 and str in Python 3) + as str """ raise AttributeError("Response content isn't text") diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index 5009aab81..974abaeb1 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -90,7 +90,7 @@ def to_unicode(text, encoding=None, errors='strict'): if isinstance(text, six.text_type): return text if not isinstance(text, (bytes, six.text_type)): - raise TypeError('to_unicode must receive a bytes, str or unicode ' + raise TypeError('to_unicode must receive a bytes or str ' 'object, got %s' % type(text).__name__) if encoding is None: encoding = 'utf-8' @@ -103,16 +103,16 @@ def to_bytes(text, encoding=None, errors='strict'): if isinstance(text, bytes): return text if not isinstance(text, six.string_types): - raise TypeError('to_bytes must receive a unicode, str or bytes ' + raise TypeError('to_bytes must receive a str or bytes ' 'object, got %s' % type(text).__name__) if encoding is None: encoding = 'utf-8' return text.encode(encoding, errors) +@deprecated('to_unicode') def to_native_str(text, encoding=None, errors='strict'): - """ Return str representation of ``text`` - (bytes in Python 2.x and unicode in Python 3.x). """ + """ Return str representation of ``text``. """ return to_unicode(text, encoding, errors) From 3ac4b430ae0165f25d8cd2fb1ac25ac9f307df8b Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 31 Oct 2019 15:20:28 +0500 Subject: [PATCH 179/387] Remove an unused six import. --- tests/test_downloader_handlers.py | 1 - 1 file changed, 1 deletion(-) diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 9f74577a5..e6856945c 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -1,5 +1,4 @@ import os -import six import shutil import tempfile import contextlib From 75b1d051d99cda17558a210326e77046eec851a7 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 20 Aug 2019 21:27:52 +0500 Subject: [PATCH 180/387] Simplify some more imports. --- scrapy/downloadermiddlewares/httpproxy.py | 5 +---- scrapy/loader/processors.py | 5 +---- tests/__init__.py | 5 ----- tests/test_downloader_handlers.py | 5 +---- tests/test_downloadermiddleware.py | 3 ++- tests/test_downloadermiddleware_robotstxt.py | 4 +++- tests/test_extension_telnet.py | 5 ----- tests/test_feedexport.py | 2 +- tests/test_http_request.py | 3 +-- tests/test_item.py | 2 +- tests/test_pipeline_files.py | 6 +----- tests/test_settings/__init__.py | 3 +-- tests/test_spider.py | 3 +-- tests/test_spidermiddleware.py | 3 ++- tests/test_stats.py | 6 +----- tests/test_utils_deprecate.py | 3 +-- tests/test_utils_misc/__init__.py | 2 +- tests/test_utils_trackref.py | 2 +- 18 files changed, 20 insertions(+), 47 deletions(-) diff --git a/scrapy/downloadermiddlewares/httpproxy.py b/scrapy/downloadermiddlewares/httpproxy.py index 2c35d1b90..2212d9688 100644 --- a/scrapy/downloadermiddlewares/httpproxy.py +++ b/scrapy/downloadermiddlewares/httpproxy.py @@ -1,10 +1,7 @@ import base64 from six.moves.urllib.parse import unquote, urlunparse from six.moves.urllib.request import getproxies, proxy_bypass -try: - from urllib2 import _parse_proxy -except ImportError: - from urllib.request import _parse_proxy +from urllib.request import _parse_proxy from scrapy.exceptions import NotConfigured from scrapy.utils.httpobj import urlparse_cached diff --git a/scrapy/loader/processors.py b/scrapy/loader/processors.py index 2acdc8093..02c625acc 100644 --- a/scrapy/loader/processors.py +++ b/scrapy/loader/processors.py @@ -3,10 +3,7 @@ This module provides some commonly used processors for Item Loaders. See documentation in docs/topics/loaders.rst """ -try: - from collections import ChainMap -except ImportError: - from scrapy.utils.datatypes import MergeDict as ChainMap +from collections import ChainMap from scrapy.utils.misc import arg_to_iter from scrapy.loader.common import wrap_loader_context diff --git a/tests/__init__.py b/tests/__init__.py index a54367f8c..12ce79fa9 100644 --- a/tests/__init__.py +++ b/tests/__init__.py @@ -21,11 +21,6 @@ if 'COV_CORE_CONFIG' in os.environ: os.environ['COV_CORE_CONFIG'] = os.path.join(_sourceroot, os.environ['COV_CORE_CONFIG']) -try: - import unittest.mock as mock -except ImportError: - import mock - tests_datadir = os.path.join(os.path.abspath(os.path.dirname(__file__)), 'sample_data') diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 109469503..6090998d4 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -2,11 +2,8 @@ import os import six import shutil import tempfile +from unittest import mock import contextlib -try: - from unittest import mock -except ImportError: - import mock from testfixtures import LogCapture from twisted.trial import unittest diff --git a/tests/test_downloadermiddleware.py b/tests/test_downloadermiddleware.py index 03564e748..6b9a5bee8 100644 --- a/tests/test_downloadermiddleware.py +++ b/tests/test_downloadermiddleware.py @@ -1,3 +1,5 @@ +from unittest import mock + from twisted.trial.unittest import TestCase from twisted.python.failure import Failure @@ -7,7 +9,6 @@ from scrapy.exceptions import _InvalidOutput from scrapy.core.downloader.middleware import DownloaderMiddlewareManager from scrapy.utils.test import get_crawler from scrapy.utils.python import to_bytes -from tests import mock class ManagerTestCase(TestCase): diff --git a/tests/test_downloadermiddleware_robotstxt.py b/tests/test_downloadermiddleware_robotstxt.py index fbc46cba4..8266bf35f 100644 --- a/tests/test_downloadermiddleware_robotstxt.py +++ b/tests/test_downloadermiddleware_robotstxt.py @@ -1,5 +1,8 @@ # -*- coding: utf-8 -*- from __future__ import absolute_import + +from unittest import mock + from twisted.internet import reactor, error from twisted.internet.defer import Deferred, DeferredList, maybeDeferred from twisted.python import failure @@ -9,7 +12,6 @@ from scrapy.downloadermiddlewares.robotstxt import (RobotsTxtMiddleware, from scrapy.exceptions import IgnoreRequest, NotConfigured from scrapy.http import Request, Response, TextResponse from scrapy.settings import Settings -from tests import mock from tests.test_robotstxt_interface import rerp_available, reppy_available diff --git a/tests/test_extension_telnet.py b/tests/test_extension_telnet.py index 4f389e5cb..875ceb83c 100644 --- a/tests/test_extension_telnet.py +++ b/tests/test_extension_telnet.py @@ -1,8 +1,3 @@ -try: - import unittest.mock as mock -except ImportError: - import mock - from twisted.trial import unittest from twisted.conch.telnet import ITelnetProtocol from twisted.cred import credentials diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index e1436fbe5..7431f921f 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -7,6 +7,7 @@ from io import BytesIO import tempfile import shutil import string +from unittest import mock from six.moves.urllib.parse import urljoin, urlparse, quote from six.moves.urllib.request import pathname2url @@ -15,7 +16,6 @@ from twisted.trial import unittest from twisted.internet import defer from scrapy.crawler import CrawlerRunner from scrapy.settings import Settings -from tests import mock from tests.mockserver import MockServer from w3lib.url import path_to_file_uri diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 828902b99..effb9e53b 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -3,6 +3,7 @@ import cgi import unittest import re import json +from unittest import mock from urllib.parse import unquote_to_bytes import warnings @@ -13,8 +14,6 @@ from six.moves.urllib.parse import urlparse, parse_qs, unquote from scrapy.http import Request, FormRequest, XmlRpcRequest, JsonRequest, Headers, HtmlResponse from scrapy.utils.python import to_bytes, to_native_str -from tests import mock - class RequestTest(unittest.TestCase): diff --git a/tests/test_item.py b/tests/test_item.py index 0ad278701..d98c63ddd 100644 --- a/tests/test_item.py +++ b/tests/test_item.py @@ -1,12 +1,12 @@ import sys import unittest +from unittest import mock from warnings import catch_warnings import six from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.item import ABCMeta, DictItem, Field, Item, ItemMeta -from tests import mock PY36_PLUS = (sys.version_info.major >= 3) and (sys.version_info.minor >= 6) diff --git a/tests/test_pipeline_files.py b/tests/test_pipeline_files.py index cb8f8da18..bd40e4103 100644 --- a/tests/test_pipeline_files.py +++ b/tests/test_pipeline_files.py @@ -1,10 +1,9 @@ import os import random import time -import hashlib -import warnings from tempfile import mkdtemp from shutil import rmtree +from unittest import mock from six.moves.urllib.parse import urlparse from six import BytesIO @@ -15,13 +14,10 @@ from scrapy.pipelines.files import FilesPipeline, FSFilesStore, S3FilesStore, GC from scrapy.item import Item, Field from scrapy.http import Request, Response from scrapy.settings import Settings -from scrapy.utils.python import to_bytes from scrapy.utils.test import assert_aws_environ, get_s3_content_and_delete from scrapy.utils.test import assert_gcs_environ, get_gcs_content_and_delete from scrapy.utils.boto import is_botocore -from tests import mock - def _mocked_download_func(request, info): response = request.meta.get('response') diff --git a/tests/test_settings/__init__.py b/tests/test_settings/__init__.py index 08286ff02..32e65bed5 100644 --- a/tests/test_settings/__init__.py +++ b/tests/test_settings/__init__.py @@ -1,10 +1,9 @@ import six import unittest -import warnings +from unittest import mock from scrapy.settings import (BaseSettings, Settings, SettingsAttribute, SETTINGS_PRIORITIES, get_settings_priority) -from tests import mock from . import default_settings diff --git a/tests/test_spider.py b/tests/test_spider.py index 2220b8ffc..b913a56b7 100644 --- a/tests/test_spider.py +++ b/tests/test_spider.py @@ -1,5 +1,6 @@ import gzip import inspect +from unittest import mock import warnings from io import BytesIO @@ -17,8 +18,6 @@ from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.utils.trackref import object_ref from scrapy.utils.test import get_crawler -from tests import mock - class SpiderTest(unittest.TestCase): diff --git a/tests/test_spidermiddleware.py b/tests/test_spidermiddleware.py index 832fd3330..55d665e79 100644 --- a/tests/test_spidermiddleware.py +++ b/tests/test_spidermiddleware.py @@ -1,3 +1,5 @@ +from unittest import mock + from twisted.trial.unittest import TestCase from twisted.python.failure import Failure @@ -6,7 +8,6 @@ from scrapy.http import Request, Response from scrapy.exceptions import _InvalidOutput from scrapy.utils.test import get_crawler from scrapy.core.spidermw import SpiderMiddlewareManager -from tests import mock class SpiderMiddlewareTestCase(TestCase): diff --git a/tests/test_stats.py b/tests/test_stats.py index 2033dbe07..2bbbb9e2c 100644 --- a/tests/test_stats.py +++ b/tests/test_stats.py @@ -1,10 +1,6 @@ from datetime import datetime import unittest - -try: - from unittest import mock -except ImportError: - import mock +from unittest import mock from scrapy.extensions.corestats import CoreStats from scrapy.spiders import Spider diff --git a/tests/test_utils_deprecate.py b/tests/test_utils_deprecate.py index 3e7236fb1..ce04e7f29 100644 --- a/tests/test_utils_deprecate.py +++ b/tests/test_utils_deprecate.py @@ -2,12 +2,11 @@ from __future__ import absolute_import import inspect import unittest +from unittest import mock import warnings from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.utils.deprecate import create_deprecated_class, update_classpath -from tests import mock - class MyWarning(UserWarning): pass diff --git a/tests/test_utils_misc/__init__.py b/tests/test_utils_misc/__init__.py index e109d5343..de9da9104 100644 --- a/tests/test_utils_misc/__init__.py +++ b/tests/test_utils_misc/__init__.py @@ -1,11 +1,11 @@ import sys import os import unittest +from unittest import mock from scrapy.item import Item, Field from scrapy.utils.misc import arg_to_iter, create_instance, load_object, set_environ, walk_modules -from tests import mock __doctests__ = ['scrapy.utils.misc'] diff --git a/tests/test_utils_trackref.py b/tests/test_utils_trackref.py index c6072fc0d..480a717e7 100644 --- a/tests/test_utils_trackref.py +++ b/tests/test_utils_trackref.py @@ -1,7 +1,7 @@ import six import unittest +from unittest import mock from scrapy.utils import trackref -from tests import mock class Foo(trackref.object_ref): From 397e8835564614608647c855c62c36232020f078 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 20 Aug 2019 21:35:13 +0500 Subject: [PATCH 181/387] Replace to_native_str calls with to_unicode. --- scrapy/core/downloader/handlers/http11.py | 2 +- scrapy/downloadermiddlewares/cookies.py | 6 +++--- scrapy/downloadermiddlewares/robotstxt.py | 3 --- scrapy/exporters.py | 4 ++-- scrapy/http/cookies.py | 10 +++++----- scrapy/http/response/text.py | 8 ++++---- scrapy/linkextractors/lxmlhtml.py | 4 ++-- scrapy/responsetypes.py | 6 +++--- scrapy/robotstxt.py | 8 ++++---- scrapy/spidermiddlewares/referer.py | 5 ++--- scrapy/utils/reqser.py | 4 ++-- scrapy/utils/request.py | 4 ++-- scrapy/utils/response.py | 4 ++-- scrapy/utils/ssl.py | 4 ++-- tests/test_command_parse.py | 5 ++--- tests/test_commands.py | 7 ++----- tests/test_feedexport.py | 7 +++---- tests/test_http_request.py | 7 +++---- tests/test_http_response.py | 6 +++--- tests/test_robotstxt_interface.py | 1 - 20 files changed, 47 insertions(+), 58 deletions(-) diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index 91b45a8fc..7d917cb74 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -174,7 +174,7 @@ def tunnel_request_data(host, port, proxy_auth_header=None): r""" Return binary content of a CONNECT request. - >>> from scrapy.utils.python import to_native_str as s + >>> from scrapy.utils.python import to_unicode as s >>> s(tunnel_request_data("example.com", 8080)) 'CONNECT example.com:8080 HTTP/1.1\r\nHost: example.com:8080\r\n\r\n' >>> s(tunnel_request_data("example.com", 8080, b"123")) diff --git a/scrapy/downloadermiddlewares/cookies.py b/scrapy/downloadermiddlewares/cookies.py index 321c0171b..0d2b9900c 100644 --- a/scrapy/downloadermiddlewares/cookies.py +++ b/scrapy/downloadermiddlewares/cookies.py @@ -6,7 +6,7 @@ from collections import defaultdict from scrapy.exceptions import NotConfigured from scrapy.http import Response from scrapy.http.cookies import CookieJar -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode logger = logging.getLogger(__name__) @@ -53,7 +53,7 @@ class CookiesMiddleware(object): def _debug_cookie(self, request, spider): if self.debug: - cl = [to_native_str(c, errors='replace') + cl = [to_unicode(c, errors='replace') for c in request.headers.getlist('Cookie')] if cl: cookies = "\n".join("Cookie: {}\n".format(c) for c in cl) @@ -62,7 +62,7 @@ class CookiesMiddleware(object): def _debug_set_cookie(self, response, spider): if self.debug: - cl = [to_native_str(c, errors='replace') + cl = [to_unicode(c, errors='replace') for c in response.headers.getlist('Set-Cookie')] if cl: cookies = "\n".join("Set-Cookie: {}\n".format(c) for c in cl) diff --git a/scrapy/downloadermiddlewares/robotstxt.py b/scrapy/downloadermiddlewares/robotstxt.py index 6a5dfb79c..251706c50 100644 --- a/scrapy/downloadermiddlewares/robotstxt.py +++ b/scrapy/downloadermiddlewares/robotstxt.py @@ -5,15 +5,12 @@ enable this middleware and enable the ROBOTSTXT_OBEY setting. """ import logging -import sys -import re from twisted.internet.defer import Deferred, maybeDeferred from scrapy.exceptions import NotConfigured, IgnoreRequest from scrapy.http import Request from scrapy.utils.httpobj import urlparse_cached from scrapy.utils.log import failure_to_exc_info -from scrapy.utils.python import to_native_str from scrapy.utils.misc import load_object logger = logging.getLogger(__name__) diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 40567f53b..f276c28e8 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -12,7 +12,7 @@ from six.moves import cPickle as pickle from xml.sax.saxutils import XMLGenerator from scrapy.utils.serialize import ScrapyJSONEncoder -from scrapy.utils.python import to_bytes, to_unicode, to_native_str, is_listlike +from scrapy.utils.python import to_bytes, to_unicode, is_listlike from scrapy.item import BaseItem from scrapy.exceptions import ScrapyDeprecationWarning import warnings @@ -232,7 +232,7 @@ class CsvItemExporter(BaseItemExporter): def _build_row(self, values): for s in values: try: - yield to_native_str(s, self.encoding) + yield to_unicode(s, self.encoding) except TypeError: yield s diff --git a/scrapy/http/cookies.py b/scrapy/http/cookies.py index 4e8056750..4532c3ab7 100644 --- a/scrapy/http/cookies.py +++ b/scrapy/http/cookies.py @@ -3,7 +3,7 @@ from six.moves.http_cookiejar import ( CookieJar as _CookieJar, DefaultCookiePolicy, IPV4_RE ) from scrapy.utils.httpobj import urlparse_cached -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode class CookieJar(object): @@ -165,13 +165,13 @@ class WrappedRequest(object): return name in self.request.headers def get_header(self, name, default=None): - return to_native_str(self.request.headers.get(name, default), + return to_unicode(self.request.headers.get(name, default), errors='replace') def header_items(self): return [ - (to_native_str(k, errors='replace'), - [to_native_str(x, errors='replace') for x in v]) + (to_unicode(k, errors='replace'), + [to_unicode(x, errors='replace') for x in v]) for k, v in self.request.headers.items() ] @@ -189,7 +189,7 @@ class WrappedResponse(object): # python3 cookiejars calls get_all def get_all(self, name, default=None): - return [to_native_str(v, errors='replace') + return [to_unicode(v, errors='replace') for v in self.response.headers.getlist(name)] # python2 cookiejars calls getheaders getheaders = get_all diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index a8010877c..37f450e54 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -16,7 +16,7 @@ from w3lib.html import strip_html5_whitespace from scrapy.http.request import Request from scrapy.http.response import Response from scrapy.utils.response import get_base_url -from scrapy.utils.python import memoizemethod_noargs, to_native_str +from scrapy.utils.python import memoizemethod_noargs, to_unicode class TextResponse(Response): @@ -32,7 +32,7 @@ class TextResponse(Response): def _set_url(self, url): if isinstance(url, six.text_type): - self._url = to_native_str(url, self.encoding) + self._url = to_unicode(url, self.encoding) else: super(TextResponse, self)._set_url(url) @@ -81,11 +81,11 @@ class TextResponse(Response): @memoizemethod_noargs def _headers_encoding(self): content_type = self.headers.get(b'Content-Type', b'') - return http_content_type_encoding(to_native_str(content_type)) + return http_content_type_encoding(to_unicode(content_type)) def _body_inferred_encoding(self): if self._cached_benc is None: - content_type = to_native_str(self.headers.get(b'Content-Type', b'')) + content_type = to_unicode(self.headers.get(b'Content-Type', b'')) benc, ubody = html_to_unicode(content_type, self.body, auto_detect_fun=self._auto_detect_fun, default_encoding=self._DEFAULT_ENCODING) diff --git a/scrapy/linkextractors/lxmlhtml.py b/scrapy/linkextractors/lxmlhtml.py index 8f6f93a44..890c019c8 100644 --- a/scrapy/linkextractors/lxmlhtml.py +++ b/scrapy/linkextractors/lxmlhtml.py @@ -10,7 +10,7 @@ from w3lib.url import canonicalize_url from scrapy.link import Link from scrapy.utils.misc import arg_to_iter, rel_has_nofollow -from scrapy.utils.python import unique as unique_list, to_native_str +from scrapy.utils.python import unique as unique_list, to_unicode from scrapy.utils.response import get_base_url from scrapy.linkextractors import FilteringLinkExtractor @@ -67,7 +67,7 @@ class LxmlParserLinkExtractor(object): url = self.process_attr(attr_val) if url is None: continue - url = to_native_str(url, encoding=response_encoding) + url = to_unicode(url, encoding=response_encoding) # to fix relative links after process_value url = urljoin(response_url, url) link = Link(url, _collect_string_content(el) or u'', diff --git a/scrapy/responsetypes.py b/scrapy/responsetypes.py index 4a2d5bf52..de62276c8 100644 --- a/scrapy/responsetypes.py +++ b/scrapy/responsetypes.py @@ -10,7 +10,7 @@ import six from scrapy.http import Response from scrapy.utils.misc import load_object -from scrapy.utils.python import binary_is_text, to_bytes, to_native_str +from scrapy.utils.python import binary_is_text, to_bytes, to_unicode class ResponseTypes(object): @@ -55,12 +55,12 @@ class ResponseTypes(object): header """ if content_encoding: return Response - mimetype = to_native_str(content_type).split(';')[0].strip().lower() + mimetype = to_unicode(content_type).split(';')[0].strip().lower() return self.from_mimetype(mimetype) def from_content_disposition(self, content_disposition): try: - filename = to_native_str(content_disposition, + filename = to_unicode(content_disposition, encoding='latin-1', errors='replace').split(';')[1].split('=')[1] filename = filename.strip('"\'') return self.from_filename(filename) diff --git a/scrapy/robotstxt.py b/scrapy/robotstxt.py index 189f165d1..95a8c09b8 100644 --- a/scrapy/robotstxt.py +++ b/scrapy/robotstxt.py @@ -3,14 +3,14 @@ import logging from abc import ABCMeta, abstractmethod from six import with_metaclass -from scrapy.utils.python import to_native_str, to_unicode +from scrapy.utils.python import to_unicode logger = logging.getLogger(__name__) def decode_robotstxt(robotstxt_body, spider, to_native_str_type=False): try: if to_native_str_type: - robotstxt_body = to_native_str(robotstxt_body) + robotstxt_body = to_unicode(robotstxt_body) else: robotstxt_body = robotstxt_body.decode('utf-8') except UnicodeDecodeError: @@ -66,8 +66,8 @@ class PythonRobotParser(RobotParser): return o def allowed(self, url, user_agent): - user_agent = to_native_str(user_agent) - url = to_native_str(url) + user_agent = to_unicode(user_agent) + url = to_unicode(url) return self.rp.can_fetch(user_agent, url) diff --git a/scrapy/spidermiddlewares/referer.py b/scrapy/spidermiddlewares/referer.py index 1ddfb37f4..c76e4d5a2 100644 --- a/scrapy/spidermiddlewares/referer.py +++ b/scrapy/spidermiddlewares/referer.py @@ -10,8 +10,7 @@ from w3lib.url import safe_url_string from scrapy.http import Request, Response from scrapy.exceptions import NotConfigured from scrapy import signals -from scrapy.utils.python import to_native_str -from scrapy.utils.httpobj import urlparse_cached +from scrapy.utils.python import to_unicode from scrapy.utils.misc import load_object from scrapy.utils.url import strip_url @@ -322,7 +321,7 @@ class RefererMiddleware(object): if isinstance(resp_or_url, Response): policy_header = resp_or_url.headers.get('Referrer-Policy') if policy_header is not None: - policy_name = to_native_str(policy_header.decode('latin1')) + policy_name = to_unicode(policy_header.decode('latin1')) if policy_name is None: return self.default_policy() diff --git a/scrapy/utils/reqser.py b/scrapy/utils/reqser.py index c7ea7b425..495564ac0 100644 --- a/scrapy/utils/reqser.py +++ b/scrapy/utils/reqser.py @@ -4,7 +4,7 @@ Helper functions for serializing (and deserializing) requests. import six from scrapy.http import Request -from scrapy.utils.python import to_unicode, to_native_str +from scrapy.utils.python import to_unicode from scrapy.utils.misc import load_object @@ -54,7 +54,7 @@ def request_from_dict(d, spider=None): eb = _get_method(spider, eb) request_cls = load_object(d['_class']) if '_class' in d else Request return request_cls( - url=to_native_str(d['url']), + url=to_unicode(d['url']), callback=cb, errback=eb, method=d['method'], diff --git a/scrapy/utils/request.py b/scrapy/utils/request.py index fb5af66a2..63d0ae772 100644 --- a/scrapy/utils/request.py +++ b/scrapy/utils/request.py @@ -9,7 +9,7 @@ import weakref from six.moves.urllib.parse import urlunparse from w3lib.http import basic_auth_header -from scrapy.utils.python import to_bytes, to_native_str +from scrapy.utils.python import to_bytes, to_unicode from w3lib.url import canonicalize_url from scrapy.utils.httpobj import urlparse_cached @@ -97,4 +97,4 @@ def referer_str(request): referrer = request.headers.get('Referer') if referrer is None: return referrer - return to_native_str(referrer, errors='replace') + return to_unicode(referrer, errors='replace') diff --git a/scrapy/utils/response.py b/scrapy/utils/response.py index c3236afd4..feab07431 100644 --- a/scrapy/utils/response.py +++ b/scrapy/utils/response.py @@ -8,7 +8,7 @@ import webbrowser import tempfile from twisted.web import http -from scrapy.utils.python import to_bytes, to_native_str +from scrapy.utils.python import to_bytes, to_unicode from w3lib import html @@ -36,7 +36,7 @@ def response_status_message(status): """Return status code plus status text descriptive message """ message = http.RESPONSES.get(int(status), "Unknown Status") - return '%s %s' % (status, to_native_str(message)) + return '%s %s' % (status, to_unicode(message)) def response_httprepr(response): diff --git a/scrapy/utils/ssl.py b/scrapy/utils/ssl.py index 02aed60ee..6e81b33ff 100644 --- a/scrapy/utils/ssl.py +++ b/scrapy/utils/ssl.py @@ -3,7 +3,7 @@ import OpenSSL import OpenSSL._util as pyOpenSSLutil -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode # The OpenSSL symbol is present since 1.1.1 but it's not currently supported in any version of pyOpenSSL. @@ -12,7 +12,7 @@ SSL_OP_NO_TLSv1_3 = getattr(pyOpenSSLutil.lib, 'SSL_OP_NO_TLSv1_3', 0) def ffi_buf_to_string(buf): - return to_native_str(pyOpenSSLutil.ffi.string(buf)) + return to_unicode(pyOpenSSLutil.ffi.string(buf)) def x509name_to_string(x509name): diff --git a/tests/test_command_parse.py b/tests/test_command_parse.py index 62d5d76b4..b134beb88 100644 --- a/tests/test_command_parse.py +++ b/tests/test_command_parse.py @@ -1,17 +1,16 @@ import os from os.path import join, abspath -from twisted.trial import unittest from twisted.internet import defer from scrapy.utils.testsite import SiteTest from scrapy.utils.testproc import ProcessTest -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode from tests.test_commands import CommandTest def _textmode(bstr): """Normalize input the same as writing to a file and reading from it in text mode""" - return to_native_str(bstr).replace(os.linesep, '\n') + return to_unicode(bstr).replace(os.linesep, '\n') class ParseCommandTest(ProcessTest, SiteTest, CommandTest): command = 'parse' diff --git a/tests/test_commands.py b/tests/test_commands.py index b8445ae6c..536379170 100644 --- a/tests/test_commands.py +++ b/tests/test_commands.py @@ -10,13 +10,10 @@ from contextlib import contextmanager from threading import Timer from twisted.trial import unittest -from twisted.internet import defer import scrapy -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode from scrapy.utils.test import get_testenv -from scrapy.utils.testsite import SiteTest -from scrapy.utils.testproc import ProcessTest from tests.test_crawler import ExceptionSpider, NoRequestsSpider @@ -56,7 +53,7 @@ class ProjectTest(unittest.TestCase): finally: timer.cancel() - return p, to_native_str(stdout), to_native_str(stderr) + return p, to_unicode(stdout), to_unicode(stderr) class StartprojectTest(ProjectTest): diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 7431f921f..abe2ab557 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -26,8 +26,7 @@ from scrapy.extensions.feedexport import ( S3FeedStorage, StdoutFeedStorage, BlockingFeedStorage) from scrapy.utils.test import assert_aws_environ, get_s3_content_and_delete, get_crawler -from scrapy.utils.python import to_native_str -from scrapy.utils.project import get_project_settings +from scrapy.utils.python import to_unicode class FileFeedStorageTest(unittest.TestCase): @@ -456,7 +455,7 @@ class FeedExportTest(unittest.TestCase): settings.update({'FEED_FORMAT': 'csv'}) data = yield self.exported_data(items, settings) - reader = csv.DictReader(to_native_str(data).splitlines()) + reader = csv.DictReader(to_unicode(data).splitlines()) got_rows = list(reader) if ordered: self.assertEqual(reader.fieldnames, header) @@ -470,7 +469,7 @@ class FeedExportTest(unittest.TestCase): settings = settings or {} settings.update({'FEED_FORMAT': 'jl'}) data = yield self.exported_data(items, settings) - parsed = [json.loads(to_native_str(line)) for line in data.splitlines()] + parsed = [json.loads(to_unicode(line)) for line in data.splitlines()] rows = [{k: v for k, v in row.items() if v} for row in rows] self.assertEqual(rows, parsed) diff --git a/tests/test_http_request.py b/tests/test_http_request.py index effb9e53b..807265981 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -7,12 +7,11 @@ from unittest import mock from urllib.parse import unquote_to_bytes import warnings -import six from six.moves import xmlrpc_client as xmlrpclib from six.moves.urllib.parse import urlparse, parse_qs, unquote from scrapy.http import Request, FormRequest, XmlRpcRequest, JsonRequest, Headers, HtmlResponse -from scrapy.utils.python import to_bytes, to_native_str +from scrapy.utils.python import to_bytes, to_unicode class RequestTest(unittest.TestCase): @@ -349,8 +348,8 @@ class FormRequestTest(RequestTest): request_class = FormRequest def assertQueryEqual(self, first, second, msg=None): - first = to_native_str(first).split("&") - second = to_native_str(second).split("&") + first = to_unicode(first).split("&") + second = to_unicode(second).split("&") return self.assertEqual(sorted(first), sorted(second), msg) def test_empty_formdata(self): diff --git a/tests/test_http_response.py b/tests/test_http_response.py index d6e77d6b8..80bf51647 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -7,7 +7,7 @@ from w3lib.encoding import resolve_encoding from scrapy.http import (Request, Response, TextResponse, HtmlResponse, XmlResponse, Headers) from scrapy.selector import Selector -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode from scrapy.exceptions import NotSupported from scrapy.link import Link from tests import get_testdata @@ -204,11 +204,11 @@ class TextResponseTest(BaseResponseTest): assert isinstance(resp.url, str) resp = self.response_class(url=u"http://www.example.com/price/\xa3", encoding='utf-8') - self.assertEqual(resp.url, to_native_str(b'http://www.example.com/price/\xc2\xa3')) + self.assertEqual(resp.url, to_unicode(b'http://www.example.com/price/\xc2\xa3')) resp = self.response_class(url=u"http://www.example.com/price/\xa3", encoding='latin-1') self.assertEqual(resp.url, 'http://www.example.com/price/\xa3') resp = self.response_class(u"http://www.example.com/price/\xa3", headers={"Content-type": ["text/html; charset=utf-8"]}) - self.assertEqual(resp.url, to_native_str(b'http://www.example.com/price/\xc2\xa3')) + self.assertEqual(resp.url, to_unicode(b'http://www.example.com/price/\xc2\xa3')) resp = self.response_class(u"http://www.example.com/price/\xa3", headers={"Content-type": ["text/html; charset=iso-8859-1"]}) self.assertEqual(resp.url, 'http://www.example.com/price/\xa3') diff --git a/tests/test_robotstxt_interface.py b/tests/test_robotstxt_interface.py index 9aaab560a..cd7480e33 100644 --- a/tests/test_robotstxt_interface.py +++ b/tests/test_robotstxt_interface.py @@ -1,6 +1,5 @@ # coding=utf-8 from twisted.trial import unittest -from scrapy.utils.python import to_native_str def reppy_available(): From 7299e91b1f1b231095e2a7ddbce5ad980ae7da8a Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Wed, 28 Aug 2019 16:29:53 +0500 Subject: [PATCH 182/387] Remove Py2-only code that checks sys.version_info. --- tests/test_utils_reqser.py | 7 ++----- 1 file changed, 2 insertions(+), 5 deletions(-) diff --git a/tests/test_utils_reqser.py b/tests/test_utils_reqser.py index 11ac56897..92cd16de7 100644 --- a/tests/test_utils_reqser.py +++ b/tests/test_utils_reqser.py @@ -80,8 +80,6 @@ class RequestSerializationTest(unittest.TestCase): self._assert_serializes_ok(r, spider=self.spider) def test_mixin_private_callback_serialization(self): - if sys.version_info[0] < 3: - return r = Request("http://www.example.com", callback=self.spider._TestSpiderMixin__mixin_callback, errback=self.spider.handle_error) @@ -119,9 +117,8 @@ class RequestSerializationTest(unittest.TestCase): def test_private_name_mangling(self): self._assert_mangles_to( self.spider, '_TestSpider__parse_item_private') - if sys.version_info[0] >= 3: - self._assert_mangles_to( - self.spider, '_TestSpiderMixin__mixin_callback') + self._assert_mangles_to( + self.spider, '_TestSpiderMixin__mixin_callback') def test_unserializable_callback1(self): r = Request("http://www.example.com", callback=lambda x: x) From f02c3d1dcf3e4880388d19e961e7911be5dc54ff Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 31 Oct 2019 13:31:33 +0100 Subject: [PATCH 183/387] Use communicate() instead of wait() after killing the mock server (#4095) --- tests/mockserver.py | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/tests/mockserver.py b/tests/mockserver.py index 77908284b..b766bb653 100644 --- a/tests/mockserver.py +++ b/tests/mockserver.py @@ -206,8 +206,7 @@ class MockServer(): def __exit__(self, exc_type, exc_value, traceback): self.proc.kill() - self.proc.wait() - time.sleep(0.2) + self.proc.communicate() def url(self, path, is_secure=False): host = self.http_address.replace('0.0.0.0', '127.0.0.1') From 864123132a99dda61feb407a59367c6c06133517 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 31 Oct 2019 22:55:58 +0500 Subject: [PATCH 184/387] Fix a duplicate ref name in docs. --- docs/topics/downloader-middleware.rst | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 366b95510..e93645077 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -474,7 +474,7 @@ DBM storage backend A DBM_ storage backend is also available for the HTTP cache middleware. - By default, it uses the dbm_ module, but you can change it with the + By default, it uses the `dbm module`_, but you can change it with the :setting:`HTTPCACHE_DBM_MODULE` setting. .. _httpcache-storage-custom: @@ -1202,4 +1202,4 @@ The default encoding for proxy authentication on :class:`HttpProxyMiddleware`. .. _DBM: https://en.wikipedia.org/wiki/Dbm -.. _dbm: https://docs.python.org/3/library/dbm.html +.. _dbm module: https://docs.python.org/3/library/dbm.html From a5eb59b92d3311f702933753d85b15e6d698faa1 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 31 Oct 2019 23:21:14 +0500 Subject: [PATCH 185/387] Fix test_proxy_connect.py for py3.5. --- tests/test_proxy_connect.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index 5e9470e39..f6381b5b1 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -117,7 +117,7 @@ class ProxyConnectTestCase(TestCase): with LogCapture() as l: yield crawler.crawl(seed=request) self._assert_got_response_code(200, l) - echo = json.loads(crawler.spider.meta['responses'][0].body) + echo = json.loads(crawler.spider.meta['responses'][0].body.decode('utf-8')) self.assertTrue('Proxy-Authorization' not in echo['headers']) @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') From 5eb01b617d207aabc705038d8a0e3b46e358bdbb Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 31 Oct 2019 23:21:30 +0500 Subject: [PATCH 186/387] Use an older mitmproxy for py3.5. --- tests/requirements-py3.txt | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/tests/requirements-py3.txt b/tests/requirements-py3.txt index f27e45a54..c2b16bec6 100644 --- a/tests/requirements-py3.txt +++ b/tests/requirements-py3.txt @@ -1,6 +1,7 @@ # Tests requirements jmespath -mitmproxy +mitmproxy; python_version >= '3.6' +mitmproxy==3.0.4; python_version < '3.6' pytest pytest-cov pytest-twisted From e0c5c724969ebc35970daaf0203969dc2eb00d56 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 1 Nov 2019 19:46:19 +0500 Subject: [PATCH 187/387] Improve the test_https_tunnel_without_leak_proxy_authorization_header change. --- tests/test_proxy_connect.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index f6381b5b1..651576c2c 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -117,7 +117,7 @@ class ProxyConnectTestCase(TestCase): with LogCapture() as l: yield crawler.crawl(seed=request) self._assert_got_response_code(200, l) - echo = json.loads(crawler.spider.meta['responses'][0].body.decode('utf-8')) + echo = json.loads(crawler.spider.meta['responses'][0].text) self.assertTrue('Proxy-Authorization' not in echo['headers']) @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') From 3c9963ab049e49a1d490d9505ad22ae9c1415421 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 1 Nov 2019 19:46:38 +0500 Subject: [PATCH 188/387] Only xfail test_https_connect_tunnel_error on 3.6+. --- tests/test_proxy_connect.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index 651576c2c..ec3f0716c 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -92,7 +92,7 @@ class ProxyConnectTestCase(TestCase): yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) self._assert_got_response_code(200, l) - @pytest.mark.xfail(reason='Python 3 fails this earlier') + @pytest.mark.xfail(reason='Python 3.6+ fails this earlier', condition=sys.version_info.minor >= 6) @defer.inlineCallbacks def test_https_connect_tunnel_error(self): crawler = get_crawler(SimpleSpider) From 4b0cdf7f3ed94e81612cceacfab6d81d54117865 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 1 Nov 2019 19:50:56 +0500 Subject: [PATCH 189/387] Use self.proc.communicate() after killing mitmdump. --- tests/test_proxy_connect.py | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index ec3f0716c..69925f80c 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -48,8 +48,7 @@ sys.exit(mitmdump()) def stop(self): self.proc.kill() - self.proc.wait() - time.sleep(0.2) + self.proc.communicate() def _wrong_credentials(proxy_url): From 350aa67c3dc8997ef9d3aac9ef3f596c83758e1c Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 1 Nov 2019 19:52:57 +0500 Subject: [PATCH 190/387] Rename tests/py3-ignores.txt to tests/ignores.txt. --- conftest.py | 2 +- tests/{py3-ignores.txt => ignores.txt} | 0 2 files changed, 1 insertion(+), 1 deletion(-) rename tests/{py3-ignores.txt => ignores.txt} (100%) diff --git a/conftest.py b/conftest.py index ede091e9f..7da4c4976 100644 --- a/conftest.py +++ b/conftest.py @@ -7,7 +7,7 @@ collect_ignore = [ ] -for line in open('tests/py3-ignores.txt'): +for line in open('tests/ignores.txt'): file_path = line.strip() if file_path and file_path[0] != '#': collect_ignore.append(file_path) diff --git a/tests/py3-ignores.txt b/tests/ignores.txt similarity index 100% rename from tests/py3-ignores.txt rename to tests/ignores.txt From 48b8ac60099c3751ad1f595a9a629be26ffb34cd Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 1 Nov 2019 20:05:37 +0500 Subject: [PATCH 191/387] Improve the dbm module ref. MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves <adrian@chaves.io> --- docs/topics/downloader-middleware.rst | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index e93645077..ae6d41809 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -474,7 +474,7 @@ DBM storage backend A DBM_ storage backend is also available for the HTTP cache middleware. - By default, it uses the `dbm module`_, but you can change it with the + By default, it uses the :mod:`dbm`, but you can change it with the :setting:`HTTPCACHE_DBM_MODULE` setting. .. _httpcache-storage-custom: @@ -1202,4 +1202,3 @@ The default encoding for proxy authentication on :class:`HttpProxyMiddleware`. .. _DBM: https://en.wikipedia.org/wiki/Dbm -.. _dbm module: https://docs.python.org/3/library/dbm.html From 415526d922b6c70bfd5872c8d67d69fcce0ee1fb Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sat, 2 Nov 2019 23:25:15 -0300 Subject: [PATCH 192/387] Remove __future__ imports --- extras/qps-bench-server.py | 1 - scrapy/cmdline.py | 1 - scrapy/commands/check.py | 1 - scrapy/commands/fetch.py | 1 - scrapy/commands/genspider.py | 1 - scrapy/commands/list.py | 1 - scrapy/commands/parse.py | 1 - scrapy/commands/settings.py | 1 - scrapy/commands/startproject.py | 1 - scrapy/commands/version.py | 2 -- scrapy/core/downloader/__init__.py | 1 - scrapy/core/downloader/handlers/http.py | 1 - scrapy/downloadermiddlewares/ajaxcrawl.py | 1 - scrapy/dupefilters.py | 1 - scrapy/extensions/httpcache.py | 2 -- scrapy/pipelines/media.py | 2 -- scrapy/responsetypes.py | 1 - scrapy/shell.py | 2 -- scrapy/signalmanager.py | 1 - scrapy/spiderloader.py | 1 - scrapy/utils/boto.py | 2 -- scrapy/utils/display.py | 1 - scrapy/utils/engine.py | 1 - scrapy/utils/ossignal.py | 2 -- scrapy/utils/request.py | 1 - scrapy/utils/test.py | 1 - scrapy/utils/testproc.py | 1 - scrapy/utils/testsite.py | 1 - scrapy/utils/trackref.py | 1 - 29 files changed, 35 deletions(-) diff --git a/extras/qps-bench-server.py b/extras/qps-bench-server.py index 3bef20bf3..da7a0022b 100755 --- a/extras/qps-bench-server.py +++ b/extras/qps-bench-server.py @@ -1,5 +1,4 @@ #!/usr/bin/env python -from __future__ import print_function from time import time from collections import deque from twisted.web.server import Site, NOT_DONE_YET diff --git a/scrapy/cmdline.py b/scrapy/cmdline.py index 418dc1ac9..69e917004 100644 --- a/scrapy/cmdline.py +++ b/scrapy/cmdline.py @@ -1,4 +1,3 @@ -from __future__ import print_function import sys import os import optparse diff --git a/scrapy/commands/check.py b/scrapy/commands/check.py index ab73e85e7..ac2a95c9b 100644 --- a/scrapy/commands/check.py +++ b/scrapy/commands/check.py @@ -1,4 +1,3 @@ -from __future__ import print_function import time import sys from collections import defaultdict diff --git a/scrapy/commands/fetch.py b/scrapy/commands/fetch.py index d45133e0e..95f6f7b9a 100644 --- a/scrapy/commands/fetch.py +++ b/scrapy/commands/fetch.py @@ -1,4 +1,3 @@ -from __future__ import print_function import sys from w3lib.url import is_url diff --git a/scrapy/commands/genspider.py b/scrapy/commands/genspider.py index d5498bb5c..adb01fa70 100644 --- a/scrapy/commands/genspider.py +++ b/scrapy/commands/genspider.py @@ -1,4 +1,3 @@ -from __future__ import print_function import os import shutil import string diff --git a/scrapy/commands/list.py b/scrapy/commands/list.py index a255b3b94..60686f109 100644 --- a/scrapy/commands/list.py +++ b/scrapy/commands/list.py @@ -1,4 +1,3 @@ -from __future__ import print_function from scrapy.commands import ScrapyCommand class Command(ScrapyCommand): diff --git a/scrapy/commands/parse.py b/scrapy/commands/parse.py index ef8acd29c..ff6f1d8cd 100644 --- a/scrapy/commands/parse.py +++ b/scrapy/commands/parse.py @@ -1,4 +1,3 @@ -from __future__ import print_function import json import logging diff --git a/scrapy/commands/settings.py b/scrapy/commands/settings.py index bee52f06a..a34338715 100644 --- a/scrapy/commands/settings.py +++ b/scrapy/commands/settings.py @@ -1,4 +1,3 @@ -from __future__ import print_function import json from scrapy.commands import ScrapyCommand diff --git a/scrapy/commands/startproject.py b/scrapy/commands/startproject.py index 67337c26e..34df25cef 100644 --- a/scrapy/commands/startproject.py +++ b/scrapy/commands/startproject.py @@ -1,4 +1,3 @@ -from __future__ import print_function import re import os import string diff --git a/scrapy/commands/version.py b/scrapy/commands/version.py index 577365c3b..494855500 100644 --- a/scrapy/commands/version.py +++ b/scrapy/commands/version.py @@ -1,5 +1,3 @@ -from __future__ import print_function - import scrapy from scrapy.commands import ScrapyCommand from scrapy.utils.versions import scrapy_components_versions diff --git a/scrapy/core/downloader/__init__.py b/scrapy/core/downloader/__init__.py index 949dacbc8..f5f8be6e8 100644 --- a/scrapy/core/downloader/__init__.py +++ b/scrapy/core/downloader/__init__.py @@ -1,4 +1,3 @@ -from __future__ import absolute_import import random import warnings from time import time diff --git a/scrapy/core/downloader/handlers/http.py b/scrapy/core/downloader/handlers/http.py index ac4b867c3..6111e132a 100644 --- a/scrapy/core/downloader/handlers/http.py +++ b/scrapy/core/downloader/handlers/http.py @@ -1,3 +1,2 @@ -from __future__ import absolute_import from .http10 import HTTP10DownloadHandler from .http11 import HTTP11DownloadHandler as HTTPDownloadHandler diff --git a/scrapy/downloadermiddlewares/ajaxcrawl.py b/scrapy/downloadermiddlewares/ajaxcrawl.py index 72715dba7..78b802673 100644 --- a/scrapy/downloadermiddlewares/ajaxcrawl.py +++ b/scrapy/downloadermiddlewares/ajaxcrawl.py @@ -1,5 +1,4 @@ # -*- coding: utf-8 -*- -from __future__ import absolute_import import re import logging diff --git a/scrapy/dupefilters.py b/scrapy/dupefilters.py index 0bcdd3495..f8802eb7d 100644 --- a/scrapy/dupefilters.py +++ b/scrapy/dupefilters.py @@ -1,4 +1,3 @@ -from __future__ import print_function import os import logging diff --git a/scrapy/extensions/httpcache.py b/scrapy/extensions/httpcache.py index f3fabf710..b1ed0b9f8 100644 --- a/scrapy/extensions/httpcache.py +++ b/scrapy/extensions/httpcache.py @@ -1,5 +1,3 @@ -from __future__ import print_function - import gzip import logging import os diff --git a/scrapy/pipelines/media.py b/scrapy/pipelines/media.py index 95dca9a3f..c174addf9 100644 --- a/scrapy/pipelines/media.py +++ b/scrapy/pipelines/media.py @@ -1,5 +1,3 @@ -from __future__ import print_function - import functools import logging from collections import defaultdict diff --git a/scrapy/responsetypes.py b/scrapy/responsetypes.py index de62276c8..b64fbbd42 100644 --- a/scrapy/responsetypes.py +++ b/scrapy/responsetypes.py @@ -2,7 +2,6 @@ This module implements a class which returns the appropriate Response class based on different criteria. """ -from __future__ import absolute_import from mimetypes import MimeTypes from pkgutil import get_data from io import StringIO diff --git a/scrapy/shell.py b/scrapy/shell.py index 80b625633..a649d555f 100644 --- a/scrapy/shell.py +++ b/scrapy/shell.py @@ -3,8 +3,6 @@ See documentation in docs/topics/shell.rst """ -from __future__ import print_function - import os import signal import warnings diff --git a/scrapy/signalmanager.py b/scrapy/signalmanager.py index 296d27ed8..d474f1806 100644 --- a/scrapy/signalmanager.py +++ b/scrapy/signalmanager.py @@ -1,4 +1,3 @@ -from __future__ import absolute_import from pydispatch import dispatcher from scrapy.utils import signal as _signal diff --git a/scrapy/spiderloader.py b/scrapy/spiderloader.py index 7478faa78..3beca4060 100644 --- a/scrapy/spiderloader.py +++ b/scrapy/spiderloader.py @@ -1,5 +1,4 @@ # -*- coding: utf-8 -*- -from __future__ import absolute_import from collections import defaultdict import traceback import warnings diff --git a/scrapy/utils/boto.py b/scrapy/utils/boto.py index c8fc911bb..46816b54d 100644 --- a/scrapy/utils/boto.py +++ b/scrapy/utils/boto.py @@ -1,7 +1,5 @@ """Boto/botocore helpers""" -from __future__ import absolute_import - from scrapy.exceptions import NotConfigured diff --git a/scrapy/utils/display.py b/scrapy/utils/display.py index f6a6c4645..536de6b88 100644 --- a/scrapy/utils/display.py +++ b/scrapy/utils/display.py @@ -2,7 +2,6 @@ pprint and pformat wrappers with colorization support """ -from __future__ import print_function import sys from pprint import pformat as pformat_ diff --git a/scrapy/utils/engine.py b/scrapy/utils/engine.py index 11dd36d91..36ef8626a 100644 --- a/scrapy/utils/engine.py +++ b/scrapy/utils/engine.py @@ -1,6 +1,5 @@ """Some debugging functions for working with the Scrapy engine""" -from __future__ import print_function from time import time # used in global tests code def get_engine_status(engine): diff --git a/scrapy/utils/ossignal.py b/scrapy/utils/ossignal.py index f87d5a803..7a7aec9be 100644 --- a/scrapy/utils/ossignal.py +++ b/scrapy/utils/ossignal.py @@ -1,5 +1,3 @@ - -from __future__ import absolute_import import signal from twisted.internet import reactor diff --git a/scrapy/utils/request.py b/scrapy/utils/request.py index 63d0ae772..45f1ef17e 100644 --- a/scrapy/utils/request.py +++ b/scrapy/utils/request.py @@ -3,7 +3,6 @@ This module provides some useful functions for working with scrapy.http.Request objects """ -from __future__ import print_function import hashlib import weakref from six.moves.urllib.parse import urlunparse diff --git a/scrapy/utils/test.py b/scrapy/utils/test.py index 4b935c51b..febaa4dcc 100644 --- a/scrapy/utils/test.py +++ b/scrapy/utils/test.py @@ -2,7 +2,6 @@ This module contains some assorted functions used in tests """ -from __future__ import absolute_import import os from importlib import import_module diff --git a/scrapy/utils/testproc.py b/scrapy/utils/testproc.py index f268e91ff..0f15cf60a 100644 --- a/scrapy/utils/testproc.py +++ b/scrapy/utils/testproc.py @@ -1,4 +1,3 @@ -from __future__ import absolute_import import sys import os diff --git a/scrapy/utils/testsite.py b/scrapy/utils/testsite.py index e50a989b3..05d06d53b 100644 --- a/scrapy/utils/testsite.py +++ b/scrapy/utils/testsite.py @@ -1,4 +1,3 @@ -from __future__ import print_function from six.moves.urllib.parse import urljoin from twisted.internet import reactor diff --git a/scrapy/utils/trackref.py b/scrapy/utils/trackref.py index eed14c5a1..78389e464 100644 --- a/scrapy/utils/trackref.py +++ b/scrapy/utils/trackref.py @@ -9,7 +9,6 @@ and no performance penalty at all when disabled (as object_ref becomes just an alias to object in that case). """ -from __future__ import print_function import weakref from time import time from operator import itemgetter From c0bfaef37abe029c98072e7b1a1bf346ab79016c Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sat, 2 Nov 2019 23:27:04 -0300 Subject: [PATCH 193/387] Remove __future__ imports from tests --- tests/mockserver.py | 1 - tests/test_downloadermiddleware_httpcache.py | 1 - tests/test_downloadermiddleware_robotstxt.py | 2 -- tests/test_engine.py | 1 - tests/test_exporters.py | 1 - tests/test_feedexport.py | 1 - tests/test_pipeline_media.py | 2 -- tests/test_utils_deprecate.py | 1 - tests/test_utils_log.py | 1 - tests/test_utils_request.py | 1 - 10 files changed, 12 deletions(-) diff --git a/tests/mockserver.py b/tests/mockserver.py index 77908284b..8aff7b525 100644 --- a/tests/mockserver.py +++ b/tests/mockserver.py @@ -1,4 +1,3 @@ -from __future__ import print_function import sys, time, random, os, json from six.moves.urllib.parse import urlencode from subprocess import Popen, PIPE diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index 950664ffe..db0843b57 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -1,4 +1,3 @@ -from __future__ import print_function import time import tempfile import shutil diff --git a/tests/test_downloadermiddleware_robotstxt.py b/tests/test_downloadermiddleware_robotstxt.py index 8266bf35f..a1645ed96 100644 --- a/tests/test_downloadermiddleware_robotstxt.py +++ b/tests/test_downloadermiddleware_robotstxt.py @@ -1,6 +1,4 @@ # -*- coding: utf-8 -*- -from __future__ import absolute_import - from unittest import mock from twisted.internet import reactor, error diff --git a/tests/test_engine.py b/tests/test_engine.py index 30150391a..d5b911a40 100644 --- a/tests/test_engine.py +++ b/tests/test_engine.py @@ -10,7 +10,6 @@ module with the ``runserver`` argument:: python test_engine.py runserver """ -from __future__ import print_function import sys, os, re from six.moves.urllib.parse import urlparse diff --git a/tests/test_exporters.py b/tests/test_exporters.py index 0046c5666..7880c5bf8 100644 --- a/tests/test_exporters.py +++ b/tests/test_exporters.py @@ -1,4 +1,3 @@ -from __future__ import absolute_import import re import json import marshal diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index abe2ab557..b13b16b41 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -1,4 +1,3 @@ -from __future__ import absolute_import import os import csv import json diff --git a/tests/test_pipeline_media.py b/tests/test_pipeline_media.py index ad2618ec9..70f11466b 100644 --- a/tests/test_pipeline_media.py +++ b/tests/test_pipeline_media.py @@ -1,5 +1,3 @@ -from __future__ import print_function - import sys from testfixtures import LogCapture diff --git a/tests/test_utils_deprecate.py b/tests/test_utils_deprecate.py index ce04e7f29..159ef8f25 100644 --- a/tests/test_utils_deprecate.py +++ b/tests/test_utils_deprecate.py @@ -1,5 +1,4 @@ # -*- coding: utf-8 -*- -from __future__ import absolute_import import inspect import unittest from unittest import mock diff --git a/tests/test_utils_log.py b/tests/test_utils_log.py index 742e04803..2c23f3616 100644 --- a/tests/test_utils_log.py +++ b/tests/test_utils_log.py @@ -1,5 +1,4 @@ # -*- coding: utf-8 -*- -from __future__ import print_function import sys import logging import unittest diff --git a/tests/test_utils_request.py b/tests/test_utils_request.py index 625a32048..446497086 100644 --- a/tests/test_utils_request.py +++ b/tests/test_utils_request.py @@ -1,4 +1,3 @@ -from __future__ import print_function import unittest from scrapy.http import Request from scrapy.utils.request import request_fingerprint, _fingerprint_cache, \ From df00389c16fc8e74c5ab7737e5bffcbde6cc011b Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sat, 2 Nov 2019 23:48:11 -0300 Subject: [PATCH 194/387] Remove six.moves occurrences --- scrapy/_monkeypatches.py | 2 +- scrapy/commands/bench.py | 3 +-- scrapy/core/downloader/handlers/ftp.py | 2 +- scrapy/core/downloader/handlers/http11.py | 4 ++-- scrapy/core/downloader/handlers/s3.py | 2 +- scrapy/core/downloader/webclient.py | 2 +- scrapy/downloadermiddlewares/httpproxy.py | 5 ++--- scrapy/downloadermiddlewares/redirect.py | 3 ++- scrapy/exporters.py | 7 ++++--- scrapy/extensions/feedexport.py | 2 +- scrapy/extensions/httpcache.py | 2 +- scrapy/extensions/spiderstate.py | 3 ++- scrapy/http/cookies.py | 5 ++--- scrapy/http/request/form.py | 4 ++-- scrapy/http/request/rpc.py | 2 +- scrapy/http/response/__init__.py | 2 +- scrapy/http/response/text.py | 4 ++-- scrapy/linkextractors/__init__.py | 2 +- scrapy/linkextractors/htmlparser.py | 6 +++--- scrapy/linkextractors/lxmlhtml.py | 4 ++-- scrapy/linkextractors/regex.py | 3 ++- scrapy/linkextractors/sgml.py | 4 ++-- scrapy/mail.py | 14 +++++++------- scrapy/pipelines/files.py | 10 +++++----- scrapy/robotstxt.py | 2 +- scrapy/spidermiddlewares/referer.py | 2 +- scrapy/squeues.py | 2 +- scrapy/utils/benchserver.py | 3 ++- scrapy/utils/curl.py | 4 ++-- scrapy/utils/httpobj.py | 2 +- scrapy/utils/project.py | 6 +++--- scrapy/utils/python.py | 2 +- scrapy/utils/request.py | 6 +++--- scrapy/utils/sitemap.py | 3 ++- scrapy/utils/testsite.py | 2 +- scrapy/utils/url.py | 2 +- 36 files changed, 68 insertions(+), 65 deletions(-) diff --git a/scrapy/_monkeypatches.py b/scrapy/_monkeypatches.py index 1f8067b35..f74f89bda 100644 --- a/scrapy/_monkeypatches.py +++ b/scrapy/_monkeypatches.py @@ -1,4 +1,4 @@ -from six.moves import copyreg +import copyreg # Undo what Twisted's perspective broker adds to pickle register diff --git a/scrapy/commands/bench.py b/scrapy/commands/bench.py index 90c8d56a2..7bbe362e7 100644 --- a/scrapy/commands/bench.py +++ b/scrapy/commands/bench.py @@ -1,8 +1,7 @@ import sys import time import subprocess - -from six.moves.urllib.parse import urlencode +from urllib.parse import urlencode import scrapy from scrapy.commands import ScrapyCommand diff --git a/scrapy/core/downloader/handlers/ftp.py b/scrapy/core/downloader/handlers/ftp.py index 806a537d4..2116a5a44 100644 --- a/scrapy/core/downloader/handlers/ftp.py +++ b/scrapy/core/downloader/handlers/ftp.py @@ -30,7 +30,7 @@ In case of status 200 request, response.headers will come with two keys: import re from io import BytesIO -from six.moves.urllib.parse import unquote +from urllib.parse import unquote from twisted.internet import reactor from twisted.protocols.ftp import FTPClient, CommandFailed diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index 7d917cb74..63dedc19b 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -2,10 +2,10 @@ import re import logging +import warnings from io import BytesIO from time import time -import warnings -from six.moves.urllib.parse import urldefrag +from urllib.parse import urldefrag from zope.interface import implementer from twisted.internet import defer, reactor, protocol diff --git a/scrapy/core/downloader/handlers/s3.py b/scrapy/core/downloader/handlers/s3.py index d8bbdd326..85b733228 100644 --- a/scrapy/core/downloader/handlers/s3.py +++ b/scrapy/core/downloader/handlers/s3.py @@ -1,4 +1,4 @@ -from six.moves.urllib.parse import unquote +from urllib.parse import unquote from scrapy.exceptions import NotConfigured from scrapy.utils.httpobj import urlparse_cached diff --git a/scrapy/core/downloader/webclient.py b/scrapy/core/downloader/webclient.py index 3a5890ed0..9699da109 100644 --- a/scrapy/core/downloader/webclient.py +++ b/scrapy/core/downloader/webclient.py @@ -1,5 +1,5 @@ from time import time -from six.moves.urllib.parse import urlparse, urlunparse, urldefrag +from urllib.parse import urlparse, urlunparse, urldefrag from twisted.web.client import HTTPClientFactory from twisted.web.http import HTTPClient diff --git a/scrapy/downloadermiddlewares/httpproxy.py b/scrapy/downloadermiddlewares/httpproxy.py index 2212d9688..814ce78fe 100644 --- a/scrapy/downloadermiddlewares/httpproxy.py +++ b/scrapy/downloadermiddlewares/httpproxy.py @@ -1,7 +1,6 @@ import base64 -from six.moves.urllib.parse import unquote, urlunparse -from six.moves.urllib.request import getproxies, proxy_bypass -from urllib.request import _parse_proxy +from urllib.parse import unquote, urlunparse +from urllib.request import getproxies, proxy_bypass, _parse_proxy from scrapy.exceptions import NotConfigured from scrapy.utils.httpobj import urlparse_cached diff --git a/scrapy/downloadermiddlewares/redirect.py b/scrapy/downloadermiddlewares/redirect.py index b73f864dd..77cb5aa94 100644 --- a/scrapy/downloadermiddlewares/redirect.py +++ b/scrapy/downloadermiddlewares/redirect.py @@ -1,5 +1,5 @@ import logging -from six.moves.urllib.parse import urljoin, urlparse +from urllib.parse import urljoin, urlparse from w3lib.url import safe_url_string @@ -7,6 +7,7 @@ from scrapy.http import HtmlResponse from scrapy.utils.response import get_meta_refresh from scrapy.exceptions import IgnoreRequest, NotConfigured + logger = logging.getLogger(__name__) diff --git a/scrapy/exporters.py b/scrapy/exporters.py index f276c28e8..d9531e67c 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -7,15 +7,16 @@ import io import sys import pprint import marshal -import six -from six.moves import cPickle as pickle +import warnings +from import pickle from xml.sax.saxutils import XMLGenerator +import six + from scrapy.utils.serialize import ScrapyJSONEncoder from scrapy.utils.python import to_bytes, to_unicode, is_listlike from scrapy.item import BaseItem from scrapy.exceptions import ScrapyDeprecationWarning -import warnings __all__ = ['BaseItemExporter', 'PprintItemExporter', 'PickleItemExporter', diff --git a/scrapy/extensions/feedexport.py b/scrapy/extensions/feedexport.py index 07ffd3476..0b854e6f3 100644 --- a/scrapy/extensions/feedexport.py +++ b/scrapy/extensions/feedexport.py @@ -10,7 +10,7 @@ import logging import posixpath from tempfile import NamedTemporaryFile from datetime import datetime -from six.moves.urllib.parse import urlparse, unquote +from urllib.parse import urlparse, unquote from ftplib import FTP from zope.interface import Interface, implementer diff --git a/scrapy/extensions/httpcache.py b/scrapy/extensions/httpcache.py index b1ed0b9f8..b98ec218f 100644 --- a/scrapy/extensions/httpcache.py +++ b/scrapy/extensions/httpcache.py @@ -1,13 +1,13 @@ import gzip import logging import os +import pickle from email.utils import mktime_tz, parsedate_tz from importlib import import_module from time import time from warnings import warn from weakref import WeakKeyDictionary -from six.moves import cPickle as pickle from w3lib.http import headers_raw_to_dict, headers_dict_to_raw from scrapy.exceptions import ScrapyDeprecationWarning diff --git a/scrapy/extensions/spiderstate.py b/scrapy/extensions/spiderstate.py index 2220cbd8f..2c8e46914 100644 --- a/scrapy/extensions/spiderstate.py +++ b/scrapy/extensions/spiderstate.py @@ -1,10 +1,11 @@ import os -from six.moves import cPickle as pickle +import pickle from scrapy import signals from scrapy.exceptions import NotConfigured from scrapy.utils.job import job_dir + class SpiderState(object): """Store and load spider state during a scraping job""" diff --git a/scrapy/http/cookies.py b/scrapy/http/cookies.py index 4532c3ab7..c39de0b52 100644 --- a/scrapy/http/cookies.py +++ b/scrapy/http/cookies.py @@ -1,7 +1,6 @@ import time -from six.moves.http_cookiejar import ( - CookieJar as _CookieJar, DefaultCookiePolicy, IPV4_RE -) +from http.cookiejar import CookieJar as _CookieJar, DefaultCookiePolicy, IPV4_RE + from scrapy.utils.httpobj import urlparse_cached from scrapy.utils.python import to_unicode diff --git a/scrapy/http/request/form.py b/scrapy/http/request/form.py index b6feede07..d2bcf77b7 100644 --- a/scrapy/http/request/form.py +++ b/scrapy/http/request/form.py @@ -5,10 +5,10 @@ This module implements the FormRequest class which is a more convenient class See documentation in docs/topics/request-response.rst """ -import six -from six.moves.urllib.parse import urljoin, urlencode +from urllib.parse import urljoin, urlencode import lxml.html +import six from parsel.selector import create_root_node from w3lib.html import strip_html5_whitespace diff --git a/scrapy/http/request/rpc.py b/scrapy/http/request/rpc.py index bd09f7534..811d3ad6b 100644 --- a/scrapy/http/request/rpc.py +++ b/scrapy/http/request/rpc.py @@ -4,7 +4,7 @@ This module implements the XmlRpcRequest class which is a more convenient class See documentation in docs/topics/request-response.rst """ -from six.moves import xmlrpc_client as xmlrpclib +import xmlrpc.client as xmlrpclib from scrapy.http.request import Request from scrapy.utils.python import get_func_args diff --git a/scrapy/http/response/__init__.py b/scrapy/http/response/__init__.py index a81404afb..64e9c6c20 100644 --- a/scrapy/http/response/__init__.py +++ b/scrapy/http/response/__init__.py @@ -4,7 +4,7 @@ responses in Scrapy. See documentation in docs/topics/request-response.rst """ -from six.moves.urllib.parse import urljoin +from urllib.parse import urljoin from scrapy.http.request import Request from scrapy.http.headers import Headers diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 37f450e54..69bcba577 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -5,10 +5,10 @@ discovering (through HTTP headers) to base Response class. See documentation in docs/topics/request-response.rst """ -import six -from six.moves.urllib.parse import urljoin +from urllib.parse import urljoin import parsel +import six from w3lib.encoding import html_to_unicode, resolve_encoding, \ html_body_declared_encoding, http_content_type_encoding from w3lib.html import strip_html5_whitespace diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index ebf3cd7d8..594c7264d 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -6,8 +6,8 @@ This package contains a collection of Link Extractors. For more info see docs/topics/link-extractors.rst """ import re +from urllib.parse import urlparse -from six.moves.urllib.parse import urlparse from parsel.csstranslator import HTMLTranslator from w3lib.url import canonicalize_url diff --git a/scrapy/linkextractors/htmlparser.py b/scrapy/linkextractors/htmlparser.py index 27978a8a1..623732cee 100644 --- a/scrapy/linkextractors/htmlparser.py +++ b/scrapy/linkextractors/htmlparser.py @@ -2,10 +2,10 @@ HTMLParser-based link extractor """ import warnings -import six -from six.moves.html_parser import HTMLParser -from six.moves.urllib.parse import urljoin +from html.parser import HTMLParser +from urllib.parse import urljoin +import six from w3lib.url import safe_url_string from w3lib.html import strip_html5_whitespace diff --git a/scrapy/linkextractors/lxmlhtml.py b/scrapy/linkextractors/lxmlhtml.py index 890c019c8..5b7e709de 100644 --- a/scrapy/linkextractors/lxmlhtml.py +++ b/scrapy/linkextractors/lxmlhtml.py @@ -1,9 +1,9 @@ """ Link extractor based on lxml.html """ -import six -from six.moves.urllib.parse import urljoin +from urllib.parse import urljoin +import six import lxml.etree as etree from w3lib.html import strip_html5_whitespace from w3lib.url import canonicalize_url diff --git a/scrapy/linkextractors/regex.py b/scrapy/linkextractors/regex.py index e689b4727..f96db256b 100644 --- a/scrapy/linkextractors/regex.py +++ b/scrapy/linkextractors/regex.py @@ -1,11 +1,12 @@ import re -from six.moves.urllib.parse import urljoin +from urllib.parse import urljoin from w3lib.html import remove_tags, replace_entities, replace_escape_chars, get_base_url from scrapy.link import Link from .sgml import SgmlLinkExtractor + linkre = re.compile( "<a\s.*?href=(\"[.#]+?\"|\'[.#]+?\'|[^\s]+?)(>|\s.*?>)(.*?)<[/ ]?a>", re.DOTALL | re.IGNORECASE) diff --git a/scrapy/linkextractors/sgml.py b/scrapy/linkextractors/sgml.py index 8940a4d77..98bed15e9 100644 --- a/scrapy/linkextractors/sgml.py +++ b/scrapy/linkextractors/sgml.py @@ -1,11 +1,11 @@ """ SGMLParser-based Link extractors """ -import six -from six.moves.urllib.parse import urljoin import warnings +from urllib.parse import urljoin from sgmllib import SGMLParser +import six from w3lib.url import safe_url_string, canonicalize_url from w3lib.html import strip_html5_whitespace diff --git a/scrapy/mail.py b/scrapy/mail.py index d24de2212..891bb5e09 100644 --- a/scrapy/mail.py +++ b/scrapy/mail.py @@ -3,21 +3,21 @@ Mail sending helpers See documentation in docs/topics/email.rst """ -from io import BytesIO import logging - -from email.utils import COMMASPACE, formatdate -from six.moves.email_mime_multipart import MIMEMultipart -from six.moves.email_mime_text import MIMEText -from six.moves.email_mime_base import MIMEBase -from email.mime.nonmultipart import MIMENonMultipart from email import encoders as Encoders +from email.mime.base import MIMEBase +from email.mime.multipart import MIMEMultipart +from email.mime.nonmultipart import MIMENonMultipart +from email.mime.text import MIMEText +from email.utils import COMMASPACE, formatdate +from io import BytesIO from twisted.internet import defer, reactor, ssl from scrapy.utils.misc import arg_to_iter from scrapy.utils.python import to_bytes + logger = logging.getLogger(__name__) diff --git a/scrapy/pipelines/files.py b/scrapy/pipelines/files.py index 8d74c5011..432d4c182 100644 --- a/scrapy/pipelines/files.py +++ b/scrapy/pipelines/files.py @@ -5,17 +5,17 @@ See documentation in topics/media-pipeline.rst """ import functools import hashlib -from io import BytesIO +import logging import mimetypes import os import os.path import time -import logging -from email.utils import parsedate_tz, mktime_tz -from six.moves.urllib.parse import urlparse from collections import defaultdict -import six +from email.utils import parsedate_tz, mktime_tz +from io import BytesIO +from urllib.parse import urlparse +import six from twisted.internet import defer, threads from scrapy.pipelines.media import MediaPipeline diff --git a/scrapy/robotstxt.py b/scrapy/robotstxt.py index 95a8c09b8..397924110 100644 --- a/scrapy/robotstxt.py +++ b/scrapy/robotstxt.py @@ -53,7 +53,7 @@ class RobotParser(with_metaclass(ABCMeta)): class PythonRobotParser(RobotParser): def __init__(self, robotstxt_body, spider): - from six.moves.urllib_robotparser import RobotFileParser + from urllib.robotparser import RobotFileParser self.spider = spider robotstxt_body = decode_robotstxt(robotstxt_body, spider, to_native_str_type=True) self.rp = RobotFileParser() diff --git a/scrapy/spidermiddlewares/referer.py b/scrapy/spidermiddlewares/referer.py index c76e4d5a2..dce2b3598 100644 --- a/scrapy/spidermiddlewares/referer.py +++ b/scrapy/spidermiddlewares/referer.py @@ -2,8 +2,8 @@ RefererMiddleware: populates Request referer field, based on the Response which originated it. """ -from six.moves.urllib.parse import urlparse import warnings +from urllib.parse import urlparse from w3lib.url import safe_url_string diff --git a/scrapy/squeues.py b/scrapy/squeues.py index 30cc926e5..d5d3be67e 100644 --- a/scrapy/squeues.py +++ b/scrapy/squeues.py @@ -3,7 +3,7 @@ Scheduler queues """ import marshal -from six.moves import cPickle as pickle +import pickle from queuelib import queue diff --git a/scrapy/utils/benchserver.py b/scrapy/utils/benchserver.py index 5bbda6e27..cdbe21942 100644 --- a/scrapy/utils/benchserver.py +++ b/scrapy/utils/benchserver.py @@ -1,5 +1,6 @@ import random -from six.moves.urllib.parse import urlencode +from urllib.parse import urlencode + from twisted.web.server import Site from twisted.web.resource import Resource from twisted.internet import reactor diff --git a/scrapy/utils/curl.py b/scrapy/utils/curl.py index b3fd0a497..a0a47e473 100644 --- a/scrapy/utils/curl.py +++ b/scrapy/utils/curl.py @@ -1,9 +1,9 @@ import argparse import warnings from shlex import split +from http.cookies import SimpleCookie +from urllib.parse import urlparse -from six.moves.http_cookies import SimpleCookie -from six.moves.urllib.parse import urlparse from six import string_types, iteritems from w3lib.http import basic_auth_header diff --git a/scrapy/utils/httpobj.py b/scrapy/utils/httpobj.py index b4c929b0e..54ffe086d 100644 --- a/scrapy/utils/httpobj.py +++ b/scrapy/utils/httpobj.py @@ -1,8 +1,8 @@ """Helper functions for scrapy.http objects (Request, Response)""" import weakref +from urllib.parse import urlparse -from six.moves.urllib.parse import urlparse _urlparse_cache = weakref.WeakKeyDictionary() def urlparse_cached(request_or_response): diff --git a/scrapy/utils/project.py b/scrapy/utils/project.py index 1cbda141a..f28c2eaa1 100644 --- a/scrapy/utils/project.py +++ b/scrapy/utils/project.py @@ -1,5 +1,5 @@ import os -from six.moves import cPickle as pickle +import pickle import warnings from importlib import import_module @@ -7,8 +7,8 @@ from os.path import join, dirname, abspath, isabs, exists from scrapy.utils.conf import closest_scrapy_cfg, get_config, init_env from scrapy.settings import Settings -from scrapy.exceptions import NotConfigured -from scrapy.exceptions import ScrapyDeprecationWarning +from scrapy.exceptions import NotConfigured, ScrapyDeprecationWarning + ENVVAR = 'SCRAPY_SETTINGS_MODULE' DATADIR_CFG_SECTION = 'datadir' diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index 974abaeb1..cb8ac3c8e 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -65,7 +65,7 @@ def is_listlike(x): True >>> is_listlike((x for x in range(3))) True - >>> is_listlike(six.moves.xrange(5)) + >>> is_listlike(range(5)) True """ return hasattr(x, "__iter__") and not isinstance(x, (six.text_type, bytes)) diff --git a/scrapy/utils/request.py b/scrapy/utils/request.py index 45f1ef17e..c5e877acd 100644 --- a/scrapy/utils/request.py +++ b/scrapy/utils/request.py @@ -5,13 +5,13 @@ scrapy.http.Request objects import hashlib import weakref -from six.moves.urllib.parse import urlunparse +from urllib.parse import urlunparse from w3lib.http import basic_auth_header -from scrapy.utils.python import to_bytes, to_unicode - from w3lib.url import canonicalize_url + from scrapy.utils.httpobj import urlparse_cached +from scrapy.utils.python import to_bytes, to_unicode _fingerprint_cache = weakref.WeakKeyDictionary() diff --git a/scrapy/utils/sitemap.py b/scrapy/utils/sitemap.py index 4742b3e13..2f10cf4de 100644 --- a/scrapy/utils/sitemap.py +++ b/scrapy/utils/sitemap.py @@ -5,8 +5,9 @@ Note: The main purpose of this module is to provide support for the SitemapSpider, its API is subject to change without notice. """ +from urllib.parse import urljoin + import lxml.etree -from six.moves.urllib.parse import urljoin class Sitemap(object): diff --git a/scrapy/utils/testsite.py b/scrapy/utils/testsite.py index 05d06d53b..6f5c21624 100644 --- a/scrapy/utils/testsite.py +++ b/scrapy/utils/testsite.py @@ -1,4 +1,4 @@ -from six.moves.urllib.parse import urljoin +from urllib.parse import urljoin from twisted.internet import reactor from twisted.web import server, resource, static, util diff --git a/scrapy/utils/url.py b/scrapy/utils/url.py index b3a4be007..4b48868fe 100644 --- a/scrapy/utils/url.py +++ b/scrapy/utils/url.py @@ -7,7 +7,7 @@ to the w3lib.url module. Always import those from there instead. """ import posixpath import re -from six.moves.urllib.parse import (ParseResult, urldefrag, urlparse, urlunparse) +from urllib.parse import ParseResult, urldefrag, urlparse, urlunparse # scrapy.utils.url was moved to w3lib.url and import * ensures this # move doesn't break old code From 5ab0f189ce33f235b847f96b66a243d808b5ce1f Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 00:01:09 -0300 Subject: [PATCH 195/387] Remove six.moves occurrences from tests --- tests/mockserver.py | 8 ++++++-- tests/spiders.py | 2 +- tests/test_crawl.py | 2 +- tests/test_engine.py | 6 ++++-- tests/test_exporters.py | 2 +- tests/test_feedexport.py | 6 +++--- tests/test_http_cookies.py | 2 +- tests/test_http_request.py | 8 +++----- tests/test_pipeline_files.py | 4 ++-- tests/test_proxy_connect.py | 6 +++--- tests/test_spidermiddleware_offsite.py | 9 ++++----- tests/test_spidermiddleware_referer.py | 2 +- tests/test_urlparse_monkeypatches.py | 2 +- tests/test_utils_datatypes.py | 8 -------- tests/test_utils_defer.py | 6 ++---- tests/test_utils_httpobj.py | 3 ++- tests/test_utils_response.py | 3 ++- tests/test_utils_url.py | 3 ++- 18 files changed, 39 insertions(+), 43 deletions(-) diff --git a/tests/mockserver.py b/tests/mockserver.py index 8aff7b525..15b1b24f5 100644 --- a/tests/mockserver.py +++ b/tests/mockserver.py @@ -1,6 +1,10 @@ -import sys, time, random, os, json -from six.moves.urllib.parse import urlencode +import json +import os +import random +import sys +import time from subprocess import Popen, PIPE +from urllib.parse import urlencode from OpenSSL import SSL from twisted.web.server import Site, NOT_DONE_YET diff --git a/tests/spiders.py b/tests/spiders.py index 7816bf7c7..72a428c50 100644 --- a/tests/spiders.py +++ b/tests/spiders.py @@ -3,7 +3,7 @@ Some spiders used for testing and benchmarking """ import time -from six.moves.urllib.parse import urlencode +from urllib.parse import urlencode from scrapy.spiders import Spider from scrapy.http import Request diff --git a/tests/test_crawl.py b/tests/test_crawl.py index 3fc13eeb7..3307899b7 100644 --- a/tests/test_crawl.py +++ b/tests/test_crawl.py @@ -140,7 +140,7 @@ class CrawlTestCase(TestCase): def test_unbounded_response(self): # Completeness of responses without Content-Length or Transfer-Encoding # can not be determined, we treat them as valid but flagged as "partial" - from six.moves.urllib.parse import urlencode + from urllib.parse import urlencode query = urlencode({'raw': '''\ HTTP/1.1 200 OK Server: Apache-Coyote/1.1 diff --git a/tests/test_engine.py b/tests/test_engine.py index d5b911a40..002c4e6bb 100644 --- a/tests/test_engine.py +++ b/tests/test_engine.py @@ -10,8 +10,10 @@ module with the ``runserver`` argument:: python test_engine.py runserver """ -import sys, os, re -from six.moves.urllib.parse import urlparse +import os +import re +import sys +from urllib.parse import urlparse from twisted.internet import reactor, defer from twisted.web import server, static, util diff --git a/tests/test_exporters.py b/tests/test_exporters.py index 7880c5bf8..f151a1285 100644 --- a/tests/test_exporters.py +++ b/tests/test_exporters.py @@ -1,11 +1,11 @@ import re import json import marshal +import pickle import tempfile import unittest from io import BytesIO from datetime import datetime -from six.moves import cPickle as pickle import lxml.etree import six diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index b13b16b41..7a71669c3 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -2,13 +2,13 @@ import os import csv import json import warnings -from io import BytesIO import tempfile import shutil import string +from io import BytesIO from unittest import mock -from six.moves.urllib.parse import urljoin, urlparse, quote -from six.moves.urllib.request import pathname2url +from urllib.parse import urljoin, urlparse, quote +from urllib.request import pathname2url from zope.interface.verify import verifyObject from twisted.trial import unittest diff --git a/tests/test_http_cookies.py b/tests/test_http_cookies.py index 0a9ed500a..45ddb42ba 100644 --- a/tests/test_http_cookies.py +++ b/tests/test_http_cookies.py @@ -1,4 +1,4 @@ -from six.moves.urllib.parse import urlparse +from urllib.parse import urlparse from unittest import TestCase from scrapy.http import Request, Response diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 807265981..62d2847d7 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -3,12 +3,10 @@ import cgi import unittest import re import json -from unittest import mock -from urllib.parse import unquote_to_bytes +import xmlrpc.client as xmlrpclib import warnings - -from six.moves import xmlrpc_client as xmlrpclib -from six.moves.urllib.parse import urlparse, parse_qs, unquote +from unittest import mock +from urllib.parse import parse_qs, unquote, unquote_to_bytes, urlparse from scrapy.http import Request, FormRequest, XmlRpcRequest, JsonRequest, Headers, HtmlResponse from scrapy.utils.python import to_bytes, to_unicode diff --git a/tests/test_pipeline_files.py b/tests/test_pipeline_files.py index bd40e4103..dede4bf12 100644 --- a/tests/test_pipeline_files.py +++ b/tests/test_pipeline_files.py @@ -4,9 +4,9 @@ import time from tempfile import mkdtemp from shutil import rmtree from unittest import mock -from six.moves.urllib.parse import urlparse -from six import BytesIO +from urllib.parse import urlparse +from six import BytesIO from twisted.trial import unittest from twisted.internet import defer diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index ae1236bcb..2b62a378d 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -1,15 +1,15 @@ import json import os import time - -from six.moves.urllib.parse import urlsplit, urlunsplit +from urllib.parse import urlsplit, urlunsplit from threading import Thread + from libmproxy import controller, proxy from netlib import http_auth from testfixtures import LogCapture - from twisted.internet import defer from twisted.trial.unittest import TestCase + from scrapy.utils.test import get_crawler from scrapy.http import Request from tests.spiders import SimpleSpider, SingleRequestSpider diff --git a/tests/test_spidermiddleware_offsite.py b/tests/test_spidermiddleware_offsite.py index 7e4af0d4c..5833591af 100644 --- a/tests/test_spidermiddleware_offsite.py +++ b/tests/test_spidermiddleware_offsite.py @@ -1,13 +1,12 @@ from unittest import TestCase - -from six.moves.urllib.parse import urlparse +from urllib.parse import urlparse +import warnings from scrapy.http import Response, Request from scrapy.spiders import Spider -from scrapy.spidermiddlewares.offsite import OffsiteMiddleware -from scrapy.spidermiddlewares.offsite import URLWarning +from scrapy.spidermiddlewares.offsite import OffsiteMiddleware, URLWarning from scrapy.utils.test import get_crawler -import warnings + class TestOffsiteMiddleware(TestCase): diff --git a/tests/test_spidermiddleware_referer.py b/tests/test_spidermiddleware_referer.py index 21439c20e..23c38d9b1 100644 --- a/tests/test_spidermiddleware_referer.py +++ b/tests/test_spidermiddleware_referer.py @@ -1,4 +1,4 @@ -from six.moves.urllib.parse import urlparse +from urllib.parse import urlparse from unittest import TestCase import warnings diff --git a/tests/test_urlparse_monkeypatches.py b/tests/test_urlparse_monkeypatches.py index 22e39821c..bea0cf3e5 100644 --- a/tests/test_urlparse_monkeypatches.py +++ b/tests/test_urlparse_monkeypatches.py @@ -1,4 +1,4 @@ -from six.moves.urllib.parse import urlparse +from urllib.parse import urlparse import unittest diff --git a/tests/test_utils_datatypes.py b/tests/test_utils_datatypes.py index 47877f555..8782e4c53 100644 --- a/tests/test_utils_datatypes.py +++ b/tests/test_utils_datatypes.py @@ -192,14 +192,6 @@ class SequenceExcludeTest(unittest.TestCase): self.assertIn(20, d) self.assertNotIn(15, d) - def test_six_range(self): - import six.moves - seq = six.moves.range(10**3, 10**6) - d = SequenceExclude(seq) - self.assertIn(10**2, d) - self.assertIn(10**7, d) - self.assertNotIn(10**4, d) - def test_range_step(self): seq = range(10, 20, 3) d = SequenceExclude(seq) diff --git a/tests/test_utils_defer.py b/tests/test_utils_defer.py index 003bb9b02..0d8c46657 100644 --- a/tests/test_utils_defer.py +++ b/tests/test_utils_defer.py @@ -5,8 +5,6 @@ from twisted.python.failure import Failure from scrapy.utils.defer import mustbe_deferred, process_chain, \ process_chain_both, process_parallel, iter_errback -from six.moves import xrange - class MustbeDeferredTest(unittest.TestCase): def test_success_function(self): @@ -83,7 +81,7 @@ class IterErrbackTest(unittest.TestCase): def test_iter_errback_good(self): def itergood(): - for x in xrange(10): + for x in range(10): yield x errors = [] @@ -93,7 +91,7 @@ class IterErrbackTest(unittest.TestCase): def test_iter_errback_bad(self): def iterbad(): - for x in xrange(10): + for x in range(10): if x == 5: a = 1/0 yield x diff --git a/tests/test_utils_httpobj.py b/tests/test_utils_httpobj.py index 4f9f7a370..cf8ad1f23 100644 --- a/tests/test_utils_httpobj.py +++ b/tests/test_utils_httpobj.py @@ -1,9 +1,10 @@ import unittest -from six.moves.urllib.parse import urlparse +from urllib.parse import urlparse from scrapy.http import Request from scrapy.utils.httpobj import urlparse_cached + class HttpobjUtilsTest(unittest.TestCase): def test_urlparse_cached(self): diff --git a/tests/test_utils_response.py b/tests/test_utils_response.py index bea4dade3..6ebf290c0 100644 --- a/tests/test_utils_response.py +++ b/tests/test_utils_response.py @@ -1,12 +1,13 @@ import os import unittest -from six.moves.urllib.parse import urlparse +from urllib.parse import urlparse from scrapy.http import Response, TextResponse, HtmlResponse from scrapy.utils.python import to_bytes from scrapy.utils.response import (response_httprepr, open_in_browser, get_meta_refresh, get_base_url, response_status_message) + __doctests__ = ['scrapy.utils.response'] diff --git a/tests/test_utils_url.py b/tests/test_utils_url.py index e6588055c..a8e37d7b8 100644 --- a/tests/test_utils_url.py +++ b/tests/test_utils_url.py @@ -1,14 +1,15 @@ # -*- coding: utf-8 -*- import unittest +from urllib.parse import urlparse import six -from six.moves.urllib.parse import urlparse from scrapy.spiders import Spider from scrapy.utils.url import (url_is_from_any_domain, url_is_from_spider, add_http_if_no_scheme, guess_scheme, parse_url, strip_url) + __doctests__ = ['scrapy.utils.url'] From 1aba5136939ff26503a59ca9fb9320a030aa77de Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 00:26:44 -0300 Subject: [PATCH 196/387] Remove six.iter* occurrences --- scrapy/core/downloader/__init__.py | 3 +-- scrapy/core/downloader/handlers/__init__.py | 5 +++-- scrapy/downloadermiddlewares/cookies.py | 8 +++++--- scrapy/downloadermiddlewares/decompression.py | 7 +++---- scrapy/exporters.py | 6 +++--- scrapy/extensions/memdebug.py | 3 +-- scrapy/http/request/form.py | 3 +-- scrapy/item.py | 2 +- scrapy/loader/__init__.py | 6 ++---- scrapy/pipelines/files.py | 11 +++++------ scrapy/pipelines/images.py | 3 +-- scrapy/responsetypes.py | 3 +-- scrapy/settings/__init__.py | 8 ++++---- scrapy/utils/conf.py | 12 +++++------- scrapy/utils/datatypes.py | 7 ++----- scrapy/utils/python.py | 4 ++-- scrapy/utils/spider.py | 5 ++--- scrapy/utils/trackref.py | 13 ++++++------- tests/test_settings/__init__.py | 10 +++++----- tests/test_webclient.py | 2 +- 20 files changed, 54 insertions(+), 67 deletions(-) diff --git a/scrapy/core/downloader/__init__.py b/scrapy/core/downloader/__init__.py index f5f8be6e8..73d84664f 100644 --- a/scrapy/core/downloader/__init__.py +++ b/scrapy/core/downloader/__init__.py @@ -4,7 +4,6 @@ from time import time from datetime import datetime from collections import deque -import six from twisted.internet import reactor, defer, task from scrapy.utils.defer import mustbe_deferred @@ -189,7 +188,7 @@ class Downloader(object): def close(self): self._slot_gc_loop.stop() - for slot in six.itervalues(self.slots): + for slot in self.slots.values(): slot.close() def _slot_gc(self, age=60): diff --git a/scrapy/core/downloader/handlers/__init__.py b/scrapy/core/downloader/handlers/__init__.py index 0b55d32fa..39a0b1f51 100644 --- a/scrapy/core/downloader/handlers/__init__.py +++ b/scrapy/core/downloader/handlers/__init__.py @@ -1,8 +1,9 @@ """Download handlers for different schemes""" import logging + from twisted.internet import defer -import six + from scrapy.exceptions import NotSupported, NotConfigured from scrapy.utils.httpobj import urlparse_cached from scrapy.utils.misc import load_object @@ -22,7 +23,7 @@ class DownloadHandlers(object): self._notconfigured = {} # remembers failed handlers handlers = without_none_values( crawler.settings.getwithbase('DOWNLOAD_HANDLERS')) - for scheme, clspath in six.iteritems(handlers): + for scheme, clspath in handlers.items(): self._schemes[scheme] = clspath self._load_handler(scheme, skip_lazy=True) diff --git a/scrapy/downloadermiddlewares/cookies.py b/scrapy/downloadermiddlewares/cookies.py index 0d2b9900c..9deba40d6 100644 --- a/scrapy/downloadermiddlewares/cookies.py +++ b/scrapy/downloadermiddlewares/cookies.py @@ -1,5 +1,4 @@ import os -import six import logging from collections import defaultdict @@ -8,6 +7,7 @@ from scrapy.http import Response from scrapy.http.cookies import CookieJar from scrapy.utils.python import to_unicode + logger = logging.getLogger(__name__) @@ -82,8 +82,10 @@ class CookiesMiddleware(object): def _get_request_cookies(self, jar, request): if isinstance(request.cookies, dict): - cookie_list = [{'name': k, 'value': v} for k, v in \ - six.iteritems(request.cookies)] + cookie_list = [ + {'name': k, 'value': v} + for k, v in request.cookies.items() + ] else: cookie_list = request.cookies diff --git a/scrapy/downloadermiddlewares/decompression.py b/scrapy/downloadermiddlewares/decompression.py index e2d73f347..dfd64c35f 100644 --- a/scrapy/downloadermiddlewares/decompression.py +++ b/scrapy/downloadermiddlewares/decompression.py @@ -4,16 +4,15 @@ and extract the potentially compressed responses that may arrive. import bz2 import gzip -from io import BytesIO import zipfile import tarfile import logging +from io import BytesIO from tempfile import mktemp -import six - from scrapy.responsetypes import responsetypes + logger = logging.getLogger(__name__) @@ -75,7 +74,7 @@ class DecompressionMiddleware(object): if not response.body: return response - for fmt, func in six.iteritems(self._formats): + for fmt, func in self._formats.items(): new_response = func(response) if new_response: logger.debug('Decompressed response with format: %(responsefmt)s', diff --git a/scrapy/exporters.py b/scrapy/exporters.py index d9531e67c..19af2d6e4 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -62,9 +62,9 @@ class BaseItemExporter(object): include_empty = self.export_empty_fields if self.fields_to_export is None: if include_empty and not isinstance(item, dict): - field_iter = six.iterkeys(item.fields) + field_iter = item.fields.keys() else: - field_iter = six.iterkeys(item) + field_iter = item.keys() else: if include_empty: field_iter = self.fields_to_export @@ -326,7 +326,7 @@ class PythonItemExporter(BaseItemExporter): return value def _serialize_dict(self, value): - for key, val in six.iteritems(value): + for key, val in value.items(): key = to_bytes(key) if self.binary else key yield key, self._serialize_value(val) diff --git a/scrapy/extensions/memdebug.py b/scrapy/extensions/memdebug.py index 263d8ce4c..892aa8a86 100644 --- a/scrapy/extensions/memdebug.py +++ b/scrapy/extensions/memdebug.py @@ -5,7 +5,6 @@ See documentation in docs/topics/extensions.rst """ import gc -import six from scrapy import signals from scrapy.exceptions import NotConfigured @@ -28,7 +27,7 @@ class MemoryDebugger(object): def spider_closed(self, spider, reason): gc.collect() self.stats.set_value('memdebug/gc_garbage_count', len(gc.garbage), spider=spider) - for cls, wdict in six.iteritems(live_refs): + for cls, wdict in live_refs.items(): if not wdict: continue self.stats.set_value('memdebug/live_refs/%s' % cls.__name__, len(wdict), spider=spider) diff --git a/scrapy/http/request/form.py b/scrapy/http/request/form.py index d2bcf77b7..af02c8484 100644 --- a/scrapy/http/request/form.py +++ b/scrapy/http/request/form.py @@ -8,7 +8,6 @@ See documentation in docs/topics/request-response.rst from urllib.parse import urljoin, urlencode import lxml.html -import six from parsel.selector import create_root_node from w3lib.html import strip_html5_whitespace @@ -208,7 +207,7 @@ def _get_clickable(clickdata, form): # We didn't find it, so now we build an XPath expression out of the other # arguments, because they can be used as such xpath = u'.//*' + \ - u''.join(u'[@%s="%s"]' % c for c in six.iteritems(clickdata)) + u''.join(u'[@%s="%s"]' % c for c in clickdata.items()) el = form.xpath(xpath) if len(el) == 1: return (el[0].get('name'), el[0].get('value') or '') diff --git a/scrapy/item.py b/scrapy/item.py index 32f9b2ebb..4e0f0ac44 100644 --- a/scrapy/item.py +++ b/scrapy/item.py @@ -78,7 +78,7 @@ class DictItem(MutableMapping, BaseItem): def __init__(self, *args, **kwargs): self._values = {} if args or kwargs: # avoid creating dict for most common case - for k, v in six.iteritems(dict(*args, **kwargs)): + for k, v in dict(*args, **kwargs).items(): self[k] = v def __getitem__(self, key): diff --git a/scrapy/loader/__init__.py b/scrapy/loader/__init__.py index 60fd6d222..fe01a856f 100644 --- a/scrapy/loader/__init__.py +++ b/scrapy/loader/__init__.py @@ -5,8 +5,6 @@ See documentation in docs/topics/loaders.rst """ from collections import defaultdict -import six - from scrapy.item import Item from scrapy.loader.common import wrap_loader_context from scrapy.loader.processors import Identity @@ -72,7 +70,7 @@ class ItemLoader(object): if value is None: return if not field_name: - for k, v in six.iteritems(value): + for k, v in value.items(): self._add_value(k, v) else: self._add_value(field_name, value) @@ -82,7 +80,7 @@ class ItemLoader(object): if value is None: return if not field_name: - for k, v in six.iteritems(value): + for k, v in value.items(): self._replace_value(k, v) else: self._replace_value(field_name, value) diff --git a/scrapy/pipelines/files.py b/scrapy/pipelines/files.py index 432d4c182..6d55c8980 100644 --- a/scrapy/pipelines/files.py +++ b/scrapy/pipelines/files.py @@ -8,14 +8,12 @@ import hashlib import logging import mimetypes import os -import os.path import time from collections import defaultdict from email.utils import parsedate_tz, mktime_tz from io import BytesIO from urllib.parse import urlparse -import six from twisted.internet import defer, threads from scrapy.pipelines.media import MediaPipeline @@ -29,6 +27,7 @@ from scrapy.utils.request import referer_str from scrapy.utils.boto import is_botocore from scrapy.utils.datatypes import CaselessDict + logger = logging.getLogger(__name__) @@ -153,14 +152,14 @@ class S3FilesStore(object): Bucket=self.bucket, Key=key_name, Body=buf, - Metadata={k: str(v) for k, v in six.iteritems(meta or {})}, + Metadata={k: str(v) for k, v in (meta or {}).items()}, ACL=self.POLICY, **extra) else: b = self._get_boto_bucket() k = b.new_key(key_name) if meta: - for metakey, metavalue in six.iteritems(meta): + for metakey, metavalue in meta.items(): k.set_metadata(metakey, str(metavalue)) h = self.HEADERS.copy() if headers: @@ -201,7 +200,7 @@ class S3FilesStore(object): 'X-Amz-Website-Redirect-Location': 'WebsiteRedirectLocation', }) extra = {} - for key, value in six.iteritems(headers): + for key, value in headers.items(): try: kwarg = mapping[key] except KeyError: @@ -249,7 +248,7 @@ class GCSFilesStore(object): def persist_file(self, path, buf, info, meta=None, headers=None): blob = self.bucket.blob(self.prefix + path) blob.cache_control = self.CACHE_CONTROL - blob.metadata = {k: str(v) for k, v in six.iteritems(meta or {})} + blob.metadata = {k: str(v) for k, v in (meta or {}).items()} return threads.deferToThread( blob.upload_from_string, data=buf.getvalue(), diff --git a/scrapy/pipelines/images.py b/scrapy/pipelines/images.py index e77cef4ff..e9c6b759c 100644 --- a/scrapy/pipelines/images.py +++ b/scrapy/pipelines/images.py @@ -6,7 +6,6 @@ See documentation in topics/media-pipeline.rst import functools import hashlib from io import BytesIO -import six from PIL import Image @@ -126,7 +125,7 @@ class ImagesPipeline(FilesPipeline): image, buf = self.convert_image(orig_image) yield path, image, buf - for thumb_id, size in six.iteritems(self.thumbs): + for thumb_id, size in self.thumbs.items(): thumb_path = self.thumb_path(request, thumb_id, response=response, info=info) thumb_image, thumb_buf = self.convert_image(image, size) yield thumb_path, thumb_image, thumb_buf diff --git a/scrapy/responsetypes.py b/scrapy/responsetypes.py index b64fbbd42..91d309147 100644 --- a/scrapy/responsetypes.py +++ b/scrapy/responsetypes.py @@ -5,7 +5,6 @@ based on different criteria. from mimetypes import MimeTypes from pkgutil import get_data from io import StringIO -import six from scrapy.http import Response from scrapy.utils.misc import load_object @@ -36,7 +35,7 @@ class ResponseTypes(object): self.mimetypes = MimeTypes() mimedata = get_data('scrapy', 'mime.types').decode('utf8') self.mimetypes.readfp(StringIO(mimedata)) - for mimetype, cls in six.iteritems(self.CLASSES): + for mimetype, cls in self.CLASSES.items(): self.classes[mimetype] = load_object(cls) def from_mimetype(self, mimetype): diff --git a/scrapy/settings/__init__.py b/scrapy/settings/__init__.py index c871e86e0..d53b28895 100644 --- a/scrapy/settings/__init__.py +++ b/scrapy/settings/__init__.py @@ -317,10 +317,10 @@ class BaseSettings(MutableMapping): values = json.loads(values) if values is not None: if isinstance(values, BaseSettings): - for name, value in six.iteritems(values): + for name, value in values.items(): self.set(name, value, values.getpriority(name)) else: - for name, value in six.iteritems(values): + for name, value in values.items(): self.set(name, value, priority) def delete(self, name, priority='project'): @@ -377,7 +377,7 @@ class BaseSettings(MutableMapping): def _to_dict(self): return {k: (v._to_dict() if isinstance(v, BaseSettings) else v) - for k, v in six.iteritems(self)} + for k, v in self.items()} def copy_to_dict(self): """ @@ -445,7 +445,7 @@ class Settings(BaseSettings): self.setmodule(default_settings, 'default') # Promote default dictionaries to BaseSettings instances for per-key # priorities - for name, val in six.iteritems(self): + for name, val in self.items(): if isinstance(val, dict): self.set(name, BaseSettings(val, 'default'), 'default') self.update(values, priority) diff --git a/scrapy/utils/conf.py b/scrapy/utils/conf.py index 561bb72fc..7a15e77ff 100644 --- a/scrapy/utils/conf.py +++ b/scrapy/utils/conf.py @@ -1,11 +1,9 @@ -from configparser import ConfigParser import os import sys import numbers +from configparser import ConfigParser from operator import itemgetter -import six - from scrapy.settings import BaseSettings from scrapy.utils.deprecate import update_classpath from scrapy.utils.python import without_none_values @@ -22,7 +20,7 @@ def build_component_list(compdict, custom=None, convert=update_classpath): def _map_keys(compdict): if isinstance(compdict, BaseSettings): compbs = BaseSettings() - for k, v in six.iteritems(compdict): + for k, v in compdict.items(): prio = compdict.getpriority(k) if compbs.getpriority(convert(k)) == prio: raise ValueError('Some paths in {!r} convert to the same ' @@ -33,11 +31,11 @@ def build_component_list(compdict, custom=None, convert=update_classpath): return compbs else: _check_components(compdict) - return {convert(k): v for k, v in six.iteritems(compdict)} + return {convert(k): v for k, v in compdict.items()} def _validate_values(compdict): """Fail if a value in the components dict is not a real number or None.""" - for name, value in six.iteritems(compdict): + for name, value in compdict.items(): if value is not None and not isinstance(value, numbers.Real): raise ValueError('Invalid value {} for component {}, please provide ' \ 'a real number or None instead'.format(value, name)) @@ -53,7 +51,7 @@ def build_component_list(compdict, custom=None, convert=update_classpath): _validate_values(compdict) compdict = without_none_values(_map_keys(compdict)) - return [k for k, v in sorted(six.iteritems(compdict), key=itemgetter(1))] + return [k for k, v in sorted(compdict.items(), key=itemgetter(1))] def arglist_to_dict(arglist): diff --git a/scrapy/utils/datatypes.py b/scrapy/utils/datatypes.py index 6e9de47f3..56d4d1b8e 100644 --- a/scrapy/utils/datatypes.py +++ b/scrapy/utils/datatypes.py @@ -7,11 +7,8 @@ This module must not depend on any module outside the Standard Library. import copy import collections -from collections.abc import Mapping import warnings -import six - from scrapy.exceptions import ScrapyDeprecationWarning @@ -151,7 +148,7 @@ class MultiValueDict(dict): self.setlistdefault(key, []).append(value) except TypeError: raise ValueError("MultiValueDict.update() takes either a MultiValueDict or dictionary") - for key, value in six.iteritems(kwargs): + for key, value in kwargs.items(): self.setlistdefault(key, []).append(value) @@ -226,7 +223,7 @@ class CaselessDict(dict): return dict.setdefault(self, self.normkey(key), self.normvalue(def_val)) def update(self, seq): - seq = seq.items() if isinstance(seq, Mapping) else seq + seq = seq.items() if isinstance(seq, collections.abc.Mapping) else seq iseq = ((self.normkey(k), self.normvalue(v)) for k, v in seq) super(CaselessDict, self).update(iseq) diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index cb8ac3c8e..845c19fb9 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -300,7 +300,7 @@ def stringify_dict(dct_or_tuples, encoding='utf-8', keys_only=True): dict or a list of tuples, like any dict constructor supports. """ d = {} - for k, v in six.iteritems(dict(dct_or_tuples)): + for k, v in dict(dct_or_tuples).items(): k = k.encode(encoding) if isinstance(k, six.text_type) else k if not keys_only: v = v.encode(encoding) if isinstance(v, six.text_type) else v @@ -345,7 +345,7 @@ def without_none_values(iterable): value ``None`` have been removed. """ try: - return {k: v for k, v in six.iteritems(iterable) if v is not None} + return {k: v for k, v in iterable.items() if v is not None} except AttributeError: return type(iterable)((v for v in iterable if v is not None)) diff --git a/scrapy/utils/spider.py b/scrapy/utils/spider.py index 94b24f67e..48ad5041e 100644 --- a/scrapy/utils/spider.py +++ b/scrapy/utils/spider.py @@ -1,11 +1,10 @@ import logging import inspect -import six - from scrapy.spiders import Spider from scrapy.utils.misc import arg_to_iter + logger = logging.getLogger(__name__) @@ -21,7 +20,7 @@ def iter_spider_classes(module): # singleton in scrapy.spider.spiders from scrapy.spiders import Spider - for obj in six.itervalues(vars(module)): + for obj in vars(module).values(): if inspect.isclass(obj) and \ issubclass(obj, Spider) and \ obj.__module__ == module.__name__ and \ diff --git a/scrapy/utils/trackref.py b/scrapy/utils/trackref.py index 78389e464..4842b95df 100644 --- a/scrapy/utils/trackref.py +++ b/scrapy/utils/trackref.py @@ -13,7 +13,6 @@ import weakref from time import time from operator import itemgetter from collections import defaultdict -import six NoneType = type(None) @@ -36,13 +35,13 @@ def format_live_refs(ignore=NoneType): """Return a tabular representation of tracked objects""" s = "Live References\n\n" now = time() - for cls, wdict in sorted(six.iteritems(live_refs), + for cls, wdict in sorted(live_refs.items(), key=lambda x: x[0].__name__): if not wdict: continue if issubclass(cls, ignore): continue - oldest = min(six.itervalues(wdict)) + oldest = min(wdict.values()) s += "%-30s %6d oldest: %ds ago\n" % ( cls.__name__, len(wdict), now - oldest ) @@ -56,15 +55,15 @@ def print_live_refs(*a, **kw): def get_oldest(class_name): """Get the oldest object for a specific class name""" - for cls, wdict in six.iteritems(live_refs): + for cls, wdict in live_refs.items(): if cls.__name__ == class_name: if not wdict: break - return min(six.iteritems(wdict), key=itemgetter(1))[0] + return min(wdict.items(), key=itemgetter(1))[0] def iter_all(class_name): """Iterate over all objects of the same class by its class name""" - for cls, wdict in six.iteritems(live_refs): + for cls, wdict in live_refs.items(): if cls.__name__ == class_name: - return six.iterkeys(wdict) + return wdict.keys() diff --git a/tests/test_settings/__init__.py b/tests/test_settings/__init__.py index 32e65bed5..d5cbef6f5 100644 --- a/tests/test_settings/__init__.py +++ b/tests/test_settings/__init__.py @@ -10,7 +10,7 @@ from . import default_settings class SettingsGlobalFuncsTest(unittest.TestCase): def test_get_settings_priority(self): - for prio_str, prio_num in six.iteritems(SETTINGS_PRIORITIES): + for prio_str, prio_num in SETTINGS_PRIORITIES.items(): self.assertEqual(get_settings_priority(prio_str), prio_num) self.assertEqual(get_settings_priority(99), 99) @@ -148,10 +148,10 @@ class BaseSettingsTest(unittest.TestCase): self.settings.setmodule( 'tests.test_settings.default_settings', 10) - self.assertCountEqual(six.iterkeys(self.settings.attributes), - six.iterkeys(ctrl_attributes)) + self.assertCountEqual(self.settings.attributes.keys(), + ctrl_attributes.keys()) - for key in six.iterkeys(ctrl_attributes): + for key in ctrl_attributes.keys(): attr = self.settings.attributes[key] ctrl_attr = ctrl_attributes[key] self.assertEqual(attr.value, ctrl_attr.value) @@ -227,7 +227,7 @@ class BaseSettingsTest(unittest.TestCase): } settings = self.settings settings.attributes = {key: SettingsAttribute(value, 0) for key, value - in six.iteritems(test_configuration)} + in test_configuration.items()} self.assertTrue(settings.getbool('TEST_ENABLED1')) self.assertTrue(settings.getbool('TEST_ENABLED2')) diff --git a/tests/test_webclient.py b/tests/test_webclient.py index a81946490..0c04e7114 100644 --- a/tests/test_webclient.py +++ b/tests/test_webclient.py @@ -318,7 +318,7 @@ class WebClientTestCase(unittest.TestCase): def cleanup(passthrough): # Clean up the server which is hanging around not doing # anything. - connected = list(six.iterkeys(self.wrapper.protocols)) + connected = list(self.wrapper.protocols.keys()) # There might be nothing here if the server managed to already see # that the connection was lost. if connected: From 68bf192172e548da49613064086a6e02cf0bae54 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 00:32:07 -0300 Subject: [PATCH 197/387] Fix bad import --- scrapy/exporters.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 19af2d6e4..2ff089ce9 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -8,7 +8,7 @@ import sys import pprint import marshal import warnings -from import pickle +import pickle from xml.sax.saxutils import XMLGenerator import six From ce8e515fa8960d0229132ac33fc5f8a6e43bbd6f Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 00:36:25 -0300 Subject: [PATCH 198/387] Remove six type wrappers --- scrapy/crawler.py | 2 +- scrapy/exporters.py | 4 ++-- scrapy/http/headers.py | 6 +++--- scrapy/http/request/__init__.py | 2 +- scrapy/http/response/text.py | 6 +++--- scrapy/linkextractors/htmlparser.py | 2 +- scrapy/linkextractors/lxmlhtml.py | 2 +- scrapy/linkextractors/sgml.py | 2 +- scrapy/settings/__init__.py | 10 +++++----- scrapy/spiders/crawl.py | 2 +- scrapy/spiders/sitemap.py | 4 ++-- scrapy/utils/iterators.py | 6 +++--- scrapy/utils/misc.py | 6 +++--- scrapy/utils/python.py | 14 +++++++------- 14 files changed, 34 insertions(+), 34 deletions(-) diff --git a/scrapy/crawler.py b/scrapy/crawler.py index 19b998e0d..b88df8f64 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -205,7 +205,7 @@ class CrawlerRunner(object): return self._create_crawler(crawler_or_spidercls) def _create_crawler(self, spidercls): - if isinstance(spidercls, six.string_types): + if isinstance(spidercls, str): spidercls = self.spider_loader.load(spidercls) return Crawler(spidercls, self.settings) diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 2ff089ce9..e45317491 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -182,7 +182,7 @@ class XmlItemExporter(BaseItemExporter): for value in serialized_value: self._export_xml_field('value', value, depth=depth+1) self._beautify_indent(depth=depth) - elif isinstance(serialized_value, six.text_type): + elif isinstance(serialized_value, str): self.xg.characters(serialized_value) else: self.xg.characters(str(serialized_value)) @@ -321,7 +321,7 @@ class PythonItemExporter(BaseItemExporter): if is_listlike(value): return [self._serialize_value(v) for v in value] encode_func = to_bytes if self.binary else to_unicode - if isinstance(value, (six.text_type, bytes)): + if isinstance(value, (str, bytes)): return encode_func(value, encoding=self.encoding) return value diff --git a/scrapy/http/headers.py b/scrapy/http/headers.py index 62507eb19..5bbe7d72a 100644 --- a/scrapy/http/headers.py +++ b/scrapy/http/headers.py @@ -19,7 +19,7 @@ class Headers(CaselessDict): """Normalize values to bytes""" if value is None: value = [] - elif isinstance(value, (six.text_type, bytes)): + elif isinstance(value, (str, bytes)): value = [value] elif not hasattr(value, '__iter__'): value = [value] @@ -29,10 +29,10 @@ class Headers(CaselessDict): def _tobytes(self, x): if isinstance(x, bytes): return x - elif isinstance(x, six.text_type): + elif isinstance(x, str): return x.encode(self.encoding) elif isinstance(x, int): - return six.text_type(x).encode(self.encoding) + return str(x).encode(self.encoding) else: raise TypeError('Unsupported value type: {}'.format(type(x))) diff --git a/scrapy/http/request/__init__.py b/scrapy/http/request/__init__.py index d09eaf849..ff7a44545 100644 --- a/scrapy/http/request/__init__.py +++ b/scrapy/http/request/__init__.py @@ -60,7 +60,7 @@ class Request(object_ref): return self._url def _set_url(self, url): - if not isinstance(url, six.string_types): + if not isinstance(url, str): raise TypeError('Request url must be str or unicode, got %s:' % type(url).__name__) s = safe_url_string(url, self.encoding) diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 69bcba577..65400803f 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -31,14 +31,14 @@ class TextResponse(Response): super(TextResponse, self).__init__(*args, **kwargs) def _set_url(self, url): - if isinstance(url, six.text_type): + if isinstance(url, str): self._url = to_unicode(url, self.encoding) else: super(TextResponse, self)._set_url(url) def _set_body(self, body): self._body = b'' # used by encoding detection - if isinstance(body, six.text_type): + if isinstance(body, str): if self._encoding is None: raise TypeError('Cannot convert unicode body - %s has no encoding' % type(self).__name__) @@ -158,7 +158,7 @@ class TextResponse(Response): def _url_from_selector(sel): # type: (parsel.Selector) -> str - if isinstance(sel.root, six.string_types): + if isinstance(sel.root, str): # e.g. ::attr(href) result return strip_html5_whitespace(sel.root) if not hasattr(sel.root, 'tag'): diff --git a/scrapy/linkextractors/htmlparser.py b/scrapy/linkextractors/htmlparser.py index 623732cee..2fec35799 100644 --- a/scrapy/linkextractors/htmlparser.py +++ b/scrapy/linkextractors/htmlparser.py @@ -42,7 +42,7 @@ class HtmlParserLinkExtractor(HTMLParser): ret = [] base_url = urljoin(response_url, self.base_url) if self.base_url else response_url for link in links: - if isinstance(link.url, six.text_type): + if isinstance(link.url, str): link.url = link.url.encode(response_encoding) try: link.url = urljoin(base_url, link.url) diff --git a/scrapy/linkextractors/lxmlhtml.py b/scrapy/linkextractors/lxmlhtml.py index 5b7e709de..496ad3053 100644 --- a/scrapy/linkextractors/lxmlhtml.py +++ b/scrapy/linkextractors/lxmlhtml.py @@ -22,7 +22,7 @@ _collect_string_content = etree.XPath("string()") def _nons(tag): - if isinstance(tag, six.string_types): + if isinstance(tag, str): if tag[0] == '{' and tag[1:len(XHTML_NAMESPACE)+1] == XHTML_NAMESPACE: return tag.split('}')[-1] return tag diff --git a/scrapy/linkextractors/sgml.py b/scrapy/linkextractors/sgml.py index 98bed15e9..a9dffcad3 100644 --- a/scrapy/linkextractors/sgml.py +++ b/scrapy/linkextractors/sgml.py @@ -49,7 +49,7 @@ class BaseSgmlLinkExtractor(SGMLParser): if base_url is None: base_url = urljoin(response_url, self.base_url) if self.base_url else response_url for link in self.links: - if isinstance(link.url, six.text_type): + if isinstance(link.url, str): link.url = link.url.encode(response_encoding) try: link.url = urljoin(base_url, link.url) diff --git a/scrapy/settings/__init__.py b/scrapy/settings/__init__.py index d53b28895..c88c9c0e2 100644 --- a/scrapy/settings/__init__.py +++ b/scrapy/settings/__init__.py @@ -23,7 +23,7 @@ def get_settings_priority(priority): :attr:`~scrapy.settings.SETTINGS_PRIORITIES` dictionary and returns its numerical value, or directly returns a given numerical priority. """ - if isinstance(priority, six.string_types): + if isinstance(priority, str): return SETTINGS_PRIORITIES[priority] else: return priority @@ -173,7 +173,7 @@ class BaseSettings(MutableMapping): :type default: any """ value = self.get(name, default or []) - if isinstance(value, six.string_types): + if isinstance(value, str): value = value.split(',') return list(value) @@ -194,7 +194,7 @@ class BaseSettings(MutableMapping): :type default: any """ value = self.get(name, default or {}) - if isinstance(value, six.string_types): + if isinstance(value, str): value = json.loads(value) return dict(value) @@ -284,7 +284,7 @@ class BaseSettings(MutableMapping): :type priority: string or int """ self._assert_mutability() - if isinstance(module, six.string_types): + if isinstance(module, str): module = import_module(module) for key in dir(module): if key.isupper(): @@ -313,7 +313,7 @@ class BaseSettings(MutableMapping): :type priority: string or int """ self._assert_mutability() - if isinstance(values, six.string_types): + if isinstance(values, str): values = json.loads(values) if values is not None: if isinstance(values, BaseSettings): diff --git a/scrapy/spiders/crawl.py b/scrapy/spiders/crawl.py index 03000ce54..59e2c5661 100644 --- a/scrapy/spiders/crawl.py +++ b/scrapy/spiders/crawl.py @@ -25,7 +25,7 @@ def _identity(request, response): def _get_method(method, spider): if callable(method): return method - elif isinstance(method, six.string_types): + elif isinstance(method, str): return getattr(spider, method, None) diff --git a/scrapy/spiders/sitemap.py b/scrapy/spiders/sitemap.py index 534c45c70..2917daf57 100644 --- a/scrapy/spiders/sitemap.py +++ b/scrapy/spiders/sitemap.py @@ -22,7 +22,7 @@ class SitemapSpider(Spider): super(SitemapSpider, self).__init__(*a, **kw) self._cbs = [] for r, c in self.sitemap_rules: - if isinstance(c, six.string_types): + if isinstance(c, str): c = getattr(self, c) self._cbs.append((regex(r), c)) self._follow = [regex(x) for x in self.sitemap_follow] @@ -86,7 +86,7 @@ class SitemapSpider(Spider): def regex(x): - if isinstance(x, six.string_types): + if isinstance(x, str): return re.compile(x) return x diff --git a/scrapy/utils/iterators.py b/scrapy/utils/iterators.py index 9693ba768..10481fe8d 100644 --- a/scrapy/utils/iterators.py +++ b/scrapy/utils/iterators.py @@ -60,7 +60,7 @@ class _StreamReader(object): self._text, self.encoding = obj.body, obj.encoding else: self._text, self.encoding = obj, 'utf-8' - self._is_unicode = isinstance(self._text, six.text_type) + self._is_unicode = isinstance(self._text, str) def read(self, n=65535): self.read = self._read_unicode if self._is_unicode else self._read_string @@ -125,7 +125,7 @@ def csviter(obj, delimiter=None, headers=None, encoding=None, quotechar=None): def _body_or_str(obj, unicode=True): - expected_types = (Response, six.text_type, six.binary_type) + expected_types = (Response, str, bytes) assert isinstance(obj, expected_types), \ "obj must be %s, not %s" % ( " or ".join(t.__name__ for t in expected_types), @@ -137,7 +137,7 @@ def _body_or_str(obj, unicode=True): return obj.text else: return obj.body.decode('utf-8') - elif isinstance(obj, six.text_type): + elif isinstance(obj, str): return obj if unicode else obj.encode('utf-8') else: return obj.decode('utf-8') if unicode else obj diff --git a/scrapy/utils/misc.py b/scrapy/utils/misc.py index f638adb25..9a44f3576 100644 --- a/scrapy/utils/misc.py +++ b/scrapy/utils/misc.py @@ -13,7 +13,7 @@ from scrapy.utils.python import flatten, to_unicode from scrapy.item import BaseItem -_ITERABLE_SINGLE_VALUES = dict, BaseItem, six.text_type, bytes +_ITERABLE_SINGLE_VALUES = dict, BaseItem, str, bytes def arg_to_iter(arg): @@ -83,7 +83,7 @@ def extract_regex(regex, text, encoding='utf-8'): * if the regex doesn't contain any group the entire regex matching is returned """ - if isinstance(regex, six.string_types): + if isinstance(regex, str): regex = re.compile(regex, re.UNICODE) try: @@ -92,7 +92,7 @@ def extract_regex(regex, text, encoding='utf-8'): strings = regex.findall(text) # full regex or numbered groups strings = flatten(strings) - if isinstance(text, six.text_type): + if isinstance(text, str): return [replace_entities(s, keep=['lt', 'amp']) for s in strings] else: return [replace_entities(to_unicode(s, encoding), keep=['lt', 'amp']) diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index 845c19fb9..18fee1964 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -68,7 +68,7 @@ def is_listlike(x): >>> is_listlike(range(5)) True """ - return hasattr(x, "__iter__") and not isinstance(x, (six.text_type, bytes)) + return hasattr(x, "__iter__") and not isinstance(x, (str, bytes)) def unique(list_, key=lambda x: x): @@ -87,9 +87,9 @@ def unique(list_, key=lambda x: x): def to_unicode(text, encoding=None, errors='strict'): """Return the unicode representation of a bytes object ``text``. If ``text`` is already an unicode object, return it as-is.""" - if isinstance(text, six.text_type): + if isinstance(text, str): return text - if not isinstance(text, (bytes, six.text_type)): + if not isinstance(text, (bytes, str)): raise TypeError('to_unicode must receive a bytes or str ' 'object, got %s' % type(text).__name__) if encoding is None: @@ -102,7 +102,7 @@ def to_bytes(text, encoding=None, errors='strict'): is already a bytes object, return it as-is.""" if isinstance(text, bytes): return text - if not isinstance(text, six.string_types): + if not isinstance(text, str): raise TypeError('to_bytes must receive a str or bytes ' 'object, got %s' % type(text).__name__) if encoding is None: @@ -138,7 +138,7 @@ def re_rsearch(pattern, text, chunk_size=1024): yield (text[offset:], offset) yield (text, 0) - if isinstance(pattern, six.string_types): + if isinstance(pattern, str): pattern = re.compile(pattern) for chunk, offset in _chunk_iter(): @@ -301,9 +301,9 @@ def stringify_dict(dct_or_tuples, encoding='utf-8', keys_only=True): """ d = {} for k, v in dict(dct_or_tuples).items(): - k = k.encode(encoding) if isinstance(k, six.text_type) else k + k = k.encode(encoding) if isinstance(k, str) else k if not keys_only: - v = v.encode(encoding) if isinstance(v, six.text_type) else v + v = v.encode(encoding) if isinstance(v, str) else v d[k] = v return d From 54a786b102a3862c94fcdb123a26b54a39892697 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 00:58:47 -0300 Subject: [PATCH 199/387] Remove six imports --- scrapy/crawler.py | 6 +++--- scrapy/exporters.py | 2 -- scrapy/http/headers.py | 4 ---- scrapy/http/request/__init__.py | 1 - scrapy/http/response/text.py | 1 - scrapy/linkextractors/htmlparser.py | 1 - scrapy/linkextractors/lxmlhtml.py | 1 - scrapy/linkextractors/sgml.py | 1 - scrapy/settings/__init__.py | 1 - scrapy/spiders/crawl.py | 2 -- scrapy/spiders/sitemap.py | 1 - scrapy/utils/curl.py | 3 +-- scrapy/utils/iterators.py | 6 +++--- scrapy/utils/misc.py | 1 - tests/test_http_headers.py | 3 --- tests/test_http_request.py | 2 +- 16 files changed, 8 insertions(+), 28 deletions(-) diff --git a/scrapy/crawler.py b/scrapy/crawler.py index b88df8f64..4d7d9bac4 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -1,9 +1,8 @@ -import six -import signal import logging +import signal +import sys import warnings -import sys from twisted.internet import reactor, defer from zope.interface.verify import verifyClass, DoesNotImplement @@ -22,6 +21,7 @@ from scrapy.utils.log import ( get_scrapy_root_handler, install_scrapy_root_handler) from scrapy import signals + logger = logging.getLogger(__name__) diff --git a/scrapy/exporters.py b/scrapy/exporters.py index e45317491..f2999daea 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -11,8 +11,6 @@ import warnings import pickle from xml.sax.saxutils import XMLGenerator -import six - from scrapy.utils.serialize import ScrapyJSONEncoder from scrapy.utils.python import to_bytes, to_unicode, is_listlike from scrapy.item import BaseItem diff --git a/scrapy/http/headers.py b/scrapy/http/headers.py index 5bbe7d72a..860a5c9c6 100644 --- a/scrapy/http/headers.py +++ b/scrapy/http/headers.py @@ -1,4 +1,3 @@ -import six from w3lib.http import headers_dict_to_raw from scrapy.utils.datatypes import CaselessDict from scrapy.utils.python import to_unicode @@ -68,9 +67,6 @@ class Headers(CaselessDict): self[key] = lst def items(self): - return list(self.iteritems()) - - def iteritems(self): return ((k, self.getlist(k)) for k in self.keys()) def values(self): diff --git a/scrapy/http/request/__init__.py b/scrapy/http/request/__init__.py index ff7a44545..1e4a9c166 100644 --- a/scrapy/http/request/__init__.py +++ b/scrapy/http/request/__init__.py @@ -4,7 +4,6 @@ requests in Scrapy. See documentation in docs/topics/request-response.rst """ -import six from w3lib.url import safe_url_string from scrapy.http.headers import Headers diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 65400803f..1079fd6e8 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -8,7 +8,6 @@ See documentation in docs/topics/request-response.rst from urllib.parse import urljoin import parsel -import six from w3lib.encoding import html_to_unicode, resolve_encoding, \ html_body_declared_encoding, http_content_type_encoding from w3lib.html import strip_html5_whitespace diff --git a/scrapy/linkextractors/htmlparser.py b/scrapy/linkextractors/htmlparser.py index 2fec35799..0425d4340 100644 --- a/scrapy/linkextractors/htmlparser.py +++ b/scrapy/linkextractors/htmlparser.py @@ -5,7 +5,6 @@ import warnings from html.parser import HTMLParser from urllib.parse import urljoin -import six from w3lib.url import safe_url_string from w3lib.html import strip_html5_whitespace diff --git a/scrapy/linkextractors/lxmlhtml.py b/scrapy/linkextractors/lxmlhtml.py index 496ad3053..cb55e805a 100644 --- a/scrapy/linkextractors/lxmlhtml.py +++ b/scrapy/linkextractors/lxmlhtml.py @@ -3,7 +3,6 @@ Link extractor based on lxml.html """ from urllib.parse import urljoin -import six import lxml.etree as etree from w3lib.html import strip_html5_whitespace from w3lib.url import canonicalize_url diff --git a/scrapy/linkextractors/sgml.py b/scrapy/linkextractors/sgml.py index a9dffcad3..2ba6bca45 100644 --- a/scrapy/linkextractors/sgml.py +++ b/scrapy/linkextractors/sgml.py @@ -5,7 +5,6 @@ import warnings from urllib.parse import urljoin from sgmllib import SGMLParser -import six from w3lib.url import safe_url_string, canonicalize_url from w3lib.html import strip_html5_whitespace diff --git a/scrapy/settings/__init__.py b/scrapy/settings/__init__.py index c88c9c0e2..b6133619c 100644 --- a/scrapy/settings/__init__.py +++ b/scrapy/settings/__init__.py @@ -1,4 +1,3 @@ -import six import json import copy from collections.abc import MutableMapping diff --git a/scrapy/spiders/crawl.py b/scrapy/spiders/crawl.py index 59e2c5661..a5eb1a518 100644 --- a/scrapy/spiders/crawl.py +++ b/scrapy/spiders/crawl.py @@ -8,8 +8,6 @@ See documentation in docs/topics/spiders.rst import copy import warnings -import six - from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.http import Request, HtmlResponse from scrapy.linkextractors import LinkExtractor diff --git a/scrapy/spiders/sitemap.py b/scrapy/spiders/sitemap.py index 2917daf57..d368c7108 100644 --- a/scrapy/spiders/sitemap.py +++ b/scrapy/spiders/sitemap.py @@ -1,6 +1,5 @@ import re import logging -import six from scrapy.spiders import Spider from scrapy.http import Request, XmlResponse diff --git a/scrapy/utils/curl.py b/scrapy/utils/curl.py index a0a47e473..16639356e 100644 --- a/scrapy/utils/curl.py +++ b/scrapy/utils/curl.py @@ -4,7 +4,6 @@ from shlex import split from http.cookies import SimpleCookie from urllib.parse import urlparse -from six import string_types, iteritems from w3lib.http import basic_auth_header @@ -76,7 +75,7 @@ def curl_to_request_kwargs(curl_command, ignore_unknown_options=True): name = name.strip() val = val.strip() if name.title() == 'Cookie': - for name, morsel in iteritems(SimpleCookie(val)): + for name, morsel in SimpleCookie(val).items(): cookies[name] = morsel.value else: headers.append((name, val)) diff --git a/scrapy/utils/iterators.py b/scrapy/utils/iterators.py index 10481fe8d..3c0cb68c3 100644 --- a/scrapy/utils/iterators.py +++ b/scrapy/utils/iterators.py @@ -1,13 +1,13 @@ -import re import csv -from io import StringIO import logging -import six +import re +from io import StringIO from scrapy.http import TextResponse, Response from scrapy.selector import Selector from scrapy.utils.python import re_rsearch, to_unicode + logger = logging.getLogger(__name__) diff --git a/scrapy/utils/misc.py b/scrapy/utils/misc.py index 9a44f3576..0de3f18b3 100644 --- a/scrapy/utils/misc.py +++ b/scrapy/utils/misc.py @@ -6,7 +6,6 @@ from contextlib import contextmanager from importlib import import_module from pkgutil import iter_modules -import six from w3lib.html import replace_entities from scrapy.utils.python import flatten, to_unicode diff --git a/tests/test_http_headers.py b/tests/test_http_headers.py index 69d906fbf..50763d8f7 100644 --- a/tests/test_http_headers.py +++ b/tests/test_http_headers.py @@ -85,9 +85,6 @@ class HeadersTest(unittest.TestCase): self.assertSortedEqual(h.items(), [(b'X-Forwarded-For', [b'ip1', b'ip2']), (b'Content-Type', [b'text/html'])]) - self.assertSortedEqual(h.iteritems(), - [(b'X-Forwarded-For', [b'ip1', b'ip2']), - (b'Content-Type', [b'text/html'])]) self.assertSortedEqual(h.values(), [b'ip2', b'text/html']) def test_update(self): diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 62d2847d7..988c8a811 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -62,7 +62,7 @@ class RequestTest(unittest.TestCase): # headers must not be unicode h = Headers({'key1': u'val1', u'key2': 'val2'}) h[u'newkey'] = u'newval' - for k, v in h.iteritems(): + for k, v in h.items(): self.assertIsInstance(k, bytes) for s in v: self.assertIsInstance(s, bytes) From ac62524824c590e9dd2d323bf418f71a683ecbf4 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 01:00:54 -0300 Subject: [PATCH 200/387] Remove six.get_method_* --- scrapy/core/downloader/middleware.py | 6 +++--- scrapy/core/spidermw.py | 4 ++-- scrapy/utils/reqser.py | 4 ++-- 3 files changed, 7 insertions(+), 7 deletions(-) diff --git a/scrapy/core/downloader/middleware.py b/scrapy/core/downloader/middleware.py index 7a6a4dfac..72432558a 100644 --- a/scrapy/core/downloader/middleware.py +++ b/scrapy/core/downloader/middleware.py @@ -38,7 +38,7 @@ class DownloaderMiddlewareManager(MiddlewareManager): response = yield method(request=request, spider=spider) if response is not None and not isinstance(response, (Response, Request)): raise _InvalidOutput('Middleware %s.process_request must return None, Response or Request, got %s' % \ - (six.get_method_self(method).__class__.__name__, response.__class__.__name__)) + (method.__self__.__class__.__name__, response.__class__.__name__)) if response: defer.returnValue(response) defer.returnValue((yield download_func(request=request, spider=spider))) @@ -53,7 +53,7 @@ class DownloaderMiddlewareManager(MiddlewareManager): response = yield method(request=request, response=response, spider=spider) if not isinstance(response, (Response, Request)): raise _InvalidOutput('Middleware %s.process_response must return Response or Request, got %s' % \ - (six.get_method_self(method).__class__.__name__, type(response))) + (method.__self__.__class__.__name__, type(response))) if isinstance(response, Request): defer.returnValue(response) defer.returnValue(response) @@ -65,7 +65,7 @@ class DownloaderMiddlewareManager(MiddlewareManager): response = yield method(request=request, exception=exception, spider=spider) if response is not None and not isinstance(response, (Response, Request)): raise _InvalidOutput('Middleware %s.process_exception must return None, Response or Request, got %s' % \ - (six.get_method_self(method).__class__.__name__, type(response))) + (method.__self__.__class__.__name__, type(response))) if response: defer.returnValue(response) defer.returnValue(_failure) diff --git a/scrapy/core/spidermw.py b/scrapy/core/spidermw.py index b5f9837ff..57436d2f6 100644 --- a/scrapy/core/spidermw.py +++ b/scrapy/core/spidermw.py @@ -37,8 +37,8 @@ class SpiderMiddlewareManager(MiddlewareManager): def scrape_response(self, scrape_func, response, request, spider): fname = lambda f:'%s.%s' % ( - six.get_method_self(f).__class__.__name__, - six.get_method_function(f).__name__) + f.__self__.__class__.__name__, + f.__func__.__name__) def process_spider_input(response): for method in self.methods['process_spider_input']: diff --git a/scrapy/utils/reqser.py b/scrapy/utils/reqser.py index 495564ac0..e961ffca9 100644 --- a/scrapy/utils/reqser.py +++ b/scrapy/utils/reqser.py @@ -87,12 +87,12 @@ def _mangle_private_name(obj, func, name): def _find_method(obj, func): if obj: try: - func_self = six.get_method_self(func) + func_self = func.__self__ except AttributeError: # func has no __self__ pass else: if func_self is obj: - name = six.get_method_function(func).__name__ + name = func.__func__.__name__ if _is_private_method(name): return _mangle_private_name(obj, func, name) return name From 5d8abdde59e7501e8bda25b27c8926fda66b9af1 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 01:01:10 -0300 Subject: [PATCH 201/387] Remove six.text_type from tests --- tests/test_exporters.py | 2 +- tests/test_http_response.py | 8 ++++---- tests/test_loader.py | 14 +++++++------- tests/test_logformatter.py | 4 ++-- tests/test_toplevel.py | 2 +- tests/test_utils_iterators.py | 4 ++-- tests/test_utils_python.py | 14 +++++++------- 7 files changed, 24 insertions(+), 24 deletions(-) diff --git a/tests/test_exporters.py b/tests/test_exporters.py index f151a1285..8433fa4db 100644 --- a/tests/test_exporters.py +++ b/tests/test_exporters.py @@ -79,7 +79,7 @@ class BaseItemExporterTest(unittest.TestCase): ie = self._get_exporter(fields_to_export=['name'], encoding='latin-1') _, name = list(ie._get_serialized_fields(self.i))[0] - assert isinstance(name, six.text_type) + assert isinstance(name, str) self.assertEqual(name, u'John\xa3') def test_field_custom_serializer(self): diff --git a/tests/test_http_response.py b/tests/test_http_response.py index 80bf51647..1d121cc83 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -102,7 +102,7 @@ class BaseResponseTest(unittest.TestCase): self.assertEqual(r4.flags, []) def _assert_response_values(self, response, encoding, body): - if isinstance(body, six.text_type): + if isinstance(body, str): body_unicode = body body_bytes = body.encode(encoding) else: @@ -110,7 +110,7 @@ class BaseResponseTest(unittest.TestCase): body_bytes = body assert isinstance(response.body, bytes) - assert isinstance(response.text, six.text_type) + assert isinstance(response.text, str) self._assert_response_encoding(response, encoding) self.assertEqual(response.body, body_bytes) self.assertEqual(response.body_as_unicode(), body_unicode) @@ -220,11 +220,11 @@ class TextResponseTest(BaseResponseTest): r1 = self.response_class('http://www.example.com', body=original_string, encoding='cp1251') # check body_as_unicode - self.assertTrue(isinstance(r1.body_as_unicode(), six.text_type)) + self.assertTrue(isinstance(r1.body_as_unicode(), str)) self.assertEqual(r1.body_as_unicode(), unicode_string) # check response.text - self.assertTrue(isinstance(r1.text, six.text_type)) + self.assertTrue(isinstance(r1.text, str)) self.assertEqual(r1.text, unicode_string) def test_encoding(self): diff --git a/tests/test_loader.py b/tests/test_loader.py index 4a4264a2a..f1cf0114f 100644 --- a/tests/test_loader.py +++ b/tests/test_loader.py @@ -157,7 +157,7 @@ class BasicItemLoaderTest(unittest.TestCase): def test_get_value(self): il = NameItemLoader() - self.assertEqual(u'FOO', il.get_value([u'foo', u'bar'], TakeFirst(), six.text_type.upper)) + self.assertEqual(u'FOO', il.get_value([u'foo', u'bar'], TakeFirst(), str.upper)) self.assertEqual([u'foo', u'bar'], il.get_value([u'name:foo', u'name:bar'], re=u'name:(.*)$')) self.assertEqual(u'foo', il.get_value([u'name:foo', u'name:bar'], TakeFirst(), re=u'name:(.*)$')) @@ -258,7 +258,7 @@ class BasicItemLoaderTest(unittest.TestCase): def test_extend_custom_input_processors(self): class ChildItemLoader(TestItemLoader): - name_in = MapCompose(TestItemLoader.name_in, six.text_type.swapcase) + name_in = MapCompose(TestItemLoader.name_in, str.swapcase) il = ChildItemLoader() il.add_value('name', u'marta') @@ -266,7 +266,7 @@ class BasicItemLoaderTest(unittest.TestCase): def test_extend_default_input_processors(self): class ChildDefaultedItemLoader(DefaultedItemLoader): - name_in = MapCompose(DefaultedItemLoader.default_input_processor, six.text_type.swapcase) + name_in = MapCompose(DefaultedItemLoader.default_input_processor, str.swapcase) il = ChildDefaultedItemLoader() il.add_value('name', u'marta') @@ -689,7 +689,7 @@ class ProcessorsTest(unittest.TestCase): self.assertRaises(TypeError, proc, [None, '', 'hello', 'world']) self.assertEqual(proc(['', 'hello', 'world']), u' hello world') self.assertEqual(proc(['hello', 'world']), u'hello world') - self.assertIsInstance(proc(['hello', 'world']), six.text_type) + self.assertIsInstance(proc(['hello', 'world']), str) def test_compose(self): proc = Compose(lambda v: v[0], str.upper) @@ -704,12 +704,12 @@ class ProcessorsTest(unittest.TestCase): def test_mapcompose(self): def filter_world(x): return None if x == 'world' else x - proc = MapCompose(filter_world, six.text_type.upper) + proc = MapCompose(filter_world, str.upper) self.assertEqual(proc([u'hello', u'world', u'this', u'is', u'scrapy']), [u'HELLO', u'THIS', u'IS', u'SCRAPY']) - proc = MapCompose(filter_world, six.text_type.upper) + proc = MapCompose(filter_world, str.upper) self.assertEqual(proc(None), []) - proc = MapCompose(filter_world, six.text_type.upper) + proc = MapCompose(filter_world, str.upper) self.assertRaises(ValueError, proc, [1]) proc = MapCompose(filter_world, lambda x: x + 1) self.assertRaises(ValueError, proc, 'hello') diff --git a/tests/test_logformatter.py b/tests/test_logformatter.py index eb9c4a561..ca90de9d2 100644 --- a/tests/test_logformatter.py +++ b/tests/test_logformatter.py @@ -59,7 +59,7 @@ class LoggingContribTest(unittest.TestCase): logkws = self.formatter.dropped(item, exception, response, self.spider) logline = logkws['msg'] % logkws['args'] lines = logline.splitlines() - assert all(isinstance(x, six.text_type) for x in lines) + assert all(isinstance(x, str) for x in lines) self.assertEqual(lines, [u"Dropped: \u2018", '{}']) def test_scraped(self): @@ -69,7 +69,7 @@ class LoggingContribTest(unittest.TestCase): logkws = self.formatter.scraped(item, response, self.spider) logline = logkws['msg'] % logkws['args'] lines = logline.splitlines() - assert all(isinstance(x, six.text_type) for x in lines) + assert all(isinstance(x, str) for x in lines) self.assertEqual(lines, [u"Scraped from <200 http://www.example.com>", u'name: \xa3']) diff --git a/tests/test_toplevel.py b/tests/test_toplevel.py index 91bbe43bc..6d305249a 100644 --- a/tests/test_toplevel.py +++ b/tests/test_toplevel.py @@ -6,7 +6,7 @@ import scrapy class ToplevelTestCase(TestCase): def test_version(self): - self.assertIs(type(scrapy.__version__), six.text_type) + self.assertIs(type(scrapy.__version__), str) def test_version_info(self): self.assertIs(type(scrapy.version_info), tuple) diff --git a/tests/test_utils_iterators.py b/tests/test_utils_iterators.py index 2d845697e..0910bd560 100644 --- a/tests/test_utils_iterators.py +++ b/tests/test_utils_iterators.py @@ -255,8 +255,8 @@ class UtilsCsvTestCase(unittest.TestCase): # explicit type check cuz' we no like stinkin' autocasting! yarrr for result_row in result: - self.assertTrue(all((isinstance(k, six.text_type) for k in result_row.keys()))) - self.assertTrue(all((isinstance(v, six.text_type) for v in result_row.values()))) + self.assertTrue(all((isinstance(k, str) for k in result_row.keys()))) + self.assertTrue(all((isinstance(v, str) for v in result_row.values()))) def test_csviter_delimiter(self): body = get_testdata('feeds', 'feed-sample3.csv').replace(b',', b'\t') diff --git a/tests/test_utils_python.py b/tests/test_utils_python.py index 096aa50b7..bea431969 100644 --- a/tests/test_utils_python.py +++ b/tests/test_utils_python.py @@ -169,8 +169,8 @@ class UtilsPythonTestCase(unittest.TestCase): d2 = stringify_dict(d, keys_only=False) self.assertEqual(d, d2) self.assertIsNot(d, d2) # shouldn't modify in place - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.keys())) - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.values())) + self.assertFalse(any(isinstance(x, str) for x in d2.keys())) + self.assertFalse(any(isinstance(x, str) for x in d2.values())) @unittest.skipUnless(six.PY2, "deprecated function") def test_stringify_dict_tuples(self): @@ -179,8 +179,8 @@ class UtilsPythonTestCase(unittest.TestCase): d2 = stringify_dict(tuples, keys_only=False) self.assertEqual(d, d2) self.assertIsNot(d, d2) # shouldn't modify in place - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.keys()), d2.keys()) - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.values())) + self.assertFalse(any(isinstance(x, str) for x in d2.keys()), d2.keys()) + self.assertFalse(any(isinstance(x, str) for x in d2.values())) @unittest.skipUnless(six.PY2, "deprecated function") def test_stringify_dict_keys_only(self): @@ -188,7 +188,7 @@ class UtilsPythonTestCase(unittest.TestCase): d2 = stringify_dict(d) self.assertEqual(d, d2) self.assertIsNot(d, d2) # shouldn't modify in place - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.keys())) + self.assertFalse(any(isinstance(x, str) for x in d2.keys())) def test_get_func_args(self): def f1(a, b, c): @@ -227,12 +227,12 @@ class UtilsPythonTestCase(unittest.TestCase): if platform.python_implementation() == 'CPython': # TODO: how do we fix this to return the actual argument names? - self.assertEqual(get_func_args(six.text_type.split), []) + self.assertEqual(get_func_args(str.split), []) self.assertEqual(get_func_args(" ".join), []) self.assertEqual(get_func_args(operator.itemgetter(2)), []) else: self.assertEqual( - get_func_args(six.text_type.split, True), ['sep', 'maxsplit']) + get_func_args(str.split, True), ['sep', 'maxsplit']) self.assertEqual(get_func_args(" ".join, True), ['list']) self.assertEqual( get_func_args(operator.itemgetter(2), True), ['obj']) From eaeaa40b991bcb2113ba63146fbe4b2dc9e93016 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 01:08:08 -0300 Subject: [PATCH 202/387] Remove six.PY* checks --- conftest.py | 12 +++++------- tests/test_webclient.py | 20 -------------------- 2 files changed, 5 insertions(+), 27 deletions(-) diff --git a/conftest.py b/conftest.py index 06d65ba1d..5e6a42977 100644 --- a/conftest.py +++ b/conftest.py @@ -1,4 +1,3 @@ -import six import pytest @@ -7,12 +6,11 @@ collect_ignore = [ "scrapy/utils/testsite.py", ] - -if six.PY3: - for line in open('tests/py3-ignores.txt'): - file_path = line.strip() - if file_path and file_path[0] != '#': - collect_ignore.append(file_path) +# FIXME: fix or delete these tests +for line in open('tests/py3-ignores.txt'): + file_path = line.strip() + if file_path and file_path[0] != '#': + collect_ignore.append(file_path) @pytest.fixture() diff --git a/tests/test_webclient.py b/tests/test_webclient.py index 0c04e7114..608cfe597 100644 --- a/tests/test_webclient.py +++ b/tests/test_webclient.py @@ -78,26 +78,6 @@ class ParseUrlTestCase(unittest.TestCase): to_bytes(x) if not isinstance(x, int) else x for x in test) self.assertEqual(client._parse(url), test, url) - def test_externalUnicodeInterference(self): - """ - L{client._parse} should return C{str} for the scheme, host, and path - elements of its return tuple, even when passed an URL which has - previously been passed to L{urlparse} as a C{unicode} string. - """ - if not six.PY2: - raise unittest.SkipTest( - "Applies only to Py2, as urls can be ONLY unicode on Py3") - badInput = u'http://example.com/path' - goodInput = badInput.encode('ascii') - self._parse(badInput) # cache badInput in urlparse_cached - scheme, netloc, host, port, path = self._parse(goodInput) - self.assertTrue(isinstance(scheme, str)) - self.assertTrue(isinstance(netloc, str)) - self.assertTrue(isinstance(host, str)) - self.assertTrue(isinstance(path, str)) - self.assertTrue(isinstance(port, int)) - - class ScrapyHTTPPageGetterTests(unittest.TestCase): From d72444b9c8db94dd5567e763ee53b2c70a423c9a Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 01:11:23 -0300 Subject: [PATCH 203/387] Remove more six imports --- scrapy/core/downloader/middleware.py | 2 -- scrapy/core/spidermw.py | 1 - scrapy/utils/reqser.py | 2 -- tests/test_exporters.py | 1 - tests/test_http_response.py | 1 - tests/test_loader.py | 2 -- tests/test_logformatter.py | 1 - tests/test_toplevel.py | 2 +- tests/test_utils_iterators.py | 3 ++- tests/test_utils_reqser.py | 2 -- tests/test_utils_url.py | 2 -- tests/test_webclient.py | 1 - 12 files changed, 3 insertions(+), 17 deletions(-) diff --git a/scrapy/core/downloader/middleware.py b/scrapy/core/downloader/middleware.py index 72432558a..38608a429 100644 --- a/scrapy/core/downloader/middleware.py +++ b/scrapy/core/downloader/middleware.py @@ -3,8 +3,6 @@ Downloader Middleware manager See documentation in docs/topics/downloader-middleware.rst """ -import six - from twisted.internet import defer from scrapy.exceptions import _InvalidOutput diff --git a/scrapy/core/spidermw.py b/scrapy/core/spidermw.py index 57436d2f6..e4c6df8c7 100644 --- a/scrapy/core/spidermw.py +++ b/scrapy/core/spidermw.py @@ -5,7 +5,6 @@ See documentation in docs/topics/spider-middleware.rst """ from itertools import chain, islice -import six from twisted.python.failure import Failure from scrapy.exceptions import _InvalidOutput from scrapy.middleware import MiddlewareManager diff --git a/scrapy/utils/reqser.py b/scrapy/utils/reqser.py index e961ffca9..749bbc387 100644 --- a/scrapy/utils/reqser.py +++ b/scrapy/utils/reqser.py @@ -1,8 +1,6 @@ """ Helper functions for serializing (and deserializing) requests. """ -import six - from scrapy.http import Request from scrapy.utils.python import to_unicode from scrapy.utils.misc import load_object diff --git a/tests/test_exporters.py b/tests/test_exporters.py index 8433fa4db..5d1f5c182 100644 --- a/tests/test_exporters.py +++ b/tests/test_exporters.py @@ -8,7 +8,6 @@ from io import BytesIO from datetime import datetime import lxml.etree -import six from scrapy.item import Item, Field from scrapy.utils.python import to_unicode diff --git a/tests/test_http_response.py b/tests/test_http_response.py index 1d121cc83..ee7177ceb 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -1,7 +1,6 @@ # -*- coding: utf-8 -*- import unittest -import six from w3lib.encoding import resolve_encoding from scrapy.http import (Request, Response, TextResponse, HtmlResponse, diff --git a/tests/test_loader.py b/tests/test_loader.py index f1cf0114f..69ded4d50 100644 --- a/tests/test_loader.py +++ b/tests/test_loader.py @@ -1,8 +1,6 @@ from functools import partial import unittest -import six - from scrapy.http import HtmlResponse from scrapy.item import Item, Field from scrapy.loader import ItemLoader diff --git a/tests/test_logformatter.py b/tests/test_logformatter.py index ca90de9d2..5dc077c5b 100644 --- a/tests/test_logformatter.py +++ b/tests/test_logformatter.py @@ -3,7 +3,6 @@ import unittest from testfixtures import LogCapture from twisted.internet import defer from twisted.trial.unittest import TestCase as TwistedTestCase -import six from scrapy.crawler import CrawlerRunner from scrapy.exceptions import DropItem diff --git a/tests/test_toplevel.py b/tests/test_toplevel.py index 6d305249a..fdc5df166 100644 --- a/tests/test_toplevel.py +++ b/tests/test_toplevel.py @@ -1,5 +1,5 @@ from unittest import TestCase -import six + import scrapy diff --git a/tests/test_utils_iterators.py b/tests/test_utils_iterators.py index 0910bd560..4d69edb31 100644 --- a/tests/test_utils_iterators.py +++ b/tests/test_utils_iterators.py @@ -1,12 +1,13 @@ # -*- coding: utf-8 -*- import os -import six + from twisted.trial import unittest from scrapy.utils.iterators import csviter, xmliter, _body_or_str, xmliter_lxml from scrapy.http import XmlResponse, TextResponse, Response from tests import get_testdata + FOOBAR_NL = u"foo\nbar" diff --git a/tests/test_utils_reqser.py b/tests/test_utils_reqser.py index 92cd16de7..b5729b086 100644 --- a/tests/test_utils_reqser.py +++ b/tests/test_utils_reqser.py @@ -2,8 +2,6 @@ import unittest import sys -import six - from scrapy.http import Request, FormRequest from scrapy.spiders import Spider from scrapy.utils.reqser import request_to_dict, request_from_dict, _is_private_method, _mangle_private_name diff --git a/tests/test_utils_url.py b/tests/test_utils_url.py index a8e37d7b8..c5fdc752b 100644 --- a/tests/test_utils_url.py +++ b/tests/test_utils_url.py @@ -2,8 +2,6 @@ import unittest from urllib.parse import urlparse -import six - from scrapy.spiders import Spider from scrapy.utils.url import (url_is_from_any_domain, url_is_from_spider, add_http_if_no_scheme, guess_scheme, diff --git a/tests/test_webclient.py b/tests/test_webclient.py index 608cfe597..746367b41 100644 --- a/tests/test_webclient.py +++ b/tests/test_webclient.py @@ -3,7 +3,6 @@ from twisted.internet import defer Tests borrowed from the twisted.web.client tests. """ import os -import six import shutil import OpenSSL.SSL From e461570f991cf70d17a784aa628496521c241873 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sat, 2 Nov 2019 23:13:54 -0300 Subject: [PATCH 204/387] Remove protego from requirements file --- requirements-py3.txt | 1 - 1 file changed, 1 deletion(-) diff --git a/requirements-py3.txt b/requirements-py3.txt index 2c98e6f6d..cd183a525 100644 --- a/requirements-py3.txt +++ b/requirements-py3.txt @@ -2,7 +2,6 @@ parsel>=1.5.0 PyDispatcher>=2.0.5 Twisted>=17.9.0 w3lib>=1.17.0 -protego>=0.1.15 pyOpenSSL>=16.2.0 # Earlier versions fail with "AttributeError: module 'lib' has no attribute 'SSL_ST_INIT'" queuelib>=1.4.2 # Earlier versions fail with "AttributeError: '...QueueTest' object has no attribute 'qpath'" From 7f3cb05d8e20048eb3e37c5e83bcf495d28c6064 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 12:03:02 -0300 Subject: [PATCH 205/387] Remove metaclass-related six code --- scrapy/item.py | 5 +---- tests/test_item.py | 4 +--- 2 files changed, 2 insertions(+), 7 deletions(-) diff --git a/scrapy/item.py b/scrapy/item.py index 4e0f0ac44..1d39b48b2 100644 --- a/scrapy/item.py +++ b/scrapy/item.py @@ -10,8 +10,6 @@ from copy import deepcopy from pprint import pformat from warnings import warn -import six - from scrapy.utils.deprecate import ScrapyDeprecationWarning from scrapy.utils.trackref import object_ref @@ -130,6 +128,5 @@ class DictItem(MutableMapping, BaseItem): return deepcopy(self) -@six.add_metaclass(ItemMeta) -class Item(DictItem): +class Item(DictItem, metaclass=ItemMeta): pass diff --git a/tests/test_item.py b/tests/test_item.py index d98c63ddd..0da8fa1ac 100644 --- a/tests/test_item.py +++ b/tests/test_item.py @@ -3,8 +3,6 @@ import unittest from unittest import mock from warnings import catch_warnings -import six - from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.item import ABCMeta, DictItem, Field, Item, ItemMeta @@ -302,7 +300,7 @@ class ItemMetaTest(unittest.TestCase): class ItemMetaClassCellRegression(unittest.TestCase): def test_item_meta_classcell_regression(self): - class MyItem(six.with_metaclass(ItemMeta, Item)): + class MyItem(Item, metaclass=ItemMeta): def __init__(self, *args, **kwargs): # This call to super() trigger the __classcell__ propagation # requirement. When not done properly raises an error: From 586b25d27e1641433d505338901fa9a1409ccdd4 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 12:10:37 -0300 Subject: [PATCH 206/387] Remove six types --- scrapy/downloadermiddlewares/ajaxcrawl.py | 3 +-- scrapy/utils/python.py | 3 +-- tests/test_utils_trackref.py | 7 ++++--- 3 files changed, 6 insertions(+), 7 deletions(-) diff --git a/scrapy/downloadermiddlewares/ajaxcrawl.py b/scrapy/downloadermiddlewares/ajaxcrawl.py index 78b802673..c618e9ffc 100644 --- a/scrapy/downloadermiddlewares/ajaxcrawl.py +++ b/scrapy/downloadermiddlewares/ajaxcrawl.py @@ -2,7 +2,6 @@ import re import logging -import six from w3lib import html from scrapy.exceptions import NotConfigured @@ -66,7 +65,7 @@ class AjaxCrawlMiddleware(object): # XXX: move it to w3lib? -_ajax_crawlable_re = re.compile(six.u(r'<meta\s+name=["\']fragment["\']\s+content=["\']!["\']/?>')) +_ajax_crawlable_re = re.compile(r'<meta\s+name=["\']fragment["\']\s+content=["\']!["\']/?>') def _has_ajaxcrawlable_meta(text): """ >>> _has_ajaxcrawlable_meta('<html><head><meta name="fragment" content="!"/></head><body></body></html>') diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index 18fee1964..d32ee5a3a 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -7,7 +7,6 @@ import re import inspect import weakref import errno -import six from functools import partial, wraps from itertools import chain import sys @@ -162,7 +161,7 @@ def memoizemethod_noargs(method): return new_method -_BINARYCHARS = {six.b(chr(i)) for i in range(32)} - {b"\0", b"\t", b"\n", b"\r"} +_BINARYCHARS = {to_bytes(chr(i)) for i in range(32)} - {b"\0", b"\t", b"\n", b"\r"} _BINARYCHARS |= {ord(ch) for ch in _BINARYCHARS} @deprecated("scrapy.utils.python.binary_is_text") diff --git a/tests/test_utils_trackref.py b/tests/test_utils_trackref.py index 480a717e7..16e02f919 100644 --- a/tests/test_utils_trackref.py +++ b/tests/test_utils_trackref.py @@ -1,6 +1,7 @@ -import six import unittest +from io import StringIO from unittest import mock + from scrapy.utils import trackref @@ -38,12 +39,12 @@ Live References Bar 1 oldest: 0s ago ''') - @mock.patch('sys.stdout', new_callable=six.StringIO) + @mock.patch('sys.stdout', new_callable=StringIO) def test_print_live_refs_empty(self, stdout): trackref.print_live_refs() self.assertEqual(stdout.getvalue(), 'Live References\n\n\n') - @mock.patch('sys.stdout', new_callable=six.StringIO) + @mock.patch('sys.stdout', new_callable=StringIO) def test_print_live_refs_with_objects(self, stdout): o1 = Foo() # NOQA trackref.print_live_refs() From 5797aefd4c561666c2c047494b7424d77eebb469 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 12:18:35 -0300 Subject: [PATCH 207/387] Remove six.assertCountEqual --- tests/test_cmdline/__init__.py | 7 +++---- tests/test_settings/__init__.py | 14 ++++++-------- 2 files changed, 9 insertions(+), 12 deletions(-) diff --git a/tests/test_cmdline/__init__.py b/tests/test_cmdline/__init__.py index 56cfe642a..909ea90e0 100644 --- a/tests/test_cmdline/__init__.py +++ b/tests/test_cmdline/__init__.py @@ -1,13 +1,12 @@ -from io import StringIO import json import os import pstats import shutil -import six -from subprocess import Popen, PIPE import sys import tempfile import unittest +from io import StringIO +from subprocess import Popen, PIPE from scrapy.utils.test import get_testenv @@ -65,5 +64,5 @@ class CmdlineTest(unittest.TestCase): for char in ("'", "<", ">", 'u"'): settingsstr = settingsstr.replace(char, '"') settingsdict = json.loads(settingsstr) - six.assertCountEqual(self, settingsdict.keys(), EXTENSIONS.keys()) + self.assertCountEqual(settingsdict.keys(), EXTENSIONS.keys()) self.assertEqual(200, settingsdict[EXT_PATH]) diff --git a/tests/test_settings/__init__.py b/tests/test_settings/__init__.py index d5cbef6f5..fda44653a 100644 --- a/tests/test_settings/__init__.py +++ b/tests/test_settings/__init__.py @@ -1,4 +1,3 @@ -import six import unittest from unittest import mock @@ -43,14 +42,14 @@ class SettingsAttributeTest(unittest.TestCase): new_dict = {'three': 11, 'four': 21} attribute.set(new_dict, 10) self.assertIsInstance(attribute.value, BaseSettings) - six.assertCountEqual(self, attribute.value, new_dict) - six.assertCountEqual(self, original_settings, original_dict) + self.assertCountEqual(attribute.value, new_dict) + self.assertCountEqual(original_settings, original_dict) new_settings = BaseSettings({'five': 12}, 0) attribute.set(new_settings, 0) # Insufficient priority - six.assertCountEqual(self, attribute.value, new_dict) + self.assertCountEqual(attribute.value, new_dict) attribute.set(new_settings, 10) - six.assertCountEqual(self, attribute.value, new_settings) + self.assertCountEqual(attribute.value, new_settings) def test_repr(self): self.assertEqual(repr(self.attribute), @@ -276,9 +275,8 @@ class BaseSettingsTest(unittest.TestCase): 'TEST': BaseSettings({1: 10, 3: 30}, 'default'), 'HASNOBASE': BaseSettings({3: 3000}, 'default')}) s['TEST'].set(2, 200, 'cmdline') - six.assertCountEqual(self, s.getwithbase('TEST'), - {1: 1, 2: 200, 3: 30}) - six.assertCountEqual(self, s.getwithbase('HASNOBASE'), s['HASNOBASE']) + self.assertCountEqual(s.getwithbase('TEST'), {1: 1, 2: 200, 3: 30}) + self.assertCountEqual(s.getwithbase('HASNOBASE'), s['HASNOBASE']) self.assertEqual(s.getwithbase('NONEXISTENT'), {}) def test_maxpriority(self): From 00b793dc59a32d13c89ac5c0d54985677b228801 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 12:26:38 -0300 Subject: [PATCH 208/387] Remove elluding six occurrences --- scrapy/robotstxt.py | 6 ++++-- tests/test_contracts.py | 5 ++--- tests/test_pipeline_files.py | 2 +- tests/test_utils_curl.py | 7 ++----- 4 files changed, 9 insertions(+), 11 deletions(-) diff --git a/scrapy/robotstxt.py b/scrapy/robotstxt.py index 397924110..0a9af3a62 100644 --- a/scrapy/robotstxt.py +++ b/scrapy/robotstxt.py @@ -1,12 +1,13 @@ import sys import logging from abc import ABCMeta, abstractmethod -from six import with_metaclass from scrapy.utils.python import to_unicode + logger = logging.getLogger(__name__) + def decode_robotstxt(robotstxt_body, spider, to_native_str_type=False): try: if to_native_str_type: @@ -23,7 +24,8 @@ def decode_robotstxt(robotstxt_body, spider, to_native_str_type=False): robotstxt_body = '' return robotstxt_body -class RobotParser(with_metaclass(ABCMeta)): + +class RobotParser(metaclass=ABCMeta): @classmethod @abstractmethod def from_crawler(cls, crawler, robotstxt_body): diff --git a/tests/test_contracts.py b/tests/test_contracts.py index b2e358700..582e3d052 100644 --- a/tests/test_contracts.py +++ b/tests/test_contracts.py @@ -1,6 +1,5 @@ from unittest import TextTestResult -from six import get_unbound_function from twisted.internet import defer from twisted.python import failure from twisted.trial import unittest @@ -395,8 +394,8 @@ class ContractsManagerTest(unittest.TestCase): with MockServer() as mockserver: contract_doc = '@url {}'.format(mockserver.url('/status?n=200')) - get_unbound_function(TestSameUrlSpider.parse_first).__doc__ = contract_doc - get_unbound_function(TestSameUrlSpider.parse_second).__doc__ = contract_doc + TestSameUrlSpider.parse_first.__doc__ = contract_doc + TestSameUrlSpider.parse_second.__doc__ = contract_doc crawler = CrawlerRunner().create_crawler(TestSameUrlSpider) yield crawler.crawl() diff --git a/tests/test_pipeline_files.py b/tests/test_pipeline_files.py index dede4bf12..52f2b554e 100644 --- a/tests/test_pipeline_files.py +++ b/tests/test_pipeline_files.py @@ -1,12 +1,12 @@ import os import random import time +from io import BytesIO from tempfile import mkdtemp from shutil import rmtree from unittest import mock from urllib.parse import urlparse -from six import BytesIO from twisted.trial import unittest from twisted.internet import defer diff --git a/tests/test_utils_curl.py b/tests/test_utils_curl.py index c5655df7e..50e1bfd5f 100644 --- a/tests/test_utils_curl.py +++ b/tests/test_utils_curl.py @@ -1,7 +1,6 @@ import unittest import warnings -from six import assertRaisesRegex from w3lib.http import basic_auth_header from scrapy import Request @@ -177,8 +176,7 @@ class CurlToRequestKwargsTest(unittest.TestCase): self.assertEqual(curl_to_request_kwargs(curl_command), expected_result) def test_too_few_arguments_error(self): - assertRaisesRegex( - self, + self.assertRaisesRegex( ValueError, r"too few arguments|the following arguments are required:\s*url", lambda: curl_to_request_kwargs("curl"), @@ -194,8 +192,7 @@ class CurlToRequestKwargsTest(unittest.TestCase): self.assertEqual(curl_to_request_kwargs(curl_command), expected_result) # case 2: ignore_unknown_options=False (raise exception): - assertRaisesRegex( - self, + self.assertRaisesRegex( ValueError, "Unrecognized options:.*--bar.*--baz", lambda: curl_to_request_kwargs( From 0c4e5b68ea0e077033fde5da28c6dc4fdd859d92 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Sun, 3 Nov 2019 12:30:34 -0300 Subject: [PATCH 209/387] Remove six from requirements and setup files --- requirements-py3.txt | 1 - setup.py | 1 - 2 files changed, 2 deletions(-) diff --git a/requirements-py3.txt b/requirements-py3.txt index cd183a525..28c649e28 100644 --- a/requirements-py3.txt +++ b/requirements-py3.txt @@ -13,5 +13,4 @@ cryptography>=2.0 # Earlier versions would fail to install cssselect>=0.9.1 lxml>=3.5.0 service_identity>=16.0.0 -six>=1.10.0 zope.interface>=4.1.3 diff --git a/setup.py b/setup.py index 8f5f14f0d..85d797f88 100644 --- a/setup.py +++ b/setup.py @@ -72,7 +72,6 @@ setup( 'pyOpenSSL>=16.2.0', 'queuelib>=1.4.2', 'service_identity>=16.0.0', - 'six>=1.10.0', 'w3lib>=1.17.0', 'zope.interface>=4.1.3', 'protego>=0.1.15', From 439a3e59b8e858441f8d97dbc32f398db392330d Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Mon, 4 Nov 2019 10:35:58 -0300 Subject: [PATCH 210/387] Fix scrapy.utils.datatypes.LocalCache limit issue --- scrapy/utils/datatypes.py | 5 +++-- tests/test_utils_datatypes.py | 29 +++++++++++++++++++++++++++-- 2 files changed, 30 insertions(+), 4 deletions(-) diff --git a/scrapy/utils/datatypes.py b/scrapy/utils/datatypes.py index df2b99c28..f7e3240c1 100644 --- a/scrapy/utils/datatypes.py +++ b/scrapy/utils/datatypes.py @@ -315,8 +315,9 @@ class LocalCache(collections.OrderedDict): self.limit = limit def __setitem__(self, key, value): - while len(self) >= self.limit: - self.popitem(last=False) + if self.limit: + while len(self) >= self.limit: + self.popitem(last=False) super(LocalCache, self).__setitem__(key, value) diff --git a/tests/test_utils_datatypes.py b/tests/test_utils_datatypes.py index 535095b8d..6ffd7c73c 100644 --- a/tests/test_utils_datatypes.py +++ b/tests/test_utils_datatypes.py @@ -7,7 +7,7 @@ if six.PY2: else: from collections.abc import Mapping, MutableMapping -from scrapy.utils.datatypes import CaselessDict, SequenceExclude +from scrapy.utils.datatypes import CaselessDict, SequenceExclude, LocalCache __doctests__ = ['scrapy.utils.datatypes'] @@ -242,6 +242,31 @@ class SequenceExcludeTest(unittest.TestCase): for v in [-3, "test", 1.1]: self.assertNotIn(v, d) + +class LocalCacheTest(unittest.TestCase): + + def test_cache_with_limit(self): + cache = LocalCache(limit=2) + cache['a'] = 1 + cache['b'] = 2 + cache['c'] = 3 + self.assertEqual(len(cache), 2) + self.assertNotIn('a', cache) + self.assertIn('b', cache) + self.assertIn('c', cache) + self.assertEqual(cache['b'], 2) + self.assertEqual(cache['c'], 3) + + def test_cache_without_limit(self): + max = 10**4 + cache = LocalCache() + for x in range(max): + cache[str(x)] = x + self.assertEqual(len(cache), max) + for x in range(max): + self.assertIn(str(x), cache) + self.assertEqual(cache[str(x)], x) + + if __name__ == "__main__": unittest.main() - From fed9fbe62d54175d70660498f2fba6b5b7f68a92 Mon Sep 17 00:00:00 2001 From: elacuesta <elacuesta@users.noreply.github.com> Date: Mon, 4 Nov 2019 15:34:27 -0300 Subject: [PATCH 211/387] Update tests/test_utils_datatypes.py MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves <adrian@chaves.io> --- tests/test_utils_datatypes.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_utils_datatypes.py b/tests/test_utils_datatypes.py index 6ffd7c73c..fb2362829 100644 --- a/tests/test_utils_datatypes.py +++ b/tests/test_utils_datatypes.py @@ -7,7 +7,7 @@ if six.PY2: else: from collections.abc import Mapping, MutableMapping -from scrapy.utils.datatypes import CaselessDict, SequenceExclude, LocalCache +from scrapy.utils.datatypes import CaselessDict, LocalCache, SequenceExclude __doctests__ = ['scrapy.utils.datatypes'] From 613c66a034cebc885b90c8ba2a76aea2109ef72e Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Tue, 5 Nov 2019 09:45:51 -0300 Subject: [PATCH 212/387] Do not override built-in max function --- tests/test_utils_datatypes.py | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/tests/test_utils_datatypes.py b/tests/test_utils_datatypes.py index fb2362829..7e671f627 100644 --- a/tests/test_utils_datatypes.py +++ b/tests/test_utils_datatypes.py @@ -258,12 +258,12 @@ class LocalCacheTest(unittest.TestCase): self.assertEqual(cache['c'], 3) def test_cache_without_limit(self): - max = 10**4 + maximum = 10**4 cache = LocalCache() - for x in range(max): + for x in range(maximum): cache[str(x)] = x - self.assertEqual(len(cache), max) - for x in range(max): + self.assertEqual(len(cache), maximum) + for x in range(maximum): self.assertIn(str(x), cache) self.assertEqual(cache[str(x)], x) From 698aa704b98e206cb755190f74d73a6ba47c5fac Mon Sep 17 00:00:00 2001 From: seregaxvm <seregaxvm.main@gmail.com> Date: Tue, 5 Nov 2019 18:30:01 +0300 Subject: [PATCH 213/387] Fix zsh completion file extension (#4122) --- extras/scrapy_zsh_completion | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/extras/scrapy_zsh_completion b/extras/scrapy_zsh_completion index 86c52c36c..e995947cb 100644 --- a/extras/scrapy_zsh_completion +++ b/extras/scrapy_zsh_completion @@ -61,7 +61,7 @@ _scrapy() { '-c[evaluate the code in the shell, print the result and exit]:code:(CODE)' '--no-redirect[do not handle HTTP 3xx status codes and print response as-is]' '--spider[use this spider]:spider:_scrapy_spiders' - '::file:_files -g \*.http' + '::file:_files -g \*.html' '::URL:_httpie_urls' ) _scrapy_glb_opts $options From fe31695ba0266deaa94222fc01885b7270af4294 Mon Sep 17 00:00:00 2001 From: elacuesta <elacuesta@users.noreply.github.com> Date: Tue, 5 Nov 2019 15:36:19 -0300 Subject: [PATCH 214/387] Remove unused import (urllib.parse.unquote) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves <adrian@chaves.io> --- tests/test_http_request.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 988c8a811..05cac617c 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -6,7 +6,7 @@ import json import xmlrpc.client as xmlrpclib import warnings from unittest import mock -from urllib.parse import parse_qs, unquote, unquote_to_bytes, urlparse +from urllib.parse import parse_qs, unquote_to_bytes, urlparse from scrapy.http import Request, FormRequest, XmlRpcRequest, JsonRequest, Headers, HtmlResponse from scrapy.utils.python import to_bytes, to_unicode From 98caf055b5a343ff663f2c1150b301a3f4546a16 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Wed, 6 Nov 2019 11:53:46 +0100 Subject: [PATCH 215/387] =?UTF-8?q?Fix=20a=20typo:=20specifiy=20=E2=86=92?= =?UTF-8?q?=20specify=20(#4128)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- scrapy/utils/misc.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/utils/misc.py b/scrapy/utils/misc.py index f638adb25..b74f34451 100644 --- a/scrapy/utils/misc.py +++ b/scrapy/utils/misc.py @@ -136,7 +136,7 @@ def create_instance(objcls, settings, crawler, *args, **kwargs): """ if settings is None: if crawler is None: - raise ValueError("Specifiy at least one of settings and crawler.") + raise ValueError("Specify at least one of settings and crawler.") settings = crawler.settings if crawler and hasattr(objcls, 'from_crawler'): return objcls.from_crawler(crawler, *args, **kwargs) From e8b1e46e85fbcdf22408d320f3cc61b6e802c5e2 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Marc=20Hern=C3=A1ndez?= <noviluni@gmail.com> Date: Thu, 7 Nov 2019 14:05:01 +0100 Subject: [PATCH 216/387] Add pytest-flake8 (#3945) --- .travis.yml | 2 + conftest.py | 10 +++ pytest.ini | 253 ++++++++++++++++++++++++++++++++++++++++++++++++++++ tox.ini | 8 ++ 4 files changed, 273 insertions(+) diff --git a/.travis.yml b/.travis.yml index 2352ef124..f5b9d3f72 100644 --- a/.travis.yml +++ b/.travis.yml @@ -7,6 +7,8 @@ branches: - /^\d\.\d+\.\d+(rc\d+|\.dev\d+)?$/ matrix: include: + - env: TOXENV=flake8 + python: 3.8 - env: TOXENV=pypy3 python: 3.5 - env: TOXENV=py35 diff --git a/conftest.py b/conftest.py index 06d65ba1d..d5d61ddd3 100644 --- a/conftest.py +++ b/conftest.py @@ -19,3 +19,13 @@ if six.PY3: def chdir(tmpdir): """Change to pytest-provided temporary directory""" tmpdir.chdir() + + +def pytest_collection_modifyitems(session, config, items): + # Avoid executing tests when executing `--flake8` flag (pytest-flake8) + try: + from pytest_flake8 import Flake8Item + if config.getoption('--flake8'): + items[:] = [item for item in items if isinstance(item, Flake8Item)] + except ImportError: + pass diff --git a/pytest.ini b/pytest.ini index 73d169601..0e4ee9ed1 100644 --- a/pytest.ini +++ b/pytest.ini @@ -4,3 +4,256 @@ python_files=test_*.py __init__.py python_classes= addopts = --doctest-modules --assert=plain twisted = 1 +flake8-ignore = + # extras + extras/qps-bench-server.py E261 E501 + extras/qpsclient.py E501 E261 E501 + # scrapy/commands + scrapy/commands/__init__.py E128 E501 + scrapy/commands/check.py F401 E501 W391 + scrapy/commands/crawl.py E501 + scrapy/commands/edit.py E501 + scrapy/commands/fetch.py E401 E302 E501 E128 E502 E731 + scrapy/commands/genspider.py E128 E501 E502 + scrapy/commands/list.py E302 + scrapy/commands/parse.py E128 E501 E731 E226 + scrapy/commands/runspider.py E501 + scrapy/commands/settings.py E302 E128 + scrapy/commands/shell.py E128 E501 E502 + scrapy/commands/startproject.py E502 E127 E501 E128 W391 + scrapy/commands/version.py E501 E128 W391 + scrapy/commands/view.py F401 E302 + # scrapy/contracts + scrapy/contracts/__init__.py E501 W504 + scrapy/contracts/default.py E502 E128 + # scrapy/core + scrapy/core/engine.py E261 E501 E128 E127 E306 E502 + scrapy/core/scheduler.py E501 + scrapy/core/scraper.py E501 E306 E261 E128 W391 W504 + scrapy/core/spidermw.py E501 E731 E502 E231 E126 E226 + scrapy/core/downloader/__init__.py F401 E501 + scrapy/core/downloader/contextfactory.py E501 E128 E126 + scrapy/core/downloader/middleware.py E501 E502 + scrapy/core/downloader/tls.py E501 E305 E241 + scrapy/core/downloader/webclient.py E731 E501 E261 E502 E128 W391 E126 E226 + scrapy/core/downloader/handlers/__init__.py E501 + scrapy/core/downloader/handlers/ftp.py E501 E305 E128 E127 W391 + scrapy/core/downloader/handlers/http.py F401 + scrapy/core/downloader/handlers/http10.py E501 + scrapy/core/downloader/handlers/http11.py E501 + scrapy/core/downloader/handlers/s3.py E501 F401 E502 E128 E126 + # scrapy/downloadermiddlewares + scrapy/downloadermiddlewares/ajaxcrawl.py E302 E501 E226 + scrapy/downloadermiddlewares/decompression.py E501 + scrapy/downloadermiddlewares/defaultheaders.py E501 + scrapy/downloadermiddlewares/httpcache.py E501 E126 + scrapy/downloadermiddlewares/httpcompression.py E502 E128 + scrapy/downloadermiddlewares/httpproxy.py E501 + scrapy/downloadermiddlewares/redirect.py E501 W504 + scrapy/downloadermiddlewares/retry.py E501 E126 + scrapy/downloadermiddlewares/robotstxt.py F401 E501 + scrapy/downloadermiddlewares/stats.py E501 + # scrapy/extensions + scrapy/extensions/closespider.py E501 E502 E128 E123 + scrapy/extensions/corestats.py E302 E501 + scrapy/extensions/feedexport.py E128 E501 + scrapy/extensions/httpcache.py E128 E501 E303 F401 + scrapy/extensions/memdebug.py E501 + scrapy/extensions/spiderstate.py E302 E501 + scrapy/extensions/telnet.py E501 W504 + scrapy/extensions/throttle.py E501 + # scrapy/http + scrapy/http/__init__.py F401 + scrapy/http/common.py E501 + scrapy/http/cookies.py E501 + scrapy/http/headers.py W391 + scrapy/http/request/__init__.py E501 + scrapy/http/request/form.py E501 E123 + scrapy/http/request/json_request.py E501 + scrapy/http/response/__init__.py E501 E128 W293 W291 + scrapy/http/response/html.py E302 + scrapy/http/response/text.py E501 W293 E128 E124 + scrapy/http/response/xml.py E302 + # scrapy/linkextractors + scrapy/linkextractors/__init__.py E731 E502 E501 E402 F401 + scrapy/linkextractors/lxmlhtml.py E501 E731 E226 + # scrapy/loader + scrapy/loader/__init__.py E501 E502 E128 + scrapy/loader/common.py E302 + scrapy/loader/processors.py E501 + # scrapy/pipelines + scrapy/pipelines/__init__.py E302 + scrapy/pipelines/files.py E116 E501 E266 + scrapy/pipelines/images.py E265 E501 + scrapy/pipelines/media.py E125 E501 E266 + # scrapy/selector + scrapy/selector/__init__.py F403 F401 + scrapy/selector/unified.py F401 E501 E111 + # scrapy/settings + scrapy/settings/__init__.py E501 + scrapy/settings/default_settings.py E501 E261 E114 E116 E226 + scrapy/settings/deprecated.py E501 + # scrapy/spidermiddlewares + scrapy/spidermiddlewares/httperror.py E501 + scrapy/spidermiddlewares/offsite.py E501 + scrapy/spidermiddlewares/referer.py F401 E501 E129 W503 W504 + scrapy/spidermiddlewares/urllength.py E501 + # scrapy/spiders + scrapy/spiders/__init__.py F401 E501 E402 + scrapy/spiders/crawl.py E501 + scrapy/spiders/feed.py E501 E261 W391 + scrapy/spiders/init.py W391 + scrapy/spiders/sitemap.py E501 + # scrapy/utils + scrapy/utils/benchserver.py E501 + scrapy/utils/boto.py F401 + scrapy/utils/conf.py E402 E502 E501 + scrapy/utils/console.py E302 E261 F401 E306 E305 + scrapy/utils/curl.py F401 + scrapy/utils/datatypes.py E501 E226 + scrapy/utils/decorators.py E501 E302 + scrapy/utils/defer.py E501 E302 E128 + scrapy/utils/deprecate.py E128 E501 E127 E502 + scrapy/utils/display.py E302 + scrapy/utils/engine.py F401 E261 E302 + scrapy/utils/ftp.py E302 + scrapy/utils/gz.py E305 E501 E302 W504 + scrapy/utils/http.py F403 F401 W391 E226 + scrapy/utils/httpobj.py E302 E501 + scrapy/utils/iterators.py E501 E701 + scrapy/utils/job.py E302 + scrapy/utils/log.py E128 W503 + scrapy/utils/markup.py F403 F401 W292 + scrapy/utils/misc.py E501 E226 + scrapy/utils/multipart.py F403 F401 W292 + scrapy/utils/project.py E501 + scrapy/utils/python.py E501 E302 + scrapy/utils/reactor.py E302 E226 + scrapy/utils/reqser.py E501 + scrapy/utils/request.py E302 E127 E501 + scrapy/utils/response.py E501 E302 E128 + scrapy/utils/signal.py E501 E128 + scrapy/utils/sitemap.py E501 + scrapy/utils/spider.py E271 E302 E501 + scrapy/utils/ssl.py E501 + scrapy/utils/template.py E302 + scrapy/utils/test.py E302 E501 + scrapy/utils/url.py E501 F403 F401 E128 F405 + # scrapy + scrapy/__init__.py E402 E501 + scrapy/_monkeypatches.py W293 + scrapy/cmdline.py E502 E501 + scrapy/crawler.py E501 + scrapy/dupefilters.py E302 E501 E202 + scrapy/exceptions.py E302 E501 + scrapy/exporters.py E501 E261 E226 + scrapy/extension.py E302 + scrapy/interfaces.py E302 E501 W391 + scrapy/item.py E501 E128 + scrapy/link.py E501 W391 + scrapy/logformatter.py E501 W293 + scrapy/mail.py E402 E128 E501 E502 + scrapy/middleware.py E502 E128 E501 + scrapy/pqueues.py E501 + scrapy/resolver.py E302 + scrapy/responsetypes.py E128 E501 E305 + scrapy/robotstxt.py E302 E501 + scrapy/shell.py E501 + scrapy/signalmanager.py E501 + scrapy/spiderloader.py E225 F841 E501 E126 + scrapy/squeues.py E128 + scrapy/statscollectors.py E501 W391 + # tests + tests/__init__.py F401 E402 E501 + tests/mockserver.py E401 E501 E126 E123 F401 + tests/pipelines.py E302 F841 E226 + tests/spiders.py E302 E501 E127 + tests/test_closespider.py E501 E127 + tests/test_command_fetch.py E501 E261 + tests/test_command_parse.py F401 E302 E501 E128 E303 E226 + tests/test_command_shell.py E501 E128 + tests/test_commands.py F401 E128 E501 + tests/test_contracts.py E501 E128 W293 + tests/test_crawl.py E501 E741 E265 + tests/test_crawler.py F841 E306 E501 + tests/test_dependencies.py E302 F841 E501 E305 + tests/test_downloader_handlers.py E124 E127 E128 E225 E261 E265 F401 E501 E502 E701 E711 E126 E226 E123 + tests/test_downloadermiddleware.py E501 + tests/test_downloadermiddleware_ajaxcrawlable.py E302 E501 + tests/test_downloadermiddleware_cookies.py E731 E741 E501 E128 E303 E265 E126 + tests/test_downloadermiddleware_decompression.py E127 + tests/test_downloadermiddleware_defaultheaders.py E501 + tests/test_downloadermiddleware_downloadtimeout.py E501 + tests/test_downloadermiddleware_httpcache.py E713 E501 E302 E305 F401 + tests/test_downloadermiddleware_httpcompression.py E501 F401 E251 E126 E123 + tests/test_downloadermiddleware_httpproxy.py F401 E501 E128 + tests/test_downloadermiddleware_redirect.py E501 E303 E128 E306 E127 E305 + tests/test_downloadermiddleware_retry.py E501 E128 W293 E251 E502 E303 E126 + tests/test_downloadermiddleware_robotstxt.py E501 + tests/test_downloadermiddleware_stats.py E501 + tests/test_dupefilters.py E302 E221 E501 E741 W293 W291 E128 E124 + tests/test_engine.py E401 E501 E502 E128 E261 + tests/test_exporters.py E501 E731 E306 E128 E124 + tests/test_extension_telnet.py F401 F841 + tests/test_feedexport.py E501 F401 F841 E241 + tests/test_http_cookies.py E501 + tests/test_http_headers.py E302 E501 + tests/test_http_request.py F401 E402 E501 E231 E261 E127 E128 W293 E502 E128 E502 E126 E123 + tests/test_http_response.py E501 E301 E502 E128 E265 + tests/test_item.py E701 E128 E231 F841 E306 + tests/test_link.py E501 + tests/test_linkextractors.py E501 E128 E231 E124 + tests/test_loader.py E302 E501 E731 E303 E741 E128 E117 E241 + tests/test_logformatter.py E128 E501 E231 E122 E302 + tests/test_mail.py E302 E128 E501 E305 + tests/test_middleware.py E302 E501 E128 + tests/test_pipeline_crawl.py E131 E501 E128 E126 + tests/test_pipeline_files.py F401 E501 W293 E303 E272 E226 + tests/test_pipeline_images.py F401 F841 E501 E303 + tests/test_pipeline_media.py E501 E741 E731 E128 E261 E306 E502 + tests/test_request_cb_kwargs.py E501 + tests/test_responsetypes.py E501 E302 E305 + tests/test_robotstxt_interface.py F401 E302 E501 W291 E501 + tests/test_scheduler.py E501 E126 E123 + tests/test_selector.py F401 E501 E127 + tests/test_spider.py E501 F401 + tests/test_spidermiddleware.py E501 E226 + tests/test_spidermiddleware_depth.py W391 + tests/test_spidermiddleware_httperror.py E128 E501 E127 E121 + tests/test_spidermiddleware_offsite.py E302 E501 E128 E111 W293 + tests/test_spidermiddleware_output_chain.py F401 E501 E302 W293 E226 + tests/test_spidermiddleware_referer.py F401 E501 E302 F841 E125 E201 E261 E124 E501 W391 E241 E121 + tests/test_spidermiddleware_urllength.py W391 + tests/test_squeues.py E501 E302 E701 E741 + tests/test_utils_conf.py E501 E231 E303 E128 + tests/test_utils_console.py E302 E231 + tests/test_utils_curl.py E501 + tests/test_utils_datatypes.py E402 E501 E305 W391 + tests/test_utils_defer.py E306 E261 E501 E302 F841 E226 + tests/test_utils_deprecate.py F841 E306 E501 + tests/test_utils_http.py E302 E501 E502 E128 W391 W504 + tests/test_utils_httpobj.py E302 + tests/test_utils_iterators.py E501 E128 E129 E302 E303 E241 + tests/test_utils_log.py E741 E226 + tests/test_utils_python.py E501 E303 E731 E701 E305 + tests/test_utils_reqser.py F401 E501 E128 + tests/test_utils_request.py E302 E501 E128 E305 + tests/test_utils_response.py E501 + tests/test_utils_signal.py E741 F841 E302 E731 E226 + tests/test_utils_sitemap.py E302 E128 E501 E124 + tests/test_utils_spider.py E261 E302 E305 W391 + tests/test_utils_template.py E305 + tests/test_utils_url.py F401 E501 E127 E302 E305 E211 E125 E501 E226 E241 E126 E123 + tests/test_webclient.py E501 E128 E122 E303 E402 E306 E226 E241 E123 E126 + tests/mocks/dummydbm.py E302 + tests/test_cmdline/__init__.py E502 E501 + tests/test_cmdline/extensions.py E302 W391 + tests/test_settings/__init__.py F401 E501 E128 + tests/test_settings/default_settings.py W391 + tests/test_spiderloader/__init__.py E128 E501 E302 + tests/test_spiderloader/test_spiders/spider0.py E302 + tests/test_spiderloader/test_spiders/spider1.py E302 + tests/test_spiderloader/test_spiders/spider2.py E302 + tests/test_spiderloader/test_spiders/spider3.py E302 + tests/test_spiderloader/test_spiders/nested/spider4.py E302 + tests/test_utils_misc/__init__.py E501 E231 diff --git a/tox.ini b/tox.ini index 821144381..cc3463a4d 100644 --- a/tox.ini +++ b/tox.ini @@ -62,6 +62,14 @@ basepython = pypy3 commands = py.test {posargs:scrapy tests} +[testenv:flake8] +basepython = python3.8 +deps = + {[testenv]deps} + pytest-flake8 +commands = + py.test --flake8 {posargs:scrapy tests} + [docs] changedir = docs deps = From c377c14e3263ce3c2bffa446cd0965006e845664 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Marc=20Hern=C3=A1ndez?= <noviluni@gmail.com> Date: Thu, 7 Nov 2019 17:47:35 +0100 Subject: [PATCH 217/387] Fix W391 Blank line at end of file (#4137) --- pytest.ini | 37 ++++++++++-------------- scrapy/commands/check.py | 1 - scrapy/commands/startproject.py | 1 - scrapy/commands/version.py | 1 - scrapy/core/downloader/handlers/ftp.py | 1 - scrapy/core/downloader/webclient.py | 1 - scrapy/core/scraper.py | 1 - scrapy/http/headers.py | 2 -- scrapy/interfaces.py | 1 - scrapy/link.py | 1 - scrapy/spiders/feed.py | 1 - scrapy/spiders/init.py | 1 - scrapy/statscollectors.py | 2 -- scrapy/utils/http.py | 1 - tests/test_cmdline/extensions.py | 1 - tests/test_settings/default_settings.py | 1 - tests/test_spidermiddleware_depth.py | 1 - tests/test_spidermiddleware_urllength.py | 1 - tests/test_utils_datatypes.py | 1 - tests/test_utils_http.py | 2 -- tests/test_utils_spider.py | 1 - 21 files changed, 16 insertions(+), 44 deletions(-) diff --git a/pytest.ini b/pytest.ini index 0e4ee9ed1..db5bee228 100644 --- a/pytest.ini +++ b/pytest.ini @@ -10,7 +10,7 @@ flake8-ignore = extras/qpsclient.py E501 E261 E501 # scrapy/commands scrapy/commands/__init__.py E128 E501 - scrapy/commands/check.py F401 E501 W391 + scrapy/commands/check.py F401 E501 scrapy/commands/crawl.py E501 scrapy/commands/edit.py E501 scrapy/commands/fetch.py E401 E302 E501 E128 E502 E731 @@ -20,8 +20,8 @@ flake8-ignore = scrapy/commands/runspider.py E501 scrapy/commands/settings.py E302 E128 scrapy/commands/shell.py E128 E501 E502 - scrapy/commands/startproject.py E502 E127 E501 E128 W391 - scrapy/commands/version.py E501 E128 W391 + scrapy/commands/startproject.py E502 E127 E501 E128 + scrapy/commands/version.py E501 E128 scrapy/commands/view.py F401 E302 # scrapy/contracts scrapy/contracts/__init__.py E501 W504 @@ -29,15 +29,15 @@ flake8-ignore = # scrapy/core scrapy/core/engine.py E261 E501 E128 E127 E306 E502 scrapy/core/scheduler.py E501 - scrapy/core/scraper.py E501 E306 E261 E128 W391 W504 + scrapy/core/scraper.py E501 E306 E261 E128 W504 scrapy/core/spidermw.py E501 E731 E502 E231 E126 E226 scrapy/core/downloader/__init__.py F401 E501 scrapy/core/downloader/contextfactory.py E501 E128 E126 scrapy/core/downloader/middleware.py E501 E502 scrapy/core/downloader/tls.py E501 E305 E241 - scrapy/core/downloader/webclient.py E731 E501 E261 E502 E128 W391 E126 E226 + scrapy/core/downloader/webclient.py E731 E501 E261 E502 E128 E126 E226 scrapy/core/downloader/handlers/__init__.py E501 - scrapy/core/downloader/handlers/ftp.py E501 E305 E128 E127 W391 + scrapy/core/downloader/handlers/ftp.py E501 E305 E128 E127 scrapy/core/downloader/handlers/http.py F401 scrapy/core/downloader/handlers/http10.py E501 scrapy/core/downloader/handlers/http11.py E501 @@ -66,7 +66,6 @@ flake8-ignore = scrapy/http/__init__.py F401 scrapy/http/common.py E501 scrapy/http/cookies.py E501 - scrapy/http/headers.py W391 scrapy/http/request/__init__.py E501 scrapy/http/request/form.py E501 E123 scrapy/http/request/json_request.py E501 @@ -101,8 +100,7 @@ flake8-ignore = # scrapy/spiders scrapy/spiders/__init__.py F401 E501 E402 scrapy/spiders/crawl.py E501 - scrapy/spiders/feed.py E501 E261 W391 - scrapy/spiders/init.py W391 + scrapy/spiders/feed.py E501 E261 scrapy/spiders/sitemap.py E501 # scrapy/utils scrapy/utils/benchserver.py E501 @@ -118,7 +116,7 @@ flake8-ignore = scrapy/utils/engine.py F401 E261 E302 scrapy/utils/ftp.py E302 scrapy/utils/gz.py E305 E501 E302 W504 - scrapy/utils/http.py F403 F401 W391 E226 + scrapy/utils/http.py F403 F401 E226 scrapy/utils/httpobj.py E302 E501 scrapy/utils/iterators.py E501 E701 scrapy/utils/job.py E302 @@ -148,9 +146,9 @@ flake8-ignore = scrapy/exceptions.py E302 E501 scrapy/exporters.py E501 E261 E226 scrapy/extension.py E302 - scrapy/interfaces.py E302 E501 W391 + scrapy/interfaces.py E302 E501 scrapy/item.py E501 E128 - scrapy/link.py E501 W391 + scrapy/link.py E501 scrapy/logformatter.py E501 W293 scrapy/mail.py E402 E128 E501 E502 scrapy/middleware.py E502 E128 E501 @@ -162,7 +160,7 @@ flake8-ignore = scrapy/signalmanager.py E501 scrapy/spiderloader.py E225 F841 E501 E126 scrapy/squeues.py E128 - scrapy/statscollectors.py E501 W391 + scrapy/statscollectors.py E501 # tests tests/__init__.py F401 E402 E501 tests/mockserver.py E401 E501 E126 E123 F401 @@ -218,20 +216,18 @@ flake8-ignore = tests/test_selector.py F401 E501 E127 tests/test_spider.py E501 F401 tests/test_spidermiddleware.py E501 E226 - tests/test_spidermiddleware_depth.py W391 tests/test_spidermiddleware_httperror.py E128 E501 E127 E121 tests/test_spidermiddleware_offsite.py E302 E501 E128 E111 W293 tests/test_spidermiddleware_output_chain.py F401 E501 E302 W293 E226 - tests/test_spidermiddleware_referer.py F401 E501 E302 F841 E125 E201 E261 E124 E501 W391 E241 E121 - tests/test_spidermiddleware_urllength.py W391 + tests/test_spidermiddleware_referer.py F401 E501 E302 F841 E125 E201 E261 E124 E501 E241 E121 tests/test_squeues.py E501 E302 E701 E741 tests/test_utils_conf.py E501 E231 E303 E128 tests/test_utils_console.py E302 E231 tests/test_utils_curl.py E501 - tests/test_utils_datatypes.py E402 E501 E305 W391 + tests/test_utils_datatypes.py E402 E501 E305 tests/test_utils_defer.py E306 E261 E501 E302 F841 E226 tests/test_utils_deprecate.py F841 E306 E501 - tests/test_utils_http.py E302 E501 E502 E128 W391 W504 + tests/test_utils_http.py E302 E501 E502 E128 W504 tests/test_utils_httpobj.py E302 tests/test_utils_iterators.py E501 E128 E129 E302 E303 E241 tests/test_utils_log.py E741 E226 @@ -241,15 +237,14 @@ flake8-ignore = tests/test_utils_response.py E501 tests/test_utils_signal.py E741 F841 E302 E731 E226 tests/test_utils_sitemap.py E302 E128 E501 E124 - tests/test_utils_spider.py E261 E302 E305 W391 + tests/test_utils_spider.py E261 E302 E305 tests/test_utils_template.py E305 tests/test_utils_url.py F401 E501 E127 E302 E305 E211 E125 E501 E226 E241 E126 E123 tests/test_webclient.py E501 E128 E122 E303 E402 E306 E226 E241 E123 E126 tests/mocks/dummydbm.py E302 tests/test_cmdline/__init__.py E502 E501 - tests/test_cmdline/extensions.py E302 W391 + tests/test_cmdline/extensions.py E302 tests/test_settings/__init__.py F401 E501 E128 - tests/test_settings/default_settings.py W391 tests/test_spiderloader/__init__.py E128 E501 E302 tests/test_spiderloader/test_spiders/spider0.py E302 tests/test_spiderloader/test_spiders/spider1.py E302 diff --git a/scrapy/commands/check.py b/scrapy/commands/check.py index ab73e85e7..3e6c11b7d 100644 --- a/scrapy/commands/check.py +++ b/scrapy/commands/check.py @@ -96,4 +96,3 @@ class Command(ScrapyCommand): result.printErrors() result.printSummary(start, stop) self.exitcode = int(not result.wasSuccessful()) - diff --git a/scrapy/commands/startproject.py b/scrapy/commands/startproject.py index 67337c26e..3b9f6eabb 100644 --- a/scrapy/commands/startproject.py +++ b/scrapy/commands/startproject.py @@ -119,4 +119,3 @@ class Command(ScrapyCommand): _templates_base_dir = self.settings['TEMPLATES_DIR'] or \ join(scrapy.__path__[0], 'templates') return join(_templates_base_dir, 'project') - diff --git a/scrapy/commands/version.py b/scrapy/commands/version.py index 577365c3b..8651948f7 100644 --- a/scrapy/commands/version.py +++ b/scrapy/commands/version.py @@ -30,4 +30,3 @@ class Command(ScrapyCommand): print(patt % (name, version)) else: print("Scrapy %s" % scrapy.__version__) - diff --git a/scrapy/core/downloader/handlers/ftp.py b/scrapy/core/downloader/handlers/ftp.py index 806a537d4..39ed67a1a 100644 --- a/scrapy/core/downloader/handlers/ftp.py +++ b/scrapy/core/downloader/handlers/ftp.py @@ -112,4 +112,3 @@ class FTPDownloadHandler(object): httpcode = self.CODE_MAPPING.get(ftpcode, self.CODE_MAPPING["default"]) return Response(url=request.url, status=httpcode, body=to_bytes(message)) raise result.type(result.value) - diff --git a/scrapy/core/downloader/webclient.py b/scrapy/core/downloader/webclient.py index 3a5890ed0..3fe13414a 100644 --- a/scrapy/core/downloader/webclient.py +++ b/scrapy/core/downloader/webclient.py @@ -157,4 +157,3 @@ class ScrapyHTTPClientFactory(HTTPClientFactory): def gotHeaders(self, headers): self.headers_time = time() self.response_headers = headers - diff --git a/scrapy/core/scraper.py b/scrapy/core/scraper.py index 1f389cf2e..40de6b87a 100644 --- a/scrapy/core/scraper.py +++ b/scrapy/core/scraper.py @@ -244,4 +244,3 @@ class Scraper(object): return self.signals.send_catch_log_deferred( signal=signals.item_scraped, item=output, response=response, spider=spider) - diff --git a/scrapy/http/headers.py b/scrapy/http/headers.py index 62507eb19..f3b46b994 100644 --- a/scrapy/http/headers.py +++ b/scrapy/http/headers.py @@ -91,5 +91,3 @@ class Headers(CaselessDict): def __copy__(self): return self.__class__(self) copy = __copy__ - - diff --git a/scrapy/interfaces.py b/scrapy/interfaces.py index 89ad2b14f..d48babc3c 100644 --- a/scrapy/interfaces.py +++ b/scrapy/interfaces.py @@ -15,4 +15,3 @@ class ISpiderLoader(Interface): def find_by_request(request): """Return the list of spiders names that can handle the given request""" - diff --git a/scrapy/link.py b/scrapy/link.py index 2c8301680..f0638ced2 100644 --- a/scrapy/link.py +++ b/scrapy/link.py @@ -39,4 +39,3 @@ class Link(object): def __repr__(self): return 'Link(url=%r, text=%r, fragment=%r, nofollow=%r)' % \ (self.url, self.text, self.fragment, self.nofollow) - diff --git a/scrapy/spiders/feed.py b/scrapy/spiders/feed.py index 06e212e1c..197812a26 100644 --- a/scrapy/spiders/feed.py +++ b/scrapy/spiders/feed.py @@ -133,4 +133,3 @@ class CSVFeedSpider(Spider): raise NotConfigured('You must define parse_row method in order to scrape this CSV feed') response = self.adapt_response(response) return self.parse_rows(response) - diff --git a/scrapy/spiders/init.py b/scrapy/spiders/init.py index 2efb1a869..fd41133ea 100644 --- a/scrapy/spiders/init.py +++ b/scrapy/spiders/init.py @@ -29,4 +29,3 @@ class InitSpider(Spider): spider """ return self.initialized() - diff --git a/scrapy/statscollectors.py b/scrapy/statscollectors.py index 6da9ddcd2..f0bfaed34 100644 --- a/scrapy/statscollectors.py +++ b/scrapy/statscollectors.py @@ -80,5 +80,3 @@ class DummyStatsCollector(StatsCollector): def min_value(self, key, value, spider=None): pass - - diff --git a/scrapy/utils/http.py b/scrapy/utils/http.py index b6e05c862..ad49ef3e9 100644 --- a/scrapy/utils/http.py +++ b/scrapy/utils/http.py @@ -34,4 +34,3 @@ def decode_chunked_transfer(chunked_body): body += t[:size] t = t[size+2:] return body - diff --git a/tests/test_cmdline/extensions.py b/tests/test_cmdline/extensions.py index 72867eb56..28456b55d 100644 --- a/tests/test_cmdline/extensions.py +++ b/tests/test_cmdline/extensions.py @@ -12,4 +12,3 @@ class TestExtension(object): class DummyExtension(object): pass - diff --git a/tests/test_settings/default_settings.py b/tests/test_settings/default_settings.py index c24b5a9b9..26a555275 100644 --- a/tests/test_settings/default_settings.py +++ b/tests/test_settings/default_settings.py @@ -2,4 +2,3 @@ TEST_DEFAULT = 'defvalue' TEST_DICT = {'key': 'val'} - diff --git a/tests/test_spidermiddleware_depth.py b/tests/test_spidermiddleware_depth.py index 3685d5a6f..71cca2472 100644 --- a/tests/test_spidermiddleware_depth.py +++ b/tests/test_spidermiddleware_depth.py @@ -40,4 +40,3 @@ class TestDepthMiddleware(TestCase): def tearDown(self): self.stats.close_spider(self.spider, '') - diff --git a/tests/test_spidermiddleware_urllength.py b/tests/test_spidermiddleware_urllength.py index a0aae0fdd..5ef2b23fd 100644 --- a/tests/test_spidermiddleware_urllength.py +++ b/tests/test_spidermiddleware_urllength.py @@ -18,4 +18,3 @@ class TestUrlLengthMiddleware(TestCase): spider = Spider('foo') out = list(mw.process_spider_output(res, reqs, spider)) self.assertEqual(out, [short_url_req]) - diff --git a/tests/test_utils_datatypes.py b/tests/test_utils_datatypes.py index 535095b8d..9455172fa 100644 --- a/tests/test_utils_datatypes.py +++ b/tests/test_utils_datatypes.py @@ -244,4 +244,3 @@ class SequenceExcludeTest(unittest.TestCase): if __name__ == "__main__": unittest.main() - diff --git a/tests/test_utils_http.py b/tests/test_utils_http.py index 583105673..2524153ea 100644 --- a/tests/test_utils_http.py +++ b/tests/test_utils_http.py @@ -16,5 +16,3 @@ class ChunkedTest(unittest.TestCase): "This is the data in the first chunk\r\n" + "and this is the second one\r\n" + "consequence") - - diff --git a/tests/test_utils_spider.py b/tests/test_utils_spider.py index 045e72117..d9de1ce77 100644 --- a/tests/test_utils_spider.py +++ b/tests/test_utils_spider.py @@ -34,4 +34,3 @@ class UtilsSpidersTestCase(unittest.TestCase): if __name__ == "__main__": unittest.main() - From d874c4d90bcf96c7e5b507babaa2a45a233da506 Mon Sep 17 00:00:00 2001 From: Andrey Rahmatullin <wrar@wrar.name> Date: Thu, 7 Nov 2019 22:02:17 +0500 Subject: [PATCH 218/387] Remove the old Python 2 PyPy installation code from .travis.yml (#4138) --- .travis.yml | 7 ------- 1 file changed, 7 deletions(-) diff --git a/.travis.yml b/.travis.yml index f5b9d3f72..0e77af9fd 100644 --- a/.travis.yml +++ b/.travis.yml @@ -27,13 +27,6 @@ matrix: python: 3.6 install: - | - if [ "$TOXENV" = "pypy" ]; then - export PYPY_VERSION="pypy-6.0.0-linux_x86_64-portable" - wget "https://bitbucket.org/squeaky/portable-pypy/downloads/${PYPY_VERSION}.tar.bz2" - tar -jxf ${PYPY_VERSION}.tar.bz2 - virtualenv --python="$PYPY_VERSION/bin/pypy" "$HOME/virtualenvs/$PYPY_VERSION" - source "$HOME/virtualenvs/$PYPY_VERSION/bin/activate" - fi if [ "$TOXENV" = "pypy3" ]; then export PYPY_VERSION="pypy3.5-5.9-beta-linux_x86_64-portable" wget "https://bitbucket.org/squeaky/portable-pypy/downloads/${PYPY_VERSION}.tar.bz2" From aef98188facfc79dc574d8a86200b4e95b96b880 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 7 Nov 2019 18:06:55 +0100 Subject: [PATCH 219/387] Improve the details about request serialization requirements for JOBDIR --- docs/topics/jobs.rst | 31 ++++--------------------------- 1 file changed, 4 insertions(+), 27 deletions(-) diff --git a/docs/topics/jobs.rst b/docs/topics/jobs.rst index 9fd311c69..f5542495b 100644 --- a/docs/topics/jobs.rst +++ b/docs/topics/jobs.rst @@ -71,34 +71,11 @@ on cookies. Request serialization --------------------- -Requests must be serializable by the ``pickle`` module, in order for persistence -to work, so you should make sure that your requests are serializable. - -The most common issue here is to use ``lambda`` functions on request callbacks that -can't be persisted. - -So, for example, this won't work:: - - def some_callback(self, response): - somearg = 'test' - return scrapy.Request('http://www.example.com', - callback=lambda r: self.other_callback(r, somearg)) - - def other_callback(self, response, somearg): - print("the argument passed is: %s" % somearg) - -But this will:: - - def some_callback(self, response): - somearg = 'test' - return scrapy.Request('http://www.example.com', - callback=self.other_callback, cb_kwargs={'somearg': somearg}) - - def other_callback(self, response, somearg): - print("the argument passed is: %s" % somearg) +For persistence to work, :class:`~scrapy.http.Request` objects must be +serializable with :mod:`pickle`, except for the ``callback`` and ``errback`` +values passed to their ``__init__`` method, which must be methods of the +runnning :class:`~scrapy.spiders.Spider` class. If you wish to log the requests that couldn't be serialized, you can set the :setting:`SCHEDULER_DEBUG` setting to ``True`` in the project's settings page. It is ``False`` by default. - -.. _pickle: https://docs.python.org/library/pickle.html From 44f19df3119d553aa5c002321bd901a424c4bb2c Mon Sep 17 00:00:00 2001 From: elacuesta <elacuesta@users.noreply.github.com> Date: Fri, 8 Nov 2019 11:32:50 -0300 Subject: [PATCH 220/387] [test] Update mitmproxy version MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves <adrian@chaves.io> --- tests/requirements-py3.txt | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/requirements-py3.txt b/tests/requirements-py3.txt index c2b16bec6..7abb66b9c 100644 --- a/tests/requirements-py3.txt +++ b/tests/requirements-py3.txt @@ -1,7 +1,7 @@ # Tests requirements jmespath mitmproxy; python_version >= '3.6' -mitmproxy==3.0.4; python_version < '3.6' +mitmproxy<4.0.0; python_version < '3.6' pytest pytest-cov pytest-twisted From 1df5755699eac5a98239ae73dfb82908706bf03b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Fri, 8 Nov 2019 16:00:10 +0100 Subject: [PATCH 221/387] Set the bases for testing code examples from the documentation --- docs/conftest.py | 16 ++++++++++++++++ pytest.ini | 20 +++++++++++++++++++- tests/requirements-py3.txt | 1 + tox.ini | 6 +++--- 4 files changed, 39 insertions(+), 4 deletions(-) create mode 100644 docs/conftest.py diff --git a/docs/conftest.py b/docs/conftest.py new file mode 100644 index 000000000..91c1d4428 --- /dev/null +++ b/docs/conftest.py @@ -0,0 +1,16 @@ +from doctest import ELLIPSIS + +from sybil import Sybil +from sybil.parsers.codeblock import CodeBlockParser +from sybil.parsers.doctest import DocTestParser +from sybil.parsers.skip import skip + + +pytest_collect_file = Sybil( + parsers=[ + DocTestParser(optionflags=ELLIPSIS), + CodeBlockParser(future_imports=['print_function']), + skip, + ], + pattern='*.rst', +).pytest() diff --git a/pytest.ini b/pytest.ini index db5bee228..8c5a2cd54 100644 --- a/pytest.ini +++ b/pytest.ini @@ -2,7 +2,25 @@ usefixtures = chdir python_files=test_*.py __init__.py python_classes= -addopts = --doctest-modules --assert=plain +addopts = + --assert=plain + --doctest-modules + --ignore=docs/_ext + --ignore=docs/conf.py + --ignore=docs/intro/tutorial.rst + --ignore=docs/news.rst + --ignore=docs/topics/commands.rst + --ignore=docs/topics/debug.rst + --ignore=docs/topics/developer-tools.rst + --ignore=docs/topics/dynamic-content.rst + --ignore=docs/topics/items.rst + --ignore=docs/topics/leaks.rst + --ignore=docs/topics/loaders.rst + --ignore=docs/topics/selectors.rst + --ignore=docs/topics/shell.rst + --ignore=docs/topics/stats.rst + --ignore=docs/topics/telnetconsole.rst + --ignore=docs/utils twisted = 1 flake8-ignore = # extras diff --git a/tests/requirements-py3.txt b/tests/requirements-py3.txt index dd5b23cc3..2e8d319d2 100644 --- a/tests/requirements-py3.txt +++ b/tests/requirements-py3.txt @@ -4,6 +4,7 @@ pytest pytest-cov pytest-twisted pytest-xdist +sybil testfixtures # optional for shell wrapper tests diff --git a/tox.ini b/tox.ini index cc3463a4d..3668058c3 100644 --- a/tox.ini +++ b/tox.ini @@ -21,7 +21,7 @@ passenv = GCS_TEST_FILE_URI GCS_PROJECT_ID commands = - py.test --cov=scrapy --cov-report= {posargs:scrapy tests} + py.test --cov=scrapy --cov-report= {posargs:docs scrapy tests} [testenv:py35] basepython = python3.5 @@ -60,7 +60,7 @@ basepython = python3.8 [testenv:pypy3] basepython = pypy3 commands = - py.test {posargs:scrapy tests} + py.test {posargs:docs scrapy tests} [testenv:flake8] basepython = python3.8 @@ -68,7 +68,7 @@ deps = {[testenv]deps} pytest-flake8 commands = - py.test --flake8 {posargs:scrapy tests} + py.test --flake8 {posargs:docs scrapy tests} [docs] changedir = docs From 6cde428af43a5c8268208c6e4e239ab5ce507af4 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Fri, 8 Nov 2019 12:26:40 -0300 Subject: [PATCH 222/387] Remove deprecated MergeDict class --- scrapy/utils/datatypes.py | 59 --------------------------------------- 1 file changed, 59 deletions(-) diff --git a/scrapy/utils/datatypes.py b/scrapy/utils/datatypes.py index 56d4d1b8e..e194a7613 100644 --- a/scrapy/utils/datatypes.py +++ b/scrapy/utils/datatypes.py @@ -235,65 +235,6 @@ class CaselessDict(dict): return dict.pop(self, self.normkey(key), *args) -class MergeDict(object): - """ - A simple class for creating new "virtual" dictionaries that actually look - up values in more than one dictionary, passed in the constructor. - - If a key appears in more than one of the given dictionaries, only the - first occurrence will be used. - """ - def __init__(self, *dicts): - warnings.warn( - "scrapy.utils.datatypes.MergeDict is deprecated in favor " - "of collections.ChainMap (introduced in Python 3.3)", - category=ScrapyDeprecationWarning, - stacklevel=2, - ) - self.dicts = dicts - - def __getitem__(self, key): - for dict_ in self.dicts: - try: - return dict_[key] - except KeyError: - pass - raise KeyError - - def __copy__(self): - return self.__class__(*self.dicts) - - def get(self, key, default=None): - try: - return self[key] - except KeyError: - return default - - def getlist(self, key): - for dict_ in self.dicts: - if key in dict_.keys(): - return dict_.getlist(key) - return [] - - def items(self): - item_list = [] - for dict_ in self.dicts: - item_list.extend(dict_.items()) - return item_list - - def has_key(self, key): - for dict_ in self.dicts: - if key in dict_: - return True - return False - - __contains__ = has_key - - def copy(self): - """Returns a copy of this object.""" - return self.__copy__() - - class LocalCache(collections.OrderedDict): """Dictionary with a finite number of keys. From b6bbb2819707a2202c87abd6b3dba6af13a7cc85 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Fri, 8 Nov 2019 22:13:03 -0300 Subject: [PATCH 223/387] PEP8 adjustments --- scrapy/crawler.py | 1 - scrapy/exporters.py | 1 - scrapy/http/cookies.py | 2 +- scrapy/link.py | 2 ++ tests/test_pipeline_media.py | 2 -- tests/test_proxy_connect.py | 53 ++++++++++++++++++++---------------- tests/test_utils_python.py | 2 +- 7 files changed, 34 insertions(+), 29 deletions(-) diff --git a/scrapy/crawler.py b/scrapy/crawler.py index 4d7d9bac4..ab62c678c 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -1,6 +1,5 @@ import logging import signal -import sys import warnings from twisted.internet import reactor, defer diff --git a/scrapy/exporters.py b/scrapy/exporters.py index f2999daea..e31ab1780 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -4,7 +4,6 @@ Item Exporters are used to export/serialize items into different formats. import csv import io -import sys import pprint import marshal import warnings diff --git a/scrapy/http/cookies.py b/scrapy/http/cookies.py index c39de0b52..0903fd4f8 100644 --- a/scrapy/http/cookies.py +++ b/scrapy/http/cookies.py @@ -165,7 +165,7 @@ class WrappedRequest(object): def get_header(self, name, default=None): return to_unicode(self.request.headers.get(name, default), - errors='replace') + errors='replace') def header_items(self): return [ diff --git a/scrapy/link.py b/scrapy/link.py index be1888ef0..a809c5ca4 100644 --- a/scrapy/link.py +++ b/scrapy/link.py @@ -4,6 +4,8 @@ This module defines the Link object used in Link extractors. For actual link extractors implementation see scrapy.linkextractors, or its documentation in: docs/topics/link-extractors.rst """ + + class Link(object): """Link objects represent an extracted link by the LinkExtractor.""" diff --git a/tests/test_pipeline_media.py b/tests/test_pipeline_media.py index 70f11466b..0d23f51cc 100644 --- a/tests/test_pipeline_media.py +++ b/tests/test_pipeline_media.py @@ -1,5 +1,3 @@ -import sys - from testfixtures import LogCapture from twisted.trial import unittest from twisted.python.failure import Failure diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index bf56136b1..4147dc944 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -28,17 +28,24 @@ from mitmproxy.tools.main import mitmdump sys.argv[0] = "mitmdump" sys.exit(mitmdump()) """ - cert_path = os.path.join(os.path.abspath(os.path.dirname(__file__)), - 'keys', 'mitmproxy-ca.pem') - self.proc = Popen([sys.executable, - '-c', script, - '--listen-host', '127.0.0.1', - '--listen-port', '0', - '--proxyauth', '%s:%s' % (self.auth_user, self.auth_pass), - '--certs', cert_path, - '--ssl-insecure', - ], - stdout=PIPE, env=get_testenv()) + cert_path = os.path.join( + os.path.abspath(os.path.dirname(__file__)), + 'keys', + 'mitmproxy-ca.pem' + ) + self.proc = Popen( + [ + sys.executable, + '-c', script, + '--listen-host', '127.0.0.1', + '--listen-port', '0', + '--proxyauth', '%s:%s' % (self.auth_user, self.auth_pass), + '--certs', cert_path, + '--ssl-insecure', + ], + stdout=PIPE, + env=get_testenv() + ) line = self.proc.stdout.readline().decode('utf-8') host_port = re.search(r'listening at http://([^:]+:\d+)', line).group(1) address = 'http://%s:%s@%s' % (self.auth_user, self.auth_pass, host_port) @@ -75,9 +82,9 @@ class ProxyConnectTestCase(TestCase): @defer.inlineCallbacks def test_https_connect_tunnel(self): crawler = get_crawler(SimpleSpider) - with LogCapture() as l: + with LogCapture() as logs: yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(200, l) + self._assert_got_response_code(200, logs) @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') @defer.inlineCallbacks @@ -85,35 +92,35 @@ class ProxyConnectTestCase(TestCase): proxy = os.environ['https_proxy'] os.environ['https_proxy'] = proxy + '?noconnect' crawler = get_crawler(SimpleSpider) - with LogCapture() as l: + with LogCapture() as logs: yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(200, l) + self._assert_got_response_code(200, logs) @pytest.mark.xfail(reason='Python 3.6+ fails this earlier', condition=sys.version_info.minor >= 6) @defer.inlineCallbacks def test_https_connect_tunnel_error(self): crawler = get_crawler(SimpleSpider) - with LogCapture() as l: + with LogCapture() as logs: yield crawler.crawl("https://localhost:99999/status?n=200") - self._assert_got_tunnel_error(l) + self._assert_got_tunnel_error(logs) @defer.inlineCallbacks def test_https_tunnel_auth_error(self): os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) crawler = get_crawler(SimpleSpider) - with LogCapture() as l: + with LogCapture() as logs: yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) # The proxy returns a 407 error code but it does not reach the client; # he just sees a TunnelError. - self._assert_got_tunnel_error(l) + self._assert_got_tunnel_error(logs) @defer.inlineCallbacks def test_https_tunnel_without_leak_proxy_authorization_header(self): request = Request(self.mockserver.url("/echo", is_secure=True)) crawler = get_crawler(SingleRequestSpider) - with LogCapture() as l: + with LogCapture() as logs: yield crawler.crawl(seed=request) - self._assert_got_response_code(200, l) + self._assert_got_response_code(200, logs) echo = json.loads(crawler.spider.meta['responses'][0].text) self.assertTrue('Proxy-Authorization' not in echo['headers']) @@ -122,9 +129,9 @@ class ProxyConnectTestCase(TestCase): def test_https_noconnect_auth_error(self): os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) + '?noconnect' crawler = get_crawler(SimpleSpider) - with LogCapture() as l: + with LogCapture() as logs: yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(407, l) + self._assert_got_response_code(407, logs) def _assert_got_response_code(self, code, log): print(log) diff --git a/tests/test_utils_python.py b/tests/test_utils_python.py index faf0d4b73..b36c2a5e3 100644 --- a/tests/test_utils_python.py +++ b/tests/test_utils_python.py @@ -7,7 +7,7 @@ from itertools import count from scrapy.utils.python import ( memoizemethod_noargs, binary_is_text, equal_attributes, - WeakKeyCache, stringify_dict, get_func_args, to_bytes, to_unicode, + WeakKeyCache, get_func_args, to_bytes, to_unicode, without_none_values, MutableChain) From 084a1cda6dfd94a3671d49c362db2ee6ea88a10d Mon Sep 17 00:00:00 2001 From: purvaudai <purvaudai99@gmail.com> Date: Mon, 11 Nov 2019 15:41:00 +0530 Subject: [PATCH 224/387] Adding test --- tests/test_feedexport.py | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 8c0e5cd3d..16916f728 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -415,7 +415,7 @@ class FeedExportTest(unittest.TestCase): spider_cls.start_urls = [s.url('/')] yield runner.crawl(spider_cls) - with open(res_path, 'rb') as f: + with open(res_uri, 'rb') as f: content = f.read() finally: @@ -850,12 +850,12 @@ class FeedExportTest(unittest.TestCase): def test_pathlib_uri(self): tmpdir = tempfile.mkdtemp() feed_uri = Path(tmpdir) / 'res' - feed_uri=str(feed_uri) res_uri = urljoin('file:', pathname2url(feed_uri)) settings = { 'FEED_FORMAT': 'csv', 'FEED_STORE_EMPTY': True, - 'FEED_URI': res_uri, + 'FEED_URI': feed_uri, + 'FEED_URI_ISPATH' : True } data = yield self.exported_no_data(settings) From 0042c389eb2d44d017bc8af069c7dd7ebd9319bd Mon Sep 17 00:00:00 2001 From: purvaudai <purvaudai99@gmail.com> Date: Mon, 11 Nov 2019 15:57:58 +0530 Subject: [PATCH 225/387] Adding test --- tests/test_feedexport.py | 1 - 1 file changed, 1 deletion(-) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 16916f728..275579959 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -855,7 +855,6 @@ class FeedExportTest(unittest.TestCase): 'FEED_FORMAT': 'csv', 'FEED_STORE_EMPTY': True, 'FEED_URI': feed_uri, - 'FEED_URI_ISPATH' : True } data = yield self.exported_no_data(settings) From 9e6e2dde2b7736278075d6a8268511a3bc44b8b5 Mon Sep 17 00:00:00 2001 From: purvaudai <purvaudai99@gmail.com> Date: Mon, 11 Nov 2019 16:10:37 +0530 Subject: [PATCH 226/387] Adding test --- tests/test_feedexport.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 275579959..b6e4a5449 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -415,7 +415,7 @@ class FeedExportTest(unittest.TestCase): spider_cls.start_urls = [s.url('/')] yield runner.crawl(spider_cls) - with open(res_uri, 'rb') as f: + with open(defaults['FEED_URI'], 'rb') as f: content = f.read() finally: From 970c3be1603483a61637b94afdb965eb24342744 Mon Sep 17 00:00:00 2001 From: purvaudai <purvaudai99@gmail.com> Date: Mon, 11 Nov 2019 18:34:15 +0530 Subject: [PATCH 227/387] Added Test --- scrapy/extensions/feedexport.py | 2 +- tests/test_feedexport.py | 7 ++++--- 2 files changed, 5 insertions(+), 4 deletions(-) diff --git a/scrapy/extensions/feedexport.py b/scrapy/extensions/feedexport.py index 981efee55..8dacce459 100644 --- a/scrapy/extensions/feedexport.py +++ b/scrapy/extensions/feedexport.py @@ -201,7 +201,7 @@ class FeedExporter(object): self.settings = settings if not settings['FEED_URI']: raise NotConfigured - self.urifmt=str(settings['FEED_URI']) + self.urifmt = str(settings['FEED_URI']) self.format = settings['FEED_FORMAT'].lower() self.export_encoding = settings['FEED_EXPORT_ENCODING'] self.storages = self._load_components('FEED_STORAGES') diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index b6e4a5449..11d32bd14 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -407,6 +407,7 @@ class FeedExportTest(unittest.TestCase): defaults = { 'FEED_URI': res_uri, 'FEED_FORMAT': 'csv', + 'FEED_PATH': res_path } defaults.update(settings or {}) try: @@ -415,7 +416,7 @@ class FeedExportTest(unittest.TestCase): spider_cls.start_urls = [s.url('/')] yield runner.crawl(spider_cls) - with open(defaults['FEED_URI'], 'rb') as f: + with open(defaults['FEED_PATH'], 'rb') as f: content = f.read() finally: @@ -855,8 +856,8 @@ class FeedExportTest(unittest.TestCase): 'FEED_FORMAT': 'csv', 'FEED_STORE_EMPTY': True, 'FEED_URI': feed_uri, + 'FEED_PATH': feed_uri } - data = yield self.exported_no_data(settings) self.assertEqual(data, b'') - shutil.rmtree(tmpdir, ignore_errors=True) \ No newline at end of file + shutil.rmtree(tmpdir, ignore_errors=True) From 0c2dcd5092eccf08399f0838d86d598329cc3a28 Mon Sep 17 00:00:00 2001 From: purvaudai <purvaudai99@gmail.com> Date: Mon, 11 Nov 2019 18:35:50 +0530 Subject: [PATCH 228/387] Added Test --- requirements-py2.txt | 19 ------------------- 1 file changed, 19 deletions(-) delete mode 100644 requirements-py2.txt diff --git a/requirements-py2.txt b/requirements-py2.txt deleted file mode 100644 index 42e057417..000000000 --- a/requirements-py2.txt +++ /dev/null @@ -1,19 +0,0 @@ -parsel>=1.5.0 -PyDispatcher>=2.0.5 -w3lib>=1.17.0 -protego>=0.1.15 - -pyOpenSSL>=16.2.0 # Earlier versions fail with "AttributeError: module 'lib' has no attribute 'SSL_ST_INIT'" -queuelib>=1.4.2 # Earlier versions fail with "AttributeError: '...QueueTest' object has no attribute 'qpath'" -cryptography>=2.0 # Earlier versions would fail to install - -# Reference versions taken from -# https://packages.ubuntu.com/xenial/python/ -# https://packages.ubuntu.com/xenial/zope/ -cssselect>=0.9.1 -lxml>=3.5.0 -service_identity>=16.0.0 -six>=1.10.0 -Twisted>=16.0.0 -zope.interface>=4.1.3 -pathlib2>=2.0 From f39ff4945854fb5d98389fabfa2d2e5a059c0643 Mon Sep 17 00:00:00 2001 From: purvaudai <purvaudai99@gmail.com> Date: Mon, 11 Nov 2019 18:54:21 +0530 Subject: [PATCH 229/387] Added Test --- tests/test_feedexport.py | 1 - 1 file changed, 1 deletion(-) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 11d32bd14..9b07c2051 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -851,7 +851,6 @@ class FeedExportTest(unittest.TestCase): def test_pathlib_uri(self): tmpdir = tempfile.mkdtemp() feed_uri = Path(tmpdir) / 'res' - res_uri = urljoin('file:', pathname2url(feed_uri)) settings = { 'FEED_FORMAT': 'csv', 'FEED_STORE_EMPTY': True, From 50eaabe1fc540218f9f04197810367154eb3e102 Mon Sep 17 00:00:00 2001 From: purvaudai <purvaudai99@gmail.com> Date: Mon, 11 Nov 2019 20:00:26 +0530 Subject: [PATCH 230/387] Added Test --- tests/test_feedexport.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 9b07c2051..2819f8f0b 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -416,7 +416,7 @@ class FeedExportTest(unittest.TestCase): spider_cls.start_urls = [s.url('/')] yield runner.crawl(spider_cls) - with open(defaults['FEED_PATH'], 'rb') as f: + with open(str(defaults['FEED_PATH']), 'rb') as f: content = f.read() finally: From 79d2f99995a12ffc19ab7cd2fda10112cf5e9b65 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Tue, 12 Nov 2019 08:08:50 +0100 Subject: [PATCH 231/387] Make tutorial doctests pass --- docs/_tests/quotes1.html | 281 +++++++++++++++++++++++++++++++++++++++ docs/conftest.py | 17 ++- docs/intro/tutorial.rst | 23 ++-- pytest.ini | 1 - 4 files changed, 310 insertions(+), 12 deletions(-) create mode 100644 docs/_tests/quotes1.html diff --git a/docs/_tests/quotes1.html b/docs/_tests/quotes1.html new file mode 100644 index 000000000..71aff8847 --- /dev/null +++ b/docs/_tests/quotes1.html @@ -0,0 +1,281 @@ +<!DOCTYPE html> +<html lang="en"> +<head> + <meta charset="UTF-8"> + <title>Quotes to Scrape + + + + +
+
+ +
+

+ + Login + +

+
+
+ + +
+
+ +
+ “The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.” + by + (about) + +
+ Tags: + + + change + + deep-thoughts + + thinking + + world + +
+
+ +
+ “It is our choices, Harry, that show what we truly are, far more than our abilities.” + by + (about) + +
+ Tags: + + + abilities + + choices + +
+
+ +
+ “There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.” + by + (about) + +
+ Tags: + + + inspirational + + life + + live + + miracle + + miracles + +
+
+ +
+ “The person, be it gentleman or lady, who has not pleasure in a good novel, must be intolerably stupid.” + by + (about) + +
+ Tags: + + + aliteracy + + books + + classic + + humor + +
+
+ +
+ “Imperfection is beauty, madness is genius and it's better to be absolutely ridiculous than absolutely boring.” + by + (about) + +
+ Tags: + + + be-yourself + + inspirational + +
+
+ +
+ “Try not to become a man of success. Rather become a man of value.” + by + (about) + +
+ Tags: + + + adulthood + + success + + value + +
+
+ +
+ “It is better to be hated for what you are than to be loved for what you are not.” + by + (about) + +
+ Tags: + + + life + + love + +
+
+ +
+ “I have not failed. I've just found 10,000 ways that won't work.” + by + (about) + +
+ Tags: + + + edison + + failure + + inspirational + + paraphrased + +
+
+ +
+ “A woman is like a tea bag; you never know how strong it is until it's in hot water.” + by + (about) + + +
+ +
+ “A day without sunshine is like, you know, night.” + by + (about) + +
+ Tags: + + + humor + + obvious + + simile + +
+
+ + +
+
+ +

Top Ten tags

+ + + love + + + + inspirational + + + + life + + + + humor + + + + books + + + + reading + + + + friendship + + + + friends + + + + truth + + + + simile + + + +
+
+ +
+ + + \ No newline at end of file diff --git a/docs/conftest.py b/docs/conftest.py index 91c1d4428..8c735e838 100644 --- a/docs/conftest.py +++ b/docs/conftest.py @@ -1,16 +1,29 @@ -from doctest import ELLIPSIS +import os +from doctest import ELLIPSIS, NORMALIZE_WHITESPACE +from scrapy.http.response.html import HtmlResponse from sybil import Sybil from sybil.parsers.codeblock import CodeBlockParser from sybil.parsers.doctest import DocTestParser from sybil.parsers.skip import skip +def load_response(url, filename): + input_path = os.path.join(os.path.dirname(__file__), '_tests', filename) + with open(input_path, 'rb') as input_file: + return HtmlResponse(url, body=input_file.read()) + + +def setup(namespace): + namespace['load_response'] = load_response + + pytest_collect_file = Sybil( parsers=[ - DocTestParser(optionflags=ELLIPSIS), + DocTestParser(optionflags=ELLIPSIS | NORMALIZE_WHITESPACE), CodeBlockParser(future_imports=['print_function']), skip, ], pattern='*.rst', + setup=setup, ).pytest() diff --git a/docs/intro/tutorial.rst b/docs/intro/tutorial.rst index 0629b9e19..996e3b475 100644 --- a/docs/intro/tutorial.rst +++ b/docs/intro/tutorial.rst @@ -235,13 +235,16 @@ You will see something like:: [s] shelp() Shell help (print this help) [s] fetch(req_or_url) Fetch request (or URL) and update local objects [s] view(response) View response in a browser - >>> Using the shell, you can try selecting elements using `CSS`_ with the response -object:: +object: - >>> response.css('title') - [] +.. invisible-code-block: python + + response = load_response('http://quotes.toscrape.com/page/1/', 'quotes1.html') + +>>> response.css('title') +[] The result of running ``response.css('title')`` is a list-like object called :class:`~scrapy.selector.SelectorList`, which represents a list of @@ -372,6 +375,9 @@ we want:: We get a list of selectors for the quote HTML elements with:: >>> response.css("div.quote") + [, + , + ...] Each of the selectors returned by the query above allows us to run further queries over their sub-elements. Let's assign the first selector to a @@ -404,10 +410,9 @@ quotes elements and put them together into a Python dictionary:: ... author = quote.css("small.author::text").get() ... tags = quote.css("div.tags a.tag::text").getall() ... print(dict(text=text, author=author, tags=tags)) - {'tags': ['change', 'deep-thoughts', 'thinking', 'world'], 'author': 'Albert Einstein', 'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”'} - {'tags': ['abilities', 'choices'], 'author': 'J.K. Rowling', 'text': '“It is our choices, Harry, that show what we truly are, far more than our abilities.”'} - ... a few more of these, omitted for brevity - >>> + {'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein', 'tags': ['change', 'deep-thoughts', 'thinking', 'world']} + {'text': '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', 'author': 'J.K. Rowling', 'tags': ['abilities', 'choices']} + ... Extracting data in our spider ----------------------------- @@ -521,7 +526,7 @@ There is also an ``attrib`` property available (see :ref:`selecting-attributes` for more):: >>> response.css('li.next a').attrib['href'] - '/page/2' + '/page/2/' Let's see now our spider modified to recursively follow the link to the next page, extracting data from it:: diff --git a/pytest.ini b/pytest.ini index 8c5a2cd54..3f1cc5800 100644 --- a/pytest.ini +++ b/pytest.ini @@ -7,7 +7,6 @@ addopts = --doctest-modules --ignore=docs/_ext --ignore=docs/conf.py - --ignore=docs/intro/tutorial.rst --ignore=docs/news.rst --ignore=docs/topics/commands.rst --ignore=docs/topics/debug.rst From 7b7bb028f45a06b3444950521ea582d0e83691ba Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Tue, 12 Nov 2019 08:49:06 +0100 Subject: [PATCH 232/387] Use intersphinx for links to the Sphinx documentation --- docs/conf.py | 1 + docs/contributing.rst | 17 ++++++++--------- 2 files changed, 9 insertions(+), 9 deletions(-) diff --git a/docs/conf.py b/docs/conf.py index 34dd5bcb7..0cc1dc22b 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -273,4 +273,5 @@ coverage_ignore_pyobjects = [ intersphinx_mapping = { 'python': ('https://docs.python.org/3', None), + 'sphinx': ('https://www.sphinx-doc.org/en/stable', None), } diff --git a/docs/contributing.rst b/docs/contributing.rst index 28dea74de..f084bd23d 100644 --- a/docs/contributing.rst +++ b/docs/contributing.rst @@ -177,20 +177,19 @@ Documentation policies ====================== For reference documentation of API members (classes, methods, etc.) use -docstrings and make sure that the Sphinx documentation uses the autodoc_ -extension to pull the docstrings. API reference documentation should follow -docstring conventions (`PEP 257`_) and be IDE-friendly: short, to the point, -and it may provide short examples. +docstrings and make sure that the Sphinx documentation uses the +:mod:`~sphinx.ext.autodoc` extension to pull the docstrings. API reference +documentation should follow docstring conventions (`PEP 257`_) and be +IDE-friendly: short, to the point, and it may provide short examples. Other types of documentation, such as tutorials or topics, should be covered in files within the ``docs/`` directory. This includes documentation that is specific to an API member, but goes beyond API reference documentation. -In any case, if something is covered in a docstring, use the autodoc_ -extension to pull the docstring into the documentation instead of duplicating -the docstring in files within the ``docs/`` directory. - -.. _autodoc: http://www.sphinx-doc.org/en/stable/ext/autodoc.html +In any case, if something is covered in a docstring, use the +:mod:`~sphinx.ext.autodoc` extension to pull the docstring into the +documentation instead of duplicating the docstring in files within the +``docs/`` directory. Tests ===== From 8a6a063778d45e8ba68a59558d2930332c1e9a83 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Tue, 12 Nov 2019 10:23:19 +0100 Subject: [PATCH 233/387] Allow opening the source code from the API documentation --- docs/conf.py | 1 + 1 file changed, 1 insertion(+) diff --git a/docs/conf.py b/docs/conf.py index 34dd5bcb7..b09000de0 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -31,6 +31,7 @@ extensions = [ 'sphinx.ext.autodoc', 'sphinx.ext.coverage', 'sphinx.ext.intersphinx', + 'sphinx.ext.viewcode', ] # Add any paths that contain templates here, relative to this directory. From 4b8b0345e58ee5990bd0c28205b5eb0b892680d1 Mon Sep 17 00:00:00 2001 From: purvaudai Date: Tue, 12 Nov 2019 18:17:15 +0530 Subject: [PATCH 234/387] Mades Changes as per review --- requirements-py3.txt | 1 - tests/test_feedexport.py | 2 +- 2 files changed, 1 insertion(+), 2 deletions(-) diff --git a/requirements-py3.txt b/requirements-py3.txt index 77296b91b..2c98e6f6d 100644 --- a/requirements-py3.txt +++ b/requirements-py3.txt @@ -16,4 +16,3 @@ lxml>=3.5.0 service_identity>=16.0.0 six>=1.10.0 zope.interface>=4.1.3 -pathlib2>=2.0 diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 2819f8f0b..1f46ac04a 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -29,7 +29,7 @@ from scrapy.utils.test import assert_aws_environ, get_s3_content_and_delete, get from scrapy.utils.python import to_native_str from scrapy.utils.project import get_project_settings -from pathlib2 import Path +from pathlib import Path class FileFeedStorageTest(unittest.TestCase): From 414e6e2fd568e0dfae699873d3c1ccd865261d09 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Wed, 13 Nov 2019 07:56:45 +0100 Subject: [PATCH 235/387] Skip a doctest in Python 3.5- because of dictionary changes --- docs/intro/tutorial.rst | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/intro/tutorial.rst b/docs/intro/tutorial.rst index 996e3b475..30b1ddeab 100644 --- a/docs/intro/tutorial.rst +++ b/docs/intro/tutorial.rst @@ -402,6 +402,12 @@ to get all of them:: >>> tags ['change', 'deep-thoughts', 'thinking', 'world'] +.. invisible-code-block: python + + from sys import version_info + +.. skip: next if(version_info <= (3, 5), reason="Only Python 3.6+ dictionaries match the output") + Having figured out how to extract each bit, we can now iterate over all the quotes elements and put them together into a Python dictionary:: From b642a1fca29852adf0ba3ddd62c2ecfdbaf9610e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Wed, 13 Nov 2019 09:14:20 +0100 Subject: [PATCH 236/387] Fix doctest skipping based on the running Python version --- docs/intro/tutorial.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/intro/tutorial.rst b/docs/intro/tutorial.rst index 30b1ddeab..6b15a5fbd 100644 --- a/docs/intro/tutorial.rst +++ b/docs/intro/tutorial.rst @@ -406,7 +406,7 @@ to get all of them:: from sys import version_info -.. skip: next if(version_info <= (3, 5), reason="Only Python 3.6+ dictionaries match the output") +.. skip: next if(version_info < (3, 6), reason="Only Python 3.6+ dictionaries match the output") Having figured out how to extract each bit, we can now iterate over all the quotes elements and put them together into a Python dictionary:: From 76c31094dff2920778b42ed746c6a9cd5b08f4e4 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Wed, 13 Nov 2019 09:28:48 +0100 Subject: [PATCH 237/387] Install the sphinx-notfound-page Sphinx extension --- docs/conf.py | 1 + docs/requirements.txt | 3 ++- 2 files changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/conf.py b/docs/conf.py index 6ab5959d5..935c3c9a1 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -27,6 +27,7 @@ sys.path.insert(0, path.dirname(path.dirname(__file__))) # Add any Sphinx extension module names here, as strings. They can be extensions # coming with Sphinx (named 'sphinx.ext.*') or your custom ones. extensions = [ + 'notfound.extension', 'scrapydocs', 'sphinx.ext.autodoc', 'sphinx.ext.coverage', diff --git a/docs/requirements.txt b/docs/requirements.txt index 379da9994..f9db85146 100644 --- a/docs/requirements.txt +++ b/docs/requirements.txt @@ -1,2 +1,3 @@ Sphinx>=2.1 -sphinx_rtd_theme \ No newline at end of file +sphinx-notfound-page +sphinx_rtd_theme From a3a3107bc45483e0d7c77e45412121fdb8d539b3 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Wed, 13 Nov 2019 09:46:54 +0100 Subject: [PATCH 238/387] MutableChain: return self from __iter__ --- scrapy/utils/python.py | 4 +--- 1 file changed, 1 insertion(+), 3 deletions(-) diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index ea5193f12..64402a2bb 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -388,9 +388,7 @@ class MutableChain(object): self.data = chain(self.data, *iterables) def __iter__(self): - return self.data.__iter__() + return self def __next__(self): return next(self.data) - - next = __next__ From 33ef24c757c797c804c8fd242b8ab5705219b452 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Wed, 13 Nov 2019 10:52:05 +0100 Subject: [PATCH 239/387] =?UTF-8?q?Add=20missing=20whitespace=20after=20?= =?UTF-8?q?=E2=80=98,=E2=80=99,=20=E2=80=98;=E2=80=99=20or=20=E2=80=98:?= =?UTF-8?q?=E2=80=99?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- pytest.ini | 16 ++++++++-------- scrapy/core/spidermw.py | 2 +- tests/test_http_request.py | 4 ++-- tests/test_item.py | 2 +- tests/test_linkextractors.py | 4 ++-- tests/test_logformatter.py | 2 +- tests/test_utils_conf.py | 2 +- tests/test_utils_console.py | 2 +- tests/test_utils_misc/__init__.py | 2 +- 9 files changed, 18 insertions(+), 18 deletions(-) diff --git a/pytest.ini b/pytest.ini index 8c5a2cd54..984930459 100644 --- a/pytest.ini +++ b/pytest.ini @@ -48,7 +48,7 @@ flake8-ignore = scrapy/core/engine.py E261 E501 E128 E127 E306 E502 scrapy/core/scheduler.py E501 scrapy/core/scraper.py E501 E306 E261 E128 W504 - scrapy/core/spidermw.py E501 E731 E502 E231 E126 E226 + scrapy/core/spidermw.py E501 E731 E502 E126 E226 scrapy/core/downloader/__init__.py F401 E501 scrapy/core/downloader/contextfactory.py E501 E128 E126 scrapy/core/downloader/middleware.py E501 E502 @@ -214,13 +214,13 @@ flake8-ignore = tests/test_feedexport.py E501 F401 F841 E241 tests/test_http_cookies.py E501 tests/test_http_headers.py E302 E501 - tests/test_http_request.py F401 E402 E501 E231 E261 E127 E128 W293 E502 E128 E502 E126 E123 + tests/test_http_request.py F401 E402 E501 E261 E127 E128 W293 E502 E128 E502 E126 E123 tests/test_http_response.py E501 E301 E502 E128 E265 - tests/test_item.py E701 E128 E231 F841 E306 + tests/test_item.py E701 E128 F841 E306 tests/test_link.py E501 - tests/test_linkextractors.py E501 E128 E231 E124 + tests/test_linkextractors.py E501 E128 E124 tests/test_loader.py E302 E501 E731 E303 E741 E128 E117 E241 - tests/test_logformatter.py E128 E501 E231 E122 E302 + tests/test_logformatter.py E128 E501 E122 E302 tests/test_mail.py E302 E128 E501 E305 tests/test_middleware.py E302 E501 E128 tests/test_pipeline_crawl.py E131 E501 E128 E126 @@ -239,8 +239,8 @@ flake8-ignore = tests/test_spidermiddleware_output_chain.py F401 E501 E302 W293 E226 tests/test_spidermiddleware_referer.py F401 E501 E302 F841 E125 E201 E261 E124 E501 E241 E121 tests/test_squeues.py E501 E302 E701 E741 - tests/test_utils_conf.py E501 E231 E303 E128 - tests/test_utils_console.py E302 E231 + tests/test_utils_conf.py E501 E303 E128 + tests/test_utils_console.py E302 tests/test_utils_curl.py E501 tests/test_utils_datatypes.py E402 E501 E305 tests/test_utils_defer.py E306 E261 E501 E302 F841 E226 @@ -269,4 +269,4 @@ flake8-ignore = tests/test_spiderloader/test_spiders/spider2.py E302 tests/test_spiderloader/test_spiders/spider3.py E302 tests/test_spiderloader/test_spiders/nested/spider4.py E302 - tests/test_utils_misc/__init__.py E501 E231 + tests/test_utils_misc/__init__.py E501 diff --git a/scrapy/core/spidermw.py b/scrapy/core/spidermw.py index b5f9837ff..00cee3ada 100644 --- a/scrapy/core/spidermw.py +++ b/scrapy/core/spidermw.py @@ -36,7 +36,7 @@ class SpiderMiddlewareManager(MiddlewareManager): self.methods['process_spider_exception'].appendleft(getattr(mw, 'process_spider_exception', None)) def scrape_response(self, scrape_func, response, request, spider): - fname = lambda f:'%s.%s' % ( + fname = lambda f: '%s.%s' % ( six.get_method_self(f).__class__.__name__, six.get_method_function(f).__name__) diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 96a4fb141..45a547f40 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -56,7 +56,7 @@ class RequestTest(unittest.TestCase): def test_headers(self): # Different ways of setting headers attribute url = 'http://www.scrapy.org' - headers = {b'Accept':'gzip', b'Custom-Header':'nothing to tell you'} + headers = {b'Accept': 'gzip', b'Custom-Header': 'nothing to tell you'} r = self.request_class(url=url, headers=headers) p = self.request_class(url=url, headers=r.headers) @@ -816,7 +816,7 @@ class FormRequestTest(RequestTest): """) - r1 = self.request_class.from_response(response, formdata={'two':'3'}) + r1 = self.request_class.from_response(response, formdata={'two': '3'}) self.assertEqual(r1.method, 'POST') self.assertEqual(r1.headers['Content-type'], b'application/x-www-form-urlencoded') fs = _qs(r1) diff --git a/tests/test_item.py b/tests/test_item.py index 947566686..7c9468f65 100644 --- a/tests/test_item.py +++ b/tests/test_item.py @@ -245,7 +245,7 @@ class ItemTest(unittest.TestCase): def test_copy(self): class TestItem(Item): name = Field() - item = TestItem({'name':'lower'}) + item = TestItem({'name': 'lower'}) copied_item = item.copy() self.assertNotEqual(id(item), id(copied_item)) copied_item['name'] = copied_item['name'].upper() diff --git a/tests/test_linkextractors.py b/tests/test_linkextractors.py index d96e259f6..57ef1694a 100644 --- a/tests/test_linkextractors.py +++ b/tests/test_linkextractors.py @@ -322,7 +322,7 @@ class Base: Link(url=page4_url, text=u'href with whitespaces'), ]) - lx = self.extractor_cls(attrs=("href","src"), tags=("a","area","img"), deny_extensions=()) + lx = self.extractor_cls(attrs=("href", "src"), tags=("a", "area", "img"), deny_extensions=()) self.assertEqual(lx.extract_links(self.response), [ Link(url='http://example.com/sample1.html', text=u''), Link(url='http://example.com/sample2.html', text=u'sample 2'), @@ -360,7 +360,7 @@ class Base: Link(url='http://example.com/sample2.html', text=u'sample 2'), ]) - lx = self.extractor_cls(tags=("a","img"), attrs=("href", "src"), deny_extensions=()) + lx = self.extractor_cls(tags=("a", "img"), attrs=("href", "src"), deny_extensions=()) self.assertEqual(lx.extract_links(response), [ Link(url='http://example.com/sample2.html', text=u'sample 2'), Link(url='http://example.com/sample2.jpg', text=u''), diff --git a/tests/test_logformatter.py b/tests/test_logformatter.py index eb9c4a561..b4ea30bb7 100644 --- a/tests/test_logformatter.py +++ b/tests/test_logformatter.py @@ -45,7 +45,7 @@ class LoggingContribTest(unittest.TestCase): "Crawled (200) (referer: http://example.com) ['cached']") def test_flags_in_request(self): - req = Request("http://www.example.com", flags=['test','flag']) + req = Request("http://www.example.com", flags=['test', 'flag']) res = Response("http://www.example.com") logkws = self.formatter.crawled(req, res, self.spider) logline = logkws['msg'] % logkws['args'] diff --git a/tests/test_utils_conf.py b/tests/test_utils_conf.py index 29937c189..02d8ba51e 100644 --- a/tests/test_utils_conf.py +++ b/tests/test_utils_conf.py @@ -79,7 +79,7 @@ class BuildComponentListTest(unittest.TestCase): self.assertRaises(ValueError, build_component_list, {}, d, convert=lambda x: x) d = {'one': {'a': 'a', 'b': 2}} self.assertRaises(ValueError, build_component_list, {}, d, convert=lambda x: x) - d = {'one': 'lorem ipsum',} + d = {'one': 'lorem ipsum'} self.assertRaises(ValueError, build_component_list, {}, d, convert=lambda x: x) diff --git a/tests/test_utils_console.py b/tests/test_utils_console.py index 65782747b..c2211848c 100644 --- a/tests/test_utils_console.py +++ b/tests/test_utils_console.py @@ -21,7 +21,7 @@ class UtilsConsoleTestCase(unittest.TestCase): shell = get_shell_embed_func(['invalid']) self.assertEqual(shell, None) - shell = get_shell_embed_func(['invalid','python']) + shell = get_shell_embed_func(['invalid', 'python']) self.assertTrue(callable(shell)) self.assertEqual(shell.__name__, '_embed_standard_shell') diff --git a/tests/test_utils_misc/__init__.py b/tests/test_utils_misc/__init__.py index e109d5343..de6f173e0 100644 --- a/tests/test_utils_misc/__init__.py +++ b/tests/test_utils_misc/__init__.py @@ -74,7 +74,7 @@ class UtilsMiscTestCase(unittest.TestCase): self.assertEqual(list(arg_to_iter(100)), [100]) self.assertEqual(list(arg_to_iter(l for l in 'abc')), ['a', 'b', 'c']) self.assertEqual(list(arg_to_iter([1, 2, 3])), [1, 2, 3]) - self.assertEqual(list(arg_to_iter({'a':1})), [{'a': 1}]) + self.assertEqual(list(arg_to_iter({'a': 1})), [{'a': 1}]) self.assertEqual(list(arg_to_iter(TestItem(name="john"))), [TestItem(name="john")]) def test_create_instance(self): From 1d7c8cb0b1d3aa225fd396f263938fd2f171fb73 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Mon, 22 Jul 2019 22:27:29 +0500 Subject: [PATCH 240/387] Remove six.PY2 and six.PY3 conditionals. --- docs/topics/downloader-middleware.rst | 6 +++--- scrapy/_monkeypatches.py | 10 ---------- scrapy/commands/fetch.py | 5 ++--- scrapy/crawler.py | 11 ----------- scrapy/exporters.py | 2 +- scrapy/extensions/feedexport.py | 3 +-- scrapy/http/request/form.py | 3 +-- scrapy/http/response/text.py | 3 --- scrapy/item.py | 8 +------- scrapy/link.py | 15 ++------------- scrapy/mail.py | 9 ++------- scrapy/settings/__init__.py | 8 +------- scrapy/settings/default_settings.py | 4 +--- scrapy/utils/boto.py | 10 +--------- scrapy/utils/conf.py | 5 +---- scrapy/utils/datatypes.py | 20 +++++++------------- scrapy/utils/gz.py | 13 +++---------- scrapy/utils/iterators.py | 6 +----- scrapy/utils/python.py | 10 ++-------- tests/__init__.py | 15 --------------- tests/test_http_request.py | 11 +++-------- tests/test_http_response.py | 3 +-- tests/test_item.py | 8 ++------ tests/test_link.py | 13 ++----------- tests/test_middleware.py | 21 ++++++--------------- tests/test_request_cb_kwargs.py | 8 +------- tests/test_settings/__init__.py | 8 +------- tests/test_utils_datatypes.py | 7 +------ tests/test_utils_python.py | 7 +++---- 29 files changed, 50 insertions(+), 202 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 8048e1c86..366b95510 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -474,7 +474,7 @@ DBM storage backend A DBM_ storage backend is also available for the HTTP cache middleware. - By default, it uses the anydbm_ module, but you can change it with the + By default, it uses the dbm_ module, but you can change it with the :setting:`HTTPCACHE_DBM_MODULE` setting. .. _httpcache-storage-custom: @@ -626,7 +626,7 @@ HTTPCACHE_DBM_MODULE .. versionadded:: 0.13 -Default: ``'anydbm'`` +Default: ``'dbm'`` The database module to use in the :ref:`DBM storage backend `. This setting is specific to the DBM backend. @@ -1202,4 +1202,4 @@ The default encoding for proxy authentication on :class:`HttpProxyMiddleware`. .. _DBM: https://en.wikipedia.org/wiki/Dbm -.. _anydbm: https://docs.python.org/2/library/anydbm.html +.. _dbm: https://docs.python.org/3/library/dbm.html diff --git a/scrapy/_monkeypatches.py b/scrapy/_monkeypatches.py index b68099cad..1f8067b35 100644 --- a/scrapy/_monkeypatches.py +++ b/scrapy/_monkeypatches.py @@ -1,16 +1,6 @@ -import six from six.moves import copyreg -if six.PY2: - from urlparse import urlparse - - # workaround for https://bugs.python.org/issue9374 - Python < 2.7.4 - if urlparse('s3://bucket/key?key=value').query != 'key=value': - from urlparse import uses_query - uses_query.append('s3') - - # Undo what Twisted's perspective broker adds to pickle register # to prevent bugs like Twisted#7989 while serializing requests import twisted.persisted.styles # NOQA diff --git a/scrapy/commands/fetch.py b/scrapy/commands/fetch.py index 7d4840529..d45133e0e 100644 --- a/scrapy/commands/fetch.py +++ b/scrapy/commands/fetch.py @@ -1,5 +1,5 @@ from __future__ import print_function -import sys, six +import sys from w3lib.url import is_url from scrapy.commands import ScrapyCommand @@ -45,8 +45,7 @@ class Command(ScrapyCommand): self._print_bytes(response.body) def _print_bytes(self, bytes_): - bytes_writer = sys.stdout if six.PY2 else sys.stdout.buffer - bytes_writer.write(bytes_ + b'\n') + sys.stdout.buffer.write(bytes_ + b'\n') def run(self, args, opts): if len(args) != 1 or not is_url(args[0]): diff --git a/scrapy/crawler.py b/scrapy/crawler.py index ded3c082b..19b998e0d 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -88,20 +88,9 @@ class Crawler(object): yield self.engine.open_spider(self.spider, start_requests) yield defer.maybeDeferred(self.engine.start) except Exception: - # In Python 2 reraising an exception after yield discards - # the original traceback (see https://bugs.python.org/issue7563), - # so sys.exc_info() workaround is used. - # This workaround also works in Python 3, but it is not needed, - # and it is slower, so in Python 3 we use native `raise`. - if six.PY2: - exc_info = sys.exc_info() - self.crawling = False if self.engine is not None: yield self.engine.close() - - if six.PY2: - six.reraise(*exc_info) raise def _create_spider(self, *args, **kwargs): diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 8ed8d55f1..0d9c35654 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -216,7 +216,7 @@ class CsvItemExporter(BaseItemExporter): write_through=True, encoding=self.encoding, newline='' # Windows needs this https://github.com/scrapy/scrapy/issues/3034 - ) if six.PY3 else file + ) self.csv_writer = csv.writer(self.stream, **kwargs) self._headers_not_written = True self._join_multivalued = join_multivalued diff --git a/scrapy/extensions/feedexport.py b/scrapy/extensions/feedexport.py index 41d68fb14..e2492d506 100644 --- a/scrapy/extensions/feedexport.py +++ b/scrapy/extensions/feedexport.py @@ -10,7 +10,6 @@ import logging import posixpath from tempfile import NamedTemporaryFile from datetime import datetime -import six from six.moves.urllib.parse import urlparse, unquote from ftplib import FTP @@ -65,7 +64,7 @@ class StdoutFeedStorage(object): def __init__(self, uri, _stdout=None): if not _stdout: - _stdout = sys.stdout if six.PY2 else sys.stdout.buffer + _stdout = sys.stdout.buffer self._stdout = _stdout def open(self, spider): diff --git a/scrapy/http/request/form.py b/scrapy/http/request/form.py index 3ce8fc48e..b6feede07 100644 --- a/scrapy/http/request/form.py +++ b/scrapy/http/request/form.py @@ -104,8 +104,7 @@ def _get_form(response, formname, formid, formnumber, formxpath): el = el.getparent() if el is None: break - encoded = formxpath if six.PY3 else formxpath.encode('unicode_escape') - raise ValueError('No
element found with %s' % encoded) + raise ValueError('No element found with %s' % formxpath) # If we get here, it means that either formname was None # or invalid diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 339913d4e..a8010877c 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -32,9 +32,6 @@ class TextResponse(Response): def _set_url(self, url): if isinstance(url, six.text_type): - if six.PY2 and self.encoding is None: - raise TypeError("Cannot convert unicode url - %s " - "has no encoding" % type(self).__name__) self._url = to_native_str(url, self.encoding) else: super(TextResponse, self)._set_url(url) diff --git a/scrapy/item.py b/scrapy/item.py index 73b8f54b0..32f9b2ebb 100644 --- a/scrapy/item.py +++ b/scrapy/item.py @@ -4,8 +4,8 @@ Scrapy Item See documentation in docs/topics/item.rst """ -import collections from abc import ABCMeta +from collections.abc import MutableMapping from copy import deepcopy from pprint import pformat from warnings import warn @@ -16,12 +16,6 @@ from scrapy.utils.deprecate import ScrapyDeprecationWarning from scrapy.utils.trackref import object_ref -if six.PY2: - MutableMapping = collections.MutableMapping -else: - MutableMapping = collections.abc.MutableMapping - - class BaseItem(object_ref): """Base class for all scraped items. diff --git a/scrapy/link.py b/scrapy/link.py index f0638ced2..be1888ef0 100644 --- a/scrapy/link.py +++ b/scrapy/link.py @@ -4,12 +4,6 @@ This module defines the Link object used in Link extractors. For actual link extractors implementation see scrapy.linkextractors, or its documentation in: docs/topics/link-extractors.rst """ -import warnings -import six - -from scrapy.utils.python import to_bytes - - class Link(object): """Link objects represent an extracted link by the LinkExtractor.""" @@ -17,13 +11,8 @@ class Link(object): def __init__(self, url, text='', fragment='', nofollow=False): if not isinstance(url, str): - if six.PY2: - warnings.warn("Link urls must be str objects. " - "Assuming utf-8 encoding (which could be wrong)") - url = to_bytes(url, encoding='utf8') - else: - got = url.__class__.__name__ - raise TypeError("Link urls must be str objects, got %s" % got) + got = url.__class__.__name__ + raise TypeError("Link urls must be str objects, got %s" % got) self.url = url self.text = text self.fragment = fragment diff --git a/scrapy/mail.py b/scrapy/mail.py index 5b944e1c4..746468e25 100644 --- a/scrapy/mail.py +++ b/scrapy/mail.py @@ -9,18 +9,13 @@ try: from cStringIO import StringIO as BytesIO except ImportError: from io import BytesIO -import six from email.utils import COMMASPACE, formatdate from six.moves.email_mime_multipart import MIMEMultipart from six.moves.email_mime_text import MIMEText from six.moves.email_mime_base import MIMEBase -if six.PY2: - from email.MIMENonMultipart import MIMENonMultipart - from email import Encoders -else: - from email.mime.nonmultipart import MIMENonMultipart - from email import encoders as Encoders +from email.mime.nonmultipart import MIMENonMultipart +from email import encoders as Encoders from twisted.internet import defer, reactor, ssl diff --git a/scrapy/settings/__init__.py b/scrapy/settings/__init__.py index f28c7940d..c871e86e0 100644 --- a/scrapy/settings/__init__.py +++ b/scrapy/settings/__init__.py @@ -1,19 +1,13 @@ import six import json import copy -import collections +from collections.abc import MutableMapping from importlib import import_module from pprint import pformat from scrapy.settings import default_settings -if six.PY2: - MutableMapping = collections.MutableMapping -else: - MutableMapping = collections.abc.MutableMapping - - SETTINGS_PRIORITIES = { 'default': 0, 'command': 10, diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 9c22999cb..5c9678c01 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -17,8 +17,6 @@ import sys from importlib import import_module from os.path import join, abspath, dirname -import six - AJAXCRAWL_ENABLED = False AUTOTHROTTLE_ENABLED = False @@ -179,7 +177,7 @@ HTTPCACHE_ALWAYS_STORE = False HTTPCACHE_IGNORE_HTTP_CODES = [] HTTPCACHE_IGNORE_SCHEMES = ['file'] HTTPCACHE_IGNORE_RESPONSE_CACHE_CONTROLS = [] -HTTPCACHE_DBM_MODULE = 'anydbm' if six.PY2 else 'dbm' +HTTPCACHE_DBM_MODULE = 'dbm' HTTPCACHE_POLICY = 'scrapy.extensions.httpcache.DummyPolicy' HTTPCACHE_GZIP = False diff --git a/scrapy/utils/boto.py b/scrapy/utils/boto.py index 421ab2f7e..c8fc911bb 100644 --- a/scrapy/utils/boto.py +++ b/scrapy/utils/boto.py @@ -1,7 +1,6 @@ """Boto/botocore helpers""" from __future__ import absolute_import -import six from scrapy.exceptions import NotConfigured @@ -11,11 +10,4 @@ def is_botocore(): import botocore return True except ImportError: - if six.PY2: - try: - import boto - return False - except ImportError: - raise NotConfigured('missing botocore or boto library') - else: - raise NotConfigured('missing botocore library') + raise NotConfigured('missing botocore library') diff --git a/scrapy/utils/conf.py b/scrapy/utils/conf.py index fb7ca3310..561bb72fc 100644 --- a/scrapy/utils/conf.py +++ b/scrapy/utils/conf.py @@ -1,13 +1,10 @@ +from configparser import ConfigParser import os import sys import numbers from operator import itemgetter import six -if six.PY2: - from ConfigParser import SafeConfigParser as ConfigParser -else: - from configparser import ConfigParser from scrapy.settings import BaseSettings from scrapy.utils.deprecate import update_classpath diff --git a/scrapy/utils/datatypes.py b/scrapy/utils/datatypes.py index 87536e9d7..39d389fa6 100644 --- a/scrapy/utils/datatypes.py +++ b/scrapy/utils/datatypes.py @@ -7,6 +7,7 @@ This module must not depend on any module outside the Standard Library. import copy import collections +from collections.abc import Mapping import warnings import six @@ -14,12 +15,6 @@ import six from scrapy.exceptions import ScrapyDeprecationWarning -if six.PY2: - Mapping = collections.Mapping -else: - Mapping = collections.abc.Mapping - - class MultiValueDictKeyError(KeyError): def __init__(self, *args, **kwargs): warnings.warn( @@ -252,13 +247,12 @@ class MergeDict(object): first occurrence will be used. """ def __init__(self, *dicts): - if not six.PY2: - warnings.warn( - "scrapy.utils.datatypes.MergeDict is deprecated in favor " - "of collections.ChainMap (introduced in Python 3.3)", - category=ScrapyDeprecationWarning, - stacklevel=2, - ) + warnings.warn( + "scrapy.utils.datatypes.MergeDict is deprecated in favor " + "of collections.ChainMap (introduced in Python 3.3)", + category=ScrapyDeprecationWarning, + stacklevel=2, + ) self.dicts = dicts def __getitem__(self, key): diff --git a/scrapy/utils/gz.py b/scrapy/utils/gz.py index b3fb16b1e..9984492f0 100644 --- a/scrapy/utils/gz.py +++ b/scrapy/utils/gz.py @@ -6,7 +6,6 @@ except ImportError: from io import BytesIO from gzip import GzipFile -import six import re from scrapy.utils.decorators import deprecated @@ -17,14 +16,8 @@ from scrapy.utils.decorators import deprecated # (regression or bug-fix compared to Python 3.4) # - read1(), which fetches data before raising EOFError on next call # works here but is only available from Python>=3.3 -# - scrapy does not support Python 3.2 -# - Python 2.7 GzipFile works fine with standard read() + extrabuf -if six.PY2: - def read1(gzf, size=-1): - return gzf.read(size) -else: - def read1(gzf, size=-1): - return gzf.read1(size) +def read1(gzf, size=-1): + return gzf.read1(size) def gunzip(data): @@ -37,7 +30,7 @@ def gunzip(data): chunk = b'.' while chunk: try: - chunk = read1(f, 8196) + chunk = f.read1(8196) output_list.append(chunk) except (IOError, EOFError, struct.error): # complete only if there is some data, otherwise re-raise diff --git a/scrapy/utils/iterators.py b/scrapy/utils/iterators.py index a12e14005..dbc1e0d20 100644 --- a/scrapy/utils/iterators.py +++ b/scrapy/utils/iterators.py @@ -102,11 +102,7 @@ def csviter(obj, delimiter=None, headers=None, encoding=None, quotechar=None): def row_to_unicode(row_): return [to_unicode(field, encoding) for field in row_] - # Python 3 csv reader input object needs to return strings - if six.PY3: - lines = StringIO(_body_or_str(obj, unicode=True)) - else: - lines = BytesIO(_body_or_str(obj, unicode=False)) + lines = StringIO(_body_or_str(obj, unicode=True)) kwargs = {} if delimiter: kwargs["delimiter"] = delimiter diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index ea5193f12..37e6be868 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -113,10 +113,7 @@ def to_bytes(text, encoding=None, errors='strict'): def to_native_str(text, encoding=None, errors='strict'): """ Return str representation of ``text`` (bytes in Python 2.x and unicode in Python 3.x). """ - if six.PY2: - return to_bytes(text, encoding, errors) - else: - return to_unicode(text, encoding, errors) + return to_unicode(text, encoding, errors) def re_rsearch(pattern, text, chunk_size=1024): @@ -189,7 +186,7 @@ def _getargspec_py23(func): """_getargspec_py23(function) -> named tuple ArgSpec(args, varargs, keywords, defaults) - Identical to inspect.getargspec() in python2, but uses + Was identical to inspect.getargspec() in python2, but uses inspect.getfullargspec() for python3 behind the scenes to avoid DeprecationWarning. @@ -199,9 +196,6 @@ def _getargspec_py23(func): >>> _getargspec_py23(f) ArgSpec(args=['a', 'b'], varargs='ar', keywords='kw', defaults=(2,)) """ - if six.PY2: - return inspect.getargspec(func) - return inspect.ArgSpec(*inspect.getfullargspec(func)[:4]) diff --git a/tests/__init__.py b/tests/__init__.py index 9c9e35c35..a54367f8c 100644 --- a/tests/__init__.py +++ b/tests/__init__.py @@ -35,18 +35,3 @@ def get_testdata(*paths): path = os.path.join(tests_datadir, *paths) with open(path, 'rb') as f: return f.read() - - -# FIXME: delete after dropping py2 support -# Monkey patch the unittest module to prevent the -# DeprecationWarning about assertRaisesRegexp -> assertRaisesRegex -import six -if six.PY2: - import unittest - import twisted.trial.unittest - if not getattr(unittest.TestCase, 'assertRegex', None): - unittest.TestCase.assertRegex = unittest.TestCase.assertRegexpMatches - if not getattr(unittest.TestCase, 'assertRaisesRegex', None): - unittest.TestCase.assertRaisesRegex = unittest.TestCase.assertRaisesRegexp - if not getattr(twisted.trial.unittest.TestCase, 'assertRaisesRegex', None): - twisted.trial.unittest.TestCase.assertRaisesRegex = twisted.trial.unittest.TestCase.assertRaisesRegexp diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 96a4fb141..bb451b5f4 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -3,13 +3,12 @@ import cgi import unittest import re import json +from urllib.parse import unquote_to_bytes import warnings import six from six.moves import xmlrpc_client as xmlrpclib from six.moves.urllib.parse import urlparse, parse_qs, unquote -if six.PY3: - from urllib.parse import unquote_to_bytes from scrapy.http import Request, FormRequest, XmlRpcRequest, JsonRequest, Headers, HtmlResponse from scrapy.utils.python import to_bytes, to_native_str @@ -1064,8 +1063,7 @@ class FormRequestTest(RequestTest): self.assertEqual(fs, {}) xpath = u"//form[@name='\u03b1']" - encoded = xpath if six.PY3 else xpath.encode('unicode_escape') - self.assertRaisesRegex(ValueError, re.escape(encoded), + self.assertRaisesRegex(ValueError, re.escape(xpath), self.request_class.from_response, response, formxpath=xpath) @@ -1208,10 +1206,7 @@ def _qs(req, encoding='utf-8', to_unicode=False): qs = req.body else: qs = req.url.partition('?')[2] - if six.PY2: - uqs = unquote(to_native_str(qs, encoding)) - elif six.PY3: - uqs = unquote_to_bytes(qs) + uqs = unquote_to_bytes(qs) if to_unicode: uqs = uqs.decode(encoding) return parse_qs(uqs, True) diff --git a/tests/test_http_response.py b/tests/test_http_response.py index dfc8562f3..ec3b51086 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -21,8 +21,7 @@ class BaseResponseTest(unittest.TestCase): # Response requires url in the consturctor self.assertRaises(Exception, self.response_class) self.assertTrue(isinstance(self.response_class('http://example.com/'), self.response_class)) - if not six.PY2: - self.assertRaises(TypeError, self.response_class, b"http://example.com") + self.assertRaises(TypeError, self.response_class, b"http://example.com") # body can be str or None self.assertTrue(isinstance(self.response_class('http://example.com/', body=b''), self.response_class)) self.assertTrue(isinstance(self.response_class('http://example.com/', body=b'body'), self.response_class)) diff --git a/tests/test_item.py b/tests/test_item.py index 947566686..0ad278701 100644 --- a/tests/test_item.py +++ b/tests/test_item.py @@ -62,12 +62,8 @@ class ItemTest(unittest.TestCase): i['number'] = 123 itemrepr = repr(i) - if six.PY2: - self.assertEqual(itemrepr, - "{'name': u'John Doe', 'number': 123}") - else: - self.assertEqual(itemrepr, - "{'name': 'John Doe', 'number': 123}") + self.assertEqual(itemrepr, + "{'name': 'John Doe', 'number': 123}") i2 = eval(itemrepr) self.assertEqual(i2['name'], 'John Doe') diff --git a/tests/test_link.py b/tests/test_link.py index 955430b37..5e2ce5eeb 100644 --- a/tests/test_link.py +++ b/tests/test_link.py @@ -1,6 +1,4 @@ import unittest -import warnings -import six from scrapy.link import Link @@ -46,12 +44,5 @@ class LinkTest(unittest.TestCase): self._assert_same_links(l1, l2) def test_non_str_url_py2(self): - if six.PY2: - with warnings.catch_warnings(record=True) as w: - link = Link(u"http://www.example.com/\xa3") - self.assertIsInstance(link.url, str) - self.assertEqual(link.url, b'http://www.example.com/\xc2\xa3') - assert len(w) == 1, "warning not issued" - else: - with self.assertRaises(TypeError): - Link(b"http://www.example.com/\xc2\xa3") + with self.assertRaises(TypeError): + Link(b"http://www.example.com/\xc2\xa3") diff --git a/tests/test_middleware.py b/tests/test_middleware.py index aea0be825..af9b43d61 100644 --- a/tests/test_middleware.py +++ b/tests/test_middleware.py @@ -3,7 +3,6 @@ from twisted.trial import unittest from scrapy.settings import Settings from scrapy.exceptions import NotConfigured from scrapy.middleware import MiddlewareManager -import six class M1(object): @@ -66,20 +65,12 @@ class MiddlewareManagerTest(unittest.TestCase): def test_methods(self): mwman = TestMiddlewareManager(M1(), M2(), M3()) - if six.PY2: - self.assertEqual([x.im_class for x in mwman.methods['open_spider']], - [M1, M2]) - self.assertEqual([x.im_class for x in mwman.methods['close_spider']], - [M2, M1]) - self.assertEqual([x.im_class for x in mwman.methods['process']], - [M1, M3]) - else: - self.assertEqual([x.__self__.__class__ for x in mwman.methods['open_spider']], - [M1, M2]) - self.assertEqual([x.__self__.__class__ for x in mwman.methods['close_spider']], - [M2, M1]) - self.assertEqual([x.__self__.__class__ for x in mwman.methods['process']], - [M1, M3]) + self.assertEqual([x.__self__.__class__ for x in mwman.methods['open_spider']], + [M1, M2]) + self.assertEqual([x.__self__.__class__ for x in mwman.methods['close_spider']], + [M2, M1]) + self.assertEqual([x.__self__.__class__ for x in mwman.methods['process']], + [M1, M3]) def test_enabled(self): m1, m2, m3 = M1(), M2(), M3() diff --git a/tests/test_request_cb_kwargs.py b/tests/test_request_cb_kwargs.py index c9943faa8..a5cdc0de0 100644 --- a/tests/test_request_cb_kwargs.py +++ b/tests/test_request_cb_kwargs.py @@ -1,7 +1,6 @@ from testfixtures import LogCapture from twisted.internet import defer from twisted.trial.unittest import TestCase -import six from scrapy.http import Request from scrapy.crawler import CrawlerRunner @@ -161,9 +160,4 @@ class CallbackKeywordArgumentsTestCase(TestCase): self.assertEqual(exceptions['takes_less'].exc_info[0], TypeError) self.assertEqual(str(exceptions['takes_less'].exc_info[1]), "parse_takes_less() got an unexpected keyword argument 'number'") self.assertEqual(exceptions['takes_more'].exc_info[0], TypeError) - # py2 and py3 messages are different - exc_message = str(exceptions['takes_more'].exc_info[1]) - if six.PY2: - self.assertEqual(exc_message, "parse_takes_more() takes exactly 5 arguments (4 given)") - elif six.PY3: - self.assertEqual(exc_message, "parse_takes_more() missing 1 required positional argument: 'other'") + self.assertEqual(str(exceptions['takes_more'].exc_info[1]), "parse_takes_more() missing 1 required positional argument: 'other'") diff --git a/tests/test_settings/__init__.py b/tests/test_settings/__init__.py index 1dbacbea3..08286ff02 100644 --- a/tests/test_settings/__init__.py +++ b/tests/test_settings/__init__.py @@ -60,9 +60,6 @@ class SettingsAttributeTest(unittest.TestCase): class BaseSettingsTest(unittest.TestCase): - if six.PY3: - assertItemsEqual = unittest.TestCase.assertCountEqual - def setUp(self): self.settings = BaseSettings() @@ -152,7 +149,7 @@ class BaseSettingsTest(unittest.TestCase): self.settings.setmodule( 'tests.test_settings.default_settings', 10) - self.assertItemsEqual(six.iterkeys(self.settings.attributes), + self.assertCountEqual(six.iterkeys(self.settings.attributes), six.iterkeys(ctrl_attributes)) for key in six.iterkeys(ctrl_attributes): @@ -343,9 +340,6 @@ class BaseSettingsTest(unittest.TestCase): class SettingsTest(unittest.TestCase): - if six.PY3: - assertItemsEqual = unittest.TestCase.assertCountEqual - def setUp(self): self.settings = Settings() diff --git a/tests/test_utils_datatypes.py b/tests/test_utils_datatypes.py index 7e671f627..53228fc6e 100644 --- a/tests/test_utils_datatypes.py +++ b/tests/test_utils_datatypes.py @@ -1,11 +1,6 @@ import copy import unittest - -import six -if six.PY2: - from collections import Mapping, MutableMapping -else: - from collections.abc import Mapping, MutableMapping +from collections.abc import Mapping, MutableMapping from scrapy.utils.datatypes import CaselessDict, LocalCache, SequenceExclude diff --git a/tests/test_utils_python.py b/tests/test_utils_python.py index 3e1148354..096aa50b7 100644 --- a/tests/test_utils_python.py +++ b/tests/test_utils_python.py @@ -231,12 +231,11 @@ class UtilsPythonTestCase(unittest.TestCase): self.assertEqual(get_func_args(" ".join), []) self.assertEqual(get_func_args(operator.itemgetter(2)), []) else: - stripself = not six.PY2 # PyPy3 exposes them as methods self.assertEqual( - get_func_args(six.text_type.split, stripself), ['sep', 'maxsplit']) - self.assertEqual(get_func_args(" ".join, stripself), ['list']) + get_func_args(six.text_type.split, True), ['sep', 'maxsplit']) + self.assertEqual(get_func_args(" ".join, True), ['list']) self.assertEqual( - get_func_args(operator.itemgetter(2), stripself), ['obj']) + get_func_args(operator.itemgetter(2), True), ['obj']) def test_without_none_values(self): From 0e696ed06d2975ee86e7ce3a2d3892588c420a04 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Tue, 20 Aug 2019 20:56:26 +0500 Subject: [PATCH 241/387] Remove unneeded and unused code from XmlItemExporter. --- scrapy/exporters.py | 22 ++++------------------ 1 file changed, 4 insertions(+), 18 deletions(-) diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 0d9c35654..1aa195b0b 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -143,11 +143,11 @@ class XmlItemExporter(BaseItemExporter): def _beautify_newline(self, new_item=False): if self.indent is not None and (self.indent > 0 or new_item): - self._xg_characters('\n') + self.xg.characters('\n') def _beautify_indent(self, depth=1): if self.indent: - self._xg_characters(' ' * self.indent * depth) + self.xg.characters(' ' * self.indent * depth) def start_exporting(self): self.xg.startDocument() @@ -182,26 +182,12 @@ class XmlItemExporter(BaseItemExporter): self._export_xml_field('value', value, depth=depth+1) self._beautify_indent(depth=depth) elif isinstance(serialized_value, six.text_type): - self._xg_characters(serialized_value) + self.xg.characters(serialized_value) else: - self._xg_characters(str(serialized_value)) + self.xg.characters(str(serialized_value)) self.xg.endElement(name) self._beautify_newline() - # Workaround for https://bugs.python.org/issue17606 - # Before Python 2.7.4 xml.sax.saxutils required bytes; - # since 2.7.4 it requires unicode. The bug is likely to be - # fixed in 2.7.6, but 2.7.6 will still support unicode, - # and Python 3.x will require unicode, so ">= 2.7.4" should be fine. - if sys.version_info[:3] >= (2, 7, 4): - def _xg_characters(self, serialized_value): - if not isinstance(serialized_value, six.text_type): - serialized_value = serialized_value.decode(self.encoding) - return self.xg.characters(serialized_value) - else: # pragma: no cover - def _xg_characters(self, serialized_value): - return self.xg.characters(serialized_value) - class CsvItemExporter(BaseItemExporter): From 065fe29d3cc6634893dfead320779498fb061cd4 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Tue, 20 Aug 2019 21:06:52 +0500 Subject: [PATCH 242/387] Deprecate scrapy.utils.gz.read1. --- scrapy/utils/gz.py | 1 + 1 file changed, 1 insertion(+) diff --git a/scrapy/utils/gz.py b/scrapy/utils/gz.py index 9984492f0..dc8316d8c 100644 --- a/scrapy/utils/gz.py +++ b/scrapy/utils/gz.py @@ -16,6 +16,7 @@ from scrapy.utils.decorators import deprecated # (regression or bug-fix compared to Python 3.4) # - read1(), which fetches data before raising EOFError on next call # works here but is only available from Python>=3.3 +@deprecated('GzipFile.read1') def read1(gzf, size=-1): return gzf.read1(size) From 85e79ae792752353ea60bff3d93f9e77bea300fb Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Tue, 20 Aug 2019 21:09:49 +0500 Subject: [PATCH 243/387] Remove cStringIO imports. --- scrapy/downloadermiddlewares/decompression.py | 6 +----- scrapy/mail.py | 6 +----- scrapy/pipelines/files.py | 7 +------ scrapy/pipelines/images.py | 6 +----- scrapy/utils/gz.py | 9 ++------- scrapy/utils/iterators.py | 6 +----- tests/test_cmdline/__init__.py | 5 +---- tests/test_pipeline_media.py | 6 ++---- 8 files changed, 10 insertions(+), 41 deletions(-) diff --git a/scrapy/downloadermiddlewares/decompression.py b/scrapy/downloadermiddlewares/decompression.py index 49313cc04..e2d73f347 100644 --- a/scrapy/downloadermiddlewares/decompression.py +++ b/scrapy/downloadermiddlewares/decompression.py @@ -4,6 +4,7 @@ and extract the potentially compressed responses that may arrive. import bz2 import gzip +from io import BytesIO import zipfile import tarfile import logging @@ -11,11 +12,6 @@ from tempfile import mktemp import six -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO - from scrapy.responsetypes import responsetypes logger = logging.getLogger(__name__) diff --git a/scrapy/mail.py b/scrapy/mail.py index 746468e25..d24de2212 100644 --- a/scrapy/mail.py +++ b/scrapy/mail.py @@ -3,13 +3,9 @@ Mail sending helpers See documentation in docs/topics/email.rst """ +from io import BytesIO import logging -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO - from email.utils import COMMASPACE, formatdate from six.moves.email_mime_multipart import MIMEMultipart from six.moves.email_mime_text import MIMEText diff --git a/scrapy/pipelines/files.py b/scrapy/pipelines/files.py index cc3d10b63..8d74c5011 100644 --- a/scrapy/pipelines/files.py +++ b/scrapy/pipelines/files.py @@ -5,6 +5,7 @@ See documentation in topics/media-pipeline.rst """ import functools import hashlib +from io import BytesIO import mimetypes import os import os.path @@ -15,12 +16,6 @@ from six.moves.urllib.parse import urlparse from collections import defaultdict import six - -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO - from twisted.internet import defer, threads from scrapy.pipelines.media import MediaPipeline diff --git a/scrapy/pipelines/images.py b/scrapy/pipelines/images.py index fa4d12ad1..e77cef4ff 100644 --- a/scrapy/pipelines/images.py +++ b/scrapy/pipelines/images.py @@ -5,13 +5,9 @@ See documentation in topics/media-pipeline.rst """ import functools import hashlib +from io import BytesIO import six -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO - from PIL import Image from scrapy.utils.misc import md5sum diff --git a/scrapy/utils/gz.py b/scrapy/utils/gz.py index dc8316d8c..f41e62fe3 100644 --- a/scrapy/utils/gz.py +++ b/scrapy/utils/gz.py @@ -1,12 +1,7 @@ -import struct - -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO from gzip import GzipFile - +from io import BytesIO import re +import struct from scrapy.utils.decorators import deprecated diff --git a/scrapy/utils/iterators.py b/scrapy/utils/iterators.py index dbc1e0d20..9693ba768 100644 --- a/scrapy/utils/iterators.py +++ b/scrapy/utils/iterators.py @@ -1,11 +1,7 @@ import re import csv -import logging -try: - from cStringIO import StringIO as BytesIO -except ImportError: - from io import BytesIO from io import StringIO +import logging import six from scrapy.http import TextResponse, Response diff --git a/tests/test_cmdline/__init__.py b/tests/test_cmdline/__init__.py index 68dfb1cca..56cfe642a 100644 --- a/tests/test_cmdline/__init__.py +++ b/tests/test_cmdline/__init__.py @@ -1,3 +1,4 @@ +from io import StringIO import json import os import pstats @@ -7,10 +8,6 @@ from subprocess import Popen, PIPE import sys import tempfile import unittest -try: - from cStringIO import StringIO -except ImportError: - from io import StringIO from scrapy.utils.test import get_testenv diff --git a/tests/test_pipeline_media.py b/tests/test_pipeline_media.py index 28e39cefa..ad2618ec9 100644 --- a/tests/test_pipeline_media.py +++ b/tests/test_pipeline_media.py @@ -144,10 +144,8 @@ class BaseMediaPipelineTestCase(unittest.TestCase): # The Failure should encapsulate a FileException ... self.assertEqual(failure.value, file_exc) - # ... and if we're running on Python 3 ... - if sys.version_info.major >= 3: - # ... it should have the returnValue exception set as its context - self.assertEqual(failure.value.__context__, def_gen_return_exc) + # ... and it should have the returnValue exception set as its context + self.assertEqual(failure.value.__context__, def_gen_return_exc) # Let's calculate the request fingerprint and fake some runtime data... fp = request_fingerprint(request) From cfa633f5e865fa491257215cc912d7321cc6da78 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Tue, 20 Aug 2019 21:13:01 +0500 Subject: [PATCH 244/387] Some text function messages cleanup, deprecate to_native_str. --- scrapy/http/response/__init__.py | 2 +- scrapy/utils/python.py | 8 ++++---- 2 files changed, 5 insertions(+), 5 deletions(-) diff --git a/scrapy/http/response/__init__.py b/scrapy/http/response/__init__.py index b0a526b72..a81404afb 100644 --- a/scrapy/http/response/__init__.py +++ b/scrapy/http/response/__init__.py @@ -88,7 +88,7 @@ class Response(object_ref): @property def text(self): """For subclasses of TextResponse, this will return the body - as text (unicode object in Python 2 and str in Python 3) + as str """ raise AttributeError("Response content isn't text") diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index 37e6be868..663a8ebaa 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -90,7 +90,7 @@ def to_unicode(text, encoding=None, errors='strict'): if isinstance(text, six.text_type): return text if not isinstance(text, (bytes, six.text_type)): - raise TypeError('to_unicode must receive a bytes, str or unicode ' + raise TypeError('to_unicode must receive a bytes or str ' 'object, got %s' % type(text).__name__) if encoding is None: encoding = 'utf-8' @@ -103,16 +103,16 @@ def to_bytes(text, encoding=None, errors='strict'): if isinstance(text, bytes): return text if not isinstance(text, six.string_types): - raise TypeError('to_bytes must receive a unicode, str or bytes ' + raise TypeError('to_bytes must receive a str or bytes ' 'object, got %s' % type(text).__name__) if encoding is None: encoding = 'utf-8' return text.encode(encoding, errors) +@deprecated('to_unicode') def to_native_str(text, encoding=None, errors='strict'): - """ Return str representation of ``text`` - (bytes in Python 2.x and unicode in Python 3.x). """ + """ Return str representation of ``text``. """ return to_unicode(text, encoding, errors) From 92ffd2f249ec276436277385e525397a2fb8b0a0 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Tue, 20 Aug 2019 21:27:52 +0500 Subject: [PATCH 245/387] Simplify some more imports. --- scrapy/downloadermiddlewares/httpproxy.py | 5 +---- scrapy/loader/processors.py | 5 +---- tests/__init__.py | 5 ----- tests/test_downloader_handlers.py | 5 +---- tests/test_downloadermiddleware.py | 3 ++- tests/test_downloadermiddleware_robotstxt.py | 4 +++- tests/test_extension_telnet.py | 5 ----- tests/test_feedexport.py | 2 +- tests/test_http_request.py | 3 +-- tests/test_item.py | 2 +- tests/test_pipeline_files.py | 6 +----- tests/test_settings/__init__.py | 3 +-- tests/test_spider.py | 3 +-- tests/test_spidermiddleware.py | 3 ++- tests/test_stats.py | 6 +----- tests/test_utils_deprecate.py | 3 +-- tests/test_utils_misc/__init__.py | 2 +- tests/test_utils_trackref.py | 2 +- 18 files changed, 20 insertions(+), 47 deletions(-) diff --git a/scrapy/downloadermiddlewares/httpproxy.py b/scrapy/downloadermiddlewares/httpproxy.py index 2c35d1b90..2212d9688 100644 --- a/scrapy/downloadermiddlewares/httpproxy.py +++ b/scrapy/downloadermiddlewares/httpproxy.py @@ -1,10 +1,7 @@ import base64 from six.moves.urllib.parse import unquote, urlunparse from six.moves.urllib.request import getproxies, proxy_bypass -try: - from urllib2 import _parse_proxy -except ImportError: - from urllib.request import _parse_proxy +from urllib.request import _parse_proxy from scrapy.exceptions import NotConfigured from scrapy.utils.httpobj import urlparse_cached diff --git a/scrapy/loader/processors.py b/scrapy/loader/processors.py index 2acdc8093..02c625acc 100644 --- a/scrapy/loader/processors.py +++ b/scrapy/loader/processors.py @@ -3,10 +3,7 @@ This module provides some commonly used processors for Item Loaders. See documentation in docs/topics/loaders.rst """ -try: - from collections import ChainMap -except ImportError: - from scrapy.utils.datatypes import MergeDict as ChainMap +from collections import ChainMap from scrapy.utils.misc import arg_to_iter from scrapy.loader.common import wrap_loader_context diff --git a/tests/__init__.py b/tests/__init__.py index a54367f8c..12ce79fa9 100644 --- a/tests/__init__.py +++ b/tests/__init__.py @@ -21,11 +21,6 @@ if 'COV_CORE_CONFIG' in os.environ: os.environ['COV_CORE_CONFIG'] = os.path.join(_sourceroot, os.environ['COV_CORE_CONFIG']) -try: - import unittest.mock as mock -except ImportError: - import mock - tests_datadir = os.path.join(os.path.abspath(os.path.dirname(__file__)), 'sample_data') diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 109469503..6090998d4 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -2,11 +2,8 @@ import os import six import shutil import tempfile +from unittest import mock import contextlib -try: - from unittest import mock -except ImportError: - import mock from testfixtures import LogCapture from twisted.trial import unittest diff --git a/tests/test_downloadermiddleware.py b/tests/test_downloadermiddleware.py index 03564e748..6b9a5bee8 100644 --- a/tests/test_downloadermiddleware.py +++ b/tests/test_downloadermiddleware.py @@ -1,3 +1,5 @@ +from unittest import mock + from twisted.trial.unittest import TestCase from twisted.python.failure import Failure @@ -7,7 +9,6 @@ from scrapy.exceptions import _InvalidOutput from scrapy.core.downloader.middleware import DownloaderMiddlewareManager from scrapy.utils.test import get_crawler from scrapy.utils.python import to_bytes -from tests import mock class ManagerTestCase(TestCase): diff --git a/tests/test_downloadermiddleware_robotstxt.py b/tests/test_downloadermiddleware_robotstxt.py index fbc46cba4..8266bf35f 100644 --- a/tests/test_downloadermiddleware_robotstxt.py +++ b/tests/test_downloadermiddleware_robotstxt.py @@ -1,5 +1,8 @@ # -*- coding: utf-8 -*- from __future__ import absolute_import + +from unittest import mock + from twisted.internet import reactor, error from twisted.internet.defer import Deferred, DeferredList, maybeDeferred from twisted.python import failure @@ -9,7 +12,6 @@ from scrapy.downloadermiddlewares.robotstxt import (RobotsTxtMiddleware, from scrapy.exceptions import IgnoreRequest, NotConfigured from scrapy.http import Request, Response, TextResponse from scrapy.settings import Settings -from tests import mock from tests.test_robotstxt_interface import rerp_available, reppy_available diff --git a/tests/test_extension_telnet.py b/tests/test_extension_telnet.py index 4f389e5cb..875ceb83c 100644 --- a/tests/test_extension_telnet.py +++ b/tests/test_extension_telnet.py @@ -1,8 +1,3 @@ -try: - import unittest.mock as mock -except ImportError: - import mock - from twisted.trial import unittest from twisted.conch.telnet import ITelnetProtocol from twisted.cred import credentials diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 11a5a8279..0c70bf80e 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -7,6 +7,7 @@ from io import BytesIO import tempfile import shutil import string +from unittest import mock from six.moves.urllib.parse import urljoin, urlparse, quote from six.moves.urllib.request import pathname2url @@ -15,7 +16,6 @@ from twisted.trial import unittest from twisted.internet import defer from scrapy.crawler import CrawlerRunner from scrapy.settings import Settings -from tests import mock from tests.mockserver import MockServer from w3lib.url import path_to_file_uri diff --git a/tests/test_http_request.py b/tests/test_http_request.py index bb451b5f4..9fe201579 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -3,6 +3,7 @@ import cgi import unittest import re import json +from unittest import mock from urllib.parse import unquote_to_bytes import warnings @@ -13,8 +14,6 @@ from six.moves.urllib.parse import urlparse, parse_qs, unquote from scrapy.http import Request, FormRequest, XmlRpcRequest, JsonRequest, Headers, HtmlResponse from scrapy.utils.python import to_bytes, to_native_str -from tests import mock - class RequestTest(unittest.TestCase): diff --git a/tests/test_item.py b/tests/test_item.py index 0ad278701..d98c63ddd 100644 --- a/tests/test_item.py +++ b/tests/test_item.py @@ -1,12 +1,12 @@ import sys import unittest +from unittest import mock from warnings import catch_warnings import six from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.item import ABCMeta, DictItem, Field, Item, ItemMeta -from tests import mock PY36_PLUS = (sys.version_info.major >= 3) and (sys.version_info.minor >= 6) diff --git a/tests/test_pipeline_files.py b/tests/test_pipeline_files.py index cb8f8da18..bd40e4103 100644 --- a/tests/test_pipeline_files.py +++ b/tests/test_pipeline_files.py @@ -1,10 +1,9 @@ import os import random import time -import hashlib -import warnings from tempfile import mkdtemp from shutil import rmtree +from unittest import mock from six.moves.urllib.parse import urlparse from six import BytesIO @@ -15,13 +14,10 @@ from scrapy.pipelines.files import FilesPipeline, FSFilesStore, S3FilesStore, GC from scrapy.item import Item, Field from scrapy.http import Request, Response from scrapy.settings import Settings -from scrapy.utils.python import to_bytes from scrapy.utils.test import assert_aws_environ, get_s3_content_and_delete from scrapy.utils.test import assert_gcs_environ, get_gcs_content_and_delete from scrapy.utils.boto import is_botocore -from tests import mock - def _mocked_download_func(request, info): response = request.meta.get('response') diff --git a/tests/test_settings/__init__.py b/tests/test_settings/__init__.py index 08286ff02..32e65bed5 100644 --- a/tests/test_settings/__init__.py +++ b/tests/test_settings/__init__.py @@ -1,10 +1,9 @@ import six import unittest -import warnings +from unittest import mock from scrapy.settings import (BaseSettings, Settings, SettingsAttribute, SETTINGS_PRIORITIES, get_settings_priority) -from tests import mock from . import default_settings diff --git a/tests/test_spider.py b/tests/test_spider.py index 6f6cdb8ff..c0fccfdd6 100644 --- a/tests/test_spider.py +++ b/tests/test_spider.py @@ -1,5 +1,6 @@ import gzip import inspect +from unittest import mock import warnings from io import BytesIO @@ -17,8 +18,6 @@ from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.utils.trackref import object_ref from scrapy.utils.test import get_crawler -from tests import mock - class SpiderTest(unittest.TestCase): diff --git a/tests/test_spidermiddleware.py b/tests/test_spidermiddleware.py index 832fd3330..55d665e79 100644 --- a/tests/test_spidermiddleware.py +++ b/tests/test_spidermiddleware.py @@ -1,3 +1,5 @@ +from unittest import mock + from twisted.trial.unittest import TestCase from twisted.python.failure import Failure @@ -6,7 +8,6 @@ from scrapy.http import Request, Response from scrapy.exceptions import _InvalidOutput from scrapy.utils.test import get_crawler from scrapy.core.spidermw import SpiderMiddlewareManager -from tests import mock class SpiderMiddlewareTestCase(TestCase): diff --git a/tests/test_stats.py b/tests/test_stats.py index 2033dbe07..2bbbb9e2c 100644 --- a/tests/test_stats.py +++ b/tests/test_stats.py @@ -1,10 +1,6 @@ from datetime import datetime import unittest - -try: - from unittest import mock -except ImportError: - import mock +from unittest import mock from scrapy.extensions.corestats import CoreStats from scrapy.spiders import Spider diff --git a/tests/test_utils_deprecate.py b/tests/test_utils_deprecate.py index 3e7236fb1..ce04e7f29 100644 --- a/tests/test_utils_deprecate.py +++ b/tests/test_utils_deprecate.py @@ -2,12 +2,11 @@ from __future__ import absolute_import import inspect import unittest +from unittest import mock import warnings from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.utils.deprecate import create_deprecated_class, update_classpath -from tests import mock - class MyWarning(UserWarning): pass diff --git a/tests/test_utils_misc/__init__.py b/tests/test_utils_misc/__init__.py index e109d5343..de9da9104 100644 --- a/tests/test_utils_misc/__init__.py +++ b/tests/test_utils_misc/__init__.py @@ -1,11 +1,11 @@ import sys import os import unittest +from unittest import mock from scrapy.item import Item, Field from scrapy.utils.misc import arg_to_iter, create_instance, load_object, set_environ, walk_modules -from tests import mock __doctests__ = ['scrapy.utils.misc'] diff --git a/tests/test_utils_trackref.py b/tests/test_utils_trackref.py index c6072fc0d..480a717e7 100644 --- a/tests/test_utils_trackref.py +++ b/tests/test_utils_trackref.py @@ -1,7 +1,7 @@ import six import unittest +from unittest import mock from scrapy.utils import trackref -from tests import mock class Foo(trackref.object_ref): From a138fb05d4f0d90e2002e85a348a5be34904d3d8 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Tue, 20 Aug 2019 21:35:13 +0500 Subject: [PATCH 246/387] Replace to_native_str calls with to_unicode. --- scrapy/core/downloader/handlers/http11.py | 2 +- scrapy/downloadermiddlewares/cookies.py | 6 +++--- scrapy/downloadermiddlewares/robotstxt.py | 3 --- scrapy/exporters.py | 4 ++-- scrapy/http/cookies.py | 10 +++++----- scrapy/http/response/text.py | 8 ++++---- scrapy/linkextractors/lxmlhtml.py | 4 ++-- scrapy/responsetypes.py | 6 +++--- scrapy/robotstxt.py | 8 ++++---- scrapy/spidermiddlewares/referer.py | 5 ++--- scrapy/utils/reqser.py | 4 ++-- scrapy/utils/request.py | 4 ++-- scrapy/utils/response.py | 4 ++-- scrapy/utils/ssl.py | 4 ++-- tests/test_command_parse.py | 5 ++--- tests/test_commands.py | 7 ++----- tests/test_feedexport.py | 7 +++---- tests/test_http_request.py | 7 +++---- tests/test_http_response.py | 6 +++--- tests/test_robotstxt_interface.py | 1 - 20 files changed, 47 insertions(+), 58 deletions(-) diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index 91b45a8fc..7d917cb74 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -174,7 +174,7 @@ def tunnel_request_data(host, port, proxy_auth_header=None): r""" Return binary content of a CONNECT request. - >>> from scrapy.utils.python import to_native_str as s + >>> from scrapy.utils.python import to_unicode as s >>> s(tunnel_request_data("example.com", 8080)) 'CONNECT example.com:8080 HTTP/1.1\r\nHost: example.com:8080\r\n\r\n' >>> s(tunnel_request_data("example.com", 8080, b"123")) diff --git a/scrapy/downloadermiddlewares/cookies.py b/scrapy/downloadermiddlewares/cookies.py index 321c0171b..0d2b9900c 100644 --- a/scrapy/downloadermiddlewares/cookies.py +++ b/scrapy/downloadermiddlewares/cookies.py @@ -6,7 +6,7 @@ from collections import defaultdict from scrapy.exceptions import NotConfigured from scrapy.http import Response from scrapy.http.cookies import CookieJar -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode logger = logging.getLogger(__name__) @@ -53,7 +53,7 @@ class CookiesMiddleware(object): def _debug_cookie(self, request, spider): if self.debug: - cl = [to_native_str(c, errors='replace') + cl = [to_unicode(c, errors='replace') for c in request.headers.getlist('Cookie')] if cl: cookies = "\n".join("Cookie: {}\n".format(c) for c in cl) @@ -62,7 +62,7 @@ class CookiesMiddleware(object): def _debug_set_cookie(self, response, spider): if self.debug: - cl = [to_native_str(c, errors='replace') + cl = [to_unicode(c, errors='replace') for c in response.headers.getlist('Set-Cookie')] if cl: cookies = "\n".join("Set-Cookie: {}\n".format(c) for c in cl) diff --git a/scrapy/downloadermiddlewares/robotstxt.py b/scrapy/downloadermiddlewares/robotstxt.py index 6a5dfb79c..251706c50 100644 --- a/scrapy/downloadermiddlewares/robotstxt.py +++ b/scrapy/downloadermiddlewares/robotstxt.py @@ -5,15 +5,12 @@ enable this middleware and enable the ROBOTSTXT_OBEY setting. """ import logging -import sys -import re from twisted.internet.defer import Deferred, maybeDeferred from scrapy.exceptions import NotConfigured, IgnoreRequest from scrapy.http import Request from scrapy.utils.httpobj import urlparse_cached from scrapy.utils.log import failure_to_exc_info -from scrapy.utils.python import to_native_str from scrapy.utils.misc import load_object logger = logging.getLogger(__name__) diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 1aa195b0b..8eb52995e 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -12,7 +12,7 @@ from six.moves import cPickle as pickle from xml.sax.saxutils import XMLGenerator from scrapy.utils.serialize import ScrapyJSONEncoder -from scrapy.utils.python import to_bytes, to_unicode, to_native_str, is_listlike +from scrapy.utils.python import to_bytes, to_unicode, is_listlike from scrapy.item import BaseItem from scrapy.exceptions import ScrapyDeprecationWarning import warnings @@ -232,7 +232,7 @@ class CsvItemExporter(BaseItemExporter): def _build_row(self, values): for s in values: try: - yield to_native_str(s, self.encoding) + yield to_unicode(s, self.encoding) except TypeError: yield s diff --git a/scrapy/http/cookies.py b/scrapy/http/cookies.py index 4e8056750..4532c3ab7 100644 --- a/scrapy/http/cookies.py +++ b/scrapy/http/cookies.py @@ -3,7 +3,7 @@ from six.moves.http_cookiejar import ( CookieJar as _CookieJar, DefaultCookiePolicy, IPV4_RE ) from scrapy.utils.httpobj import urlparse_cached -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode class CookieJar(object): @@ -165,13 +165,13 @@ class WrappedRequest(object): return name in self.request.headers def get_header(self, name, default=None): - return to_native_str(self.request.headers.get(name, default), + return to_unicode(self.request.headers.get(name, default), errors='replace') def header_items(self): return [ - (to_native_str(k, errors='replace'), - [to_native_str(x, errors='replace') for x in v]) + (to_unicode(k, errors='replace'), + [to_unicode(x, errors='replace') for x in v]) for k, v in self.request.headers.items() ] @@ -189,7 +189,7 @@ class WrappedResponse(object): # python3 cookiejars calls get_all def get_all(self, name, default=None): - return [to_native_str(v, errors='replace') + return [to_unicode(v, errors='replace') for v in self.response.headers.getlist(name)] # python2 cookiejars calls getheaders getheaders = get_all diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index a8010877c..37f450e54 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -16,7 +16,7 @@ from w3lib.html import strip_html5_whitespace from scrapy.http.request import Request from scrapy.http.response import Response from scrapy.utils.response import get_base_url -from scrapy.utils.python import memoizemethod_noargs, to_native_str +from scrapy.utils.python import memoizemethod_noargs, to_unicode class TextResponse(Response): @@ -32,7 +32,7 @@ class TextResponse(Response): def _set_url(self, url): if isinstance(url, six.text_type): - self._url = to_native_str(url, self.encoding) + self._url = to_unicode(url, self.encoding) else: super(TextResponse, self)._set_url(url) @@ -81,11 +81,11 @@ class TextResponse(Response): @memoizemethod_noargs def _headers_encoding(self): content_type = self.headers.get(b'Content-Type', b'') - return http_content_type_encoding(to_native_str(content_type)) + return http_content_type_encoding(to_unicode(content_type)) def _body_inferred_encoding(self): if self._cached_benc is None: - content_type = to_native_str(self.headers.get(b'Content-Type', b'')) + content_type = to_unicode(self.headers.get(b'Content-Type', b'')) benc, ubody = html_to_unicode(content_type, self.body, auto_detect_fun=self._auto_detect_fun, default_encoding=self._DEFAULT_ENCODING) diff --git a/scrapy/linkextractors/lxmlhtml.py b/scrapy/linkextractors/lxmlhtml.py index 8f6f93a44..890c019c8 100644 --- a/scrapy/linkextractors/lxmlhtml.py +++ b/scrapy/linkextractors/lxmlhtml.py @@ -10,7 +10,7 @@ from w3lib.url import canonicalize_url from scrapy.link import Link from scrapy.utils.misc import arg_to_iter, rel_has_nofollow -from scrapy.utils.python import unique as unique_list, to_native_str +from scrapy.utils.python import unique as unique_list, to_unicode from scrapy.utils.response import get_base_url from scrapy.linkextractors import FilteringLinkExtractor @@ -67,7 +67,7 @@ class LxmlParserLinkExtractor(object): url = self.process_attr(attr_val) if url is None: continue - url = to_native_str(url, encoding=response_encoding) + url = to_unicode(url, encoding=response_encoding) # to fix relative links after process_value url = urljoin(response_url, url) link = Link(url, _collect_string_content(el) or u'', diff --git a/scrapy/responsetypes.py b/scrapy/responsetypes.py index 4a2d5bf52..de62276c8 100644 --- a/scrapy/responsetypes.py +++ b/scrapy/responsetypes.py @@ -10,7 +10,7 @@ import six from scrapy.http import Response from scrapy.utils.misc import load_object -from scrapy.utils.python import binary_is_text, to_bytes, to_native_str +from scrapy.utils.python import binary_is_text, to_bytes, to_unicode class ResponseTypes(object): @@ -55,12 +55,12 @@ class ResponseTypes(object): header """ if content_encoding: return Response - mimetype = to_native_str(content_type).split(';')[0].strip().lower() + mimetype = to_unicode(content_type).split(';')[0].strip().lower() return self.from_mimetype(mimetype) def from_content_disposition(self, content_disposition): try: - filename = to_native_str(content_disposition, + filename = to_unicode(content_disposition, encoding='latin-1', errors='replace').split(';')[1].split('=')[1] filename = filename.strip('"\'') return self.from_filename(filename) diff --git a/scrapy/robotstxt.py b/scrapy/robotstxt.py index 189f165d1..95a8c09b8 100644 --- a/scrapy/robotstxt.py +++ b/scrapy/robotstxt.py @@ -3,14 +3,14 @@ import logging from abc import ABCMeta, abstractmethod from six import with_metaclass -from scrapy.utils.python import to_native_str, to_unicode +from scrapy.utils.python import to_unicode logger = logging.getLogger(__name__) def decode_robotstxt(robotstxt_body, spider, to_native_str_type=False): try: if to_native_str_type: - robotstxt_body = to_native_str(robotstxt_body) + robotstxt_body = to_unicode(robotstxt_body) else: robotstxt_body = robotstxt_body.decode('utf-8') except UnicodeDecodeError: @@ -66,8 +66,8 @@ class PythonRobotParser(RobotParser): return o def allowed(self, url, user_agent): - user_agent = to_native_str(user_agent) - url = to_native_str(url) + user_agent = to_unicode(user_agent) + url = to_unicode(url) return self.rp.can_fetch(user_agent, url) diff --git a/scrapy/spidermiddlewares/referer.py b/scrapy/spidermiddlewares/referer.py index 1ddfb37f4..c76e4d5a2 100644 --- a/scrapy/spidermiddlewares/referer.py +++ b/scrapy/spidermiddlewares/referer.py @@ -10,8 +10,7 @@ from w3lib.url import safe_url_string from scrapy.http import Request, Response from scrapy.exceptions import NotConfigured from scrapy import signals -from scrapy.utils.python import to_native_str -from scrapy.utils.httpobj import urlparse_cached +from scrapy.utils.python import to_unicode from scrapy.utils.misc import load_object from scrapy.utils.url import strip_url @@ -322,7 +321,7 @@ class RefererMiddleware(object): if isinstance(resp_or_url, Response): policy_header = resp_or_url.headers.get('Referrer-Policy') if policy_header is not None: - policy_name = to_native_str(policy_header.decode('latin1')) + policy_name = to_unicode(policy_header.decode('latin1')) if policy_name is None: return self.default_policy() diff --git a/scrapy/utils/reqser.py b/scrapy/utils/reqser.py index c7ea7b425..495564ac0 100644 --- a/scrapy/utils/reqser.py +++ b/scrapy/utils/reqser.py @@ -4,7 +4,7 @@ Helper functions for serializing (and deserializing) requests. import six from scrapy.http import Request -from scrapy.utils.python import to_unicode, to_native_str +from scrapy.utils.python import to_unicode from scrapy.utils.misc import load_object @@ -54,7 +54,7 @@ def request_from_dict(d, spider=None): eb = _get_method(spider, eb) request_cls = load_object(d['_class']) if '_class' in d else Request return request_cls( - url=to_native_str(d['url']), + url=to_unicode(d['url']), callback=cb, errback=eb, method=d['method'], diff --git a/scrapy/utils/request.py b/scrapy/utils/request.py index fb5af66a2..63d0ae772 100644 --- a/scrapy/utils/request.py +++ b/scrapy/utils/request.py @@ -9,7 +9,7 @@ import weakref from six.moves.urllib.parse import urlunparse from w3lib.http import basic_auth_header -from scrapy.utils.python import to_bytes, to_native_str +from scrapy.utils.python import to_bytes, to_unicode from w3lib.url import canonicalize_url from scrapy.utils.httpobj import urlparse_cached @@ -97,4 +97,4 @@ def referer_str(request): referrer = request.headers.get('Referer') if referrer is None: return referrer - return to_native_str(referrer, errors='replace') + return to_unicode(referrer, errors='replace') diff --git a/scrapy/utils/response.py b/scrapy/utils/response.py index c3236afd4..feab07431 100644 --- a/scrapy/utils/response.py +++ b/scrapy/utils/response.py @@ -8,7 +8,7 @@ import webbrowser import tempfile from twisted.web import http -from scrapy.utils.python import to_bytes, to_native_str +from scrapy.utils.python import to_bytes, to_unicode from w3lib import html @@ -36,7 +36,7 @@ def response_status_message(status): """Return status code plus status text descriptive message """ message = http.RESPONSES.get(int(status), "Unknown Status") - return '%s %s' % (status, to_native_str(message)) + return '%s %s' % (status, to_unicode(message)) def response_httprepr(response): diff --git a/scrapy/utils/ssl.py b/scrapy/utils/ssl.py index 02aed60ee..6e81b33ff 100644 --- a/scrapy/utils/ssl.py +++ b/scrapy/utils/ssl.py @@ -3,7 +3,7 @@ import OpenSSL import OpenSSL._util as pyOpenSSLutil -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode # The OpenSSL symbol is present since 1.1.1 but it's not currently supported in any version of pyOpenSSL. @@ -12,7 +12,7 @@ SSL_OP_NO_TLSv1_3 = getattr(pyOpenSSLutil.lib, 'SSL_OP_NO_TLSv1_3', 0) def ffi_buf_to_string(buf): - return to_native_str(pyOpenSSLutil.ffi.string(buf)) + return to_unicode(pyOpenSSLutil.ffi.string(buf)) def x509name_to_string(x509name): diff --git a/tests/test_command_parse.py b/tests/test_command_parse.py index 62d5d76b4..b134beb88 100644 --- a/tests/test_command_parse.py +++ b/tests/test_command_parse.py @@ -1,17 +1,16 @@ import os from os.path import join, abspath -from twisted.trial import unittest from twisted.internet import defer from scrapy.utils.testsite import SiteTest from scrapy.utils.testproc import ProcessTest -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode from tests.test_commands import CommandTest def _textmode(bstr): """Normalize input the same as writing to a file and reading from it in text mode""" - return to_native_str(bstr).replace(os.linesep, '\n') + return to_unicode(bstr).replace(os.linesep, '\n') class ParseCommandTest(ProcessTest, SiteTest, CommandTest): command = 'parse' diff --git a/tests/test_commands.py b/tests/test_commands.py index b8445ae6c..536379170 100644 --- a/tests/test_commands.py +++ b/tests/test_commands.py @@ -10,13 +10,10 @@ from contextlib import contextmanager from threading import Timer from twisted.trial import unittest -from twisted.internet import defer import scrapy -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode from scrapy.utils.test import get_testenv -from scrapy.utils.testsite import SiteTest -from scrapy.utils.testproc import ProcessTest from tests.test_crawler import ExceptionSpider, NoRequestsSpider @@ -56,7 +53,7 @@ class ProjectTest(unittest.TestCase): finally: timer.cancel() - return p, to_native_str(stdout), to_native_str(stderr) + return p, to_unicode(stdout), to_unicode(stderr) class StartprojectTest(ProjectTest): diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 0c70bf80e..87139e81f 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -26,8 +26,7 @@ from scrapy.extensions.feedexport import ( S3FeedStorage, StdoutFeedStorage, BlockingFeedStorage) from scrapy.utils.test import assert_aws_environ, get_s3_content_and_delete, get_crawler -from scrapy.utils.python import to_native_str -from scrapy.utils.project import get_project_settings +from scrapy.utils.python import to_unicode from pathlib import Path @@ -459,7 +458,7 @@ class FeedExportTest(unittest.TestCase): settings.update({'FEED_FORMAT': 'csv'}) data = yield self.exported_data(items, settings) - reader = csv.DictReader(to_native_str(data).splitlines()) + reader = csv.DictReader(to_unicode(data).splitlines()) got_rows = list(reader) if ordered: self.assertEqual(reader.fieldnames, header) @@ -473,7 +472,7 @@ class FeedExportTest(unittest.TestCase): settings = settings or {} settings.update({'FEED_FORMAT': 'jl'}) data = yield self.exported_data(items, settings) - parsed = [json.loads(to_native_str(line)) for line in data.splitlines()] + parsed = [json.loads(to_unicode(line)) for line in data.splitlines()] rows = [{k: v for k, v in row.items() if v} for row in rows] self.assertEqual(rows, parsed) diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 9fe201579..3518da21c 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -7,12 +7,11 @@ from unittest import mock from urllib.parse import unquote_to_bytes import warnings -import six from six.moves import xmlrpc_client as xmlrpclib from six.moves.urllib.parse import urlparse, parse_qs, unquote from scrapy.http import Request, FormRequest, XmlRpcRequest, JsonRequest, Headers, HtmlResponse -from scrapy.utils.python import to_bytes, to_native_str +from scrapy.utils.python import to_bytes, to_unicode class RequestTest(unittest.TestCase): @@ -349,8 +348,8 @@ class FormRequestTest(RequestTest): request_class = FormRequest def assertQueryEqual(self, first, second, msg=None): - first = to_native_str(first).split("&") - second = to_native_str(second).split("&") + first = to_unicode(first).split("&") + second = to_unicode(second).split("&") return self.assertEqual(sorted(first), sorted(second), msg) def test_empty_formdata(self): diff --git a/tests/test_http_response.py b/tests/test_http_response.py index ec3b51086..883c943da 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -7,7 +7,7 @@ from w3lib.encoding import resolve_encoding from scrapy.http import (Request, Response, TextResponse, HtmlResponse, XmlResponse, Headers) from scrapy.selector import Selector -from scrapy.utils.python import to_native_str +from scrapy.utils.python import to_unicode from scrapy.exceptions import NotSupported from scrapy.link import Link from tests import get_testdata @@ -204,11 +204,11 @@ class TextResponseTest(BaseResponseTest): assert isinstance(resp.url, str) resp = self.response_class(url=u"http://www.example.com/price/\xa3", encoding='utf-8') - self.assertEqual(resp.url, to_native_str(b'http://www.example.com/price/\xc2\xa3')) + self.assertEqual(resp.url, to_unicode(b'http://www.example.com/price/\xc2\xa3')) resp = self.response_class(url=u"http://www.example.com/price/\xa3", encoding='latin-1') self.assertEqual(resp.url, 'http://www.example.com/price/\xa3') resp = self.response_class(u"http://www.example.com/price/\xa3", headers={"Content-type": ["text/html; charset=utf-8"]}) - self.assertEqual(resp.url, to_native_str(b'http://www.example.com/price/\xc2\xa3')) + self.assertEqual(resp.url, to_unicode(b'http://www.example.com/price/\xc2\xa3')) resp = self.response_class(u"http://www.example.com/price/\xa3", headers={"Content-type": ["text/html; charset=iso-8859-1"]}) self.assertEqual(resp.url, 'http://www.example.com/price/\xa3') diff --git a/tests/test_robotstxt_interface.py b/tests/test_robotstxt_interface.py index 9aaab560a..cd7480e33 100644 --- a/tests/test_robotstxt_interface.py +++ b/tests/test_robotstxt_interface.py @@ -1,6 +1,5 @@ # coding=utf-8 from twisted.trial import unittest -from scrapy.utils.python import to_native_str def reppy_available(): From 87c23ba22d2ef714d778b62c794333acaf232f60 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Wed, 28 Aug 2019 16:29:53 +0500 Subject: [PATCH 247/387] Remove Py2-only code that checks sys.version_info. --- tests/test_utils_reqser.py | 7 ++----- 1 file changed, 2 insertions(+), 5 deletions(-) diff --git a/tests/test_utils_reqser.py b/tests/test_utils_reqser.py index 11ac56897..92cd16de7 100644 --- a/tests/test_utils_reqser.py +++ b/tests/test_utils_reqser.py @@ -80,8 +80,6 @@ class RequestSerializationTest(unittest.TestCase): self._assert_serializes_ok(r, spider=self.spider) def test_mixin_private_callback_serialization(self): - if sys.version_info[0] < 3: - return r = Request("http://www.example.com", callback=self.spider._TestSpiderMixin__mixin_callback, errback=self.spider.handle_error) @@ -119,9 +117,8 @@ class RequestSerializationTest(unittest.TestCase): def test_private_name_mangling(self): self._assert_mangles_to( self.spider, '_TestSpider__parse_item_private') - if sys.version_info[0] >= 3: - self._assert_mangles_to( - self.spider, '_TestSpiderMixin__mixin_callback') + self._assert_mangles_to( + self.spider, '_TestSpiderMixin__mixin_callback') def test_unserializable_callback1(self): r = Request("http://www.example.com", callback=lambda x: x) From a9c891399d1bf8a888392afe80c82a8ec8f2e8e7 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 31 Oct 2019 22:55:58 +0500 Subject: [PATCH 248/387] Fix a duplicate ref name in docs. --- docs/topics/downloader-middleware.rst | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 366b95510..e93645077 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -474,7 +474,7 @@ DBM storage backend A DBM_ storage backend is also available for the HTTP cache middleware. - By default, it uses the dbm_ module, but you can change it with the + By default, it uses the `dbm module`_, but you can change it with the :setting:`HTTPCACHE_DBM_MODULE` setting. .. _httpcache-storage-custom: @@ -1202,4 +1202,4 @@ The default encoding for proxy authentication on :class:`HttpProxyMiddleware`. .. _DBM: https://en.wikipedia.org/wiki/Dbm -.. _dbm: https://docs.python.org/3/library/dbm.html +.. _dbm module: https://docs.python.org/3/library/dbm.html From dd367438fa7a7fec923b28648c4e909cbed1b47d Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Fri, 1 Nov 2019 20:05:37 +0500 Subject: [PATCH 249/387] Improve the dbm module ref. MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves --- docs/topics/downloader-middleware.rst | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index e93645077..ae6d41809 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -474,7 +474,7 @@ DBM storage backend A DBM_ storage backend is also available for the HTTP cache middleware. - By default, it uses the `dbm module`_, but you can change it with the + By default, it uses the :mod:`dbm`, but you can change it with the :setting:`HTTPCACHE_DBM_MODULE` setting. .. _httpcache-storage-custom: @@ -1202,4 +1202,3 @@ The default encoding for proxy authentication on :class:`HttpProxyMiddleware`. .. _DBM: https://en.wikipedia.org/wiki/Dbm -.. _dbm module: https://docs.python.org/3/library/dbm.html From 1a4a77d49fa580d35d1e023ab4b54a397b88088a Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 14 Nov 2019 10:24:31 +0100 Subject: [PATCH 250/387] Remove Python 2 check from MutableChainTest --- tests/test_utils_python.py | 5 ++--- 1 file changed, 2 insertions(+), 3 deletions(-) diff --git a/tests/test_utils_python.py b/tests/test_utils_python.py index 3e1148354..6cae9793d 100644 --- a/tests/test_utils_python.py +++ b/tests/test_utils_python.py @@ -21,9 +21,8 @@ class MutableChainTest(unittest.TestCase): m.extend([7, 8]) m.extend([9, 10], (11, 12)) self.assertEqual(next(m), 0) - self.assertEqual(m.next(), 1) - self.assertEqual(m.__next__(), 2) - self.assertEqual(list(m), list(range(3, 13))) + self.assertEqual(m.__next__(), 1) + self.assertEqual(list(m), list(range(2, 13))) class ToUnicodeTest(unittest.TestCase): From be6da52019990c1db45b1101dd99787752a14313 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 14 Nov 2019 10:31:55 +0100 Subject: [PATCH 251/387] Include extensions from #2067 --- scrapy/linkextractors/__init__.py | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index 3c75e683d..a2ac963fe 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -19,9 +19,12 @@ from scrapy.utils.url import ( # common file extensions that are not followed if they occur in links IGNORED_EXTENSIONS = [ + # archives + '7z', '7zip', 'bz2', 'rar', 'tar', 'tar.gz', 'xz', 'zip', + # images 'mng', 'pct', 'bmp', 'gif', 'jpg', 'jpeg', 'png', 'pst', 'psp', 'tif', - 'tiff', 'ai', 'drw', 'dxf', 'eps', 'ps', 'svg', + 'tiff', 'ai', 'drw', 'dxf', 'eps', 'ps', 'svg', 'cdr', 'ico', # audio 'mp3', 'wma', 'ogg', 'wav', 'ra', 'aac', 'mid', 'au', 'aiff', @@ -35,7 +38,7 @@ IGNORED_EXTENSIONS = [ 'odp', # other - 'css', 'pdf', 'exe', 'bin', 'rss', 'zip', 'rar', 'dmg', 'iso', 'apk', 'xz' + 'css', 'pdf', 'exe', 'bin', 'rss', 'dmg', 'iso', 'apk' ] From 3631453bfb1fb2426916919dbfd489ba9ffd9505 Mon Sep 17 00:00:00 2001 From: Andrey Rahmatullin Date: Thu, 14 Nov 2019 15:07:53 +0500 Subject: [PATCH 252/387] Remove spaces on a blank line. --- scrapy/linkextractors/__init__.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index a2ac963fe..e4c62f87b 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -21,7 +21,7 @@ from scrapy.utils.url import ( IGNORED_EXTENSIONS = [ # archives '7z', '7zip', 'bz2', 'rar', 'tar', 'tar.gz', 'xz', 'zip', - + # images 'mng', 'pct', 'bmp', 'gif', 'jpg', 'jpeg', 'png', 'pst', 'psp', 'tif', 'tiff', 'ai', 'drw', 'dxf', 'eps', 'ps', 'svg', 'cdr', 'ico', From e291460db67ef8c9e52de02a0a4a86f4bce39b9d Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 14 Nov 2019 15:24:37 +0500 Subject: [PATCH 253/387] Fix flake8-detected errors. --- scrapy/crawler.py | 1 - scrapy/exporters.py | 1 - scrapy/http/cookies.py | 2 +- scrapy/link.py | 2 ++ tests/test_pipeline_media.py | 2 -- 5 files changed, 3 insertions(+), 5 deletions(-) diff --git a/scrapy/crawler.py b/scrapy/crawler.py index 19b998e0d..8868a985b 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -3,7 +3,6 @@ import signal import logging import warnings -import sys from twisted.internet import reactor, defer from zope.interface.verify import verifyClass, DoesNotImplement diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 8eb52995e..3defafd60 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -4,7 +4,6 @@ Item Exporters are used to export/serialize items into different formats. import csv import io -import sys import pprint import marshal import six diff --git a/scrapy/http/cookies.py b/scrapy/http/cookies.py index 4532c3ab7..60a14c6f8 100644 --- a/scrapy/http/cookies.py +++ b/scrapy/http/cookies.py @@ -166,7 +166,7 @@ class WrappedRequest(object): def get_header(self, name, default=None): return to_unicode(self.request.headers.get(name, default), - errors='replace') + errors='replace') def header_items(self): return [ diff --git a/scrapy/link.py b/scrapy/link.py index be1888ef0..a809c5ca4 100644 --- a/scrapy/link.py +++ b/scrapy/link.py @@ -4,6 +4,8 @@ This module defines the Link object used in Link extractors. For actual link extractors implementation see scrapy.linkextractors, or its documentation in: docs/topics/link-extractors.rst """ + + class Link(object): """Link objects represent an extracted link by the LinkExtractor.""" diff --git a/tests/test_pipeline_media.py b/tests/test_pipeline_media.py index ad2618ec9..ad958e25f 100644 --- a/tests/test_pipeline_media.py +++ b/tests/test_pipeline_media.py @@ -1,7 +1,5 @@ from __future__ import print_function -import sys - from testfixtures import LogCapture from twisted.trial import unittest from twisted.python.failure import Failure From b8ef12cd4707f9e095abdd26cc137558009c335f Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 14 Nov 2019 12:10:25 +0100 Subject: [PATCH 254/387] Add bandit to CI --- .bandit.yml | 16 ++++++++++++++++ .travis.yml | 2 ++ tox.ini | 7 +++++++ 3 files changed, 25 insertions(+) create mode 100644 .bandit.yml diff --git a/.bandit.yml b/.bandit.yml new file mode 100644 index 000000000..00554587a --- /dev/null +++ b/.bandit.yml @@ -0,0 +1,16 @@ +skips: +- B101 +- B105 +- B303 +- B306 +- B307 +- B311 +- B320 +- B321 +- B402 +- B404 +- B406 +- B410 +- B503 +- B603 +- B605 diff --git a/.travis.yml b/.travis.yml index 0e77af9fd..9f477e860 100644 --- a/.travis.yml +++ b/.travis.yml @@ -7,6 +7,8 @@ branches: - /^\d\.\d+\.\d+(rc\d+|\.dev\d+)?$/ matrix: include: + - env: TOXENV=security + python: 3.8 - env: TOXENV=flake8 python: 3.8 - env: TOXENV=pypy3 diff --git a/tox.ini b/tox.ini index 3668058c3..cd575f3c5 100644 --- a/tox.ini +++ b/tox.ini @@ -62,6 +62,13 @@ basepython = pypy3 commands = py.test {posargs:docs scrapy tests} +[testenv:security] +basepython = python3.8 +deps = + bandit +commands = + bandit -r -c .bandit.yml {posargs:scrapy} + [testenv:flake8] basepython = python3.8 deps = From 5ee5508cc33116b3700c9b745138b56cc67951f7 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 14 Nov 2019 15:42:34 +0100 Subject: [PATCH 255/387] Have CI record the 10 slowest tests --- tox.ini | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/tox.ini b/tox.ini index 3668058c3..ec04035f5 100644 --- a/tox.ini +++ b/tox.ini @@ -21,7 +21,7 @@ passenv = GCS_TEST_FILE_URI GCS_PROJECT_ID commands = - py.test --cov=scrapy --cov-report= {posargs:docs scrapy tests} + py.test --cov=scrapy --cov-report= {posargs:--durations=10 docs scrapy tests} [testenv:py35] basepython = python3.5 @@ -60,7 +60,7 @@ basepython = python3.8 [testenv:pypy3] basepython = pypy3 commands = - py.test {posargs:docs scrapy tests} + py.test {posargs:--durations=10 docs scrapy tests} [testenv:flake8] basepython = python3.8 From 058bdda0afe967f99a3ffd59b395566b063d9c3f Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 14 Nov 2019 16:51:47 +0100 Subject: [PATCH 256/387] Improve the performance of the DOWNLOAD_DELAY test --- tests/test_crawl.py | 51 +++++++++++++++++++++++++++++++-------------- 1 file changed, 35 insertions(+), 16 deletions(-) diff --git a/tests/test_crawl.py b/tests/test_crawl.py index 3fc13eeb7..a524287eb 100644 --- a/tests/test_crawl.py +++ b/tests/test_crawl.py @@ -30,25 +30,44 @@ class CrawlTestCase(TestCase): self.assertEqual(len(crawler.spider.urls_visited), 11) # 10 + start_url @defer.inlineCallbacks - def test_delay(self): - # short to long delays - yield self._test_delay(0.2, False) - yield self._test_delay(1, False) - # randoms - yield self._test_delay(0.2, True) - yield self._test_delay(1, True) + def test_fixed_delay(self): + yield self._test_delay(total=3, delay=0.1) @defer.inlineCallbacks - def _test_delay(self, delay, randomize): - settings = {"DOWNLOAD_DELAY": delay, 'RANDOMIZE_DOWNLOAD_DELAY': randomize} + def test_randomized_delay(self): + yield self._test_delay(total=3, delay=0.1, randomize=True) + + @defer.inlineCallbacks + def _test_delay(self, total, delay, randomize=False): + crawl_kwargs = dict( + maxlatency=delay * 2, + mockserver=self.mockserver, + total=total, + ) + tolerance = (1 - (0.6 if randomize else 0.2)) + + settings = {"DOWNLOAD_DELAY": delay, + 'RANDOMIZE_DOWNLOAD_DELAY': randomize} crawler = CrawlerRunner(settings).create_crawler(FollowAllSpider) - yield crawler.crawl(maxlatency=delay * 2, mockserver=self.mockserver) - t = crawler.spider.times - totaltime = t[-1] - t[0] - avgd = totaltime / (len(t) - 1) - tolerance = 0.6 if randomize else 0.2 - self.assertTrue(avgd > delay * (1 - tolerance), - "download delay too small: %s" % avgd) + yield crawler.crawl(**crawl_kwargs) + times = crawler.spider.times + total_time = times[-1] - times[0] + average = total_time / (len(times) - 1) + self.assertTrue(average > delay * tolerance, + "download delay too small: %s" % average) + + # Ensure that the same test parameters would cause a failure if no + # download delay is set. Otherwise, it means we are using a combination + # of ``total`` and ``delay`` values that are too small for the test + # code above to have any meaning. + settings["DOWNLOAD_DELAY"] = 0 + crawler = CrawlerRunner(settings).create_crawler(FollowAllSpider) + yield crawler.crawl(**crawl_kwargs) + times = crawler.spider.times + total_time = times[-1] - times[0] + average = total_time / (len(times) - 1) + self.assertFalse(average > delay / tolerance, + "test total or delay values are too small") @defer.inlineCallbacks def test_timeout_success(self): From fe3a121f1358fc904915f5b32e276520523c553a Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 14 Nov 2019 22:50:53 +0500 Subject: [PATCH 257/387] Use kwargs when calling get_func_args. --- tests/test_utils_python.py | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/tests/test_utils_python.py b/tests/test_utils_python.py index 096aa50b7..a94398796 100644 --- a/tests/test_utils_python.py +++ b/tests/test_utils_python.py @@ -232,10 +232,10 @@ class UtilsPythonTestCase(unittest.TestCase): self.assertEqual(get_func_args(operator.itemgetter(2)), []) else: self.assertEqual( - get_func_args(six.text_type.split, True), ['sep', 'maxsplit']) - self.assertEqual(get_func_args(" ".join, True), ['list']) + get_func_args(six.text_type.split, stripself=True), ['sep', 'maxsplit']) + self.assertEqual(get_func_args(" ".join, stripself=True), ['list']) self.assertEqual( - get_func_args(operator.itemgetter(2), True), ['obj']) + get_func_args(operator.itemgetter(2), stripself=True), ['obj']) def test_without_none_values(self): From 3b2289ad012043b94b495d411316b2778bf3db35 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 14 Nov 2019 22:53:28 +0500 Subject: [PATCH 258/387] Rename test_non_str_url_py2 to test_bytes_url. --- tests/test_link.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_link.py b/tests/test_link.py index 5e2ce5eeb..e0f1efffa 100644 --- a/tests/test_link.py +++ b/tests/test_link.py @@ -43,6 +43,6 @@ class LinkTest(unittest.TestCase): l2 = eval(repr(l1)) self._assert_same_links(l1, l2) - def test_non_str_url_py2(self): + def test_bytes_url(self): with self.assertRaises(TypeError): Link(b"http://www.example.com/\xc2\xa3") From 77a84f620ffa5d001072c7122c0f38048fe15606 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Jos=C3=A9=20Alberto=20/=20Speedy?= Date: Fri, 15 Nov 2019 11:09:24 +0100 Subject: [PATCH 259/387] Fix string MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Adrián Chaves --- README.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.rst b/README.rst index 87eaac2af..4c038030f 100644 --- a/README.rst +++ b/README.rst @@ -62,7 +62,7 @@ directory. Releases ======== -You can check https://docs.scrapy.org/en/latest/news.html for release notes. +You can check https://docs.scrapy.org/en/latest/news.html for the release notes. Community (blog, twitter, mail list, IRC) ========================================= From 0e252f5a13be3195fdc3ec1d66a111ae01a0ab80 Mon Sep 17 00:00:00 2001 From: Marc Hernandez Cabot Date: Fri, 15 Nov 2019 19:12:43 +0100 Subject: [PATCH 260/387] fix E711 and E713 --- pytest.ini | 4 ++-- tests/test_downloader_handlers.py | 4 ++-- tests/test_downloadermiddleware_httpcache.py | 4 ++-- 3 files changed, 6 insertions(+), 6 deletions(-) diff --git a/pytest.ini b/pytest.ini index fa6e7287e..529ad5d27 100644 --- a/pytest.ini +++ b/pytest.ini @@ -192,14 +192,14 @@ flake8-ignore = tests/test_crawl.py E501 E741 E265 tests/test_crawler.py F841 E306 E501 tests/test_dependencies.py E302 F841 E501 E305 - tests/test_downloader_handlers.py E124 E127 E128 E225 E261 E265 F401 E501 E502 E701 E711 E126 E226 E123 + tests/test_downloader_handlers.py E124 E127 E128 E225 E261 E265 F401 E501 E502 E701 E126 E226 E123 tests/test_downloadermiddleware.py E501 tests/test_downloadermiddleware_ajaxcrawlable.py E302 E501 tests/test_downloadermiddleware_cookies.py E731 E741 E501 E128 E303 E265 E126 tests/test_downloadermiddleware_decompression.py E127 tests/test_downloadermiddleware_defaultheaders.py E501 tests/test_downloadermiddleware_downloadtimeout.py E501 - tests/test_downloadermiddleware_httpcache.py E713 E501 E302 E305 F401 + tests/test_downloadermiddleware_httpcache.py E501 E302 E305 F401 tests/test_downloadermiddleware_httpcompression.py E501 F401 E251 E126 E123 tests/test_downloadermiddleware_httpproxy.py F401 E501 E128 tests/test_downloadermiddleware_redirect.py E501 E303 E128 E306 E127 E305 diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 6090998d4..59d4a3eec 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -615,7 +615,7 @@ class Http11MockServerTestCase(unittest.TestCase): crawler = get_crawler(SingleRequestSpider) yield crawler.crawl(seed=Request(url=self.mockserver.url(''))) failure = crawler.spider.meta.get('failure') - self.assertTrue(failure == None) + self.assertTrue(failure is None) reason = crawler.spider.meta['close_reason'] self.assertTrue(reason, 'finished') @@ -636,7 +636,7 @@ class Http11MockServerTestCase(unittest.TestCase): yield crawler.crawl(seed=request) # download_maxsize = 50 is enough for the gzipped response failure = crawler.spider.meta.get('failure') - self.assertTrue(failure == None) + self.assertTrue(failure is None) reason = crawler.spider.meta['close_reason'] self.assertTrue(reason, 'finished') else: diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index 950664ffe..9d863b6e3 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -85,8 +85,8 @@ class _BaseTest(unittest.TestCase): def assertEqualRequestButWithCacheValidators(self, request1, request2): self.assertEqual(request1.url, request2.url) - assert not b'If-None-Match' in request1.headers - assert not b'If-Modified-Since' in request1.headers + assert b'If-None-Match' not in request1.headers + assert b'If-Modified-Since' not in request1.headers assert any(h in request2.headers for h in (b'If-None-Match', b'If-Modified-Since')) self.assertEqual(request1.body, request2.body) From 393a2a197251cd4ac10671fbbff11113be42d930 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 18 Nov 2019 09:15:48 +0100 Subject: [PATCH 261/387] Include /requirements-py3.txt from /docs/requirements.txt --- docs/requirements.txt | 1 + 1 file changed, 1 insertion(+) diff --git a/docs/requirements.txt b/docs/requirements.txt index f9db85146..85812be9a 100644 --- a/docs/requirements.txt +++ b/docs/requirements.txt @@ -1,3 +1,4 @@ +-r ../requirements-py3.txt Sphinx>=2.1 sphinx-notfound-page sphinx_rtd_theme From 99d8b05a0b1997033b2240aa9f945bbe659ee6dc Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 18 Nov 2019 10:58:47 +0100 Subject: [PATCH 262/387] Deprecate scrapy.utils.python.MutableChain.next --- scrapy/utils/python.py | 4 ++++ tests/test_utils_python.py | 8 +++++++- 2 files changed, 11 insertions(+), 1 deletion(-) diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index 64402a2bb..6edf7d702 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -392,3 +392,7 @@ class MutableChain(object): def __next__(self): return next(self.data) + + @deprecated("scrapy.utils.python.MutableChain.__next__") + def next(self): + return self.__next__() diff --git a/tests/test_utils_python.py b/tests/test_utils_python.py index 6cae9793d..ca2c241e4 100644 --- a/tests/test_utils_python.py +++ b/tests/test_utils_python.py @@ -5,6 +5,7 @@ import unittest from itertools import count import platform import six +from warnings import catch_warnings from scrapy.utils.python import ( memoizemethod_noargs, binary_is_text, equal_attributes, @@ -22,7 +23,12 @@ class MutableChainTest(unittest.TestCase): m.extend([9, 10], (11, 12)) self.assertEqual(next(m), 0) self.assertEqual(m.__next__(), 1) - self.assertEqual(list(m), list(range(2, 13))) + with catch_warnings(record=True) as warnings: + self.assertEqual(m.next(), 2) + self.assertEqual(len(warnings), 1) + self.assertIn('scrapy.utils.python.MutableChain.__next__', + str(warnings[0].message)) + self.assertEqual(list(m), list(range(3, 13))) class ToUnicodeTest(unittest.TestCase): From 6d1667d5b8086dd12539a185680f9e119c5fdc55 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Thu, 7 Nov 2019 18:32:33 +0100 Subject: [PATCH 263/387] Use the latest Python version to build the documentation --- .travis.yml | 2 +- tox.ini | 10 ++++++++-- 2 files changed, 9 insertions(+), 3 deletions(-) diff --git a/.travis.yml b/.travis.yml index 9f477e860..c9c64e990 100644 --- a/.travis.yml +++ b/.travis.yml @@ -26,7 +26,7 @@ matrix: - env: TOXENV=py38-extra-deps python: 3.8 - env: TOXENV=docs - python: 3.6 + python: 3.8 install: - | if [ "$TOXENV" = "pypy3" ]; then diff --git a/tox.ini b/tox.ini index fd75d18e2..195cc106a 100644 --- a/tox.ini +++ b/tox.ini @@ -6,6 +6,9 @@ [tox] envlist = py35 +[latest] +basepython = python3.8 + [testenv] deps = -ctests/constraints.txt @@ -63,14 +66,14 @@ commands = py.test {posargs:--durations=10 docs scrapy tests} [testenv:security] -basepython = python3.8 +basepython = {[latest]basepython} deps = bandit commands = bandit -r -c .bandit.yml {posargs:scrapy} [testenv:flake8] -basepython = python3.8 +basepython = {[latest]basepython} deps = {[testenv]deps} pytest-flake8 @@ -83,18 +86,21 @@ deps = -rdocs/requirements.txt [testenv:docs] +basepython = {[latest]basepython} changedir = {[docs]changedir} deps = {[docs]deps} commands = sphinx-build -W -b html . {envtmpdir}/html [testenv:docs-coverage] +basepython = {[latest]basepython} changedir = {[docs]changedir} deps = {[docs]deps} commands = sphinx-build -b coverage . {envtmpdir}/coverage [testenv:docs-links] +basepython = {[latest]basepython} changedir = {[docs]changedir} deps = {[docs]deps} commands = From e1af85619f63fe9312920f0dea7454bb76fcc23f Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 18 Nov 2019 11:06:25 +0100 Subject: [PATCH 264/387] Add a configuration file for Read the Docs --- .readthedocs.yml | 7 +++++++ 1 file changed, 7 insertions(+) create mode 100644 .readthedocs.yml diff --git a/.readthedocs.yml b/.readthedocs.yml new file mode 100644 index 000000000..3c1c3e8be --- /dev/null +++ b/.readthedocs.yml @@ -0,0 +1,7 @@ +version: 2 +sphinx: + configuration: docs/conf.py +python: + version: 3.8 + install: + - requirements: docs/requirements.txt From 74589df02f2961110166959b49002fd8e5037291 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 18 Nov 2019 14:51:44 +0100 Subject: [PATCH 265/387] Make command doctests pass --- docs/topics/commands.rst | 28 ++++++++++++++++++++++++---- pytest.ini | 1 - 2 files changed, 24 insertions(+), 5 deletions(-) diff --git a/docs/topics/commands.rst b/docs/topics/commands.rst index a93bee06b..5b3cd7e75 100644 --- a/docs/topics/commands.rst +++ b/docs/topics/commands.rst @@ -1,3 +1,5 @@ +.. highlight:: none + .. _topics-commands: ================= @@ -66,7 +68,9 @@ structure by default, similar to this:: The directory where the ``scrapy.cfg`` file resides is known as the *project root directory*. That file contains the name of the python module that defines -the project settings. Here is an example:: +the project settings. Here is an example: + +.. code-block:: ini [settings] default = myproject.settings @@ -80,7 +84,9 @@ A project root directory, the one that contains the ``scrapy.cfg``, may be shared by multiple Scrapy projects, each with its own settings module. In that case, you must define one or more aliases for those settings modules -under ``[settings]`` in your ``scrapy.cfg`` file:: +under ``[settings]`` in your ``scrapy.cfg`` file: + +.. code-block:: ini [settings] default = myproject1.settings @@ -277,6 +283,8 @@ check Run contract checks. +.. skip: start + Usage examples:: $ scrapy check -l @@ -294,6 +302,8 @@ Usage examples:: [FAILED] first_spider:parse >>> Returned 92 requests, expected 0..4 +.. skip: end + .. command:: list list @@ -481,6 +491,8 @@ Supported options: * ``--verbose`` or ``-v``: display information for each depth level +.. skip: start + Usage example:: $ scrapy parse http://www.example.com/ -c parse_item @@ -495,6 +507,8 @@ Usage example:: # Requests ----------------------------------------------------------------- [] +.. skip: end + .. command:: settings @@ -573,7 +587,9 @@ Default: ``''`` (empty string) A module to use for looking up custom Scrapy commands. This is used to add custom commands for your Scrapy project. -Example:: +Example: + +.. code-block:: python COMMANDS_MODULE = 'mybot.commands' @@ -588,7 +604,11 @@ You can also add Scrapy commands from an external library by adding a ``scrapy.commands`` section in the entry points of the library ``setup.py`` file. -The following example adds ``my_command`` command:: +The following example adds ``my_command`` command: + +.. skip: next + +.. code-block:: python from setuptools import setup, find_packages diff --git a/pytest.ini b/pytest.ini index 529ad5d27..dab91416f 100644 --- a/pytest.ini +++ b/pytest.ini @@ -8,7 +8,6 @@ addopts = --ignore=docs/_ext --ignore=docs/conf.py --ignore=docs/news.rst - --ignore=docs/topics/commands.rst --ignore=docs/topics/debug.rst --ignore=docs/topics/developer-tools.rst --ignore=docs/topics/dynamic-content.rst From e84cb18ca0b5b09c68cc76d6c48929d9ff933e5c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 18 Nov 2019 15:50:45 +0100 Subject: [PATCH 266/387] Use InterSphinx to link to the Twisted documentation --- docs/conf.py | 3 ++- docs/contributing.rst | 6 +++--- docs/topics/api.rst | 2 -- docs/topics/architecture.rst | 3 +-- docs/topics/email.rst | 13 +++++++------ docs/topics/item-pipeline.rst | 9 ++++----- docs/topics/media-pipeline.rst | 6 +++--- docs/topics/practices.rst | 3 +-- docs/topics/request-response.rst | 9 ++++----- docs/topics/signals.rst | 16 +++++++--------- scrapy/core/downloader/contextfactory.py | 17 ++++++++++------- scrapy/crawler.py | 15 ++++++++------- scrapy/signalmanager.py | 6 ++---- 13 files changed, 52 insertions(+), 56 deletions(-) diff --git a/docs/conf.py b/docs/conf.py index 6ec4582b1..6bfd2cb0e 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -275,5 +275,6 @@ coverage_ignore_pyobjects = [ intersphinx_mapping = { 'python': ('https://docs.python.org/3', None), - 'sphinx': ('https://www.sphinx-doc.org/en/stable', None), + 'sphinx': ('https://www.sphinx-doc.org/en/master', None), + 'twisted': ('https://twistedmatrix.com/documents/current', None), } diff --git a/docs/contributing.rst b/docs/contributing.rst index f084bd23d..68ae2bf3c 100644 --- a/docs/contributing.rst +++ b/docs/contributing.rst @@ -194,8 +194,9 @@ documentation instead of duplicating the docstring in files within the Tests ===== -Tests are implemented using the `Twisted unit-testing framework`_, running -tests requires `tox`_. +Tests are implemented using the :doc:`Twisted unit-testing framework +`. Running tests requires +`tox`_. .. _running-tests: @@ -269,7 +270,6 @@ And their unit-tests are in:: .. _issue tracker: https://github.com/scrapy/scrapy/issues .. _scrapy-users: https://groups.google.com/forum/#!forum/scrapy-users .. _Scrapy subreddit: https://reddit.com/r/scrapy -.. _Twisted unit-testing framework: https://twistedmatrix.com/documents/current/core/development/policy/test-standard.html .. _AUTHORS: https://github.com/scrapy/scrapy/blob/master/AUTHORS .. _tests/: https://github.com/scrapy/scrapy/tree/master/tests .. _open issues: https://github.com/scrapy/scrapy/issues diff --git a/docs/topics/api.rst b/docs/topics/api.rst index 7c8c40b5f..1c461a511 100644 --- a/docs/topics/api.rst +++ b/docs/topics/api.rst @@ -273,5 +273,3 @@ class (which they all inherit from). Close the given spider. After this is called, no more specific stats can be accessed or collected. - -.. _reactor: https://twistedmatrix.com/documents/current/core/howto/reactor-basics.html diff --git a/docs/topics/architecture.rst b/docs/topics/architecture.rst index 2effe94dc..ae25dfa2f 100644 --- a/docs/topics/architecture.rst +++ b/docs/topics/architecture.rst @@ -166,11 +166,10 @@ for concurrency. For more information about asynchronous programming and Twisted see these links: -* `Introduction to Deferreds in Twisted`_ +* :doc:`twisted:core/howto/defer-intro` * `Twisted - hello, asynchronous programming`_ * `Twisted Introduction - Krondo`_ .. _Twisted: https://twistedmatrix.com/trac/ -.. _Introduction to Deferreds in Twisted: https://twistedmatrix.com/documents/current/core/howto/defer-intro.html .. _Twisted - hello, asynchronous programming: http://jessenoller.com/blog/2009/02/11/twisted-hello-asynchronous-programming/ .. _Twisted Introduction - Krondo: http://krondo.com/an-introduction-to-asynchronous-programming-and-twisted/ diff --git a/docs/topics/email.rst b/docs/topics/email.rst index 12eedf2cd..72bf52227 100644 --- a/docs/topics/email.rst +++ b/docs/topics/email.rst @@ -9,13 +9,13 @@ Sending e-mail Although Python makes sending e-mails relatively easy via the `smtplib`_ library, Scrapy provides its own facility for sending e-mails which is very -easy to use and it's implemented using `Twisted non-blocking IO`_, to avoid -interfering with the non-blocking IO of the crawler. It also provides a -simple API for sending attachments and it's very easy to configure, with a few -:ref:`settings `. +easy to use and it's implemented using :doc:`Twisted non-blocking IO +`, to avoid interfering with the non-blocking +IO of the crawler. It also provides a simple API for sending attachments and +it's very easy to configure, with a few :ref:`settings +`. .. _smtplib: https://docs.python.org/2/library/smtplib.html -.. _Twisted non-blocking IO: https://twistedmatrix.com/documents/current/core/howto/defer-intro.html Quick example ============= @@ -39,7 +39,8 @@ MailSender class reference ========================== MailSender is the preferred class to use for sending emails from Scrapy, as it -uses `Twisted non-blocking IO`_, like the rest of the framework. +uses :doc:`Twisted non-blocking IO `, like the +rest of the framework. .. class:: MailSender(smtphost=None, mailfrom=None, smtpuser=None, smtppass=None, smtpport=None) diff --git a/docs/topics/item-pipeline.rst b/docs/topics/item-pipeline.rst index fae18200a..cdc4953c2 100644 --- a/docs/topics/item-pipeline.rst +++ b/docs/topics/item-pipeline.rst @@ -29,7 +29,8 @@ Each item pipeline component is a Python class that must implement the following This method is called for every item pipeline component. :meth:`process_item` must either: return a dict with data, return an :class:`~scrapy.item.Item` - (or any descendant class) object, return a `Twisted Deferred`_ or raise + (or any descendant class) object, return a + :class:`~twisted.internet.defer.Deferred` or raise :exc:`~scrapy.exceptions.DropItem` exception. Dropped items are no longer processed by further pipeline components. @@ -67,8 +68,6 @@ Additionally, they may also implement the following methods: :type crawler: :class:`~scrapy.crawler.Crawler` object -.. _Twisted Deferred: https://twistedmatrix.com/documents/current/core/howto/defer.html - Item pipeline example ===================== @@ -166,7 +165,8 @@ method and how to clean up the resources properly.:: Take screenshot of item ----------------------- -This example demonstrates how to return Deferred_ from :meth:`process_item` method. +This example demonstrates how to return a +:class:`~twisted.internet.defer.Deferred` from the :meth:`process_item` method. It uses Splash_ to render screenshot of item url. Pipeline makes request to locally running instance of Splash_. After request is downloaded and Deferred callback fires, it saves item to a file and adds filename to an item. @@ -209,7 +209,6 @@ and Deferred callback fires, it saves item to a file and adds filename to an ite return item .. _Splash: https://splash.readthedocs.io/en/stable/ -.. _Deferred: https://twistedmatrix.com/documents/current/core/howto/defer.html Duplicates filter ----------------- diff --git a/docs/topics/media-pipeline.rst b/docs/topics/media-pipeline.rst index 431cc6027..206e7cfa5 100644 --- a/docs/topics/media-pipeline.rst +++ b/docs/topics/media-pipeline.rst @@ -441,8 +441,9 @@ See here the methods that you can override in your custom Files Pipeline: * ``success`` is a boolean which is ``True`` if the image was downloaded successfully or ``False`` if it failed for some reason - * ``file_info_or_error`` is a dict containing the following keys (if success - is ``True``) or a `Twisted Failure`_ if there was a problem. + * ``file_info_or_error`` is a dict containing the following keys (if + success is ``True``) or a :exc:`~twisted.python.failure.Failure` if + there was a problem. * ``url`` - the url where the file was downloaded from. This is the url of the request returned from the :meth:`~get_media_requests` @@ -577,5 +578,4 @@ above:: item['image_paths'] = image_paths return item -.. _Twisted Failure: https://twistedmatrix.com/documents/current/api/twisted.python.failure.Failure.html .. _MD5 hash: https://en.wikipedia.org/wiki/MD5 diff --git a/docs/topics/practices.rst b/docs/topics/practices.rst index a6d4f0d6d..e3e8fdc72 100644 --- a/docs/topics/practices.rst +++ b/docs/topics/practices.rst @@ -101,7 +101,7 @@ reactor after ``MySpider`` has finished running. d.addBoth(lambda _: reactor.stop()) reactor.run() # the script will block here until the crawling is finished -.. seealso:: `Twisted Reactor Overview`_. +.. seealso:: :doc:`twisted:core/howto/reactor-basics` .. _run-multiple-spiders: @@ -253,6 +253,5 @@ If you are still unable to prevent your bot getting banned, consider contacting .. _ProxyMesh: https://proxymesh.com/ .. _Google cache: http://www.googleguide.com/cached_pages.html .. _testspiders: https://github.com/scrapinghub/testspiders -.. _Twisted Reactor Overview: https://twistedmatrix.com/documents/current/core/howto/reactor-basics.html .. _Crawlera: https://scrapinghub.com/crawlera .. _scrapoxy: https://scrapoxy.io/ diff --git a/docs/topics/request-response.rst b/docs/topics/request-response.rst index ee37f648e..4cf367d96 100644 --- a/docs/topics/request-response.rst +++ b/docs/topics/request-response.rst @@ -121,8 +121,8 @@ Request objects :param errback: a function that will be called if any exception was raised while processing the request. This includes pages that failed - with 404 HTTP errors and such. It receives a `Twisted Failure`_ instance - as first parameter. + with 404 HTTP errors and such. It receives a + :exc:`~twisted.python.failure.Failure` as first parameter. For more information, see :ref:`topics-request-response-ref-errbacks` below. :type errback: callable @@ -254,8 +254,8 @@ Using errbacks to catch exceptions in request processing The errback of a request is a function that will be called when an exception is raise while processing it. -It receives a `Twisted Failure`_ instance as first parameter and can be -used to track connection establishment timeouts, DNS errors etc. +It receives a :exc:`~twisted.python.failure.Failure` as first parameter and can +be used to track connection establishment timeouts, DNS errors etc. Here's an example spider logging all errors and catching some specific errors if needed:: @@ -816,5 +816,4 @@ XmlResponse objects adds encoding auto-discovering support by looking into the XML declaration line. See :attr:`TextResponse.encoding`. -.. _Twisted Failure: https://twistedmatrix.com/documents/current/api/twisted.python.failure.Failure.html .. _bug in lxml: https://bugs.launchpad.net/lxml/+bug/1665241 diff --git a/docs/topics/signals.rst b/docs/topics/signals.rst index ff07b9d55..3f29aa323 100644 --- a/docs/topics/signals.rst +++ b/docs/topics/signals.rst @@ -50,10 +50,10 @@ Here is a simple example showing how you can catch signals and perform some acti Deferred signal handlers ======================== -Some signals support returning `Twisted deferreds`_ from their handlers, see -the :ref:`topics-signals-ref` below to know which ones. +Some signals support returning :class:`~twisted.internet.defer.Deferred` +objects from their handlers, see the :ref:`topics-signals-ref` below to know +which ones. -.. _Twisted deferreds: https://twistedmatrix.com/documents/current/core/howto/defer.html .. _topics-signals-ref: @@ -155,8 +155,8 @@ item_error :param spider: the spider which raised the exception :type spider: :class:`~scrapy.spiders.Spider` object - :param failure: the exception raised as a Twisted `Failure`_ object - :type failure: `Failure`_ object + :param failure: the exception raised + :type failure: twisted.python.failure.Failure spider_closed ------------- @@ -236,8 +236,8 @@ spider_error This signal does not support returning deferreds from their handlers. - :param failure: the exception raised as a Twisted `Failure`_ object - :type failure: `Failure`_ object + :param failure: the exception raised + :type failure: twisted.python.failure.Failure :param response: the response being processed when the exception was raised :type response: :class:`~scrapy.http.Response` object @@ -333,5 +333,3 @@ response_downloaded :param spider: the spider for which the response is intended :type spider: :class:`~scrapy.spiders.Spider` object - -.. _Failure: https://twistedmatrix.com/documents/current/api/twisted.python.failure.Failure.html diff --git a/scrapy/core/downloader/contextfactory.py b/scrapy/core/downloader/contextfactory.py index 89d2776ae..6e023ebcc 100644 --- a/scrapy/core/downloader/contextfactory.py +++ b/scrapy/core/downloader/contextfactory.py @@ -67,15 +67,18 @@ class BrowserLikeContextFactory(ScrapyClientContextFactory): """ Twisted-recommended context factory for web clients. - Quoting https://twistedmatrix.com/documents/current/api/twisted.web.client.Agent.html: - "The default is to use a BrowserLikePolicyForHTTPS, - so unless you have special requirements you can leave this as-is." + Quoting the documentation of the :class:`~twisted.web.client.Agent` class: - creatorForNetloc() is the same as BrowserLikePolicyForHTTPS - except this context factory allows setting the TLS/SSL method to use. + The default is to use a + :class:`~twisted.web.client.BrowserLikePolicyForHTTPS`, so unless you + have special requirements you can leave this as-is. - Default OpenSSL method is TLS_METHOD (also called SSLv23_METHOD) - which allows TLS protocol negotiation. + :meth:`creatorForNetloc` is the same as + :class:`~twisted.web.client.BrowserLikePolicyForHTTPS` except this context + factory allows setting the TLS/SSL method to use. + + The default OpenSSL method is ``TLS_METHOD`` (also called + ``SSLv23_METHOD``) which allows TLS protocol negotiation. """ def creatorForNetloc(self, hostname, port): diff --git a/scrapy/crawler.py b/scrapy/crawler.py index 8868a985b..f8c80880a 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -110,7 +110,7 @@ class Crawler(object): class CrawlerRunner(object): """ This is a convenient helper class that keeps track of, manages and runs - crawlers inside an already setup Twisted `reactor`_. + crawlers inside an already setup :mod:`~twisted.internet.reactor`. The CrawlerRunner object must be instantiated with a :class:`~scrapy.settings.Settings` object. @@ -233,12 +233,13 @@ class CrawlerProcess(CrawlerRunner): A class to run multiple scrapy crawlers in a process simultaneously. This class extends :class:`~scrapy.crawler.CrawlerRunner` by adding support - for starting a Twisted `reactor`_ and handling shutdown signals, like the - keyboard interrupt command Ctrl-C. It also configures top-level logging. + for starting a :mod:`~twisted.internet.reactor` and handling shutdown + signals, like the keyboard interrupt command Ctrl-C. It also configures + top-level logging. This utility should be a better fit than :class:`~scrapy.crawler.CrawlerRunner` if you aren't running another - Twisted `reactor`_ within your application. + :mod:`~twisted.internet.reactor` within your application. The CrawlerProcess object must be instantiated with a :class:`~scrapy.settings.Settings` object. @@ -273,9 +274,9 @@ class CrawlerProcess(CrawlerRunner): def start(self, stop_after_crawl=True): """ - This method starts a Twisted `reactor`_, adjusts its pool size to - :setting:`REACTOR_THREADPOOL_MAXSIZE`, and installs a DNS cache based - on :setting:`DNSCACHE_ENABLED` and :setting:`DNSCACHE_SIZE`. + This method starts a :mod:`~twisted.internet.reactor`, adjusts its pool + size to :setting:`REACTOR_THREADPOOL_MAXSIZE`, and installs a DNS cache + based on :setting:`DNSCACHE_ENABLED` and :setting:`DNSCACHE_SIZE`. If ``stop_after_crawl`` is True, the reactor will be stopped after all crawlers have finished, using :meth:`join`. diff --git a/scrapy/signalmanager.py b/scrapy/signalmanager.py index 296d27ed8..9a160f62e 100644 --- a/scrapy/signalmanager.py +++ b/scrapy/signalmanager.py @@ -46,16 +46,14 @@ class SignalManager(object): def send_catch_log_deferred(self, signal, **kwargs): """ - Like :meth:`send_catch_log` but supports returning `deferreds`_ from - signal handlers. + Like :meth:`send_catch_log` but supports returning + :class:`~twisted.internet.defer.Deferred` objects from signal handlers. Returns a Deferred that gets fired once all signal handlers deferreds were fired. Send a signal, catch exceptions and log them. The keyword arguments are passed to the signal handlers (connected through the :meth:`connect` method). - - .. _deferreds: https://twistedmatrix.com/documents/current/core/howto/defer.html """ kwargs.setdefault('sender', self.sender) return _signal.send_catch_log_deferred(signal, **kwargs) From fed93515de4e306eb3262125c09eb49decdb2944 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 18 Nov 2019 16:11:03 +0100 Subject: [PATCH 267/387] Add tooltips to documentation cross-references --- docs/conf.py | 1 + docs/requirements.txt | 1 + 2 files changed, 2 insertions(+) diff --git a/docs/conf.py b/docs/conf.py index 6ec4582b1..e2784cf17 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -27,6 +27,7 @@ sys.path.insert(0, path.dirname(path.dirname(__file__))) # Add any Sphinx extension module names here, as strings. They can be extensions # coming with Sphinx (named 'sphinx.ext.*') or your custom ones. extensions = [ + 'hoverxref.extension', 'notfound.extension', 'scrapydocs', 'sphinx.ext.autodoc', diff --git a/docs/requirements.txt b/docs/requirements.txt index f9db85146..773b92cea 100644 --- a/docs/requirements.txt +++ b/docs/requirements.txt @@ -1,3 +1,4 @@ Sphinx>=2.1 +sphinx-hoverxref sphinx-notfound-page sphinx_rtd_theme From f261cf65e999573d95a575ba362c3e32b026f894 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 18 Nov 2019 17:16:09 +0100 Subject: [PATCH 268/387] Add missing blank lines between functions and classes Also fixed 2 unrelated Flake8 issues --- pytest.ini | 118 ++++++++---------- scrapy/commands/fetch.py | 1 + scrapy/commands/list.py | 1 + scrapy/commands/settings.py | 1 + scrapy/commands/view.py | 1 + scrapy/core/downloader/handlers/datauri.py | 4 +- scrapy/downloadermiddlewares/ajaxcrawl.py | 2 + scrapy/dupefilters.py | 1 + scrapy/exceptions.py | 13 ++ scrapy/extension.py | 1 + scrapy/extensions/spiderstate.py | 1 + scrapy/http/response/html.py | 1 + scrapy/http/response/xml.py | 1 + scrapy/interfaces.py | 1 + scrapy/loader/common.py | 1 + scrapy/pipelines/__init__.py | 1 + scrapy/resolver.py | 1 + scrapy/robotstxt.py | 3 + scrapy/utils/console.py | 7 ++ scrapy/utils/decorators.py | 1 + scrapy/utils/defer.py | 9 ++ scrapy/utils/display.py | 3 + scrapy/utils/engine.py | 3 + scrapy/utils/ftp.py | 1 + scrapy/utils/gz.py | 1 + scrapy/utils/httpobj.py | 3 + scrapy/utils/job.py | 1 + scrapy/utils/python.py | 1 + scrapy/utils/reactor.py | 1 + scrapy/utils/request.py | 2 + scrapy/utils/response.py | 4 + scrapy/utils/spider.py | 1 + scrapy/utils/template.py | 2 + scrapy/utils/test.py | 6 + scrapy/utils/versions.py | 2 +- tests/mocks/dummydbm.py | 1 + tests/pipelines.py | 1 + tests/spiders.py | 1 + tests/test_cmdline/extensions.py | 1 + tests/test_command_parse.py | 1 + tests/test_dependencies.py | 1 + ...test_downloadermiddleware_ajaxcrawlable.py | 2 + tests/test_downloadermiddleware_httpcache.py | 1 + tests/test_dupefilters.py | 1 + tests/test_http_headers.py | 1 + tests/test_logformatter.py | 1 + tests/test_mail.py | 1 + tests/test_middleware.py | 4 + tests/test_responsetypes.py | 1 + tests/test_robotstxt_interface.py | 2 + tests/test_spiderloader/__init__.py | 1 + .../test_spiders/nested/spider4.py | 1 + .../test_spiderloader/test_spiders/spider0.py | 1 + .../test_spiderloader/test_spiders/spider1.py | 1 + .../test_spiderloader/test_spiders/spider2.py | 1 + .../test_spiderloader/test_spiders/spider3.py | 1 + tests/test_spidermiddleware_offsite.py | 2 + tests/test_spidermiddleware_output_chain.py | 4 + tests/test_spidermiddleware_referer.py | 4 + tests/test_squeues.py | 13 ++ tests/test_utils_console.py | 1 + tests/test_utils_defer.py | 9 ++ tests/test_utils_http.py | 1 + tests/test_utils_httpobj.py | 1 + tests/test_utils_iterators.py | 1 + tests/test_utils_request.py | 1 + tests/test_utils_signal.py | 1 + tests/test_utils_sitemap.py | 1 + tests/test_utils_spider.py | 3 + tests/test_utils_url.py | 2 + 70 files changed, 199 insertions(+), 72 deletions(-) diff --git a/pytest.ini b/pytest.ini index 529ad5d27..a3693a778 100644 --- a/pytest.ini +++ b/pytest.ini @@ -30,16 +30,15 @@ flake8-ignore = scrapy/commands/check.py F401 E501 scrapy/commands/crawl.py E501 scrapy/commands/edit.py E501 - scrapy/commands/fetch.py E401 E302 E501 E128 E502 E731 + scrapy/commands/fetch.py E401 E501 E128 E502 E731 scrapy/commands/genspider.py E128 E501 E502 - scrapy/commands/list.py E302 scrapy/commands/parse.py E128 E501 E731 E226 scrapy/commands/runspider.py E501 - scrapy/commands/settings.py E302 E128 + scrapy/commands/settings.py E128 scrapy/commands/shell.py E128 E501 E502 scrapy/commands/startproject.py E502 E127 E501 E128 scrapy/commands/version.py E501 E128 - scrapy/commands/view.py F401 E302 + scrapy/commands/view.py F401 # scrapy/contracts scrapy/contracts/__init__.py E501 W504 scrapy/contracts/default.py E502 E128 @@ -60,7 +59,7 @@ flake8-ignore = scrapy/core/downloader/handlers/http11.py E501 scrapy/core/downloader/handlers/s3.py E501 F401 E502 E128 E126 # scrapy/downloadermiddlewares - scrapy/downloadermiddlewares/ajaxcrawl.py E302 E501 E226 + scrapy/downloadermiddlewares/ajaxcrawl.py E501 E226 scrapy/downloadermiddlewares/decompression.py E501 scrapy/downloadermiddlewares/defaultheaders.py E501 scrapy/downloadermiddlewares/httpcache.py E501 E126 @@ -72,11 +71,11 @@ flake8-ignore = scrapy/downloadermiddlewares/stats.py E501 # scrapy/extensions scrapy/extensions/closespider.py E501 E502 E128 E123 - scrapy/extensions/corestats.py E302 E501 + scrapy/extensions/corestats.py E501 scrapy/extensions/feedexport.py E128 E501 scrapy/extensions/httpcache.py E128 E501 E303 F401 scrapy/extensions/memdebug.py E501 - scrapy/extensions/spiderstate.py E302 E501 + scrapy/extensions/spiderstate.py E501 scrapy/extensions/telnet.py E501 W504 scrapy/extensions/throttle.py E501 # scrapy/http @@ -87,18 +86,14 @@ flake8-ignore = scrapy/http/request/form.py E501 E123 scrapy/http/request/json_request.py E501 scrapy/http/response/__init__.py E501 E128 W293 W291 - scrapy/http/response/html.py E302 scrapy/http/response/text.py E501 W293 E128 E124 - scrapy/http/response/xml.py E302 # scrapy/linkextractors scrapy/linkextractors/__init__.py E731 E502 E501 E402 F401 scrapy/linkextractors/lxmlhtml.py E501 E731 E226 # scrapy/loader scrapy/loader/__init__.py E501 E502 E128 - scrapy/loader/common.py E302 scrapy/loader/processors.py E501 # scrapy/pipelines - scrapy/pipelines/__init__.py E302 scrapy/pipelines/files.py E116 E501 E266 scrapy/pipelines/images.py E265 E501 scrapy/pipelines/media.py E125 E501 E266 @@ -123,56 +118,50 @@ flake8-ignore = scrapy/utils/benchserver.py E501 scrapy/utils/boto.py F401 scrapy/utils/conf.py E402 E502 E501 - scrapy/utils/console.py E302 E261 F401 E306 E305 + scrapy/utils/console.py E261 F401 E306 E305 scrapy/utils/curl.py F401 scrapy/utils/datatypes.py E501 E226 - scrapy/utils/decorators.py E501 E302 - scrapy/utils/defer.py E501 E302 E128 + scrapy/utils/decorators.py E501 + scrapy/utils/defer.py E501 E128 scrapy/utils/deprecate.py E128 E501 E127 E502 - scrapy/utils/display.py E302 - scrapy/utils/engine.py F401 E261 E302 - scrapy/utils/ftp.py E302 - scrapy/utils/gz.py E305 E501 E302 W504 + scrapy/utils/engine.py F401 E261 + scrapy/utils/gz.py E305 E501 W504 scrapy/utils/http.py F403 F401 E226 - scrapy/utils/httpobj.py E302 E501 + scrapy/utils/httpobj.py E501 scrapy/utils/iterators.py E501 E701 - scrapy/utils/job.py E302 scrapy/utils/log.py E128 W503 scrapy/utils/markup.py F403 F401 W292 scrapy/utils/misc.py E501 E226 scrapy/utils/multipart.py F403 F401 W292 scrapy/utils/project.py E501 - scrapy/utils/python.py E501 E302 - scrapy/utils/reactor.py E302 E226 + scrapy/utils/python.py E501 + scrapy/utils/reactor.py E226 scrapy/utils/reqser.py E501 - scrapy/utils/request.py E302 E127 E501 - scrapy/utils/response.py E501 E302 E128 + scrapy/utils/request.py E127 E501 + scrapy/utils/response.py E501 E128 scrapy/utils/signal.py E501 E128 scrapy/utils/sitemap.py E501 - scrapy/utils/spider.py E271 E302 E501 + scrapy/utils/spider.py E271 E501 scrapy/utils/ssl.py E501 - scrapy/utils/template.py E302 - scrapy/utils/test.py E302 E501 + scrapy/utils/test.py E501 scrapy/utils/url.py E501 F403 F401 E128 F405 # scrapy scrapy/__init__.py E402 E501 scrapy/_monkeypatches.py W293 scrapy/cmdline.py E502 E501 scrapy/crawler.py E501 - scrapy/dupefilters.py E302 E501 E202 - scrapy/exceptions.py E302 E501 + scrapy/dupefilters.py E501 E202 + scrapy/exceptions.py E501 scrapy/exporters.py E501 E261 E226 - scrapy/extension.py E302 - scrapy/interfaces.py E302 E501 + scrapy/interfaces.py E501 scrapy/item.py E501 E128 scrapy/link.py E501 scrapy/logformatter.py E501 W293 scrapy/mail.py E402 E128 E501 E502 scrapy/middleware.py E502 E128 E501 scrapy/pqueues.py E501 - scrapy/resolver.py E302 scrapy/responsetypes.py E128 E501 E305 - scrapy/robotstxt.py E302 E501 + scrapy/robotstxt.py E501 scrapy/shell.py E501 scrapy/signalmanager.py E501 scrapy/spiderloader.py E225 F841 E501 E126 @@ -181,91 +170,82 @@ flake8-ignore = # tests tests/__init__.py F401 E402 E501 tests/mockserver.py E401 E501 E126 E123 F401 - tests/pipelines.py E302 F841 E226 - tests/spiders.py E302 E501 E127 + tests/pipelines.py F841 E226 + tests/spiders.py E501 E127 tests/test_closespider.py E501 E127 tests/test_command_fetch.py E501 E261 - tests/test_command_parse.py F401 E302 E501 E128 E303 E226 + tests/test_command_parse.py F401 E501 E128 E303 E226 tests/test_command_shell.py E501 E128 tests/test_commands.py F401 E128 E501 tests/test_contracts.py E501 E128 W293 tests/test_crawl.py E501 E741 E265 tests/test_crawler.py F841 E306 E501 - tests/test_dependencies.py E302 F841 E501 E305 + tests/test_dependencies.py F841 E501 E305 tests/test_downloader_handlers.py E124 E127 E128 E225 E261 E265 F401 E501 E502 E701 E126 E226 E123 tests/test_downloadermiddleware.py E501 - tests/test_downloadermiddleware_ajaxcrawlable.py E302 E501 + tests/test_downloadermiddleware_ajaxcrawlable.py E501 tests/test_downloadermiddleware_cookies.py E731 E741 E501 E128 E303 E265 E126 tests/test_downloadermiddleware_decompression.py E127 tests/test_downloadermiddleware_defaultheaders.py E501 tests/test_downloadermiddleware_downloadtimeout.py E501 - tests/test_downloadermiddleware_httpcache.py E501 E302 E305 F401 + tests/test_downloadermiddleware_httpcache.py E501 E305 F401 tests/test_downloadermiddleware_httpcompression.py E501 F401 E251 E126 E123 tests/test_downloadermiddleware_httpproxy.py F401 E501 E128 tests/test_downloadermiddleware_redirect.py E501 E303 E128 E306 E127 E305 tests/test_downloadermiddleware_retry.py E501 E128 W293 E251 E502 E303 E126 tests/test_downloadermiddleware_robotstxt.py E501 tests/test_downloadermiddleware_stats.py E501 - tests/test_dupefilters.py E302 E221 E501 E741 W293 W291 E128 E124 + tests/test_dupefilters.py E221 E501 E741 W293 W291 E128 E124 tests/test_engine.py E401 E501 E502 E128 E261 tests/test_exporters.py E501 E731 E306 E128 E124 tests/test_extension_telnet.py F401 F841 tests/test_feedexport.py E501 F401 F841 E241 tests/test_http_cookies.py E501 - tests/test_http_headers.py E302 E501 + tests/test_http_headers.py E501 tests/test_http_request.py F401 E402 E501 E261 E127 E128 W293 E502 E128 E502 E126 E123 tests/test_http_response.py E501 E301 E502 E128 E265 tests/test_item.py E701 E128 F841 E306 tests/test_link.py E501 tests/test_linkextractors.py E501 E128 E124 - tests/test_loader.py E302 E501 E731 E303 E741 E128 E117 E241 - tests/test_logformatter.py E128 E501 E122 E302 - tests/test_mail.py E302 E128 E501 E305 - tests/test_middleware.py E302 E501 E128 + tests/test_loader.py E501 E731 E303 E741 E128 E117 E241 + tests/test_logformatter.py E128 E501 E122 + tests/test_mail.py E128 E501 E305 + tests/test_middleware.py E501 E128 tests/test_pipeline_crawl.py E131 E501 E128 E126 tests/test_pipeline_files.py F401 E501 W293 E303 E272 E226 tests/test_pipeline_images.py F401 F841 E501 E303 tests/test_pipeline_media.py E501 E741 E731 E128 E261 E306 E502 tests/test_request_cb_kwargs.py E501 - tests/test_responsetypes.py E501 E302 E305 - tests/test_robotstxt_interface.py F401 E302 E501 W291 E501 + tests/test_responsetypes.py E501 E305 + tests/test_robotstxt_interface.py F401 E501 W291 E501 tests/test_scheduler.py E501 E126 E123 tests/test_selector.py F401 E501 E127 tests/test_spider.py E501 F401 tests/test_spidermiddleware.py E501 E226 tests/test_spidermiddleware_httperror.py E128 E501 E127 E121 - tests/test_spidermiddleware_offsite.py E302 E501 E128 E111 W293 - tests/test_spidermiddleware_output_chain.py F401 E501 E302 W293 E226 - tests/test_spidermiddleware_referer.py F401 E501 E302 F841 E125 E201 E261 E124 E501 E241 E121 - tests/test_squeues.py E501 E302 E701 E741 + tests/test_spidermiddleware_offsite.py E501 E128 E111 W293 + tests/test_spidermiddleware_output_chain.py F401 E501 W293 E226 + tests/test_spidermiddleware_referer.py F401 E501 F841 E125 E201 E261 E124 E501 E241 E121 + tests/test_squeues.py E501 E701 E741 tests/test_utils_conf.py E501 E303 E128 - tests/test_utils_console.py E302 tests/test_utils_curl.py E501 tests/test_utils_datatypes.py E402 E501 E305 - tests/test_utils_defer.py E306 E261 E501 E302 F841 E226 + tests/test_utils_defer.py E306 E261 E501 F841 E226 tests/test_utils_deprecate.py F841 E306 E501 - tests/test_utils_http.py E302 E501 E502 E128 W504 - tests/test_utils_httpobj.py E302 - tests/test_utils_iterators.py E501 E128 E129 E302 E303 E241 + tests/test_utils_http.py E501 E502 E128 W504 + tests/test_utils_iterators.py E501 E128 E129 E303 E241 tests/test_utils_log.py E741 E226 tests/test_utils_python.py E501 E303 E731 E701 E305 tests/test_utils_reqser.py F401 E501 E128 - tests/test_utils_request.py E302 E501 E128 E305 + tests/test_utils_request.py E501 E128 E305 tests/test_utils_response.py E501 - tests/test_utils_signal.py E741 F841 E302 E731 E226 - tests/test_utils_sitemap.py E302 E128 E501 E124 - tests/test_utils_spider.py E261 E302 E305 + tests/test_utils_signal.py E741 F841 E731 E226 + tests/test_utils_sitemap.py E128 E501 E124 + tests/test_utils_spider.py E261 E305 tests/test_utils_template.py E305 - tests/test_utils_url.py F401 E501 E127 E302 E305 E211 E125 E501 E226 E241 E126 E123 + tests/test_utils_url.py F401 E501 E127 E305 E211 E125 E501 E226 E241 E126 E123 tests/test_webclient.py E501 E128 E122 E303 E402 E306 E226 E241 E123 E126 - tests/mocks/dummydbm.py E302 tests/test_cmdline/__init__.py E502 E501 - tests/test_cmdline/extensions.py E302 tests/test_settings/__init__.py F401 E501 E128 - tests/test_spiderloader/__init__.py E128 E501 E302 - tests/test_spiderloader/test_spiders/spider0.py E302 - tests/test_spiderloader/test_spiders/spider1.py E302 - tests/test_spiderloader/test_spiders/spider2.py E302 - tests/test_spiderloader/test_spiders/spider3.py E302 - tests/test_spiderloader/test_spiders/nested/spider4.py E302 + tests/test_spiderloader/__init__.py E128 E501 tests/test_utils_misc/__init__.py E501 diff --git a/scrapy/commands/fetch.py b/scrapy/commands/fetch.py index d45133e0e..724b4a1c4 100644 --- a/scrapy/commands/fetch.py +++ b/scrapy/commands/fetch.py @@ -8,6 +8,7 @@ from scrapy.exceptions import UsageError from scrapy.utils.datatypes import SequenceExclude from scrapy.utils.spider import spidercls_for_request, DefaultSpider + class Command(ScrapyCommand): requires_project = False diff --git a/scrapy/commands/list.py b/scrapy/commands/list.py index a255b3b94..422183ac1 100644 --- a/scrapy/commands/list.py +++ b/scrapy/commands/list.py @@ -1,6 +1,7 @@ from __future__ import print_function from scrapy.commands import ScrapyCommand + class Command(ScrapyCommand): requires_project = True diff --git a/scrapy/commands/settings.py b/scrapy/commands/settings.py index bee52f06a..ffe3aa2eb 100644 --- a/scrapy/commands/settings.py +++ b/scrapy/commands/settings.py @@ -4,6 +4,7 @@ import json from scrapy.commands import ScrapyCommand from scrapy.settings import BaseSettings + class Command(ScrapyCommand): requires_project = False diff --git a/scrapy/commands/view.py b/scrapy/commands/view.py index 59e665016..31c17c0ab 100644 --- a/scrapy/commands/view.py +++ b/scrapy/commands/view.py @@ -1,6 +1,7 @@ from scrapy.commands import fetch, ScrapyCommand from scrapy.utils.response import open_in_browser + class Command(fetch.Command): def short_desc(self): diff --git a/scrapy/core/downloader/handlers/datauri.py b/scrapy/core/downloader/handlers/datauri.py index ad25beb3b..9e5020753 100644 --- a/scrapy/core/downloader/handlers/datauri.py +++ b/scrapy/core/downloader/handlers/datauri.py @@ -17,8 +17,8 @@ class DataURIDownloadHandler(object): respcls = responsetypes.from_mimetype(uri.media_type) resp_kwargs = {} - if (issubclass(respcls, TextResponse) and - uri.media_type.split('/')[0] == 'text'): + if (issubclass(respcls, TextResponse) + and uri.media_type.split('/')[0] == 'text'): charset = uri.media_type_parameters.get('charset') resp_kwargs['encoding'] = charset diff --git a/scrapy/downloadermiddlewares/ajaxcrawl.py b/scrapy/downloadermiddlewares/ajaxcrawl.py index 72715dba7..ba50793bb 100644 --- a/scrapy/downloadermiddlewares/ajaxcrawl.py +++ b/scrapy/downloadermiddlewares/ajaxcrawl.py @@ -68,6 +68,8 @@ class AjaxCrawlMiddleware(object): # XXX: move it to w3lib? _ajax_crawlable_re = re.compile(six.u(r'')) + + def _has_ajaxcrawlable_meta(text): """ >>> _has_ajaxcrawlable_meta('') diff --git a/scrapy/dupefilters.py b/scrapy/dupefilters.py index 0bcdd3495..4d95eb847 100644 --- a/scrapy/dupefilters.py +++ b/scrapy/dupefilters.py @@ -5,6 +5,7 @@ import logging from scrapy.utils.job import job_dir from scrapy.utils.request import referer_str, request_fingerprint + class BaseDupeFilter(object): @classmethod diff --git a/scrapy/exceptions.py b/scrapy/exceptions.py index 96949bdd9..7c4bb3d00 100644 --- a/scrapy/exceptions.py +++ b/scrapy/exceptions.py @@ -7,10 +7,12 @@ new exceptions here without documenting them there. # Internal + class NotConfigured(Exception): """Indicates a missing configuration situation""" pass + class _InvalidOutput(TypeError): """ Indicates an invalid value has been returned by a middleware's processing method. @@ -18,15 +20,19 @@ class _InvalidOutput(TypeError): """ pass + # HTTP and crawling + class IgnoreRequest(Exception): """Indicates a decision was made not to process a request""" + class DontCloseSpider(Exception): """Request the spider not to be closed yet""" pass + class CloseSpider(Exception): """Raise this from callbacks to request the spider to be closed""" @@ -34,30 +40,37 @@ class CloseSpider(Exception): super(CloseSpider, self).__init__() self.reason = reason + # Items + class DropItem(Exception): """Drop item from the item pipeline""" pass + class NotSupported(Exception): """Indicates a feature or method is not supported""" pass + # Commands + class UsageError(Exception): """To indicate a command-line usage error""" def __init__(self, *a, **kw): self.print_help = kw.pop('print_help', True) super(UsageError, self).__init__(*a, **kw) + class ScrapyDeprecationWarning(Warning): """Warning category for deprecated features, since the default DeprecationWarning is silenced on Python 2.7+ """ pass + class ContractFail(AssertionError): """Error raised in case of a failing contract""" pass diff --git a/scrapy/extension.py b/scrapy/extension.py index e39e456fa..050b87e5f 100644 --- a/scrapy/extension.py +++ b/scrapy/extension.py @@ -6,6 +6,7 @@ See documentation in docs/topics/extensions.rst from scrapy.middleware import MiddlewareManager from scrapy.utils.conf import build_component_list + class ExtensionManager(MiddlewareManager): component_name = 'extension' diff --git a/scrapy/extensions/spiderstate.py b/scrapy/extensions/spiderstate.py index 2220cbd8f..8ba770ec0 100644 --- a/scrapy/extensions/spiderstate.py +++ b/scrapy/extensions/spiderstate.py @@ -5,6 +5,7 @@ from scrapy import signals from scrapy.exceptions import NotConfigured from scrapy.utils.job import job_dir + class SpiderState(object): """Store and load spider state during a scraping job""" diff --git a/scrapy/http/response/html.py b/scrapy/http/response/html.py index bd3559fbb..7eed052c2 100644 --- a/scrapy/http/response/html.py +++ b/scrapy/http/response/html.py @@ -7,5 +7,6 @@ See documentation in docs/topics/request-response.rst from scrapy.http.response.text import TextResponse + class HtmlResponse(TextResponse): pass diff --git a/scrapy/http/response/xml.py b/scrapy/http/response/xml.py index 1df33fee5..abf474a2f 100644 --- a/scrapy/http/response/xml.py +++ b/scrapy/http/response/xml.py @@ -7,5 +7,6 @@ See documentation in docs/topics/request-response.rst from scrapy.http.response.text import TextResponse + class XmlResponse(TextResponse): pass diff --git a/scrapy/interfaces.py b/scrapy/interfaces.py index d48babc3c..1896ec31e 100644 --- a/scrapy/interfaces.py +++ b/scrapy/interfaces.py @@ -1,5 +1,6 @@ from zope.interface import Interface + class ISpiderLoader(Interface): def from_settings(settings): diff --git a/scrapy/loader/common.py b/scrapy/loader/common.py index 916524947..42f8de636 100644 --- a/scrapy/loader/common.py +++ b/scrapy/loader/common.py @@ -3,6 +3,7 @@ from functools import partial from scrapy.utils.python import get_func_args + def wrap_loader_context(function, context): """Wrap functions that receive loader_context to contain the context "pre-loaded" and expose a interface that receives only one argument diff --git a/scrapy/pipelines/__init__.py b/scrapy/pipelines/__init__.py index 2ef8786d0..aa1bfb77f 100644 --- a/scrapy/pipelines/__init__.py +++ b/scrapy/pipelines/__init__.py @@ -7,6 +7,7 @@ See documentation in docs/item-pipeline.rst from scrapy.middleware import MiddlewareManager from scrapy.utils.conf import build_component_list + class ItemPipelineManager(MiddlewareManager): component_name = 'item pipeline' diff --git a/scrapy/resolver.py b/scrapy/resolver.py index 0aaced7e4..4df949015 100644 --- a/scrapy/resolver.py +++ b/scrapy/resolver.py @@ -7,6 +7,7 @@ from scrapy.utils.datatypes import LocalCache dnscache = LocalCache(10000) + class CachingThreadedResolver(ThreadedResolver): def __init__(self, reactor, cache_size, timeout): super(CachingThreadedResolver, self).__init__(reactor) diff --git a/scrapy/robotstxt.py b/scrapy/robotstxt.py index 95a8c09b8..f0f9c59dc 100644 --- a/scrapy/robotstxt.py +++ b/scrapy/robotstxt.py @@ -5,8 +5,10 @@ from six import with_metaclass from scrapy.utils.python import to_unicode + logger = logging.getLogger(__name__) + def decode_robotstxt(robotstxt_body, spider, to_native_str_type=False): try: if to_native_str_type: @@ -23,6 +25,7 @@ def decode_robotstxt(robotstxt_body, spider, to_native_str_type=False): robotstxt_body = '' return robotstxt_body + class RobotParser(with_metaclass(ABCMeta)): @classmethod @abstractmethod diff --git a/scrapy/utils/console.py b/scrapy/utils/console.py index 2e9981556..a26e84d38 100644 --- a/scrapy/utils/console.py +++ b/scrapy/utils/console.py @@ -1,6 +1,7 @@ from functools import wraps from collections import OrderedDict + def _embed_ipython_shell(namespace={}, banner=''): """Start an IPython Shell""" try: @@ -23,6 +24,7 @@ def _embed_ipython_shell(namespace={}, banner=''): shell() return wrapper + def _embed_bpython_shell(namespace={}, banner=''): """Start a bpython shell""" import bpython @@ -31,6 +33,7 @@ def _embed_bpython_shell(namespace={}, banner=''): bpython.embed(locals_=namespace, banner=banner) return wrapper + def _embed_ptpython_shell(namespace={}, banner=''): """Start a ptpython shell""" import ptpython.repl @@ -40,6 +43,7 @@ def _embed_ptpython_shell(namespace={}, banner=''): ptpython.repl.embed(locals=namespace) return wrapper + def _embed_standard_shell(namespace={}, banner=''): """Start a standard python shell""" import code @@ -55,6 +59,7 @@ def _embed_standard_shell(namespace={}, banner=''): code.interact(banner=banner, local=namespace) return wrapper + DEFAULT_PYTHON_SHELLS = OrderedDict([ ('ptpython', _embed_ptpython_shell), ('ipython', _embed_ipython_shell), @@ -62,6 +67,7 @@ DEFAULT_PYTHON_SHELLS = OrderedDict([ ('python', _embed_standard_shell), ]) + def get_shell_embed_func(shells=None, known_shells=None): """Return the first acceptable shell-embed function from a given list of shell names. @@ -79,6 +85,7 @@ def get_shell_embed_func(shells=None, known_shells=None): except ImportError: continue + def start_python_console(namespace=None, banner='', shells=None): """Start Python console bound to the given namespace. Readline support and tab completion will be used on Unix, if available. diff --git a/scrapy/utils/decorators.py b/scrapy/utils/decorators.py index 38bee1a6c..2e2c7adc1 100644 --- a/scrapy/utils/decorators.py +++ b/scrapy/utils/decorators.py @@ -34,6 +34,7 @@ def defers(func): return defer.maybeDeferred(func, *a, **kw) return wrapped + def inthread(func): """Decorator to call a function in a thread and return a deferred with the result diff --git a/scrapy/utils/defer.py b/scrapy/utils/defer.py index 69d621830..c5916c21c 100644 --- a/scrapy/utils/defer.py +++ b/scrapy/utils/defer.py @@ -7,6 +7,7 @@ from twisted.python import failure from scrapy.exceptions import IgnoreRequest + def defer_fail(_failure): """Same as twisted.internet.defer.fail but delay calling errback until next reactor loop @@ -18,6 +19,7 @@ def defer_fail(_failure): reactor.callLater(0.1, d.errback, _failure) return d + def defer_succeed(result): """Same as twisted.internet.defer.succeed but delay calling callback until next reactor loop @@ -29,6 +31,7 @@ def defer_succeed(result): reactor.callLater(0.1, d.callback, result) return d + def defer_result(result): if isinstance(result, defer.Deferred): return result @@ -37,6 +40,7 @@ def defer_result(result): else: return defer_succeed(result) + def mustbe_deferred(f, *args, **kw): """Same as twisted.internet.defer.maybeDeferred, but delay calling callback/errback to next reactor loop @@ -53,6 +57,7 @@ def mustbe_deferred(f, *args, **kw): else: return defer_result(result) + def parallel(iterable, count, callable, *args, **named): """Execute a callable over the objects in the given iterable, in parallel, using no more than ``count`` concurrent calls. @@ -63,6 +68,7 @@ def parallel(iterable, count, callable, *args, **named): work = (callable(elem, *args, **named) for elem in iterable) return defer.DeferredList([coop.coiterate(work) for _ in range(count)]) + def process_chain(callbacks, input, *a, **kw): """Return a Deferred built by chaining the given callbacks""" d = defer.Deferred() @@ -71,6 +77,7 @@ def process_chain(callbacks, input, *a, **kw): d.callback(input) return d + def process_chain_both(callbacks, errbacks, input, *a, **kw): """Return a Deferred built by chaining the given callbacks and errbacks""" d = defer.Deferred() @@ -83,6 +90,7 @@ def process_chain_both(callbacks, errbacks, input, *a, **kw): d.callback(input) return d + def process_parallel(callbacks, input, *a, **kw): """Return a Deferred with the output of all successful calls to the given callbacks @@ -92,6 +100,7 @@ def process_parallel(callbacks, input, *a, **kw): d.addCallbacks(lambda r: [x[1] for x in r], lambda f: f.value.subFailure) return d + def iter_errback(iterable, errback, *a, **kw): """Wraps an iterable calling an errback if an error is caught while iterating it. diff --git a/scrapy/utils/display.py b/scrapy/utils/display.py index f6a6c4645..91ebdae11 100644 --- a/scrapy/utils/display.py +++ b/scrapy/utils/display.py @@ -6,6 +6,7 @@ from __future__ import print_function import sys from pprint import pformat as pformat_ + def _colorize(text, colorize=True): if not colorize or not sys.stdout.isatty(): return text @@ -17,8 +18,10 @@ def _colorize(text, colorize=True): except ImportError: return text + def pformat(obj, *args, **kwargs): return _colorize(pformat_(obj), kwargs.pop('colorize', True)) + def pprint(obj, *args, **kwargs): print(pformat(obj, *args, **kwargs)) diff --git a/scrapy/utils/engine.py b/scrapy/utils/engine.py index 11dd36d91..2c20b5c88 100644 --- a/scrapy/utils/engine.py +++ b/scrapy/utils/engine.py @@ -3,6 +3,7 @@ from __future__ import print_function from time import time # used in global tests code + def get_engine_status(engine): """Return a report of the current engine status""" tests = [ @@ -32,6 +33,7 @@ def get_engine_status(engine): return checks + def format_engine_status(engine=None): checks = get_engine_status(engine) s = "Execution engine status\n\n" @@ -41,5 +43,6 @@ def format_engine_status(engine=None): return s + def print_engine_status(engine): print(format_engine_status(engine)) diff --git a/scrapy/utils/ftp.py b/scrapy/utils/ftp.py index 9eca6a4da..91d2439a9 100644 --- a/scrapy/utils/ftp.py +++ b/scrapy/utils/ftp.py @@ -1,6 +1,7 @@ from ftplib import error_perm from posixpath import dirname + def ftp_makedirs_cwd(ftp, path, first_call=True): """Set the current directory of the FTP connection given in the ``ftp`` argument (as a ftplib.FTP object), creating all parent directories if they diff --git a/scrapy/utils/gz.py b/scrapy/utils/gz.py index f41e62fe3..9672e28da 100644 --- a/scrapy/utils/gz.py +++ b/scrapy/utils/gz.py @@ -45,6 +45,7 @@ def gunzip(data): _is_gzipped = re.compile(br'^application/(x-)?gzip\b', re.I).search _is_octetstream = re.compile(br'^(application|binary)/octet-stream\b', re.I).search + @deprecated def is_gzipped(response): """Return True if the response is gzipped, or False otherwise""" diff --git a/scrapy/utils/httpobj.py b/scrapy/utils/httpobj.py index b4c929b0e..b2be0a901 100644 --- a/scrapy/utils/httpobj.py +++ b/scrapy/utils/httpobj.py @@ -4,7 +4,10 @@ import weakref from six.moves.urllib.parse import urlparse + _urlparse_cache = weakref.WeakKeyDictionary() + + def urlparse_cached(request_or_response): """Return urlparse.urlparse caching the result, where the argument can be a Request or Response object diff --git a/scrapy/utils/job.py b/scrapy/utils/job.py index 389fde73a..4f1e601fc 100644 --- a/scrapy/utils/job.py +++ b/scrapy/utils/job.py @@ -1,5 +1,6 @@ import os + def job_dir(settings): path = settings['JOBDIR'] if path and not os.path.exists(path): diff --git a/scrapy/utils/python.py b/scrapy/utils/python.py index 663a8ebaa..a4201bb04 100644 --- a/scrapy/utils/python.py +++ b/scrapy/utils/python.py @@ -165,6 +165,7 @@ def memoizemethod_noargs(method): _BINARYCHARS = {six.b(chr(i)) for i in range(32)} - {b"\0", b"\t", b"\n", b"\r"} _BINARYCHARS |= {ord(ch) for ch in _BINARYCHARS} + @deprecated("scrapy.utils.python.binary_is_text") def isbinarytext(text): """ This function is deprecated. diff --git a/scrapy/utils/reactor.py b/scrapy/utils/reactor.py index 83186a372..b4b5f0645 100644 --- a/scrapy/utils/reactor.py +++ b/scrapy/utils/reactor.py @@ -1,5 +1,6 @@ from twisted.internet import reactor, error + def listen_tcp(portrange, host, factory): """Like reactor.listenTCP but tries different ports in a range.""" assert len(portrange) <= 2, "invalid portrange: %s" % portrange diff --git a/scrapy/utils/request.py b/scrapy/utils/request.py index 63d0ae772..0fce5a2e1 100644 --- a/scrapy/utils/request.py +++ b/scrapy/utils/request.py @@ -16,6 +16,8 @@ from scrapy.utils.httpobj import urlparse_cached _fingerprint_cache = weakref.WeakKeyDictionary() + + def request_fingerprint(request, include_headers=None, keep_fragments=False): """ Return the request fingerprint. diff --git a/scrapy/utils/response.py b/scrapy/utils/response.py index feab07431..29fdaaf2c 100644 --- a/scrapy/utils/response.py +++ b/scrapy/utils/response.py @@ -13,6 +13,8 @@ from w3lib import html _baseurl_cache = weakref.WeakKeyDictionary() + + def get_base_url(response): """Return the base url of the given response, joined with the response url""" if response not in _baseurl_cache: @@ -23,6 +25,8 @@ def get_base_url(response): _metaref_cache = weakref.WeakKeyDictionary() + + def get_meta_refresh(response, ignore_tags=('script', 'noscript')): """Parse the http-equiv refrsh parameter from the given response""" if response not in _metaref_cache: diff --git a/scrapy/utils/spider.py b/scrapy/utils/spider.py index 94b24f67e..bf4973fbf 100644 --- a/scrapy/utils/spider.py +++ b/scrapy/utils/spider.py @@ -28,6 +28,7 @@ def iter_spider_classes(module): getattr(obj, 'name', None): yield obj + def spidercls_for_request(spider_loader, request, default_spidercls=None, log_none=False, log_multiple=False): """Return a spider class that handles the given Request. diff --git a/scrapy/utils/template.py b/scrapy/utils/template.py index 615372fc8..96ff4b09b 100644 --- a/scrapy/utils/template.py +++ b/scrapy/utils/template.py @@ -19,6 +19,8 @@ def render_templatefile(path, **kwargs): CAMELCASE_INVALID_CHARS = re.compile(r'[^a-zA-Z\d]') + + def string_camelcase(string): """ Convert a word to its CamelCase version and remove invalid chars diff --git a/scrapy/utils/test.py b/scrapy/utils/test.py index 4b935c51b..9754366df 100644 --- a/scrapy/utils/test.py +++ b/scrapy/utils/test.py @@ -32,6 +32,7 @@ def skip_if_no_boto(): except NotConfigured as e: raise SkipTest(e) + def get_s3_content_and_delete(bucket, path, with_key=False): """ Get content from s3 key, and delete key afterwards. """ @@ -51,6 +52,7 @@ def get_s3_content_and_delete(bucket, path, with_key=False): bucket.delete_key(path) return (content, key) if with_key else content + def get_gcs_content_and_delete(bucket, path): from google.cloud import storage client = storage.Client(project=os.environ.get('GCS_PROJECT_ID')) @@ -61,6 +63,7 @@ def get_gcs_content_and_delete(bucket, path): bucket.delete_blob(path) return content, acl, blob + def get_crawler(spidercls=None, settings_dict=None): """Return an unconfigured Crawler object. If settings_dict is given, it will be used to populate the crawler settings with a project level @@ -72,12 +75,14 @@ def get_crawler(spidercls=None, settings_dict=None): runner = CrawlerRunner(settings_dict) return runner.create_crawler(spidercls or Spider) + def get_pythonpath(): """Return a PYTHONPATH suitable to use in processes so that they find this installation of Scrapy""" scrapy_path = import_module('scrapy').__path__[0] return os.path.dirname(scrapy_path) + os.pathsep + os.environ.get('PYTHONPATH', '') + def get_testenv(): """Return a OS environment dict suitable to fork processes that need to import this installation of Scrapy, instead of a system installed one. @@ -86,6 +91,7 @@ def get_testenv(): env['PYTHONPATH'] = get_pythonpath() return env + def assert_samelines(testcase, text1, text2, msg=None): """Asserts text1 and text2 have the same lines, ignoring differences in line endings between platforms diff --git a/scrapy/utils/versions.py b/scrapy/utils/versions.py index 48484b303..b0737d3d5 100644 --- a/scrapy/utils/versions.py +++ b/scrapy/utils/versions.py @@ -27,5 +27,5 @@ def scrapy_components_versions(): ("Python", sys.version.replace("\n", "- ")), ("pyOpenSSL", get_openssl_version()), ("cryptography", cryptography.__version__), - ("Platform", platform.platform()), + ("Platform", platform.platform()), ] diff --git a/tests/mocks/dummydbm.py b/tests/mocks/dummydbm.py index 431428331..75c74daf5 100644 --- a/tests/mocks/dummydbm.py +++ b/tests/mocks/dummydbm.py @@ -13,6 +13,7 @@ error = KeyError _DATABASES = collections.defaultdict(DummyDB) + def open(file, flag='r', mode=0o666): """Open or create a dummy database compatible. diff --git a/tests/pipelines.py b/tests/pipelines.py index 7e2895a5c..d7d3b5259 100644 --- a/tests/pipelines.py +++ b/tests/pipelines.py @@ -2,6 +2,7 @@ Some pipelines used for testing """ + class ZeroDivisionErrorPipeline(object): def open_spider(self, spider): diff --git a/tests/spiders.py b/tests/spiders.py index 7816bf7c7..2487ecc22 100644 --- a/tests/spiders.py +++ b/tests/spiders.py @@ -16,6 +16,7 @@ class MockServerSpider(Spider): super(MockServerSpider, self).__init__(*args, **kwargs) self.mockserver = mockserver + class MetaSpider(MockServerSpider): name = 'meta' diff --git a/tests/test_cmdline/extensions.py b/tests/test_cmdline/extensions.py index 28456b55d..c64e87d81 100644 --- a/tests/test_cmdline/extensions.py +++ b/tests/test_cmdline/extensions.py @@ -1,5 +1,6 @@ """A test extension used to check the settings loading order""" + class TestExtension(object): def __init__(self, settings): diff --git a/tests/test_command_parse.py b/tests/test_command_parse.py index b134beb88..b7035fdff 100644 --- a/tests/test_command_parse.py +++ b/tests/test_command_parse.py @@ -12,6 +12,7 @@ def _textmode(bstr): and reading from it in text mode""" return to_unicode(bstr).replace(os.linesep, '\n') + class ParseCommandTest(ProcessTest, SiteTest, CommandTest): command = 'parse' diff --git a/tests/test_dependencies.py b/tests/test_dependencies.py index 03bf2ffcf..e31ccd9b5 100644 --- a/tests/test_dependencies.py +++ b/tests/test_dependencies.py @@ -1,6 +1,7 @@ from importlib import import_module from twisted.trial import unittest + class ScrapyUtilsTest(unittest.TestCase): def test_required_openssl_version(self): try: diff --git a/tests/test_downloadermiddleware_ajaxcrawlable.py b/tests/test_downloadermiddleware_ajaxcrawlable.py index 493691ea4..5a56c9db2 100644 --- a/tests/test_downloadermiddleware_ajaxcrawlable.py +++ b/tests/test_downloadermiddleware_ajaxcrawlable.py @@ -5,8 +5,10 @@ from scrapy.spiders import Spider from scrapy.http import Request, HtmlResponse, Response from scrapy.utils.test import get_crawler + __doctests__ = ['scrapy.downloadermiddlewares.ajaxcrawl'] + class AjaxCrawlMiddlewareTest(unittest.TestCase): def setUp(self): crawler = get_crawler(Spider, {'AJAXCRAWL_ENABLED': True}) diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index 9d863b6e3..00e6c685e 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -149,6 +149,7 @@ class FilesystemStorageTest(DefaultStorageTest): storage_class = 'scrapy.extensions.httpcache.FilesystemCacheStorage' + class FilesystemStorageGzipTest(FilesystemStorageTest): def _get_settings(self, **new_settings): diff --git a/tests/test_dupefilters.py b/tests/test_dupefilters.py index d7eb98c97..e4b0bdf83 100644 --- a/tests/test_dupefilters.py +++ b/tests/test_dupefilters.py @@ -12,6 +12,7 @@ from scrapy.utils.job import job_dir from scrapy.utils.test import get_crawler from tests.spiders import SimpleSpider + class FromCrawlerRFPDupeFilter(RFPDupeFilter): @classmethod diff --git a/tests/test_http_headers.py b/tests/test_http_headers.py index 69d906fbf..c83cf3b66 100644 --- a/tests/test_http_headers.py +++ b/tests/test_http_headers.py @@ -3,6 +3,7 @@ import copy from scrapy.http import Headers + class HeadersTest(unittest.TestCase): def assertSortedEqual(self, first, second, msg=None): diff --git a/tests/test_logformatter.py b/tests/test_logformatter.py index b4ea30bb7..0724d1807 100644 --- a/tests/test_logformatter.py +++ b/tests/test_logformatter.py @@ -118,6 +118,7 @@ class DropSomeItemsPipeline(object): else: self.drop = True + class ShowOrSkipMessagesTestCase(TwistedTestCase): def setUp(self): self.mockserver = MockServer() diff --git a/tests/test_mail.py b/tests/test_mail.py index b139e98d8..ddb0f1e70 100644 --- a/tests/test_mail.py +++ b/tests/test_mail.py @@ -6,6 +6,7 @@ from email.charset import Charset from scrapy.mail import MailSender + class MailSenderTest(unittest.TestCase): def test_send(self): diff --git a/tests/test_middleware.py b/tests/test_middleware.py index af9b43d61..ebf817c7e 100644 --- a/tests/test_middleware.py +++ b/tests/test_middleware.py @@ -4,6 +4,7 @@ from scrapy.settings import Settings from scrapy.exceptions import NotConfigured from scrapy.middleware import MiddlewareManager + class M1(object): def open_spider(self, spider): @@ -15,6 +16,7 @@ class M1(object): def process(self, response, request, spider): pass + class M2(object): def open_spider(self, spider): @@ -25,6 +27,7 @@ class M2(object): pass + class M3(object): def process(self, response, request, spider): @@ -54,6 +57,7 @@ class TestMiddlewareManager(MiddlewareManager): if hasattr(mw, 'process'): self.methods['process'].append(mw.process) + class MiddlewareManagerTest(unittest.TestCase): def test_init(self): diff --git a/tests/test_responsetypes.py b/tests/test_responsetypes.py index f89042b3d..d5a3371ab 100644 --- a/tests/test_responsetypes.py +++ b/tests/test_responsetypes.py @@ -4,6 +4,7 @@ from scrapy.responsetypes import responsetypes from scrapy.http import Response, TextResponse, XmlResponse, HtmlResponse, Headers + class ResponseTypesTest(unittest.TestCase): def test_from_filename(self): diff --git a/tests/test_robotstxt_interface.py b/tests/test_robotstxt_interface.py index cd7480e33..080507276 100644 --- a/tests/test_robotstxt_interface.py +++ b/tests/test_robotstxt_interface.py @@ -19,6 +19,7 @@ def rerp_available(): return False return True + def protego_available(): # check if protego parser is installed try: @@ -27,6 +28,7 @@ def protego_available(): return False return True + class BaseRobotParserTest: def _setUp(self, parser_cls): self.parser_cls = parser_cls diff --git a/tests/test_spiderloader/__init__.py b/tests/test_spiderloader/__init__.py index 106da798c..d8be6e277 100644 --- a/tests/test_spiderloader/__init__.py +++ b/tests/test_spiderloader/__init__.py @@ -109,6 +109,7 @@ class SpiderLoaderTest(unittest.TestCase): spiders = spider_loader.list() self.assertEqual(spiders, []) + class DuplicateSpiderNameLoaderTest(unittest.TestCase): def setUp(self): diff --git a/tests/test_spiderloader/test_spiders/nested/spider4.py b/tests/test_spiderloader/test_spiders/nested/spider4.py index 35b71870a..dbd1fb123 100644 --- a/tests/test_spiderloader/test_spiders/nested/spider4.py +++ b/tests/test_spiderloader/test_spiders/nested/spider4.py @@ -1,5 +1,6 @@ from scrapy.spiders import Spider + class Spider4(Spider): name = "spider4" allowed_domains = ['spider4.com'] diff --git a/tests/test_spiderloader/test_spiders/spider0.py b/tests/test_spiderloader/test_spiders/spider0.py index 75a90794e..af679dbd6 100644 --- a/tests/test_spiderloader/test_spiders/spider0.py +++ b/tests/test_spiderloader/test_spiders/spider0.py @@ -1,4 +1,5 @@ from scrapy.spiders import Spider + class Spider0(Spider): allowed_domains = ["scrapy1.org", "scrapy3.org"] diff --git a/tests/test_spiderloader/test_spiders/spider1.py b/tests/test_spiderloader/test_spiders/spider1.py index 76efddc7f..6b4317a90 100644 --- a/tests/test_spiderloader/test_spiders/spider1.py +++ b/tests/test_spiderloader/test_spiders/spider1.py @@ -1,5 +1,6 @@ from scrapy.spiders import Spider + class Spider1(Spider): name = "spider1" allowed_domains = ["scrapy1.org", "scrapy3.org"] diff --git a/tests/test_spiderloader/test_spiders/spider2.py b/tests/test_spiderloader/test_spiders/spider2.py index 0badd8437..352601863 100644 --- a/tests/test_spiderloader/test_spiders/spider2.py +++ b/tests/test_spiderloader/test_spiders/spider2.py @@ -1,5 +1,6 @@ from scrapy.spiders import Spider + class Spider2(Spider): name = "spider2" allowed_domains = ["scrapy2.org", "scrapy3.org"] diff --git a/tests/test_spiderloader/test_spiders/spider3.py b/tests/test_spiderloader/test_spiders/spider3.py index d406f2d4f..84998ba35 100644 --- a/tests/test_spiderloader/test_spiders/spider3.py +++ b/tests/test_spiderloader/test_spiders/spider3.py @@ -1,5 +1,6 @@ from scrapy.spiders import Spider + class Spider3(Spider): name = "spider3" allowed_domains = ['spider3.com'] diff --git a/tests/test_spidermiddleware_offsite.py b/tests/test_spidermiddleware_offsite.py index 7e4af0d4c..b97d9b675 100644 --- a/tests/test_spidermiddleware_offsite.py +++ b/tests/test_spidermiddleware_offsite.py @@ -9,6 +9,7 @@ from scrapy.spidermiddlewares.offsite import URLWarning from scrapy.utils.test import get_crawler import warnings + class TestOffsiteMiddleware(TestCase): def setUp(self): @@ -53,6 +54,7 @@ class TestOffsiteMiddleware2(TestOffsiteMiddleware): out = list(self.mw.process_spider_output(res, reqs, self.spider)) self.assertEqual(out, reqs) + class TestOffsiteMiddleware3(TestOffsiteMiddleware2): def _get_spider(self): diff --git a/tests/test_spidermiddleware_output_chain.py b/tests/test_spidermiddleware_output_chain.py index 6f8727a15..940e31070 100644 --- a/tests/test_spidermiddleware_output_chain.py +++ b/tests/test_spidermiddleware_output_chain.py @@ -34,6 +34,7 @@ class RecoverySpider(Spider): if not response.meta.get('dont_fail'): raise TabError() + class RecoveryMiddleware: def process_spider_exception(self, response, exception, spider): spider.logger.info('Middleware: %s exception caught', exception.__class__.__name__) @@ -50,6 +51,7 @@ class FailProcessSpiderInputMiddleware: spider.logger.info('Middleware: will raise IndexError') raise IndexError() + class ProcessSpiderInputSpiderWithoutErrback(Spider): name = 'ProcessSpiderInputSpiderWithoutErrback' custom_settings = { @@ -177,6 +179,7 @@ class GeneratorRecoverMiddleware: spider.logger.info('%s: %s caught', method, exception.__class__.__name__) yield {'processed': [method]} + class GeneratorDoNothingAfterRecoveryMiddleware(_GeneratorDoNothingMiddleware): pass @@ -247,6 +250,7 @@ class NotGeneratorRecoverMiddleware: spider.logger.info('%s: %s caught', method, exception.__class__.__name__) return [{'processed': [method]}] + class NotGeneratorDoNothingAfterRecoveryMiddleware(_NotGeneratorDoNothingMiddleware): pass diff --git a/tests/test_spidermiddleware_referer.py b/tests/test_spidermiddleware_referer.py index 21439c20e..a9c31a983 100644 --- a/tests/test_spidermiddleware_referer.py +++ b/tests/test_spidermiddleware_referer.py @@ -349,6 +349,7 @@ class TestSettingsCustomPolicy(TestRefererMiddleware): ] + # --- Tests using Request meta dict to set policy class TestRequestMetaDefault(MixinDefault, TestRefererMiddleware): req_meta = {'referrer_policy': POLICY_SCRAPY_DEFAULT} @@ -518,14 +519,17 @@ class TestPolicyHeaderPredecence001(MixinUnsafeUrl, TestRefererMiddleware): settings = {'REFERRER_POLICY': 'scrapy.spidermiddlewares.referer.SameOriginPolicy'} resp_headers = {'Referrer-Policy': POLICY_UNSAFE_URL.upper()} + class TestPolicyHeaderPredecence002(MixinNoReferrer, TestRefererMiddleware): settings = {'REFERRER_POLICY': 'scrapy.spidermiddlewares.referer.NoReferrerWhenDowngradePolicy'} resp_headers = {'Referrer-Policy': POLICY_NO_REFERRER.swapcase()} + class TestPolicyHeaderPredecence003(MixinNoReferrerWhenDowngrade, TestRefererMiddleware): settings = {'REFERRER_POLICY': 'scrapy.spidermiddlewares.referer.OriginWhenCrossOriginPolicy'} resp_headers = {'Referrer-Policy': POLICY_NO_REFERRER_WHEN_DOWNGRADE.title()} + class TestPolicyHeaderPredecence004(MixinNoReferrerWhenDowngrade, TestRefererMiddleware): """ The empty string means "no-referrer-when-downgrade" diff --git a/tests/test_squeues.py b/tests/test_squeues.py index 3ded5c027..d5fcf2f7f 100644 --- a/tests/test_squeues.py +++ b/tests/test_squeues.py @@ -7,16 +7,20 @@ from scrapy.http import Request from scrapy.loader import ItemLoader from scrapy.selector import Selector + class TestItem(Item): name = Field() + def _test_procesor(x): return x + x + class TestLoader(ItemLoader): default_item_class = TestItem name_out = staticmethod(_test_procesor) + def nonserializable_object_test(self): q = self.queue() try: @@ -35,6 +39,7 @@ def nonserializable_object_test(self): sel = Selector(text='

some text

') self.assertRaises(ValueError, q.push, sel) + class MarshalFifoDiskQueueTest(t.FifoDiskQueueTest): chunksize = 100000 @@ -53,15 +58,19 @@ class MarshalFifoDiskQueueTest(t.FifoDiskQueueTest): test_nonserializable_object = nonserializable_object_test + class ChunkSize1MarshalFifoDiskQueueTest(MarshalFifoDiskQueueTest): chunksize = 1 + class ChunkSize2MarshalFifoDiskQueueTest(MarshalFifoDiskQueueTest): chunksize = 2 + class ChunkSize3MarshalFifoDiskQueueTest(MarshalFifoDiskQueueTest): chunksize = 3 + class ChunkSize4MarshalFifoDiskQueueTest(MarshalFifoDiskQueueTest): chunksize = 4 @@ -100,15 +109,19 @@ class PickleFifoDiskQueueTest(MarshalFifoDiskQueueTest): self.assertEqual(r.url, r2.url) assert r2.meta['request'] is r2 + class ChunkSize1PickleFifoDiskQueueTest(PickleFifoDiskQueueTest): chunksize = 1 + class ChunkSize2PickleFifoDiskQueueTest(PickleFifoDiskQueueTest): chunksize = 2 + class ChunkSize3PickleFifoDiskQueueTest(PickleFifoDiskQueueTest): chunksize = 3 + class ChunkSize4PickleFifoDiskQueueTest(PickleFifoDiskQueueTest): chunksize = 4 diff --git a/tests/test_utils_console.py b/tests/test_utils_console.py index c2211848c..380c41367 100644 --- a/tests/test_utils_console.py +++ b/tests/test_utils_console.py @@ -14,6 +14,7 @@ try: except ImportError: ipy = False + class UtilsConsoleTestCase(unittest.TestCase): def test_get_shell_embed_func(self): diff --git a/tests/test_utils_defer.py b/tests/test_utils_defer.py index 003bb9b02..d642ed3ed 100644 --- a/tests/test_utils_defer.py +++ b/tests/test_utils_defer.py @@ -33,14 +33,23 @@ class MustbeDeferredTest(unittest.TestCase): steps.append(2) # add another value, that should be catched by assertEqual return dfd + def cb1(value, arg1, arg2): return "(cb1 %s %s %s)" % (value, arg1, arg2) + + def cb2(value, arg1, arg2): return defer.succeed("(cb2 %s %s %s)" % (value, arg1, arg2)) + + def cb3(value, arg1, arg2): return "(cb3 %s %s %s)" % (value, arg1, arg2) + + def cb_fail(value, arg1, arg2): return Failure(TypeError()) + + def eb1(failure, arg1, arg2): return "(eb1 %s %s %s)" % (failure.value.__class__.__name__, arg1, arg2) diff --git a/tests/test_utils_http.py b/tests/test_utils_http.py index 2524153ea..f9af4bf87 100644 --- a/tests/test_utils_http.py +++ b/tests/test_utils_http.py @@ -2,6 +2,7 @@ import unittest from scrapy.utils.http import decode_chunked_transfer + class ChunkedTest(unittest.TestCase): def test_decode_chunked_transfer(self): diff --git a/tests/test_utils_httpobj.py b/tests/test_utils_httpobj.py index 4f9f7a370..2c3965bbc 100644 --- a/tests/test_utils_httpobj.py +++ b/tests/test_utils_httpobj.py @@ -4,6 +4,7 @@ from six.moves.urllib.parse import urlparse from scrapy.http import Request from scrapy.utils.httpobj import urlparse_cached + class HttpobjUtilsTest(unittest.TestCase): def test_urlparse_cached(self): diff --git a/tests/test_utils_iterators.py b/tests/test_utils_iterators.py index 2d845697e..f16ef8110 100644 --- a/tests/test_utils_iterators.py +++ b/tests/test_utils_iterators.py @@ -235,6 +235,7 @@ class LxmlXmliterTestCase(XmliterTestCase): i = self.xmliter(42, 'product') self.assertRaises(TypeError, next, i) + class UtilsCsvTestCase(unittest.TestCase): sample_feeds_dir = os.path.join(os.path.abspath(os.path.dirname(__file__)), 'sample_data', 'feeds') sample_feed_path = os.path.join(sample_feeds_dir, 'feed-sample3.csv') diff --git a/tests/test_utils_request.py b/tests/test_utils_request.py index 625a32048..3da95b95a 100644 --- a/tests/test_utils_request.py +++ b/tests/test_utils_request.py @@ -4,6 +4,7 @@ from scrapy.http import Request from scrapy.utils.request import request_fingerprint, _fingerprint_cache, \ request_authenticate, request_httprepr + class UtilsRequestTest(unittest.TestCase): def test_request_fingerprint(self): diff --git a/tests/test_utils_signal.py b/tests/test_utils_signal.py index 62edd420d..16b7c5c68 100644 --- a/tests/test_utils_signal.py +++ b/tests/test_utils_signal.py @@ -66,6 +66,7 @@ class SendCatchLogDeferredTest2(SendCatchLogTest): def _get_result(self, signal, *a, **kw): return send_catch_log_deferred(signal, *a, **kw) + class SendCatchLogTest2(unittest.TestCase): def test_error_logged_if_deferred_not_supported(self): diff --git a/tests/test_utils_sitemap.py b/tests/test_utils_sitemap.py index 716bb44eb..db323ab31 100644 --- a/tests/test_utils_sitemap.py +++ b/tests/test_utils_sitemap.py @@ -2,6 +2,7 @@ import unittest from scrapy.utils.sitemap import Sitemap, sitemap_urls_from_robots + class SitemapTest(unittest.TestCase): def test_sitemap(self): diff --git a/tests/test_utils_spider.py b/tests/test_utils_spider.py index d9de1ce77..edeeacc80 100644 --- a/tests/test_utils_spider.py +++ b/tests/test_utils_spider.py @@ -9,12 +9,15 @@ from scrapy.spiders import CrawlSpider class MyBaseSpider(CrawlSpider): pass # abstract spider + class MySpider1(MyBaseSpider): name = 'myspider1' + class MySpider2(MyBaseSpider): name = 'myspider2' + class UtilsSpidersTestCase(unittest.TestCase): def test_iterate_spider_output(self): diff --git a/tests/test_utils_url.py b/tests/test_utils_url.py index e6588055c..93addc082 100644 --- a/tests/test_utils_url.py +++ b/tests/test_utils_url.py @@ -187,6 +187,7 @@ class AddHttpIfNoScheme(unittest.TestCase): class GuessSchemeTest(unittest.TestCase): pass + def create_guess_scheme_t(args): def do_expected(self): url = guess_scheme(args[0]) @@ -195,6 +196,7 @@ def create_guess_scheme_t(args): args[0], url, args[1]) return do_expected + def create_skipped_scheme_t(args): def do_expected(self): raise unittest.SkipTest(args[2]) From e18014d84d058d7a702c7b0f8c190ee5468abc7d Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Mon, 22 Jul 2019 20:51:03 +0500 Subject: [PATCH 269/387] Remove Python 2-only tests. --- conftest.py | 10 +- tests/py3-ignores.txt | 3 - tests/test_downloader_handlers.py | 22 +-- tests/test_linkextractors_deprecated.py | 233 ------------------------ tests/test_proxy_connect.py | 120 ------------ tests/test_utils_python.py | 27 --- tests/test_webclient.py | 20 -- 7 files changed, 14 insertions(+), 421 deletions(-) delete mode 100644 tests/test_linkextractors_deprecated.py delete mode 100644 tests/test_proxy_connect.py diff --git a/conftest.py b/conftest.py index d5d61ddd3..74fb101e9 100644 --- a/conftest.py +++ b/conftest.py @@ -1,4 +1,3 @@ -import six import pytest @@ -8,11 +7,10 @@ collect_ignore = [ ] -if six.PY3: - for line in open('tests/py3-ignores.txt'): - file_path = line.strip() - if file_path and file_path[0] != '#': - collect_ignore.append(file_path) +for line in open('tests/py3-ignores.txt'): + file_path = line.strip() + if file_path and file_path[0] != '#': + collect_ignore.append(file_path) @pytest.fixture() diff --git a/tests/py3-ignores.txt b/tests/py3-ignores.txt index 313e74ec9..45cf6fb92 100644 --- a/tests/py3-ignores.txt +++ b/tests/py3-ignores.txt @@ -1,6 +1,3 @@ -tests/test_linkextractors_deprecated.py -tests/test_proxy_connect.py - scrapy/linkextractors/sgml.py scrapy/linkextractors/regex.py scrapy/linkextractors/htmlparser.py diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 59d4a3eec..b06fcf6c3 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -630,18 +630,16 @@ class Http11MockServerTestCase(unittest.TestCase): # download_maxsize < 100, hence the CancelledError self.assertIsInstance(failure.value, defer.CancelledError) - if six.PY2: - request.headers.setdefault(b'Accept-Encoding', b'gzip,deflate') - request = request.replace(url=self.mockserver.url('/xpayload')) - yield crawler.crawl(seed=request) - # download_maxsize = 50 is enough for the gzipped response - failure = crawler.spider.meta.get('failure') - self.assertTrue(failure is None) - reason = crawler.spider.meta['close_reason'] - self.assertTrue(reason, 'finished') - else: - # See issue https://twistedmatrix.com/trac/ticket/8175 - raise unittest.SkipTest("xpayload only enabled for PY2") + # See issue https://twistedmatrix.com/trac/ticket/8175 + raise unittest.SkipTest("xpayload fails on PY3") + request.headers.setdefault(b'Accept-Encoding', b'gzip,deflate') + request = request.replace(url=self.mockserver.url('/xpayload')) + yield crawler.crawl(seed=request) + # download_maxsize = 50 is enough for the gzipped response + failure = crawler.spider.meta.get('failure') + self.assertTrue(failure is None) + reason = crawler.spider.meta['close_reason'] + self.assertTrue(reason, 'finished') class UriResource(resource.Resource): diff --git a/tests/test_linkextractors_deprecated.py b/tests/test_linkextractors_deprecated.py deleted file mode 100644 index 1366971be..000000000 --- a/tests/test_linkextractors_deprecated.py +++ /dev/null @@ -1,233 +0,0 @@ -# -*- coding: utf-8 -*- -import unittest -from scrapy.linkextractors.regex import RegexLinkExtractor -from scrapy.http import HtmlResponse -from scrapy.link import Link -from scrapy.linkextractors.htmlparser import HtmlParserLinkExtractor -from scrapy.linkextractors.sgml import SgmlLinkExtractor, BaseSgmlLinkExtractor -from tests import get_testdata - -from tests.test_linkextractors import Base - - -class BaseSgmlLinkExtractorTestCase(unittest.TestCase): - # XXX: should we move some of these tests to base link extractor tests? - - def test_basic(self): - html = """Page title<title> - <body><p><a href="item/12.html">Item 12</a></p> - <p><a href="/about.html">About us</a></p> - <img src="/logo.png" alt="Company logo (not a link)" /> - <p><a href="../othercat.html">Other category</a></p> - <p><a href="/">>></a></p> - <p><a href="/" /></p> - </body></html>""" - response = HtmlResponse("http://example.org/somepage/index.html", body=html) - - lx = BaseSgmlLinkExtractor() # default: tag=a, attr=href - self.assertEqual(lx.extract_links(response), - [Link(url='http://example.org/somepage/item/12.html', text='Item 12'), - Link(url='http://example.org/about.html', text='About us'), - Link(url='http://example.org/othercat.html', text='Other category'), - Link(url='http://example.org/', text='>>'), - Link(url='http://example.org/', text='')]) - - def test_base_url(self): - html = """<html><head><title>Page title<title><base href="http://otherdomain.com/base/" /> - <body><p><a href="item/12.html">Item 12</a></p> - </body></html>""" - response = HtmlResponse("http://example.org/somepage/index.html", body=html) - - lx = BaseSgmlLinkExtractor() # default: tag=a, attr=href - self.assertEqual(lx.extract_links(response), - [Link(url='http://otherdomain.com/base/item/12.html', text='Item 12')]) - - # base url is an absolute path and relative to host - html = """<html><head><title>Page title<title><base href="/" /> - <body><p><a href="item/12.html">Item 12</a></p></body></html>""" - response = HtmlResponse("https://example.org/somepage/index.html", body=html) - self.assertEqual(lx.extract_links(response), - [Link(url='https://example.org/item/12.html', text='Item 12')]) - - # base url has no scheme - html = """<html><head><title>Page title<title><base href="//noschemedomain.com/path/to/" /> - <body><p><a href="item/12.html">Item 12</a></p></body></html>""" - response = HtmlResponse("https://example.org/somepage/index.html", body=html) - self.assertEqual(lx.extract_links(response), - [Link(url='https://noschemedomain.com/path/to/item/12.html', text='Item 12')]) - - def test_link_text_wrong_encoding(self): - html = """<body><p><a href="item/12.html">Wrong: \xed</a></p></body></html>""" - response = HtmlResponse("http://www.example.com", body=html, encoding='utf-8') - lx = BaseSgmlLinkExtractor() - self.assertEqual(lx.extract_links(response), [ - Link(url='http://www.example.com/item/12.html', text=u'Wrong: \ufffd'), - ]) - - def test_extraction_encoding(self): - body = get_testdata('link_extractor', 'linkextractor_noenc.html') - response_utf8 = HtmlResponse(url='http://example.com/utf8', body=body, headers={'Content-Type': ['text/html; charset=utf-8']}) - response_noenc = HtmlResponse(url='http://example.com/noenc', body=body) - body = get_testdata('link_extractor', 'linkextractor_latin1.html') - response_latin1 = HtmlResponse(url='http://example.com/latin1', body=body) - - lx = BaseSgmlLinkExtractor() - self.assertEqual(lx.extract_links(response_utf8), [ - Link(url='http://example.com/sample_%C3%B1.html', text=''), - Link(url='http://example.com/sample_%E2%82%AC.html', text='sample \xe2\x82\xac text'.decode('utf-8')), - ]) - - self.assertEqual(lx.extract_links(response_noenc), [ - Link(url='http://example.com/sample_%C3%B1.html', text=''), - Link(url='http://example.com/sample_%E2%82%AC.html', text='sample \xe2\x82\xac text'.decode('utf-8')), - ]) - - # document encoding does not affect URL path component, only query part - # >>> u'sample_ñ.html'.encode('utf8') - # b'sample_\xc3\xb1.html' - # >>> u"sample_á.html".encode('utf8') - # b'sample_\xc3\xa1.html' - # >>> u"sample_ö.html".encode('utf8') - # b'sample_\xc3\xb6.html' - # >>> u"£32".encode('latin1') - # b'\xa332' - # >>> u"µ".encode('latin1') - # b'\xb5' - self.assertEqual(lx.extract_links(response_latin1), [ - Link(url='http://example.com/sample_%C3%B1.html', text=''), - Link(url='http://example.com/sample_%C3%A1.html', text='sample \xe1 text'.decode('latin1')), - Link(url='http://example.com/sample_%C3%B6.html?price=%A332&%B5=unit', text=''), - ]) - - def test_matches(self): - url1 = 'http://lotsofstuff.com/stuff1/index' - url2 = 'http://evenmorestuff.com/uglystuff/index' - - lx = BaseSgmlLinkExtractor() - self.assertEqual(lx.matches(url1), True) - self.assertEqual(lx.matches(url2), True) - - -class HtmlParserLinkExtractorTestCase(unittest.TestCase): - - def setUp(self): - body = get_testdata('link_extractor', 'sgml_linkextractor.html') - self.response = HtmlResponse(url='http://example.com/index', body=body) - - def test_extraction(self): - # Default arguments - lx = HtmlParserLinkExtractor() - self.assertEqual(lx.extract_links(self.response), [ - Link(url='http://example.com/sample2.html', text=u'sample 2'), - Link(url='http://example.com/sample3.html', text=u'sample 3 text'), - Link(url='http://example.com/sample3.html', text=u'sample 3 repetition'), - Link(url='http://example.com/sample3.html#foo', text=u'sample 3 repetition with fragment'), - Link(url='http://www.google.com/something', text=u''), - Link(url='http://example.com/innertag.html', text=u'inner tag'), - Link(url='http://example.com/page%204.html', text=u'href with whitespaces'), - ]) - - def test_link_wrong_href(self): - html = """ - <a href="http://example.org/item1.html">Item 1</a> - <a href="http://[example.org/item2.html">Item 2</a> - <a href="http://example.org/item3.html">Item 3</a> - """ - response = HtmlResponse("http://example.org/index.html", body=html) - lx = HtmlParserLinkExtractor() - self.assertEqual([link for link in lx.extract_links(response)], [ - Link(url='http://example.org/item1.html', text=u'Item 1', nofollow=False), - Link(url='http://example.org/item3.html', text=u'Item 3', nofollow=False), - ]) - - -class SgmlLinkExtractorTestCase(Base.LinkExtractorTestCase): - extractor_cls = SgmlLinkExtractor - escapes_whitespace = True - - def test_deny_extensions(self): - html = """<a href="page.html">asd</a> and <a href="photo.jpg">""" - response = HtmlResponse("http://example.org/", body=html) - lx = SgmlLinkExtractor(deny_extensions="jpg") - self.assertEqual(lx.extract_links(response), [ - Link(url='http://example.org/page.html', text=u'asd'), - ]) - - def test_attrs_sgml(self): - html = """<html><area href="sample1.html"></area> - <a ref="sample2.html">sample text 2</a></html>""" - response = HtmlResponse("http://example.com/index.html", body=html) - lx = SgmlLinkExtractor(attrs="href") - self.assertEqual(lx.extract_links(response), [ - Link(url='http://example.com/sample1.html', text=u''), - ]) - - def test_link_nofollow(self): - html = """ - <a href="page.html?action=print" rel="nofollow">Printer-friendly page</a> - <a href="about.html">About us</a> - <a href="http://google.com/something" rel="external nofollow">Something</a> - """ - response = HtmlResponse("http://example.org/page.html", body=html) - lx = SgmlLinkExtractor() - self.assertEqual([link for link in lx.extract_links(response)], [ - Link(url='http://example.org/page.html?action=print', text=u'Printer-friendly page', nofollow=True), - Link(url='http://example.org/about.html', text=u'About us', nofollow=False), - Link(url='http://google.com/something', text=u'Something', nofollow=True), - ]) - - -class RegexLinkExtractorTestCase(unittest.TestCase): - # XXX: RegexLinkExtractor is not deprecated yet, but it must be rewritten - # not to depend on SgmlLinkExractor. Its speed is also much worse - # than it should be. - - def setUp(self): - body = get_testdata('link_extractor', 'sgml_linkextractor.html') - self.response = HtmlResponse(url='http://example.com/index', body=body) - - def test_extraction(self): - # Default arguments - lx = RegexLinkExtractor() - self.assertEqual(lx.extract_links(self.response), - [Link(url='http://example.com/sample2.html', text=u'sample 2'), - Link(url='http://example.com/sample3.html', text=u'sample 3 text'), - Link(url='http://example.com/sample3.html#foo', text=u'sample 3 repetition with fragment'), - Link(url='http://www.google.com/something', text=u''), - Link(url='http://example.com/innertag.html', text=u'inner tag'),]) - - def test_link_wrong_href(self): - html = """ - <a href="http://example.org/item1.html">Item 1</a> - <a href="http://[example.org/item2.html">Item 2</a> - <a href="http://example.org/item3.html">Item 3</a> - """ - response = HtmlResponse("http://example.org/index.html", body=html) - lx = RegexLinkExtractor() - self.assertEqual([link for link in lx.extract_links(response)], [ - Link(url='http://example.org/item1.html', text=u'Item 1', nofollow=False), - Link(url='http://example.org/item3.html', text=u'Item 3', nofollow=False), - ]) - - def test_html_base_href(self): - html = """ - <html> - <head> - <base href="http://b.com/"> - </head> - <body> - <a href="test.html"></a> - </body> - </html> - """ - response = HtmlResponse("http://a.com/", body=html) - lx = RegexLinkExtractor() - self.assertEqual([link for link in lx.extract_links(response)], [ - Link(url='http://b.com/test.html', text=u'', nofollow=False), - ]) - - @unittest.expectedFailure - def test_extraction(self): - # RegexLinkExtractor doesn't parse URLs with leading/trailing - # whitespaces correctly. - super(RegexLinkExtractorTestCase, self).test_extraction() diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py deleted file mode 100644 index ae1236bcb..000000000 --- a/tests/test_proxy_connect.py +++ /dev/null @@ -1,120 +0,0 @@ -import json -import os -import time - -from six.moves.urllib.parse import urlsplit, urlunsplit -from threading import Thread -from libmproxy import controller, proxy -from netlib import http_auth -from testfixtures import LogCapture - -from twisted.internet import defer -from twisted.trial.unittest import TestCase -from scrapy.utils.test import get_crawler -from scrapy.http import Request -from tests.spiders import SimpleSpider, SingleRequestSpider -from tests.mockserver import MockServer - - -class HTTPSProxy(controller.Master, Thread): - - def __init__(self): - password_manager = http_auth.PassManSingleUser('scrapy', 'scrapy') - authenticator = http_auth.BasicProxyAuth(password_manager, "mitmproxy") - cert_path = os.path.join(os.path.abspath(os.path.dirname(__file__)), - 'keys', 'mitmproxy-ca.pem') - server = proxy.ProxyServer(proxy.ProxyConfig( - authenticator = authenticator, - cacert = cert_path), - 0) - self.server = server - Thread.__init__(self) - controller.Master.__init__(self, server) - - def http_address(self): - return 'http://scrapy:scrapy@%s:%d' % self.server.socket.getsockname() - - -def _wrong_credentials(proxy_url): - bad_auth_proxy = list(urlsplit(proxy_url)) - bad_auth_proxy[1] = bad_auth_proxy[1].replace('scrapy:scrapy@', 'wrong:wronger@') - return urlunsplit(bad_auth_proxy) - -class ProxyConnectTestCase(TestCase): - - def setUp(self): - self.mockserver = MockServer() - self.mockserver.__enter__() - self._oldenv = os.environ.copy() - - self._proxy = HTTPSProxy() - self._proxy.start() - - # Wait for the proxy to start. - time.sleep(1.0) - os.environ['https_proxy'] = self._proxy.http_address() - os.environ['http_proxy'] = self._proxy.http_address() - - def tearDown(self): - self.mockserver.__exit__(None, None, None) - self._proxy.shutdown() - os.environ = self._oldenv - - @defer.inlineCallbacks - def test_https_connect_tunnel(self): - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(200, l) - - @defer.inlineCallbacks - def test_https_noconnect(self): - proxy = os.environ['https_proxy'] - os.environ['https_proxy'] = proxy + '?noconnect' - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(200, l) - - @defer.inlineCallbacks - def test_https_connect_tunnel_error(self): - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl("https://localhost:99999/status?n=200") - self._assert_got_tunnel_error(l) - - @defer.inlineCallbacks - def test_https_tunnel_auth_error(self): - os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - # The proxy returns a 407 error code but it does not reach the client; - # he just sees a TunnelError. - self._assert_got_tunnel_error(l) - - @defer.inlineCallbacks - def test_https_tunnel_without_leak_proxy_authorization_header(self): - request = Request(self.mockserver.url("/echo", is_secure=True)) - crawler = get_crawler(SingleRequestSpider) - with LogCapture() as l: - yield crawler.crawl(seed=request) - self._assert_got_response_code(200, l) - echo = json.loads(crawler.spider.meta['responses'][0].body) - self.assertTrue('Proxy-Authorization' not in echo['headers']) - - @defer.inlineCallbacks - def test_https_noconnect_auth_error(self): - os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) + '?noconnect' - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(407, l) - - def _assert_got_response_code(self, code, log): - print(log) - self.assertEqual(str(log).count('Crawled (%d)' % code), 1) - - def _assert_got_tunnel_error(self, log): - print(log) - self.assertIn('TunnelError', str(log)) diff --git a/tests/test_utils_python.py b/tests/test_utils_python.py index a94398796..6857356f6 100644 --- a/tests/test_utils_python.py +++ b/tests/test_utils_python.py @@ -163,33 +163,6 @@ class UtilsPythonTestCase(unittest.TestCase): gc.collect() self.assertFalse(len(wk._weakdict)) - @unittest.skipUnless(six.PY2, "deprecated function") - def test_stringify_dict(self): - d = {'a': 123, u'b': b'c', u'd': u'e', object(): u'e'} - d2 = stringify_dict(d, keys_only=False) - self.assertEqual(d, d2) - self.assertIsNot(d, d2) # shouldn't modify in place - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.keys())) - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.values())) - - @unittest.skipUnless(six.PY2, "deprecated function") - def test_stringify_dict_tuples(self): - tuples = [('a', 123), (u'b', 'c'), (u'd', u'e'), (object(), u'e')] - d = dict(tuples) - d2 = stringify_dict(tuples, keys_only=False) - self.assertEqual(d, d2) - self.assertIsNot(d, d2) # shouldn't modify in place - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.keys()), d2.keys()) - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.values())) - - @unittest.skipUnless(six.PY2, "deprecated function") - def test_stringify_dict_keys_only(self): - d = {'a': 123, u'b': 'c', u'd': u'e', object(): u'e'} - d2 = stringify_dict(d) - self.assertEqual(d, d2) - self.assertIsNot(d, d2) # shouldn't modify in place - self.assertFalse(any(isinstance(x, six.text_type) for x in d2.keys())) - def test_get_func_args(self): def f1(a, b, c): pass diff --git a/tests/test_webclient.py b/tests/test_webclient.py index a81946490..7b015ff8d 100644 --- a/tests/test_webclient.py +++ b/tests/test_webclient.py @@ -78,26 +78,6 @@ class ParseUrlTestCase(unittest.TestCase): to_bytes(x) if not isinstance(x, int) else x for x in test) self.assertEqual(client._parse(url), test, url) - def test_externalUnicodeInterference(self): - """ - L{client._parse} should return C{str} for the scheme, host, and path - elements of its return tuple, even when passed an URL which has - previously been passed to L{urlparse} as a C{unicode} string. - """ - if not six.PY2: - raise unittest.SkipTest( - "Applies only to Py2, as urls can be ONLY unicode on Py3") - badInput = u'http://example.com/path' - goodInput = badInput.encode('ascii') - self._parse(badInput) # cache badInput in urlparse_cached - scheme, netloc, host, port, path = self._parse(goodInput) - self.assertTrue(isinstance(scheme, str)) - self.assertTrue(isinstance(netloc, str)) - self.assertTrue(isinstance(host, str)) - self.assertTrue(isinstance(path, str)) - self.assertTrue(isinstance(port, int)) - - class ScrapyHTTPPageGetterTests(unittest.TestCase): From f066257e95f1d025d1e581350902067c671fd800 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 3 Sep 2019 15:17:03 +0500 Subject: [PATCH 270/387] Restore tests/test_proxy_connect.py and update it to modern mitmproxy. --- tests/requirements-py3.txt | 1 + tests/test_proxy_connect.py | 134 ++++++++++++++++++++++++++++++++++++ 2 files changed, 135 insertions(+) create mode 100644 tests/test_proxy_connect.py diff --git a/tests/requirements-py3.txt b/tests/requirements-py3.txt index 2e8d319d2..8169febea 100644 --- a/tests/requirements-py3.txt +++ b/tests/requirements-py3.txt @@ -1,5 +1,6 @@ # Tests requirements jmespath +mitmproxy pytest pytest-cov pytest-twisted diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py new file mode 100644 index 000000000..8142d9a41 --- /dev/null +++ b/tests/test_proxy_connect.py @@ -0,0 +1,134 @@ +import json +import os +import re +from subprocess import Popen, PIPE +import sys +import time + +from six.moves.urllib.parse import urlsplit, urlunsplit +from testfixtures import LogCapture + +from twisted.internet import defer +from twisted.trial.unittest import TestCase + +from scrapy.utils.test import get_crawler +from scrapy.http import Request +from tests.spiders import SimpleSpider, SingleRequestSpider +from tests.mockserver import MockServer + + +class MitmProxy: + auth_user = 'scrapy' + auth_pass = 'scrapy' + + def start(self): + from scrapy.utils.test import get_testenv + script = """ +import sys +from mitmproxy.tools.main import mitmdump +sys.argv[0] = "mitmdump" +sys.exit(mitmdump()) + """ + cert_path = os.path.join(os.path.abspath(os.path.dirname(__file__)), + 'keys', 'mitmproxy-ca.pem') + self.proc = Popen([sys.executable, + '-c', script, + '--listen-host', '127.0.0.1', + '--listen-port', '0', + '--proxyauth', '%s:%s' % (self.auth_user, self.auth_pass), + '--certs', cert_path, + '--ssl-insecure', + ], + stdout=PIPE, env=get_testenv()) + line = self.proc.stdout.readline().decode('utf-8') + host_port = re.search(r'listening at http://([^:]+:\d+)', line).group(1) + address = 'http://%s:%s@%s' % (self.auth_user, self.auth_pass, host_port) + return address + + def stop(self): + self.proc.kill() + self.proc.wait() + time.sleep(0.2) + + +def _wrong_credentials(proxy_url): + bad_auth_proxy = list(urlsplit(proxy_url)) + bad_auth_proxy[1] = bad_auth_proxy[1].replace('scrapy:scrapy@', 'wrong:wronger@') + return urlunsplit(bad_auth_proxy) + + +class ProxyConnectTestCase(TestCase): + + def setUp(self): + self.mockserver = MockServer() + self.mockserver.__enter__() + self._oldenv = os.environ.copy() + + self._proxy = MitmProxy() + proxy_url = self._proxy.start() + os.environ['https_proxy'] = proxy_url + os.environ['http_proxy'] = proxy_url + + def tearDown(self): + self.mockserver.__exit__(None, None, None) + self._proxy.stop() + os.environ = self._oldenv + + @defer.inlineCallbacks + def test_https_connect_tunnel(self): + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) + self._assert_got_response_code(200, l) + + @defer.inlineCallbacks + def test_https_noconnect(self): + proxy = os.environ['https_proxy'] + os.environ['https_proxy'] = proxy + '?noconnect' + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) + self._assert_got_response_code(200, l) + + @defer.inlineCallbacks + def test_https_connect_tunnel_error(self): + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl("https://localhost:99999/status?n=200") + self._assert_got_tunnel_error(l) + + @defer.inlineCallbacks + def test_https_tunnel_auth_error(self): + os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) + # The proxy returns a 407 error code but it does not reach the client; + # he just sees a TunnelError. + self._assert_got_tunnel_error(l) + + @defer.inlineCallbacks + def test_https_tunnel_without_leak_proxy_authorization_header(self): + request = Request(self.mockserver.url("/echo", is_secure=True)) + crawler = get_crawler(SingleRequestSpider) + with LogCapture() as l: + yield crawler.crawl(seed=request) + self._assert_got_response_code(200, l) + echo = json.loads(crawler.spider.meta['responses'][0].body) + self.assertTrue('Proxy-Authorization' not in echo['headers']) + + @defer.inlineCallbacks + def test_https_noconnect_auth_error(self): + os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) + '?noconnect' + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) + self._assert_got_response_code(407, l) + + def _assert_got_response_code(self, code, log): + print(log) + self.assertEqual(str(log).count('Crawled (%d)' % code), 1) + + def _assert_got_tunnel_error(self, log): + print(log) + self.assertIn('TunnelError', str(log)) From cbb6d0c6a71c709903048597cbbdb02616aed285 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 3 Sep 2019 15:23:24 +0500 Subject: [PATCH 271/387] Mark failing proxy tests. --- tests/test_proxy_connect.py | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index 8142d9a41..5e9470e39 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -5,6 +5,7 @@ from subprocess import Popen, PIPE import sys import time +import pytest from six.moves.urllib.parse import urlsplit, urlunsplit from testfixtures import LogCapture @@ -81,6 +82,7 @@ class ProxyConnectTestCase(TestCase): yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) self._assert_got_response_code(200, l) + @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') @defer.inlineCallbacks def test_https_noconnect(self): proxy = os.environ['https_proxy'] @@ -90,6 +92,7 @@ class ProxyConnectTestCase(TestCase): yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) self._assert_got_response_code(200, l) + @pytest.mark.xfail(reason='Python 3 fails this earlier') @defer.inlineCallbacks def test_https_connect_tunnel_error(self): crawler = get_crawler(SimpleSpider) @@ -117,6 +120,7 @@ class ProxyConnectTestCase(TestCase): echo = json.loads(crawler.spider.meta['responses'][0].body) self.assertTrue('Proxy-Authorization' not in echo['headers']) + @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') @defer.inlineCallbacks def test_https_noconnect_auth_error(self): os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) + '?noconnect' From c327ad9ba6baee934b35ab9cb832e2fbb3311e13 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 31 Oct 2019 15:20:28 +0500 Subject: [PATCH 272/387] Remove an unused six import. --- tests/test_downloader_handlers.py | 1 - 1 file changed, 1 deletion(-) diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index b06fcf6c3..a78762a64 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -1,5 +1,4 @@ import os -import six import shutil import tempfile from unittest import mock From 3ec6960732591b38a21bc0af9ae042007bb849f4 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 31 Oct 2019 23:21:14 +0500 Subject: [PATCH 273/387] Fix test_proxy_connect.py for py3.5. --- tests/test_proxy_connect.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index 5e9470e39..f6381b5b1 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -117,7 +117,7 @@ class ProxyConnectTestCase(TestCase): with LogCapture() as l: yield crawler.crawl(seed=request) self._assert_got_response_code(200, l) - echo = json.loads(crawler.spider.meta['responses'][0].body) + echo = json.loads(crawler.spider.meta['responses'][0].body.decode('utf-8')) self.assertTrue('Proxy-Authorization' not in echo['headers']) @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') From c4ef950efda32d292840027e72aa5bb7844423bb Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 31 Oct 2019 23:21:30 +0500 Subject: [PATCH 274/387] Use an older mitmproxy for py3.5. --- tests/requirements-py3.txt | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/tests/requirements-py3.txt b/tests/requirements-py3.txt index 8169febea..c4bc1f278 100644 --- a/tests/requirements-py3.txt +++ b/tests/requirements-py3.txt @@ -1,6 +1,7 @@ # Tests requirements jmespath -mitmproxy +mitmproxy; python_version >= '3.6' +mitmproxy==3.0.4; python_version < '3.6' pytest pytest-cov pytest-twisted From 5080180c759e74df9e20f7770534017272a9fea3 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 1 Nov 2019 19:46:19 +0500 Subject: [PATCH 275/387] Improve the test_https_tunnel_without_leak_proxy_authorization_header change. --- tests/test_proxy_connect.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index f6381b5b1..651576c2c 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -117,7 +117,7 @@ class ProxyConnectTestCase(TestCase): with LogCapture() as l: yield crawler.crawl(seed=request) self._assert_got_response_code(200, l) - echo = json.loads(crawler.spider.meta['responses'][0].body.decode('utf-8')) + echo = json.loads(crawler.spider.meta['responses'][0].text) self.assertTrue('Proxy-Authorization' not in echo['headers']) @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') From 5970d00eb9ade62523b0f61788cab50fe8f62eb1 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 1 Nov 2019 19:46:38 +0500 Subject: [PATCH 276/387] Only xfail test_https_connect_tunnel_error on 3.6+. --- tests/test_proxy_connect.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index 651576c2c..ec3f0716c 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -92,7 +92,7 @@ class ProxyConnectTestCase(TestCase): yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) self._assert_got_response_code(200, l) - @pytest.mark.xfail(reason='Python 3 fails this earlier') + @pytest.mark.xfail(reason='Python 3.6+ fails this earlier', condition=sys.version_info.minor >= 6) @defer.inlineCallbacks def test_https_connect_tunnel_error(self): crawler = get_crawler(SimpleSpider) From 8b730a36706f9b75e1d1b168901dad9805f4da78 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 1 Nov 2019 19:50:56 +0500 Subject: [PATCH 277/387] Use self.proc.communicate() after killing mitmdump. --- tests/test_proxy_connect.py | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index ec3f0716c..69925f80c 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -48,8 +48,7 @@ sys.exit(mitmdump()) def stop(self): self.proc.kill() - self.proc.wait() - time.sleep(0.2) + self.proc.communicate() def _wrong_credentials(proxy_url): From a7b640991d527c265d581cba1b1d1119db99e4c8 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 1 Nov 2019 19:52:57 +0500 Subject: [PATCH 278/387] Rename tests/py3-ignores.txt to tests/ignores.txt. --- conftest.py | 2 +- tests/{py3-ignores.txt => ignores.txt} | 0 2 files changed, 1 insertion(+), 1 deletion(-) rename tests/{py3-ignores.txt => ignores.txt} (100%) diff --git a/conftest.py b/conftest.py index 74fb101e9..d54ce155c 100644 --- a/conftest.py +++ b/conftest.py @@ -7,7 +7,7 @@ collect_ignore = [ ] -for line in open('tests/py3-ignores.txt'): +for line in open('tests/ignores.txt'): file_path = line.strip() if file_path and file_path[0] != '#': collect_ignore.append(file_path) diff --git a/tests/py3-ignores.txt b/tests/ignores.txt similarity index 100% rename from tests/py3-ignores.txt rename to tests/ignores.txt From 922a66cf07e68a9a42d59530ce9dc2dd29a093a8 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 14 Nov 2019 22:36:58 +0500 Subject: [PATCH 279/387] Fix or ignore flake8 problems. --- pytest.ini | 1 + tests/test_proxy_connect.py | 5 ++--- tests/test_utils_python.py | 2 +- 3 files changed, 4 insertions(+), 4 deletions(-) diff --git a/pytest.ini b/pytest.ini index 529ad5d27..4f2548db2 100644 --- a/pytest.ini +++ b/pytest.ini @@ -226,6 +226,7 @@ flake8-ignore = tests/test_pipeline_files.py F401 E501 W293 E303 E272 E226 tests/test_pipeline_images.py F401 F841 E501 E303 tests/test_pipeline_media.py E501 E741 E731 E128 E261 E306 E502 + tests/test_proxy_connect.py E501 E741 tests/test_request_cb_kwargs.py E501 tests/test_responsetypes.py E501 E302 E305 tests/test_robotstxt_interface.py F401 E302 E501 W291 E501 diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index 69925f80c..2435999f9 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -3,7 +3,6 @@ import os import re from subprocess import Popen, PIPE import sys -import time import pytest from six.moves.urllib.parse import urlsplit, urlunsplit @@ -31,7 +30,7 @@ sys.argv[0] = "mitmdump" sys.exit(mitmdump()) """ cert_path = os.path.join(os.path.abspath(os.path.dirname(__file__)), - 'keys', 'mitmproxy-ca.pem') + 'keys', 'mitmproxy-ca.pem') self.proc = Popen([sys.executable, '-c', script, '--listen-host', '127.0.0.1', @@ -40,7 +39,7 @@ sys.exit(mitmdump()) '--certs', cert_path, '--ssl-insecure', ], - stdout=PIPE, env=get_testenv()) + stdout=PIPE, env=get_testenv()) line = self.proc.stdout.readline().decode('utf-8') host_port = re.search(r'listening at http://([^:]+:\d+)', line).group(1) address = 'http://%s:%s@%s' % (self.auth_user, self.auth_pass, host_port) diff --git a/tests/test_utils_python.py b/tests/test_utils_python.py index 6857356f6..326a67c2e 100644 --- a/tests/test_utils_python.py +++ b/tests/test_utils_python.py @@ -8,7 +8,7 @@ import six from scrapy.utils.python import ( memoizemethod_noargs, binary_is_text, equal_attributes, - WeakKeyCache, stringify_dict, get_func_args, to_bytes, to_unicode, + WeakKeyCache, get_func_args, to_bytes, to_unicode, without_none_values, MutableChain) __doctests__ = ['scrapy.utils.python'] From beb7d80d6a8a82baf9cb8170758704b4ff5c63cf Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 14 Nov 2019 22:47:35 +0500 Subject: [PATCH 280/387] Add a comment about the noconnect tests. --- tests/test_proxy_connect.py | 27 +++++++++++++++++---------- 1 file changed, 17 insertions(+), 10 deletions(-) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index 2435999f9..277455751 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -80,16 +80,6 @@ class ProxyConnectTestCase(TestCase): yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) self._assert_got_response_code(200, l) - @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') - @defer.inlineCallbacks - def test_https_noconnect(self): - proxy = os.environ['https_proxy'] - os.environ['https_proxy'] = proxy + '?noconnect' - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(200, l) - @pytest.mark.xfail(reason='Python 3.6+ fails this earlier', condition=sys.version_info.minor >= 6) @defer.inlineCallbacks def test_https_connect_tunnel_error(self): @@ -118,6 +108,23 @@ class ProxyConnectTestCase(TestCase): echo = json.loads(crawler.spider.meta['responses'][0].text) self.assertTrue('Proxy-Authorization' not in echo['headers']) + # The noconnect mode isn't supported by the current mitmproxy, it returns + # "Invalid request scheme: https" as it doesn't seem to support full URLs in GET at all, + # and it's not clear what behavior is intended by Scrapy and by mitmproxy here. + # https://github.com/mitmproxy/mitmproxy/issues/848 may be related. + # The Scrapy noconnect mode was required, at least in the past, to work with Crawlera, + # and https://github.com/scrapy-plugins/scrapy-crawlera/pull/44 seems to be related. + + @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') + @defer.inlineCallbacks + def test_https_noconnect(self): + proxy = os.environ['https_proxy'] + os.environ['https_proxy'] = proxy + '?noconnect' + crawler = get_crawler(SimpleSpider) + with LogCapture() as l: + yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) + self._assert_got_response_code(200, l) + @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') @defer.inlineCallbacks def test_https_noconnect_auth_error(self): From 78ad01632f16a97fb180d8c2972075b19f471380 Mon Sep 17 00:00:00 2001 From: Andrey Rahmatullin <wrar@wrar.name> Date: Tue, 19 Nov 2019 14:43:30 +0500 Subject: [PATCH 281/387] Fix flake8 problems in PR #3989 (#4176) --- tests/test_logformatter.py | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/tests/test_logformatter.py b/tests/test_logformatter.py index 5b5d68f4f..afbd25d0c 100644 --- a/tests/test_logformatter.py +++ b/tests/test_logformatter.py @@ -62,7 +62,7 @@ class LogFormatterTestCase(unittest.TestCase): lines = logline.splitlines() assert all(isinstance(x, six.text_type) for x in lines) self.assertEqual(lines, [u"Dropped: \u2018", '{}']) - + def test_error(self): # In practice, the complete traceback is shown by passing the # 'exc_info' argument to the logging function @@ -121,7 +121,7 @@ class LogformatterSubclassTest(LogFormatterTestCase): "Crawled (200) <GET http://www.example.com> (referer: http://example.com) ['cached']") def test_flags_in_request(self): - req = Request("http://www.example.com", flags=['test','flag']) + req = Request("http://www.example.com", flags=['test', 'flag']) res = Response("http://www.example.com") logkws = self.formatter.crawled(req, res, self.spider) logline = logkws['msg'] % logkws['args'] From fc3af54dbd5b7fdcbc44c82ba4ced528a104cd88 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Wed, 20 Nov 2019 07:57:09 +0100 Subject: [PATCH 282/387] Make tox configuration more user friendly --- .travis.yml | 5 ++-- docs/contributing.rst | 14 ++-------- tox.ini | 64 ++++++++++++++++++------------------------- 3 files changed, 31 insertions(+), 52 deletions(-) diff --git a/.travis.yml b/.travis.yml index 9f477e860..4e28d6f11 100644 --- a/.travis.yml +++ b/.travis.yml @@ -12,10 +12,9 @@ matrix: - env: TOXENV=flake8 python: 3.8 - env: TOXENV=pypy3 - python: 3.5 - env: TOXENV=py35 python: 3.5 - - env: TOXENV=py35-pinned + - env: TOXENV=pinned python: 3.5 - env: TOXENV=py36 python: 3.6 @@ -23,7 +22,7 @@ matrix: python: 3.7 - env: TOXENV=py38 python: 3.8 - - env: TOXENV=py38-extra-deps + - env: TOXENV=extra-deps python: 3.8 - env: TOXENV=docs python: 3.6 diff --git a/docs/contributing.rst b/docs/contributing.rst index f084bd23d..c4cb605ab 100644 --- a/docs/contributing.rst +++ b/docs/contributing.rst @@ -202,17 +202,9 @@ tests requires `tox`_. Running tests ------------- -Make sure you have a recent enough `tox`_ installation: +To run all tests:: - ``tox --version`` - -If your version is older than 1.7.0, please update it first: - - ``pip install -U tox`` - -To run all tests go to the root directory of Scrapy source code and run: - - ``tox`` + tox To run a specific test (say ``tests/test_loader.py``) use: @@ -227,7 +219,7 @@ environment name from ``tox.ini``. For example, to run the tests with Python You can also specify a comma-separated list of environmets, and use `tox’s parallel mode`_ to run the tests on multiple environments in parallel:: - tox -e py27,py36 -p auto + tox -e py37,py38 -p auto To pass command-line options to pytest_, add them after ``--`` in your call to tox_. Using ``--`` overrides the default positional arguments defined in diff --git a/tox.ini b/tox.ini index fd75d18e2..0b41f6cc3 100644 --- a/tox.ini +++ b/tox.ini @@ -4,7 +4,8 @@ # and then run "tox" from this directory. [tox] -envlist = py35 +envlist = security,flake8,py3 +minversion = 1.7.0 [testenv] deps = @@ -23,11 +24,28 @@ passenv = commands = py.test --cov=scrapy --cov-report= {posargs:--durations=10 docs scrapy tests} -[testenv:py35] -basepython = python3.5 +[testenv:security] +basepython = python3 +deps = + bandit +commands = + bandit -r -c .bandit.yml {posargs:scrapy} -[testenv:py35-pinned] -basepython = python3.5 +[testenv:flake8] +basepython = python3 +deps = + {[testenv]deps} + pytest-flake8 +commands = + py.test --flake8 {posargs:docs scrapy tests} + +[testenv:pypy3] +basepython = pypy3 +commands = + py.test {posargs:--durations=10 docs scrapy tests} + +[testenv:pinned] +basepython = python3 deps = -ctests/constraints.txt cryptography==2.0 @@ -48,34 +66,11 @@ deps = botocore==1.3.23 Pillow==3.4.2 -[testenv:py36] -basepython = python3.6 - -[testenv:py37] -basepython = python3.7 - -[testenv:py38] -basepython = python3.8 - -[testenv:pypy3] -basepython = pypy3 -commands = - py.test {posargs:--durations=10 docs scrapy tests} - -[testenv:security] -basepython = python3.8 -deps = - bandit -commands = - bandit -r -c .bandit.yml {posargs:scrapy} - -[testenv:flake8] -basepython = python3.8 +[testenv:extra-deps] deps = {[testenv]deps} - pytest-flake8 -commands = - py.test --flake8 {posargs:docs scrapy tests} + reppy + robotexclusionrulesparser [docs] changedir = docs @@ -99,10 +94,3 @@ changedir = {[docs]changedir} deps = {[docs]deps} commands = sphinx-build -W -b linkcheck . {envtmpdir}/linkcheck - -[testenv:py38-extra-deps] -basepython = python3.8 -deps = - {[testenv]deps} - reppy - robotexclusionrulesparser From e6c5292a7c4391642897ee31ea2ded889b3b88dc Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Wed, 20 Nov 2019 09:29:55 -0300 Subject: [PATCH 283/387] Response.follow_all: Specific exception for invalid selectors --- scrapy/http/response/text.py | 31 ++++++++++++++++++------------- 1 file changed, 18 insertions(+), 13 deletions(-) diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 74017b5aa..5110b4bd4 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -5,17 +5,18 @@ discovering (through HTTP headers) to base Response class. See documentation in docs/topics/request-response.rst """ -import six -from six.moves.urllib.parse import urljoin +from contextlib import suppress import parsel -from w3lib.encoding import html_to_unicode, resolve_encoding, \ - html_body_declared_encoding, http_content_type_encoding +import six +from six.moves.urllib.parse import urljoin +from w3lib.encoding import (html_body_declared_encoding, html_to_unicode, + http_content_type_encoding, resolve_encoding) from w3lib.html import strip_html5_whitespace from scrapy.http.response import Response -from scrapy.utils.response import get_base_url from scrapy.utils.python import memoizemethod_noargs, to_native_str +from scrapy.utils.response import get_base_url class TextResponse(Response): @@ -197,10 +198,8 @@ class TextResponse(Response): selector_list = self.xpath(xpath) urls = [] for selector in selector_list: - try: + with suppress(_InvalidSelector): urls.append(_url_from_selector(selector)) - except ValueError: - pass return super(TextResponse, self).follow_all( urls=urls, callback=callback, @@ -217,18 +216,24 @@ class TextResponse(Response): ) +class _InvalidSelector(ValueError): + """ + Raised when a URL cannot be obtained from a Selector + """ + + def _url_from_selector(sel): # type: (parsel.Selector) -> str if isinstance(sel.root, six.string_types): # e.g. ::attr(href) result return strip_html5_whitespace(sel.root) if not hasattr(sel.root, 'tag'): - raise ValueError("Unsupported selector: %s" % sel) + raise _InvalidSelector("Unsupported selector: %s" % sel) if sel.root.tag not in ('a', 'link'): - raise ValueError("Only <a> and <link> elements are supported; got <%s>" % - sel.root.tag) + raise _InvalidSelector("Only <a> and <link> elements are supported; got <%s>" % + sel.root.tag) href = sel.root.get('href') if href is None: - raise ValueError("<%s> element has no href attribute: %s" % - (sel.root.tag, sel)) + raise _InvalidSelector("<%s> element has no href attribute: %s" % + (sel.root.tag, sel)) return strip_html5_whitespace(href) From b602c61e1cc00e04ce435078644d5ea0c0249903 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Wed, 20 Nov 2019 09:38:54 -0300 Subject: [PATCH 284/387] [Test] Rename outdated sample files --- .../{sgml_linkextractor.html => linkextractor.html} | 0 ..._linkextractor_no_href.html => linkextractor_no_href.html} | 0 tests/test_http_response.py | 4 ++-- 3 files changed, 2 insertions(+), 2 deletions(-) rename tests/sample_data/link_extractor/{sgml_linkextractor.html => linkextractor.html} (100%) rename tests/sample_data/link_extractor/{sgml_linkextractor_no_href.html => linkextractor_no_href.html} (100%) diff --git a/tests/sample_data/link_extractor/sgml_linkextractor.html b/tests/sample_data/link_extractor/linkextractor.html similarity index 100% rename from tests/sample_data/link_extractor/sgml_linkextractor.html rename to tests/sample_data/link_extractor/linkextractor.html diff --git a/tests/sample_data/link_extractor/sgml_linkextractor_no_href.html b/tests/sample_data/link_extractor/linkextractor_no_href.html similarity index 100% rename from tests/sample_data/link_extractor/sgml_linkextractor_no_href.html rename to tests/sample_data/link_extractor/linkextractor_no_href.html diff --git a/tests/test_http_response.py b/tests/test_http_response.py index 6d3c5cb9d..0ae1612b5 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -250,12 +250,12 @@ class BaseResponseTest(unittest.TestCase): yield req def _links_response(self): - body = get_testdata('link_extractor', 'sgml_linkextractor.html') + body = get_testdata('link_extractor', 'linkextractor.html') resp = self.response_class('http://example.com/index', body=body) return resp def _links_response_no_href(self): - body = get_testdata('link_extractor', 'sgml_linkextractor_no_href.html') + body = get_testdata('link_extractor', 'linkextractor_no_href.html') resp = self.response_class('http://example.com/index', body=body) return resp From 6f4e84ecf95feabe15fda0819bc428f8e0d4d340 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Wed, 20 Nov 2019 09:55:15 -0300 Subject: [PATCH 285/387] PEP8 adjustments for scrapy.http.response module --- scrapy/http/response/__init__.py | 6 ++++-- scrapy/http/response/html.py | 1 + scrapy/http/response/text.py | 2 ++ scrapy/http/response/xml.py | 1 + 4 files changed, 8 insertions(+), 2 deletions(-) diff --git a/scrapy/http/response/__init__.py b/scrapy/http/response/__init__.py index 79a8d0ca0..b9e638551 100644 --- a/scrapy/http/response/__init__.py +++ b/scrapy/http/response/__init__.py @@ -4,6 +4,8 @@ responses in Scrapy. See documentation in docs/topics/request-response.rst """ +from typing import Generator + from six.moves.urllib.parse import urljoin from scrapy.http.request import Request @@ -41,8 +43,8 @@ class Response(object_ref): if isinstance(url, str): self._url = url else: - raise TypeError('%s url must be str, got %s:' % (type(self).__name__, - type(url).__name__)) + raise TypeError('%s url must be str, got %s:' % + (type(self).__name__, type(url).__name__)) url = property(_get_url, obsolete_setter(_set_url, 'url')) diff --git a/scrapy/http/response/html.py b/scrapy/http/response/html.py index bd3559fbb..7eed052c2 100644 --- a/scrapy/http/response/html.py +++ b/scrapy/http/response/html.py @@ -7,5 +7,6 @@ See documentation in docs/topics/request-response.rst from scrapy.http.response.text import TextResponse + class HtmlResponse(TextResponse): pass diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index b97420345..6acf1026f 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -6,6 +6,7 @@ See documentation in docs/topics/request-response.rst """ from contextlib import suppress +from typing import Generator import parsel import six @@ -14,6 +15,7 @@ from w3lib.encoding import (html_body_declared_encoding, html_to_unicode, http_content_type_encoding, resolve_encoding) from w3lib.html import strip_html5_whitespace +from scrapy.http import Request from scrapy.http.response import Response from scrapy.utils.python import memoizemethod_noargs, to_unicode from scrapy.utils.response import get_base_url diff --git a/scrapy/http/response/xml.py b/scrapy/http/response/xml.py index 1df33fee5..abf474a2f 100644 --- a/scrapy/http/response/xml.py +++ b/scrapy/http/response/xml.py @@ -7,5 +7,6 @@ See documentation in docs/topics/request-response.rst from scrapy.http.response.text import TextResponse + class XmlResponse(TextResponse): pass From 6781d2f5b261305a4f6aa41a2b485a2e5cc7bc76 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Wed, 20 Nov 2019 09:58:25 -0300 Subject: [PATCH 286/387] Update sample file references --- tests/test_linkextractors.py | 2 +- tests/test_linkextractors_deprecated.py | 4 ++-- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/tests/test_linkextractors.py b/tests/test_linkextractors.py index 57ef1694a..0b94f937f 100644 --- a/tests/test_linkextractors.py +++ b/tests/test_linkextractors.py @@ -16,7 +16,7 @@ class Base: escapes_whitespace = False def setUp(self): - body = get_testdata('link_extractor', 'sgml_linkextractor.html') + body = get_testdata('link_extractor', 'linkextractor.html') self.response = HtmlResponse(url='http://example.com/index', body=body) def test_urls_type(self): diff --git a/tests/test_linkextractors_deprecated.py b/tests/test_linkextractors_deprecated.py index 1366971be..388ed6ad4 100644 --- a/tests/test_linkextractors_deprecated.py +++ b/tests/test_linkextractors_deprecated.py @@ -111,7 +111,7 @@ class BaseSgmlLinkExtractorTestCase(unittest.TestCase): class HtmlParserLinkExtractorTestCase(unittest.TestCase): def setUp(self): - body = get_testdata('link_extractor', 'sgml_linkextractor.html') + body = get_testdata('link_extractor', 'linkextractor.html') self.response = HtmlResponse(url='http://example.com/index', body=body) def test_extraction(self): @@ -183,7 +183,7 @@ class RegexLinkExtractorTestCase(unittest.TestCase): # than it should be. def setUp(self): - body = get_testdata('link_extractor', 'sgml_linkextractor.html') + body = get_testdata('link_extractor', 'linkextractor.html') self.response = HtmlResponse(url='http://example.com/index', body=body) def test_extraction(self): From 4f80eff1e159ceb2e5f9375f6a95f69f75f6e1c2 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 21 Nov 2019 10:30:21 +0100 Subject: [PATCH 287/387] Enable sphinx-hoverxref for all references --- docs/conf.py | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/conf.py b/docs/conf.py index 0e0df274c..04472cf86 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -279,3 +279,9 @@ intersphinx_mapping = { 'sphinx': ('https://www.sphinx-doc.org/en/master', None), 'twisted': ('https://twistedmatrix.com/documents/current', None), } + + +# Options for sphinx-hoverxref options +# ------------------------------------ + +hoverxref_auto_ref = True From f251dda2687ad0cd21c126580c4ddfff59444d93 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 21 Nov 2019 11:59:10 +0100 Subject: [PATCH 288/387] Make debug doctests pass --- docs/topics/debug.rst | 8 ++++++++ pytest.ini | 1 - 2 files changed, 8 insertions(+), 1 deletion(-) diff --git a/docs/topics/debug.rst b/docs/topics/debug.rst index 0aaad0c77..4b2588518 100644 --- a/docs/topics/debug.rst +++ b/docs/topics/debug.rst @@ -48,6 +48,10 @@ The most basic way of checking the output of your spider is to use the of the spider at the method level. It has the advantage of being flexible and simple to use, but does not allow debugging code inside a method. +.. highlight:: none + +.. skip: start + In order to see the item scraped from a specific url:: $ scrapy parse --spider=myspider -c parse_item -d 2 <item_url> @@ -85,6 +89,8 @@ using:: $ scrapy parse --spider=myspider -d 3 'http://example.com/page1' +.. skip: end + Scrapy Shell ============ @@ -94,6 +100,8 @@ spider, it is of little help to check what happens inside a callback, besides showing the response received and the output. How to debug the situation when ``parse_details`` sometimes receives no item? +.. highlight:: python + Fortunately, the :command:`shell` is your bread and butter in this case (see :ref:`topics-shell-inspect-response`):: diff --git a/pytest.ini b/pytest.ini index 7be5d8572..9d99dcd22 100644 --- a/pytest.ini +++ b/pytest.ini @@ -8,7 +8,6 @@ addopts = --ignore=docs/_ext --ignore=docs/conf.py --ignore=docs/news.rst - --ignore=docs/topics/debug.rst --ignore=docs/topics/developer-tools.rst --ignore=docs/topics/dynamic-content.rst --ignore=docs/topics/items.rst From fcfcabf1bdf6fc48105c6f58a4c4eb59e26a8f6e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 21 Nov 2019 12:15:13 +0100 Subject: [PATCH 289/387] Use InterSphinx for links to the pytest and tox documentation --- docs/conf.py | 2 ++ docs/contributing.rst | 30 ++++++++++++++---------------- 2 files changed, 16 insertions(+), 16 deletions(-) diff --git a/docs/conf.py b/docs/conf.py index 0e0df274c..e37df7e47 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -275,7 +275,9 @@ coverage_ignore_pyobjects = [ # ------------------------------------- intersphinx_mapping = { + 'pytest': ('https://docs.pytest.org/en/latest', None), 'python': ('https://docs.python.org/3', None), 'sphinx': ('https://www.sphinx-doc.org/en/master', None), + 'tox': ('https://tox.readthedocs.io/en/latest', None), 'twisted': ('https://twistedmatrix.com/documents/current', None), } diff --git a/docs/contributing.rst b/docs/contributing.rst index 68ae2bf3c..81bb50a77 100644 --- a/docs/contributing.rst +++ b/docs/contributing.rst @@ -196,14 +196,14 @@ Tests Tests are implemented using the :doc:`Twisted unit-testing framework <twisted:core/development/policy/test-standard>`. Running tests requires -`tox`_. +:doc:`tox <tox:index>`. .. _running-tests: Running tests ------------- -Make sure you have a recent enough `tox`_ installation: +Make sure you have a recent enough :doc:`tox <tox:index>` installation: ``tox --version`` @@ -219,26 +219,27 @@ To run a specific test (say ``tests/test_loader.py``) use: ``tox -- tests/test_loader.py`` -To run the tests on a specific tox_ environment, use ``-e <name>`` with an -environment name from ``tox.ini``. For example, to run the tests with Python -3.6 use:: +To run the tests on a specific :doc:`tox <tox:index>` environment, use +``-e <name>`` with an environment name from ``tox.ini``. For example, to run +the tests with Python 3.6 use:: tox -e py36 -You can also specify a comma-separated list of environmets, and use `tox’s -parallel mode`_ to run the tests on multiple environments in parallel:: +You can also specify a comma-separated list of environmets, and use :ref:`tox’s +parallel mode <tox:parallel_mode>` to run the tests on multiple environments in +parallel:: tox -e py27,py36 -p auto -To pass command-line options to pytest_, add them after ``--`` in your call to -tox_. Using ``--`` overrides the default positional arguments defined in -``tox.ini``, so you must include those default positional arguments -(``scrapy tests``) after ``--`` as well:: +To pass command-line options to :doc:`pytest <pytest:index>`, add them after +``--`` in your call to :doc:`tox <tox:index>`. Using ``--`` overrides the +default positional arguments defined in ``tox.ini``, so you must include those +default positional arguments (``scrapy tests``) after ``--`` as well:: tox -- scrapy tests -x # stop after first failure You can also use the `pytest-xdist`_ plugin. For example, to run all tests on -the Python 3.6 tox_ environment using all your CPU cores:: +the Python 3.6 :doc:`tox <tox:index>` environment using all your CPU cores:: tox -e py36 -- scrapy tests -n auto @@ -275,7 +276,4 @@ And their unit-tests are in:: .. _open issues: https://github.com/scrapy/scrapy/issues .. _PEP 257: https://www.python.org/dev/peps/pep-0257/ .. _pull request: https://help.github.com/en/articles/creating-a-pull-request -.. _pytest: https://docs.pytest.org/en/latest/usage.html -.. _pytest-xdist: https://docs.pytest.org/en/3.0.0/xdist.html -.. _tox: https://pypi.python.org/pypi/tox -.. _tox’s parallel mode: https://tox.readthedocs.io/en/latest/example/basic.html#parallel-mode +.. _pytest-xdist: https://github.com/pytest-dev/pytest-xdist From a2bf340bab796704ba6846f8ed755d3ffe37bb0d Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 21 Nov 2019 14:18:49 +0100 Subject: [PATCH 290/387] Remove unused imports --- pytest.ini | 84 +++++++++---------- scrapy/commands/check.py | 2 - scrapy/commands/view.py | 2 +- scrapy/core/downloader/__init__.py | 2 - scrapy/core/downloader/handlers/http.py | 1 - scrapy/core/downloader/handlers/s3.py | 2 +- scrapy/extensions/httpcache.py | 4 +- scrapy/linkextractors/__init__.py | 2 +- scrapy/selector/__init__.py | 2 +- scrapy/selector/unified.py | 2 - scrapy/spiders/__init__.py | 7 +- scrapy/utils/boto.py | 2 +- scrapy/utils/console.py | 2 +- scrapy/utils/curl.py | 2 +- scrapy/utils/engine.py | 4 +- scrapy/utils/http.py | 2 +- scrapy/utils/markup.py | 2 +- scrapy/utils/multipart.py | 2 +- scrapy/utils/url.py | 1 - tests/mockserver.py | 8 +- tests/test_downloader_handlers.py | 10 ++- tests/test_downloadermiddleware_httpcache.py | 3 - ...st_downloadermiddleware_httpcompression.py | 2 +- tests/test_downloadermiddleware_httpproxy.py | 5 +- tests/test_extension_telnet.py | 2 +- tests/test_feedexport.py | 6 +- tests/test_http_request.py | 4 +- tests/test_pipeline_images.py | 1 - tests/test_robotstxt_interface.py | 6 +- tests/test_selector.py | 4 +- tests/test_spider.py | 1 - tests/test_spidermiddleware_output_chain.py | 1 - tests/test_spidermiddleware_referer.py | 1 - tests/test_utils_reqser.py | 4 - tests/test_utils_url.py | 6 +- tox.ini | 1 + 36 files changed, 84 insertions(+), 108 deletions(-) diff --git a/pytest.ini b/pytest.ini index 7be5d8572..0b79860f7 100644 --- a/pytest.ini +++ b/pytest.ini @@ -21,12 +21,17 @@ addopts = --ignore=docs/utils twisted = 1 flake8-ignore = + # Files that are only meant to provide top-level imports are expected not + # to use any of their imports: + scrapy/core/downloader/handlers/http.py F401 + scrapy/http/__init__.py F401 + # Issues pending a review: # extras extras/qps-bench-server.py E261 E501 extras/qpsclient.py E501 E261 E501 # scrapy/commands scrapy/commands/__init__.py E128 E501 - scrapy/commands/check.py F401 E501 + scrapy/commands/check.py E501 scrapy/commands/crawl.py E501 scrapy/commands/edit.py E501 scrapy/commands/fetch.py E401 E501 E128 E502 E731 @@ -37,7 +42,6 @@ flake8-ignore = scrapy/commands/shell.py E128 E501 E502 scrapy/commands/startproject.py E502 E127 E501 E128 scrapy/commands/version.py E501 E128 - scrapy/commands/view.py F401 # scrapy/contracts scrapy/contracts/__init__.py E501 W504 scrapy/contracts/default.py E502 E128 @@ -46,17 +50,16 @@ flake8-ignore = scrapy/core/scheduler.py E501 scrapy/core/scraper.py E501 E306 E261 E128 W504 scrapy/core/spidermw.py E501 E731 E502 E126 E226 - scrapy/core/downloader/__init__.py F401 E501 + scrapy/core/downloader/__init__.py E501 scrapy/core/downloader/contextfactory.py E501 E128 E126 scrapy/core/downloader/middleware.py E501 E502 scrapy/core/downloader/tls.py E501 E305 E241 scrapy/core/downloader/webclient.py E731 E501 E261 E502 E128 E126 E226 scrapy/core/downloader/handlers/__init__.py E501 scrapy/core/downloader/handlers/ftp.py E501 E305 E128 E127 - scrapy/core/downloader/handlers/http.py F401 scrapy/core/downloader/handlers/http10.py E501 scrapy/core/downloader/handlers/http11.py E501 - scrapy/core/downloader/handlers/s3.py E501 F401 E502 E128 E126 + scrapy/core/downloader/handlers/s3.py E501 E502 E128 E126 # scrapy/downloadermiddlewares scrapy/downloadermiddlewares/ajaxcrawl.py E501 E226 scrapy/downloadermiddlewares/decompression.py E501 @@ -66,19 +69,18 @@ flake8-ignore = scrapy/downloadermiddlewares/httpproxy.py E501 scrapy/downloadermiddlewares/redirect.py E501 W504 scrapy/downloadermiddlewares/retry.py E501 E126 - scrapy/downloadermiddlewares/robotstxt.py F401 E501 + scrapy/downloadermiddlewares/robotstxt.py E501 scrapy/downloadermiddlewares/stats.py E501 # scrapy/extensions scrapy/extensions/closespider.py E501 E502 E128 E123 scrapy/extensions/corestats.py E501 scrapy/extensions/feedexport.py E128 E501 - scrapy/extensions/httpcache.py E128 E501 E303 F401 + scrapy/extensions/httpcache.py E128 E501 E303 scrapy/extensions/memdebug.py E501 scrapy/extensions/spiderstate.py E501 scrapy/extensions/telnet.py E501 W504 scrapy/extensions/throttle.py E501 # scrapy/http - scrapy/http/__init__.py F401 scrapy/http/common.py E501 scrapy/http/cookies.py E501 scrapy/http/request/__init__.py E501 @@ -87,7 +89,7 @@ flake8-ignore = scrapy/http/response/__init__.py E501 E128 W293 W291 scrapy/http/response/text.py E501 W293 E128 E124 # scrapy/linkextractors - scrapy/linkextractors/__init__.py E731 E502 E501 E402 F401 + scrapy/linkextractors/__init__.py E731 E502 E501 E402 scrapy/linkextractors/lxmlhtml.py E501 E731 E226 # scrapy/loader scrapy/loader/__init__.py E501 E502 E128 @@ -97,8 +99,8 @@ flake8-ignore = scrapy/pipelines/images.py E265 E501 scrapy/pipelines/media.py E125 E501 E266 # scrapy/selector - scrapy/selector/__init__.py F403 F401 - scrapy/selector/unified.py F401 E501 E111 + scrapy/selector/__init__.py F403 + scrapy/selector/unified.py E501 E111 # scrapy/settings scrapy/settings/__init__.py E501 scrapy/settings/default_settings.py E501 E261 E114 E116 E226 @@ -106,32 +108,30 @@ flake8-ignore = # scrapy/spidermiddlewares scrapy/spidermiddlewares/httperror.py E501 scrapy/spidermiddlewares/offsite.py E501 - scrapy/spidermiddlewares/referer.py F401 E501 E129 W503 W504 + scrapy/spidermiddlewares/referer.py E501 E129 W503 W504 scrapy/spidermiddlewares/urllength.py E501 # scrapy/spiders - scrapy/spiders/__init__.py F401 E501 E402 + scrapy/spiders/__init__.py E501 E402 scrapy/spiders/crawl.py E501 scrapy/spiders/feed.py E501 E261 scrapy/spiders/sitemap.py E501 # scrapy/utils scrapy/utils/benchserver.py E501 - scrapy/utils/boto.py F401 scrapy/utils/conf.py E402 E502 E501 - scrapy/utils/console.py E261 F401 E306 E305 - scrapy/utils/curl.py F401 + scrapy/utils/console.py E261 E306 E305 scrapy/utils/datatypes.py E501 E226 scrapy/utils/decorators.py E501 scrapy/utils/defer.py E501 E128 scrapy/utils/deprecate.py E128 E501 E127 E502 - scrapy/utils/engine.py F401 E261 + scrapy/utils/engine.py E261 scrapy/utils/gz.py E305 E501 W504 - scrapy/utils/http.py F403 F401 E226 + scrapy/utils/http.py F403 E226 scrapy/utils/httpobj.py E501 scrapy/utils/iterators.py E501 E701 scrapy/utils/log.py E128 W503 - scrapy/utils/markup.py F403 F401 W292 + scrapy/utils/markup.py F403 W292 scrapy/utils/misc.py E501 E226 - scrapy/utils/multipart.py F403 F401 W292 + scrapy/utils/multipart.py F403 W292 scrapy/utils/project.py E501 scrapy/utils/python.py E501 scrapy/utils/reactor.py E226 @@ -143,7 +143,7 @@ flake8-ignore = scrapy/utils/spider.py E271 E501 scrapy/utils/ssl.py E501 scrapy/utils/test.py E501 - scrapy/utils/url.py E501 F403 F401 E128 F405 + scrapy/utils/url.py E501 F403 E128 F405 # scrapy scrapy/__init__.py E402 E501 scrapy/_monkeypatches.py W293 @@ -167,29 +167,29 @@ flake8-ignore = scrapy/squeues.py E128 scrapy/statscollectors.py E501 # tests - tests/__init__.py F401 E402 E501 - tests/mockserver.py E401 E501 E126 E123 F401 + tests/__init__.py E402 E501 + tests/mockserver.py E401 E501 E126 E123 tests/pipelines.py F841 E226 tests/spiders.py E501 E127 tests/test_closespider.py E501 E127 tests/test_command_fetch.py E501 E261 - tests/test_command_parse.py F401 E501 E128 E303 E226 + tests/test_command_parse.py E501 E128 E303 E226 tests/test_command_shell.py E501 E128 - tests/test_commands.py F401 E128 E501 + tests/test_commands.py E128 E501 tests/test_contracts.py E501 E128 W293 tests/test_crawl.py E501 E741 E265 tests/test_crawler.py F841 E306 E501 tests/test_dependencies.py F841 E501 E305 - tests/test_downloader_handlers.py E124 E127 E128 E225 E261 E265 F401 E501 E502 E701 E126 E226 E123 + tests/test_downloader_handlers.py E124 E127 E128 E225 E261 E265 E501 E502 E701 E126 E226 E123 tests/test_downloadermiddleware.py E501 tests/test_downloadermiddleware_ajaxcrawlable.py E501 tests/test_downloadermiddleware_cookies.py E731 E741 E501 E128 E303 E265 E126 tests/test_downloadermiddleware_decompression.py E127 tests/test_downloadermiddleware_defaultheaders.py E501 tests/test_downloadermiddleware_downloadtimeout.py E501 - tests/test_downloadermiddleware_httpcache.py E501 E305 F401 - tests/test_downloadermiddleware_httpcompression.py E501 F401 E251 E126 E123 - tests/test_downloadermiddleware_httpproxy.py F401 E501 E128 + tests/test_downloadermiddleware_httpcache.py E501 E305 + tests/test_downloadermiddleware_httpcompression.py E501 E251 E126 E123 + tests/test_downloadermiddleware_httpproxy.py E501 E128 tests/test_downloadermiddleware_redirect.py E501 E303 E128 E306 E127 E305 tests/test_downloadermiddleware_retry.py E501 E128 W293 E251 E502 E303 E126 tests/test_downloadermiddleware_robotstxt.py E501 @@ -197,11 +197,11 @@ flake8-ignore = tests/test_dupefilters.py E221 E501 E741 W293 W291 E128 E124 tests/test_engine.py E401 E501 E502 E128 E261 tests/test_exporters.py E501 E731 E306 E128 E124 - tests/test_extension_telnet.py F401 F841 - tests/test_feedexport.py E501 F401 F841 E241 + tests/test_extension_telnet.py F841 + tests/test_feedexport.py E501 F841 E241 tests/test_http_cookies.py E501 tests/test_http_headers.py E501 - tests/test_http_request.py F401 E402 E501 E261 E127 E128 W293 E502 E128 E502 E126 E123 + tests/test_http_request.py E402 E501 E261 E127 E128 W293 E502 E128 E502 E126 E123 tests/test_http_response.py E501 E301 E502 E128 E265 tests/test_item.py E701 E128 F841 E306 tests/test_link.py E501 @@ -211,20 +211,20 @@ flake8-ignore = tests/test_mail.py E128 E501 E305 tests/test_middleware.py E501 E128 tests/test_pipeline_crawl.py E131 E501 E128 E126 - tests/test_pipeline_files.py F401 E501 W293 E303 E272 E226 - tests/test_pipeline_images.py F401 F841 E501 E303 + tests/test_pipeline_files.py E501 W293 E303 E272 E226 + tests/test_pipeline_images.py F841 E501 E303 tests/test_pipeline_media.py E501 E741 E731 E128 E261 E306 E502 tests/test_request_cb_kwargs.py E501 tests/test_responsetypes.py E501 E305 - tests/test_robotstxt_interface.py F401 E501 W291 E501 + tests/test_robotstxt_interface.py E501 W291 E501 tests/test_scheduler.py E501 E126 E123 - tests/test_selector.py F401 E501 E127 - tests/test_spider.py E501 F401 + tests/test_selector.py E501 E127 + tests/test_spider.py E501 tests/test_spidermiddleware.py E501 E226 tests/test_spidermiddleware_httperror.py E128 E501 E127 E121 tests/test_spidermiddleware_offsite.py E501 E128 E111 W293 - tests/test_spidermiddleware_output_chain.py F401 E501 W293 E226 - tests/test_spidermiddleware_referer.py F401 E501 F841 E125 E201 E261 E124 E501 E241 E121 + tests/test_spidermiddleware_output_chain.py E501 W293 E226 + tests/test_spidermiddleware_referer.py E501 F841 E125 E201 E261 E124 E501 E241 E121 tests/test_squeues.py E501 E701 E741 tests/test_utils_conf.py E501 E303 E128 tests/test_utils_curl.py E501 @@ -235,16 +235,16 @@ flake8-ignore = tests/test_utils_iterators.py E501 E128 E129 E303 E241 tests/test_utils_log.py E741 E226 tests/test_utils_python.py E501 E303 E731 E701 E305 - tests/test_utils_reqser.py F401 E501 E128 + tests/test_utils_reqser.py E501 E128 tests/test_utils_request.py E501 E128 E305 tests/test_utils_response.py E501 tests/test_utils_signal.py E741 F841 E731 E226 tests/test_utils_sitemap.py E128 E501 E124 tests/test_utils_spider.py E261 E305 tests/test_utils_template.py E305 - tests/test_utils_url.py F401 E501 E127 E305 E211 E125 E501 E226 E241 E126 E123 + tests/test_utils_url.py E501 E127 E305 E211 E125 E501 E226 E241 E126 E123 tests/test_webclient.py E501 E128 E122 E303 E402 E306 E226 E241 E123 E126 tests/test_cmdline/__init__.py E502 E501 - tests/test_settings/__init__.py F401 E501 E128 + tests/test_settings/__init__.py E501 E128 tests/test_spiderloader/__init__.py E128 E501 tests/test_utils_misc/__init__.py E501 diff --git a/scrapy/commands/check.py b/scrapy/commands/check.py index 3e6c11b7d..9d4437a47 100644 --- a/scrapy/commands/check.py +++ b/scrapy/commands/check.py @@ -1,6 +1,4 @@ -from __future__ import print_function import time -import sys from collections import defaultdict from unittest import TextTestRunner, TextTestResult as _TextTestResult diff --git a/scrapy/commands/view.py b/scrapy/commands/view.py index 31c17c0ab..41e77ba3b 100644 --- a/scrapy/commands/view.py +++ b/scrapy/commands/view.py @@ -1,4 +1,4 @@ -from scrapy.commands import fetch, ScrapyCommand +from scrapy.commands import fetch from scrapy.utils.response import open_in_browser diff --git a/scrapy/core/downloader/__init__.py b/scrapy/core/downloader/__init__.py index 949dacbc8..213268741 100644 --- a/scrapy/core/downloader/__init__.py +++ b/scrapy/core/downloader/__init__.py @@ -1,6 +1,4 @@ -from __future__ import absolute_import import random -import warnings from time import time from datetime import datetime from collections import deque diff --git a/scrapy/core/downloader/handlers/http.py b/scrapy/core/downloader/handlers/http.py index ac4b867c3..6111e132a 100644 --- a/scrapy/core/downloader/handlers/http.py +++ b/scrapy/core/downloader/handlers/http.py @@ -1,3 +1,2 @@ -from __future__ import absolute_import from .http10 import HTTP10DownloadHandler from .http11 import HTTP11DownloadHandler as HTTPDownloadHandler diff --git a/scrapy/core/downloader/handlers/s3.py b/scrapy/core/downloader/handlers/s3.py index d8bbdd326..808d1bf21 100644 --- a/scrapy/core/downloader/handlers/s3.py +++ b/scrapy/core/downloader/handlers/s3.py @@ -21,7 +21,7 @@ def _get_boto_connection(): return http_request.headers try: - import boto.auth + import boto.auth # noqa: F401 except ImportError: _S3Connection = _v19_S3Connection else: diff --git a/scrapy/extensions/httpcache.py b/scrapy/extensions/httpcache.py index f3fabf710..11403957c 100644 --- a/scrapy/extensions/httpcache.py +++ b/scrapy/extensions/httpcache.py @@ -6,18 +6,16 @@ import os from email.utils import mktime_tz, parsedate_tz from importlib import import_module from time import time -from warnings import warn from weakref import WeakKeyDictionary from six.moves import cPickle as pickle from w3lib.http import headers_raw_to_dict, headers_dict_to_raw -from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.http import Headers, Response from scrapy.responsetypes import responsetypes from scrapy.utils.httpobj import urlparse_cached from scrapy.utils.project import data_path -from scrapy.utils.python import to_bytes, to_unicode, garbage_collect +from scrapy.utils.python import to_bytes, to_unicode from scrapy.utils.request import request_fingerprint diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index e4c62f87b..8c3693f04 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -118,4 +118,4 @@ class FilteringLinkExtractor(object): # Top-level imports -from .lxmlhtml import LxmlLinkExtractor as LinkExtractor +from .lxmlhtml import LxmlLinkExtractor as LinkExtractor # noqa: F401 diff --git a/scrapy/selector/__init__.py b/scrapy/selector/__init__.py index 90e96ee92..a9240c1f6 100644 --- a/scrapy/selector/__init__.py +++ b/scrapy/selector/__init__.py @@ -1,4 +1,4 @@ """ Selectors """ -from scrapy.selector.unified import * +from scrapy.selector.unified import * # noqa: F401 diff --git a/scrapy/selector/unified.py b/scrapy/selector/unified.py index 62fda40b0..a08955dc9 100644 --- a/scrapy/selector/unified.py +++ b/scrapy/selector/unified.py @@ -2,12 +2,10 @@ XPath selectors based on lxml """ -import warnings from parsel import Selector as _ParselSelector from scrapy.utils.trackref import object_ref from scrapy.utils.python import to_bytes from scrapy.http import HtmlResponse, XmlResponse -from scrapy.utils.decorators import deprecated __all__ = ['Selector', 'SelectorList'] diff --git a/scrapy/spiders/__init__.py b/scrapy/spiders/__init__.py index 94095bc27..8d15dfceb 100644 --- a/scrapy/spiders/__init__.py +++ b/scrapy/spiders/__init__.py @@ -10,7 +10,6 @@ from scrapy import signals from scrapy.http import Request from scrapy.utils.trackref import object_ref from scrapy.utils.url import url_is_from_spider -from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.utils.deprecate import method_is_overridden @@ -100,6 +99,6 @@ class Spider(object_ref): # Top-level imports -from scrapy.spiders.crawl import CrawlSpider, Rule -from scrapy.spiders.feed import XMLFeedSpider, CSVFeedSpider -from scrapy.spiders.sitemap import SitemapSpider +from scrapy.spiders.crawl import CrawlSpider, Rule # noqa: F401 +from scrapy.spiders.feed import XMLFeedSpider, CSVFeedSpider # noqa: F401 +from scrapy.spiders.sitemap import SitemapSpider # noqa: F401 diff --git a/scrapy/utils/boto.py b/scrapy/utils/boto.py index c8fc911bb..b76d5e56e 100644 --- a/scrapy/utils/boto.py +++ b/scrapy/utils/boto.py @@ -7,7 +7,7 @@ from scrapy.exceptions import NotConfigured def is_botocore(): try: - import botocore + import botocore # noqa: F401 return True except ImportError: raise NotConfigured('missing botocore library') diff --git a/scrapy/utils/console.py b/scrapy/utils/console.py index a26e84d38..688e28c34 100644 --- a/scrapy/utils/console.py +++ b/scrapy/utils/console.py @@ -52,7 +52,7 @@ def _embed_standard_shell(namespace={}, banner=''): except ImportError: pass else: - import rlcompleter + import rlcompleter # noqa: F401 readline.parse_and_bind("tab:complete") @wraps(_embed_standard_shell) def wrapper(namespace=namespace, banner=''): diff --git a/scrapy/utils/curl.py b/scrapy/utils/curl.py index b3fd0a497..7fb25a71d 100644 --- a/scrapy/utils/curl.py +++ b/scrapy/utils/curl.py @@ -4,7 +4,7 @@ from shlex import split from six.moves.http_cookies import SimpleCookie from six.moves.urllib.parse import urlparse -from six import string_types, iteritems +from six import iteritems from w3lib.http import basic_auth_header diff --git a/scrapy/utils/engine.py b/scrapy/utils/engine.py index 2c20b5c88..267c7ecd1 100644 --- a/scrapy/utils/engine.py +++ b/scrapy/utils/engine.py @@ -1,7 +1,7 @@ """Some debugging functions for working with the Scrapy engine""" -from __future__ import print_function -from time import time # used in global tests code +# used in global tests code +from time import time # noqa: F401 def get_engine_status(engine): diff --git a/scrapy/utils/http.py b/scrapy/utils/http.py index ad49ef3e9..bab262393 100644 --- a/scrapy/utils/http.py +++ b/scrapy/utils/http.py @@ -8,7 +8,7 @@ import warnings from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.utils.decorators import deprecated -from w3lib.http import * +from w3lib.http import * # noqa: F401 warnings.warn("Module `scrapy.utils.http` is deprecated, " diff --git a/scrapy/utils/markup.py b/scrapy/utils/markup.py index a18f308a3..2455fcc16 100644 --- a/scrapy/utils/markup.py +++ b/scrapy/utils/markup.py @@ -6,7 +6,7 @@ For new code, always import from w3lib.html instead of this module import warnings from scrapy.exceptions import ScrapyDeprecationWarning -from w3lib.html import * +from w3lib.html import * # noqa: F401 warnings.warn("Module `scrapy.utils.markup` is deprecated. " diff --git a/scrapy/utils/multipart.py b/scrapy/utils/multipart.py index c2d8afd07..e81f63152 100644 --- a/scrapy/utils/multipart.py +++ b/scrapy/utils/multipart.py @@ -6,7 +6,7 @@ For new code, always import from w3lib.form instead of this module import warnings from scrapy.exceptions import ScrapyDeprecationWarning -from w3lib.form import * +from w3lib.form import * # noqa: F401 warnings.warn("Module `scrapy.utils.multipart` is deprecated. " diff --git a/scrapy/utils/url.py b/scrapy/utils/url.py index b3a4be007..a3e1bde63 100644 --- a/scrapy/utils/url.py +++ b/scrapy/utils/url.py @@ -12,7 +12,6 @@ from six.moves.urllib.parse import (ParseResult, urldefrag, urlparse, urlunparse # scrapy.utils.url was moved to w3lib.url and import * ensures this # move doesn't break old code from w3lib.url import * -from w3lib.url import _safe_chars, _unquotepath from scrapy.utils.python import to_unicode diff --git a/tests/mockserver.py b/tests/mockserver.py index b766bb653..7ebb8bb62 100644 --- a/tests/mockserver.py +++ b/tests/mockserver.py @@ -1,9 +1,11 @@ -from __future__ import print_function -import sys, time, random, os, json -from six.moves.urllib.parse import urlencode +import json +import os +import random +import sys from subprocess import Popen, PIPE from OpenSSL import SSL +from six.moves.urllib.parse import urlencode from twisted.web.server import Site, NOT_DONE_YET from twisted.web.resource import Resource from twisted.web.static import File diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 59d4a3eec..0421529a6 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -543,7 +543,7 @@ class Https11InvalidDNSPattern(Https11TestCase): def setUp(self): try: - from service_identity.exceptions import CertificateError + from service_identity.exceptions import CertificateError # noqa: F401 except ImportError: raise unittest.SkipTest("cryptography lib is too old") self.tls_log_message = 'SSL connection certificate: issuer "/C=IE/O=Scrapy/CN=127.0.0.1", subject "/C=IE/O=Scrapy/CN=127.0.0.1"' @@ -778,7 +778,7 @@ class S3TestCase(unittest.TestCase): @contextlib.contextmanager def _mocked_date(self, date): try: - import botocore.auth + import botocore.auth # noqa: F401 except ImportError: yield else: @@ -843,8 +843,10 @@ class S3TestCase(unittest.TestCase): b'AWS 0PN5J17HBGZHT7JJ3X82:thdUi9VAkzhkniLj96JIrOPGi0g=') def test_request_signing5(self): - try: import botocore - except ImportError: pass + try: + import botocore # noqa: F401 + except ImportError: + pass else: raise unittest.SkipTest( 'botocore does not support overriding date with x-amz-date') diff --git a/tests/test_downloadermiddleware_httpcache.py b/tests/test_downloadermiddleware_httpcache.py index 00e6c685e..9401dd66d 100644 --- a/tests/test_downloadermiddleware_httpcache.py +++ b/tests/test_downloadermiddleware_httpcache.py @@ -1,12 +1,9 @@ -from __future__ import print_function import time import tempfile import shutil import unittest import email.utils from contextlib import contextmanager -import pytest -import sys from scrapy.http import Response, HtmlResponse, Request from scrapy.spiders import Spider diff --git a/tests/test_downloadermiddleware_httpcompression.py b/tests/test_downloadermiddleware_httpcompression.py index 0745c8dd3..c6a823b53 100644 --- a/tests/test_downloadermiddleware_httpcompression.py +++ b/tests/test_downloadermiddleware_httpcompression.py @@ -70,7 +70,7 @@ class HttpCompressionTest(TestCase): def test_process_response_br(self): try: - import brotli + import brotli # noqa: F401 except ImportError: raise SkipTest("no brotli") response = self._getresponse('br') diff --git a/tests/test_downloadermiddleware_httpproxy.py b/tests/test_downloadermiddleware_httpproxy.py index 30920b2da..36743b1de 100644 --- a/tests/test_downloadermiddleware_httpproxy.py +++ b/tests/test_downloadermiddleware_httpproxy.py @@ -1,11 +1,10 @@ import os -import sys from functools import partial -from twisted.trial.unittest import TestCase, SkipTest +from twisted.trial.unittest import TestCase from scrapy.downloadermiddlewares.httpproxy import HttpProxyMiddleware from scrapy.exceptions import NotConfigured -from scrapy.http import Response, Request +from scrapy.http import Request from scrapy.spiders import Spider from scrapy.crawler import Crawler from scrapy.settings import Settings diff --git a/tests/test_extension_telnet.py b/tests/test_extension_telnet.py index 875ceb83c..873a97248 100644 --- a/tests/test_extension_telnet.py +++ b/tests/test_extension_telnet.py @@ -3,7 +3,7 @@ from twisted.conch.telnet import ITelnetProtocol from twisted.cred import credentials from twisted.internet import defer -from scrapy.extensions.telnet import TelnetConsole, logger +from scrapy.extensions.telnet import TelnetConsole from scrapy.utils.test import get_crawler diff --git a/tests/test_feedexport.py b/tests/test_feedexport.py index 87139e81f..ce3c4f059 100644 --- a/tests/test_feedexport.py +++ b/tests/test_feedexport.py @@ -167,7 +167,7 @@ class S3FeedStorageTest(unittest.TestCase): create=True) def test_parse_credentials(self): try: - import boto + import boto # noqa: F401 except ImportError: raise unittest.SkipTest("S3FeedStorage requires boto") aws_credentials = {'AWS_ACCESS_KEY_ID': 'settings_key', @@ -268,7 +268,7 @@ class S3FeedStorageTest(unittest.TestCase): @defer.inlineCallbacks def test_store_botocore_without_acl(self): try: - import botocore + import botocore # noqa: F401 except ImportError: raise unittest.SkipTest('botocore is required') @@ -288,7 +288,7 @@ class S3FeedStorageTest(unittest.TestCase): @defer.inlineCallbacks def test_store_botocore_with_acl(self): try: - import botocore + import botocore # noqa: F401 except ImportError: raise unittest.SkipTest('botocore is required') diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 5134a03b9..9df6ff67b 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -1,5 +1,3 @@ -# -*- coding: utf-8 -*- -import cgi import unittest import re import json @@ -8,7 +6,7 @@ from urllib.parse import unquote_to_bytes import warnings from six.moves import xmlrpc_client as xmlrpclib -from six.moves.urllib.parse import urlparse, parse_qs, unquote +from six.moves.urllib.parse import urlparse, parse_qs from scrapy.http import Request, FormRequest, XmlRpcRequest, JsonRequest, Headers, HtmlResponse from scrapy.utils.python import to_bytes, to_unicode diff --git a/tests/test_pipeline_images.py b/tests/test_pipeline_images.py index 4f7265763..7f1cb4a11 100644 --- a/tests/test_pipeline_images.py +++ b/tests/test_pipeline_images.py @@ -1,7 +1,6 @@ import io import hashlib import random -import warnings from tempfile import mkdtemp from shutil import rmtree diff --git a/tests/test_robotstxt_interface.py b/tests/test_robotstxt_interface.py index 080507276..27d79437b 100644 --- a/tests/test_robotstxt_interface.py +++ b/tests/test_robotstxt_interface.py @@ -5,7 +5,7 @@ from twisted.trial import unittest def reppy_available(): # check if reppy parser is installed try: - from reppy.robots import Robots + from reppy.robots import Robots # noqa: F401 except ImportError: return False return True @@ -14,7 +14,7 @@ def reppy_available(): def rerp_available(): # check if robotexclusionrulesparser is installed try: - from robotexclusionrulesparser import RobotExclusionRulesParser + from robotexclusionrulesparser import RobotExclusionRulesParser # noqa: F401 except ImportError: return False return True @@ -23,7 +23,7 @@ def rerp_available(): def protego_available(): # check if protego parser is installed try: - from protego import Protego + from protego import Protego # noqa: F401 except ImportError: return False return True diff --git a/tests/test_selector.py b/tests/test_selector.py index b2565dd78..09c2546fb 100644 --- a/tests/test_selector.py +++ b/tests/test_selector.py @@ -1,9 +1,9 @@ -import warnings import weakref + from twisted.trial import unittest + from scrapy.http import TextResponse, HtmlResponse, XmlResponse from scrapy.selector import Selector -from lxml import etree class SelectorTestCase(unittest.TestCase): diff --git a/tests/test_spider.py b/tests/test_spider.py index c0fccfdd6..aa43e3b3a 100644 --- a/tests/test_spider.py +++ b/tests/test_spider.py @@ -15,7 +15,6 @@ from scrapy.spiders import Spider, CrawlSpider, Rule, XMLFeedSpider, \ CSVFeedSpider, SitemapSpider from scrapy.linkextractors import LinkExtractor from scrapy.exceptions import ScrapyDeprecationWarning -from scrapy.utils.trackref import object_ref from scrapy.utils.test import get_crawler diff --git a/tests/test_spidermiddleware_output_chain.py b/tests/test_spidermiddleware_output_chain.py index 940e31070..5b7b5e7aa 100644 --- a/tests/test_spidermiddleware_output_chain.py +++ b/tests/test_spidermiddleware_output_chain.py @@ -6,7 +6,6 @@ from twisted.internet import defer from scrapy import Spider, Request from scrapy.utils.test import get_crawler from tests.mockserver import MockServer -from tests.spiders import MockServerSpider class LogExceptionMiddleware: diff --git a/tests/test_spidermiddleware_referer.py b/tests/test_spidermiddleware_referer.py index a9c31a983..2be6a1cd5 100644 --- a/tests/test_spidermiddleware_referer.py +++ b/tests/test_spidermiddleware_referer.py @@ -2,7 +2,6 @@ from six.moves.urllib.parse import urlparse from unittest import TestCase import warnings -from scrapy.exceptions import NotConfigured from scrapy.http import Response, Request from scrapy.settings import Settings from scrapy.spiders import Spider diff --git a/tests/test_utils_reqser.py b/tests/test_utils_reqser.py index 92cd16de7..06d9c004c 100644 --- a/tests/test_utils_reqser.py +++ b/tests/test_utils_reqser.py @@ -1,8 +1,4 @@ -# -*- coding: utf-8 -*- import unittest -import sys - -import six from scrapy.http import Request, FormRequest from scrapy.spiders import Spider diff --git a/tests/test_utils_url.py b/tests/test_utils_url.py index 93addc082..c7bcaf88b 100644 --- a/tests/test_utils_url.py +++ b/tests/test_utils_url.py @@ -1,13 +1,9 @@ # -*- coding: utf-8 -*- import unittest -import six -from six.moves.urllib.parse import urlparse - from scrapy.spiders import Spider from scrapy.utils.url import (url_is_from_any_domain, url_is_from_spider, - add_http_if_no_scheme, guess_scheme, - parse_url, strip_url) + add_http_if_no_scheme, guess_scheme, strip_url) __doctests__ = ['scrapy.utils.url'] diff --git a/tox.ini b/tox.ini index fd75d18e2..e8672ea27 100644 --- a/tox.ini +++ b/tox.ini @@ -73,6 +73,7 @@ commands = basepython = python3.8 deps = {[testenv]deps} + -r docs/requirements.txt pytest-flake8 commands = py.test --flake8 {posargs:docs scrapy tests} From b23288135633f9e800b575d84df0a109665a840e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 21 Nov 2019 14:30:10 +0100 Subject: [PATCH 291/387] Restore intentional import of unused objects --- scrapy/utils/url.py | 1 + 1 file changed, 1 insertion(+) diff --git a/scrapy/utils/url.py b/scrapy/utils/url.py index a3e1bde63..2c7b324a1 100644 --- a/scrapy/utils/url.py +++ b/scrapy/utils/url.py @@ -12,6 +12,7 @@ from six.moves.urllib.parse import (ParseResult, urldefrag, urlparse, urlunparse # scrapy.utils.url was moved to w3lib.url and import * ensures this # move doesn't break old code from w3lib.url import * +from w3lib.url import _safe_chars, _unquotepath # noqa: F401 from scrapy.utils.python import to_unicode From 1718e450ef9549a4fc71b01dba1e6faf7a63238a Mon Sep 17 00:00:00 2001 From: Mabel Villalba <mabelvj@gmail.com> Date: Mon, 18 Nov 2019 12:33:55 +0100 Subject: [PATCH 292/387] [start_url] Fixes #4133: Raise AttributeError error when empty 'start_urls' and 'start_url' found. Added test. --- scrapy/spiders/__init__.py | 5 +++++ tests/test_spider.py | 7 +++++++ 2 files changed, 12 insertions(+) diff --git a/scrapy/spiders/__init__.py b/scrapy/spiders/__init__.py index e9c131e3b..5a35fcdb6 100644 --- a/scrapy/spiders/__init__.py +++ b/scrapy/spiders/__init__.py @@ -68,6 +68,11 @@ class Spider(object_ref): def start_requests(self): cls = self.__class__ + if not self.start_urls and hasattr(self, 'start_url'): + raise AttributeError( + "Crawling could not start: 'start_urls' not found " + "or empty (but found 'start_url' attribute instead, " + "did you miss an 's'?)") if method_is_overridden(cls, Spider, 'make_requests_from_url'): warnings.warn( "Spider.make_requests_from_url method is deprecated; it " diff --git a/tests/test_spider.py b/tests/test_spider.py index 83fb68c2f..0a6640cec 100644 --- a/tests/test_spider.py +++ b/tests/test_spider.py @@ -391,6 +391,13 @@ class CrawlSpiderTest(SpiderTest): self.assertTrue(hasattr(spider, '_follow_links')) self.assertFalse(spider._follow_links) + def test_start_url(self): + spider = self.spider_class("example.com") + spider.start_url = 'https://www.example.com' + + with self.assertRaisesRegex(AttributeError, + r'^Crawling could not start.*$'): + list(spider.start_requests()) class SitemapSpiderTest(SpiderTest): From 9b5053c564fb465cfe85d0799388b3591638c5a5 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 21 Nov 2019 22:00:34 +0100 Subject: [PATCH 293/387] Undo unintended tox.ini changes --- tox.ini | 1 - 1 file changed, 1 deletion(-) diff --git a/tox.ini b/tox.ini index e8672ea27..fd75d18e2 100644 --- a/tox.ini +++ b/tox.ini @@ -73,7 +73,6 @@ commands = basepython = python3.8 deps = {[testenv]deps} - -r docs/requirements.txt pytest-flake8 commands = py.test --flake8 {posargs:docs scrapy tests} From 55cc5c9068a7a39908f90e6403e6bd8c8ef11a4d Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Fri, 22 Nov 2019 12:41:31 -0300 Subject: [PATCH 294/387] Skip pickle in bandit check --- .bandit.yml | 2 ++ 1 file changed, 2 insertions(+) diff --git a/.bandit.yml b/.bandit.yml index 00554587a..cc7db3a66 100644 --- a/.bandit.yml +++ b/.bandit.yml @@ -1,6 +1,7 @@ skips: - B101 - B105 +- B301 - B303 - B306 - B307 @@ -8,6 +9,7 @@ skips: - B320 - B321 - B402 +- B403 - B404 - B406 - B410 From 40b5cfc0a4adbc51fa35018b09902255228e360c Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Tue, 23 Jul 2019 18:33:19 -0300 Subject: [PATCH 295/387] Item loaders: allow single-argument processors (unbound methods) --- docs/topics/loaders.rst | 16 +------------ scrapy/loader/__init__.py | 22 +++++++++++++----- tests/test_loader.py | 48 +++++++++++++++++++++++++++++++++++++++ 3 files changed, 65 insertions(+), 21 deletions(-) diff --git a/docs/topics/loaders.rst b/docs/topics/loaders.rst index 12a5e5c60..81c8dab03 100644 --- a/docs/topics/loaders.rst +++ b/docs/topics/loaders.rst @@ -142,20 +142,6 @@ accept one (and only one) positional argument, which will be an iterable. containing the collected values (for that field). The result of the output processors is the value that will be finally assigned to the item. -If you want to use a plain function as a processor, make sure it receives -``self`` as the first argument:: - - def lowercase_processor(self, values): - for v in values: - yield v.lower() - - class MyItemLoader(ItemLoader): - name_in = lowercase_processor - -This is because whenever a function is assigned as a class variable, it becomes -a method and would be passed the instance as the the first argument when being -called. See `this answer on stackoverflow`_ for more details. - The other thing you need to keep in mind is that the values returned by input processors are collected internally (in lists) and then passed to output processors to populate the fields. @@ -163,7 +149,7 @@ processors to populate the fields. Last, but not least, Scrapy comes with some :ref:`commonly used processors <topics-loaders-available-processors>` built-in for convenience. -.. _this answer on stackoverflow: https://stackoverflow.com/a/35322635 + Declaring Item Loaders ====================== diff --git a/scrapy/loader/__init__.py b/scrapy/loader/__init__.py index 60fd6d222..7cf67e29e 100644 --- a/scrapy/loader/__init__.py +++ b/scrapy/loader/__init__.py @@ -4,8 +4,7 @@ Item Loader See documentation in docs/topics/loaders.rst """ from collections import defaultdict - -import six +from contextlib import suppress from scrapy.item import Item from scrapy.loader.common import wrap_loader_context @@ -15,6 +14,17 @@ from scrapy.utils.misc import arg_to_iter, extract_regex from scrapy.utils.python import flatten +def unbound_method(method): + """ + Allow to use single-argument functions as input or output processors + (no need to define an unused first 'self' argument) + """ + with suppress(AttributeError): + if '.' not in method.__qualname__: + return method.__func__ + return method + + class ItemLoader(object): default_item_class = Item @@ -72,7 +82,7 @@ class ItemLoader(object): if value is None: return if not field_name: - for k, v in six.iteritems(value): + for k, v in value.items(): self._add_value(k, v) else: self._add_value(field_name, value) @@ -82,7 +92,7 @@ class ItemLoader(object): if value is None: return if not field_name: - for k, v in six.iteritems(value): + for k, v in value.items(): self._replace_value(k, v) else: self._replace_value(field_name, value) @@ -142,14 +152,14 @@ class ItemLoader(object): if not proc: proc = self._get_item_field_attr(field_name, 'input_processor', self.default_input_processor) - return proc + return unbound_method(proc) def get_output_processor(self, field_name): proc = getattr(self, '%s_out' % field_name, None) if not proc: proc = self._get_item_field_attr(field_name, 'output_processor', self.default_output_processor) - return proc + return unbound_method(proc) def _process_input_value(self, field_name, value): proc = self.get_input_processor(field_name) diff --git a/tests/test_loader.py b/tests/test_loader.py index b87602809..6bfc31dbf 100644 --- a/tests/test_loader.py +++ b/tests/test_loader.py @@ -994,5 +994,53 @@ class SelectJmesTestCase(unittest.TestCase): ) +# Functions as processors + +def function_processor_strip(iterable): + return [x.strip() for x in iterable] + + +def function_processor_upper(iterable): + return [x.upper() for x in iterable] + + +class FunctionProcessorItem(Item): + foo = Field( + input_processor=function_processor_strip, + output_processor=function_processor_upper, + ) + + +class FunctionProcessorItemLoader(ItemLoader): + default_item_class = FunctionProcessorItem + + +class FunctionProcessorDictLoader(ItemLoader): + default_item_class = dict + foo_in = function_processor_strip + foo_out = function_processor_upper + + +class FunctionProcessorTestCase(unittest.TestCase): + + def test_processor_defined_in_item(self): + lo = FunctionProcessorItemLoader() + lo.add_value('foo', ' bar ') + lo.add_value('foo', [' asdf ', ' qwerty ']) + self.assertEqual( + dict(lo.load_item()), + {'foo': ['BAR', 'ASDF', 'QWERTY']} + ) + + def test_processor_defined_in_item_loader(self): + lo = FunctionProcessorDictLoader() + lo.add_value('foo', ' bar ') + lo.add_value('foo', [' asdf ', ' qwerty ']) + self.assertEqual( + dict(lo.load_item()), + {'foo': ['BAR', 'ASDF', 'QWERTY']} + ) + + if __name__ == "__main__": unittest.main() From 54b056c4be045099f9f872c5455c003d68b578a7 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Mon, 25 Nov 2019 12:13:31 +0100 Subject: [PATCH 296/387] Make developer-tools doctests pass --- docs/_tests/quotes.html | 281 ++++++++++++++++++++++++++++++++ docs/topics/developer-tools.rst | 40 +++-- pytest.ini | 1 - 3 files changed, 308 insertions(+), 14 deletions(-) create mode 100644 docs/_tests/quotes.html diff --git a/docs/_tests/quotes.html b/docs/_tests/quotes.html new file mode 100644 index 000000000..71aff8847 --- /dev/null +++ b/docs/_tests/quotes.html @@ -0,0 +1,281 @@ +<!DOCTYPE html> +<html lang="en"> +<head> + <meta charset="UTF-8"> + <title>Quotes to Scrape + + + + +
+
+ +
+

+ + Login + +

+
+
+ + +
+
+ +
+ “The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.” + by + (about) + +
+ Tags: + + + change + + deep-thoughts + + thinking + + world + +
+
+ +
+ “It is our choices, Harry, that show what we truly are, far more than our abilities.” + by + (about) + +
+ Tags: + + + abilities + + choices + +
+
+ +
+ “There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.” + by + (about) + +
+ Tags: + + + inspirational + + life + + live + + miracle + + miracles + +
+
+ +
+ “The person, be it gentleman or lady, who has not pleasure in a good novel, must be intolerably stupid.” + by + (about) + +
+ Tags: + + + aliteracy + + books + + classic + + humor + +
+
+ +
+ “Imperfection is beauty, madness is genius and it's better to be absolutely ridiculous than absolutely boring.” + by + (about) + +
+ Tags: + + + be-yourself + + inspirational + +
+
+ +
+ “Try not to become a man of success. Rather become a man of value.” + by + (about) + +
+ Tags: + + + adulthood + + success + + value + +
+
+ +
+ “It is better to be hated for what you are than to be loved for what you are not.” + by + (about) + +
+ Tags: + + + life + + love + +
+
+ +
+ “I have not failed. I've just found 10,000 ways that won't work.” + by + (about) + +
+ Tags: + + + edison + + failure + + inspirational + + paraphrased + +
+
+ +
+ “A woman is like a tea bag; you never know how strong it is until it's in hot water.” + by + (about) + + +
+ +
+ “A day without sunshine is like, you know, night.” + by + (about) + +
+ Tags: + + + humor + + obvious + + simile + +
+
+ + +
+
+ +

Top Ten tags

+ + + love + + + + inspirational + + + + life + + + + humor + + + + books + + + + reading + + + + friendship + + + + friends + + + + truth + + + + simile + + + +
+
+ +
+ + + \ No newline at end of file diff --git a/docs/topics/developer-tools.rst b/docs/topics/developer-tools.rst index bf14643be..e67ce55f9 100644 --- a/docs/topics/developer-tools.rst +++ b/docs/topics/developer-tools.rst @@ -39,7 +39,7 @@ Therefore, you should keep in mind the following things: .. _topics-inspector: Inspecting a website -=================================== +==================== By far the most handy feature of the Developer Tools is the `Inspector` feature, which allows you to inspect the underlying HTML code of @@ -79,13 +79,23 @@ sections and tags of a webpage, which greatly improves readability. You can expand and collapse a tag by clicking on the arrow in front of it or by double clicking directly on the tag. If we expand the ``span`` tag with the ``class= "text"`` we will see the quote-text we clicked on. The `Inspector` lets you -copy XPaths to selected elements. Let's try it out: Right-click on the ``span`` -tag, select ``Copy > XPath`` and paste it in the scrapy shell like so:: +copy XPaths to selected elements. Let's try it out. + +First open the Scrapy shell at http://quotes.toscrape.com/ in a terminal: + +.. code-block:: none $ scrapy shell "http://quotes.toscrape.com/" - (...) - >>> response.xpath('/html/body/div/div[2]/div[1]/div[1]/span[1]/text()').getall() - ['"The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”] + +Then, back to your web browser, right-click on the ``span`` tag, select +``Copy > XPath`` and paste it in the Scrapy shell like so: + +.. invisible-code-block: python + + response = load_response('http://quotes.toscrape.com/', 'quotes.html') + +>>> response.xpath('/html/body/div/div[2]/div[1]/div[1]/span[1]/text()').getall() +['“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”'] Adding ``text()`` at the end we are able to extract the first quote with this basic selector. But this XPath is not really that clever. All it does is @@ -112,13 +122,13 @@ see each quote: With this knowledge we can refine our XPath: Instead of a path to follow, we'll simply select all ``span`` tags with the ``class="text"`` by using -the `has-class-extension`_:: +the `has-class-extension`_: - >>> response.xpath('//span[has-class("text")]/text()').getall() - ['"The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”, - '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', - '“There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.”', - (...)] +>>> response.xpath('//span[has-class("text")]/text()').getall() +['“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', +'“It is our choices, Harry, that show what we truly are, far more than our abilities.”', +'“There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.”', +...] And with one simple, cleverer XPath we are able to extract all quotes from the page. We could have constructed a loop over our first XPath to increase @@ -159,7 +169,11 @@ The page is quite similar to the basic `quotes.toscrape.com`_-page, but instead of the above-mentioned ``Next`` button, the page automatically loads new quotes when you scroll to the bottom. We could go ahead and try out different XPaths directly, but instead -we'll check another quite useful command from the scrapy shell:: +we'll check another quite useful command from the scrapy shell: + +.. skip: next + +.. code-block:: none $ scrapy shell "quotes.toscrape.com/scroll" (...) diff --git a/pytest.ini b/pytest.ini index 33c34b8e8..0e830076d 100644 --- a/pytest.ini +++ b/pytest.ini @@ -8,7 +8,6 @@ addopts = --ignore=docs/_ext --ignore=docs/conf.py --ignore=docs/news.rst - --ignore=docs/topics/developer-tools.rst --ignore=docs/topics/dynamic-content.rst --ignore=docs/topics/items.rst --ignore=docs/topics/leaks.rst From ed1e577610b77e92f9833ca96755b9f0727085f5 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 25 Nov 2019 13:35:29 +0100 Subject: [PATCH 297/387] Use super().__init__ in BaseItemExporter subclasses --- scrapy/exporters.py | 31 ++++++++++++++++--------------- 1 file changed, 16 insertions(+), 15 deletions(-) diff --git a/scrapy/exporters.py b/scrapy/exporters.py index 3defafd60..0c28c2d7f 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -24,8 +24,9 @@ __all__ = ['BaseItemExporter', 'PprintItemExporter', 'PickleItemExporter', class BaseItemExporter(object): - def __init__(self, **kwargs): - self._configure(kwargs) + def __init__(self, dont_fail=False, **kwargs): + self._kwargs = kwargs + self._configure(kwargs, dont_fail=dont_fail) def _configure(self, options, dont_fail=False): """Configure the exporter by poping options from the ``options`` dict. @@ -82,10 +83,10 @@ class BaseItemExporter(object): class JsonLinesItemExporter(BaseItemExporter): def __init__(self, file, **kwargs): - self._configure(kwargs, dont_fail=True) + super().__init__(dont_fail=True, **kwargs) self.file = file - kwargs.setdefault('ensure_ascii', not self.encoding) - self.encoder = ScrapyJSONEncoder(**kwargs) + self._kwargs.setdefault('ensure_ascii', not self.encoding) + self.encoder = ScrapyJSONEncoder(**self._kwargs) def export_item(self, item): itemdict = dict(self._get_serialized_fields(item)) @@ -96,15 +97,15 @@ class JsonLinesItemExporter(BaseItemExporter): class JsonItemExporter(BaseItemExporter): def __init__(self, file, **kwargs): - self._configure(kwargs, dont_fail=True) + super().__init__(dont_fail=True, **kwargs) self.file = file # there is a small difference between the behaviour or JsonItemExporter.indent # and ScrapyJSONEncoder.indent. ScrapyJSONEncoder.indent=None is needed to prevent # the addition of newlines everywhere json_indent = self.indent if self.indent is not None and self.indent > 0 else None - kwargs.setdefault('indent', json_indent) - kwargs.setdefault('ensure_ascii', not self.encoding) - self.encoder = ScrapyJSONEncoder(**kwargs) + self._kwargs.setdefault('indent', json_indent) + self._kwargs.setdefault('ensure_ascii', not self.encoding) + self.encoder = ScrapyJSONEncoder(**self._kwargs) self.first_item = True def _beautify_newline(self): @@ -135,7 +136,7 @@ class XmlItemExporter(BaseItemExporter): def __init__(self, file, **kwargs): self.item_element = kwargs.pop('item_element', 'item') self.root_element = kwargs.pop('root_element', 'items') - self._configure(kwargs) + super().__init__(**kwargs) if not self.encoding: self.encoding = 'utf-8' self.xg = XMLGenerator(file, encoding=self.encoding) @@ -191,7 +192,7 @@ class XmlItemExporter(BaseItemExporter): class CsvItemExporter(BaseItemExporter): def __init__(self, file, include_headers_line=True, join_multivalued=',', **kwargs): - self._configure(kwargs, dont_fail=True) + super().__init__(dont_fail=True, **kwargs) if not self.encoding: self.encoding = 'utf-8' self.include_headers_line = include_headers_line @@ -202,7 +203,7 @@ class CsvItemExporter(BaseItemExporter): encoding=self.encoding, newline='' # Windows needs this https://github.com/scrapy/scrapy/issues/3034 ) - self.csv_writer = csv.writer(self.stream, **kwargs) + self.csv_writer = csv.writer(self.stream, **self._kwargs) self._headers_not_written = True self._join_multivalued = join_multivalued @@ -251,7 +252,7 @@ class CsvItemExporter(BaseItemExporter): class PickleItemExporter(BaseItemExporter): def __init__(self, file, protocol=2, **kwargs): - self._configure(kwargs) + super().__init__(**kwargs) self.file = file self.protocol = protocol @@ -270,7 +271,7 @@ class MarshalItemExporter(BaseItemExporter): """ def __init__(self, file, **kwargs): - self._configure(kwargs) + super().__init__(**kwargs) self.file = file def export_item(self, item): @@ -280,7 +281,7 @@ class MarshalItemExporter(BaseItemExporter): class PprintItemExporter(BaseItemExporter): def __init__(self, file, **kwargs): - self._configure(kwargs) + super().__init__(**kwargs) self.file = file def export_item(self, item): From dd12f5fdcd5ce2c63e9a5d3626031f5d93c1628e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 25 Nov 2019 15:59:59 +0100 Subject: [PATCH 298/387] Use Response.follow_all in the documentation where appropiate --- docs/intro/tutorial.rst | 26 +++++++++++++------------- 1 file changed, 13 insertions(+), 13 deletions(-) diff --git a/docs/intro/tutorial.rst b/docs/intro/tutorial.rst index 6b15a5fbd..a2775e0bb 100644 --- a/docs/intro/tutorial.rst +++ b/docs/intro/tutorial.rst @@ -625,12 +625,12 @@ attribute automatically. So the code can be shortened further:: for a in response.css('li.next a'): yield response.follow(a, callback=self.parse) -.. note:: +To create multiple requests from an iterable, you can use +:meth:`response.follow_all ` instead:: + + links = response.css('li.next a') + yield from response.follow_all(links, callback=self.parse) - ``response.follow(response.css('li.next a'))`` is not valid because - ``response.css`` returns a list-like object with selectors for all results, - not a single selector. A ``for`` loop like in the example above, or - ``response.follow(response.css('li.next a')[0])`` is fine. More examples and patterns -------------------------- @@ -647,13 +647,11 @@ this time for scraping author information:: start_urls = ['http://quotes.toscrape.com/'] def parse(self, response): - # follow links to author pages - for href in response.css('.author + a::attr(href)'): - yield response.follow(href, self.parse_author) + author_page_links = response.css('.author + a') + yield from response.follow_all(author_links, self.parse_author) - # follow pagination links - for href in response.css('li.next a::attr(href)'): - yield response.follow(href, self.parse) + pagination_links = response.css('li.next a') + yield from response.follow_all(pagination_links, self.parse) def parse_author(self, response): def extract_with_css(query): @@ -669,8 +667,10 @@ This spider will start from the main page, it will follow all the links to the authors pages calling the ``parse_author`` callback for each of them, and also the pagination links with the ``parse`` callback as we saw before. -Here we're passing callbacks to ``response.follow`` as positional arguments -to make the code shorter; it also works for ``scrapy.Request``. +Here we're passing callbacks to +:meth:`response.follow_all ` as positional +arguments to make the code shorter; it also works for +:class:`~scrapy.http.Request`. The ``parse_author`` callback defines a helper function to extract and cleanup the data from a CSS query and yields the Python dict with the author data. From b73fc99b60ed83be403e9570e84f5267d35dcc9e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Tue, 26 Nov 2019 10:31:55 +0100 Subject: [PATCH 299/387] Use InterSphinx for coverage links --- docs/conf.py | 1 + docs/contributing.rst | 5 ++--- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/conf.py b/docs/conf.py index eab366efd..914d1d05f 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -275,6 +275,7 @@ coverage_ignore_pyobjects = [ # ------------------------------------- intersphinx_mapping = { + 'coverage': ('https://coverage.readthedocs.io/en/stable', None), 'pytest': ('https://docs.pytest.org/en/latest', None), 'python': ('https://docs.python.org/3', None), 'sphinx': ('https://www.sphinx-doc.org/en/master', None), diff --git a/docs/contributing.rst b/docs/contributing.rst index 81bb50a77..234c4bcee 100644 --- a/docs/contributing.rst +++ b/docs/contributing.rst @@ -243,14 +243,13 @@ the Python 3.6 :doc:`tox ` environment using all your CPU cores:: tox -e py36 -- scrapy tests -n auto -To see coverage report install `coverage`_ (``pip install coverage``) and run: +To see coverage report install :doc:`coverage ` +(``pip install coverage``) and run: ``coverage report`` see output of ``coverage --help`` for more options like html or xml report. -.. _coverage: https://pypi.python.org/pypi/coverage - Writing tests ------------- From 63546cbf3e02380819732441ad55d95725dca7c5 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Wed, 27 Nov 2019 22:42:52 +0500 Subject: [PATCH 300/387] Deprecate the HTTPS proxy noconnect mode. --- scrapy/core/downloader/handlers/http11.py | 7 ++++++ tests/test_downloader_handlers.py | 10 --------- tests/test_proxy_connect.py | 26 ----------------------- 3 files changed, 7 insertions(+), 36 deletions(-) diff --git a/scrapy/core/downloader/handlers/http11.py b/scrapy/core/downloader/handlers/http11.py index 7d917cb74..782eca89e 100644 --- a/scrapy/core/downloader/handlers/http11.py +++ b/scrapy/core/downloader/handlers/http11.py @@ -16,6 +16,7 @@ from twisted.web.http import _DataLoss, PotentialDataLoss from twisted.web.client import Agent, ResponseDone, HTTPConnectionPool, ResponseFailed, URI from twisted.internet.endpoints import TCP4ClientEndpoint +from scrapy.exceptions import ScrapyDeprecationWarning from scrapy.http import Headers from scrapy.responsetypes import responsetypes from scrapy.core.downloader.webclient import _parse @@ -285,6 +286,12 @@ class ScrapyAgent(object): scheme = _parse(request.url)[0] proxyHost = to_unicode(proxyHost) omitConnectTunnel = b'noconnect' in proxyParams + if omitConnectTunnel: + warnings.warn("Using HTTPS proxies in the noconnect mode is deprecated. " + "If you use Crawlera, it doesn't require this mode anymore, " + "so you should update scrapy-crawlera to 1.3.0+ " + "and remove '?noconnect' from the Crawlera URL.", + ScrapyDeprecationWarning) if scheme == b'https' and not omitConnectTunnel: proxyAuth = request.headers.get(b'Proxy-Authorization', None) proxyConf = (proxyHost, proxyPort, proxyAuth) diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 60124b93f..45d4aa952 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -687,16 +687,6 @@ class HttpProxyTestCase(unittest.TestCase): request = Request('http://example.com', meta={'proxy': http_proxy}) return self.download_request(request, Spider('foo')).addCallback(_test) - def test_download_with_proxy_https_noconnect(self): - def _test(response): - self.assertEqual(response.status, 200) - self.assertEqual(response.url, request.url) - self.assertEqual(response.body, b'https://example.com') - - http_proxy = '%s?noconnect' % self.getURL('') - request = Request('https://example.com', meta={'proxy': http_proxy}) - return self.download_request(request, Spider('foo')).addCallback(_test) - def test_download_without_proxy(self): def _test(response): self.assertEqual(response.status, 200) diff --git a/tests/test_proxy_connect.py b/tests/test_proxy_connect.py index 277455751..05d371b65 100644 --- a/tests/test_proxy_connect.py +++ b/tests/test_proxy_connect.py @@ -108,32 +108,6 @@ class ProxyConnectTestCase(TestCase): echo = json.loads(crawler.spider.meta['responses'][0].text) self.assertTrue('Proxy-Authorization' not in echo['headers']) - # The noconnect mode isn't supported by the current mitmproxy, it returns - # "Invalid request scheme: https" as it doesn't seem to support full URLs in GET at all, - # and it's not clear what behavior is intended by Scrapy and by mitmproxy here. - # https://github.com/mitmproxy/mitmproxy/issues/848 may be related. - # The Scrapy noconnect mode was required, at least in the past, to work with Crawlera, - # and https://github.com/scrapy-plugins/scrapy-crawlera/pull/44 seems to be related. - - @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') - @defer.inlineCallbacks - def test_https_noconnect(self): - proxy = os.environ['https_proxy'] - os.environ['https_proxy'] = proxy + '?noconnect' - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(200, l) - - @pytest.mark.xfail(reason='mitmproxy gives an error for noconnect requests') - @defer.inlineCallbacks - def test_https_noconnect_auth_error(self): - os.environ['https_proxy'] = _wrong_credentials(os.environ['https_proxy']) + '?noconnect' - crawler = get_crawler(SimpleSpider) - with LogCapture() as l: - yield crawler.crawl(self.mockserver.url("/status?n=200", is_secure=True)) - self._assert_got_response_code(407, l) - def _assert_got_response_code(self, code, log): print(log) self.assertEqual(str(log).count('Crawled (%d)' % code), 1) From d1cdfb47013330b0391a8db3b6b812697ee64b6a Mon Sep 17 00:00:00 2001 From: Grammy Jiang <719388+grammy-jiang@users.noreply.github.com> Date: Fri, 29 Nov 2019 19:13:57 +1100 Subject: [PATCH 301/387] Use pprint.pformat on overridden settings (#4199) Keeps consistency with scrapy.middleware --- scrapy/crawler.py | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/scrapy/crawler.py b/scrapy/crawler.py index f8c80880a..19b61dc7e 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -1,3 +1,4 @@ +import pprint import six import signal import logging @@ -45,7 +46,8 @@ class Crawler(object): logging.root.addHandler(handler) d = dict(overridden_settings(self.settings)) - logger.info("Overridden settings: %(settings)r", {'settings': d}) + logger.info("Overridden settings:\n%(settings)s", + {'settings': pprint.pformat(d)}) if get_scrapy_root_handler() is not None: # scrapy root handler already installed: update it with new settings From 5980b0f2840f51f8f9c7c3a266be93527999dd11 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Mon, 2 Dec 2019 16:47:44 +0100 Subject: [PATCH 302/387] =?UTF-8?q?Don=E2=80=99t=20use=20follow=5Fall=20wh?= =?UTF-8?q?ere=20a=20single=20item=20is=20expected=20(#4)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/intro/tutorial.rst | 11 +++++------ 1 file changed, 5 insertions(+), 6 deletions(-) diff --git a/docs/intro/tutorial.rst b/docs/intro/tutorial.rst index a2775e0bb..2f97017fc 100644 --- a/docs/intro/tutorial.rst +++ b/docs/intro/tutorial.rst @@ -616,20 +616,19 @@ instance; you still have to yield this Request. You can also pass a selector to ``response.follow`` instead of a string; this selector should extract necessary attributes:: - for href in response.css('li.next a::attr(href)'): - yield response.follow(href, callback=self.parse) + href = response.css('li.next a::attr(href)')[0] + yield response.follow(href, callback=self.parse) For ```` elements there is a shortcut: ``response.follow`` uses their href attribute automatically. So the code can be shortened further:: - for a in response.css('li.next a'): - yield response.follow(a, callback=self.parse) + a = response.css('li.next a')[0] + yield response.follow(a, callback=self.parse) To create multiple requests from an iterable, you can use :meth:`response.follow_all ` instead:: - links = response.css('li.next a') - yield from response.follow_all(links, callback=self.parse) + yield from response.follow_all(response.css('a'), callback=self.parse) More examples and patterns From 2a9f5a0aefae83fc0b1dc161d117a265148ad1ef Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Tue, 3 Dec 2019 15:56:50 -0300 Subject: [PATCH 303/387] Skip invalid links when passing SelectorLists to Response.follow_all --- scrapy/http/response/text.py | 19 +++++++++++-------- tests/test_http_response.py | 12 ++++++++---- 2 files changed, 19 insertions(+), 12 deletions(-) diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 6acf1026f..e3646b2d5 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -7,10 +7,10 @@ See documentation in docs/topics/request-response.rst from contextlib import suppress from typing import Generator +from urllib.parse import urljoin import parsel import six -from six.moves.urllib.parse import urljoin from w3lib.encoding import (html_body_declared_encoding, html_to_unicode, http_content_type_encoding, resolve_encoding) from w3lib.html import strip_html5_whitespace @@ -183,22 +183,25 @@ class TextResponse(Response): In addition, ``css`` and ``xpath`` arguments are accepted to perform the link extraction within the ``follow_all`` method (only one of ``urls``, ``css`` and ``xpath`` is accepted). - Note that when using the ``css`` or ``xpath`` parameters, this method will not produce - requests for selectors from which links cannot be obtained (for instance, anchor tags - without an ``href`` attribute) + Note that when passing a ``SelectorList`` as argument for the ``urls`` parameter or + using the ``css`` or ``xpath`` parameters, this method will not produce requests for + selectors from which links cannot be obtained (for instance, anchor tags without an + ``href`` attribute) """ arg_count = len(list(filter(None, (urls, css, xpath)))) if arg_count != 1: raise ValueError('Please supply exactly one of the following arguments: urls, css, xpath') if not urls: if css: - selector_list = self.css(css) + urls = self.css(css) if xpath: - selector_list = self.xpath(xpath) + urls = self.xpath(xpath) + if isinstance(urls, parsel.SelectorList): + selectors = urls urls = [] - for selector in selector_list: + for sel in selectors: with suppress(_InvalidSelector): - urls.append(_url_from_selector(selector)) + urls.append(_url_from_selector(sel)) return super(TextResponse, self).follow_all( urls=urls, callback=callback, diff --git a/tests/test_http_response.py b/tests/test_http_response.py index 36ccdfa1f..ce13650ce 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -579,8 +579,10 @@ class TextResponseTest(BaseResponseTest): 'http://example.com/page/4/', ] response = self._links_response_no_href() - extracted = [r.url for r in response.follow_all(css='.pagination a')] - self.assertEqual(expected, extracted) + extracted1 = [r.url for r in response.follow_all(css='.pagination a')] + self.assertEqual(expected, extracted1) + extracted2 = [r.url for r in response.follow_all(response.css('.pagination a'))] + self.assertEqual(expected, extracted2) def test_follow_all_xpath(self): expected = [ @@ -598,8 +600,10 @@ class TextResponseTest(BaseResponseTest): 'http://example.com/page/4/', ] response = self._links_response_no_href() - extracted = [r.url for r in response.follow_all(xpath='//div[@id="pagination"]/a')] - self.assertEqual(expected, extracted) + extracted1 = [r.url for r in response.follow_all(xpath='//div[@id="pagination"]/a')] + self.assertEqual(expected, extracted1) + extracted2 = [r.url for r in response.follow_all(response.xpath('//div[@id="pagination"]/a'))] + self.assertEqual(expected, extracted2) def test_follow_all_too_many_arguments(self): response = self._links_response() From 62778cf23f6e0cad5839f31f9da8b3c5778dcbdb Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta Date: Wed, 11 Sep 2019 15:40:44 -0300 Subject: [PATCH 304/387] Request: remove restriction about errback without callback --- scrapy/http/request/__init__.py | 1 - tests/test_http_request.py | 46 ++++++++++++++++++++++----------- 2 files changed, 31 insertions(+), 16 deletions(-) diff --git a/scrapy/http/request/__init__.py b/scrapy/http/request/__init__.py index 76a428199..61c0a4c9e 100644 --- a/scrapy/http/request/__init__.py +++ b/scrapy/http/request/__init__.py @@ -32,7 +32,6 @@ class Request(object_ref): raise TypeError('callback must be a callable, got %s' % type(callback).__name__) if errback is not None and not callable(errback): raise TypeError('errback must be a callable, got %s' % type(errback).__name__) - assert callback or not errback, "Cannot use errback without a callback" self.callback = callback self.errback = errback diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 9df6ff67b..3449c7a40 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -246,25 +246,41 @@ class RequestTest(unittest.TestCase): self.assertRaises(AttributeError, setattr, r, 'url', 'http://example2.com') self.assertRaises(AttributeError, setattr, r, 'body', 'xxx') - def test_callback_is_callable(self): + def test_callback_and_errback(self): def a_function(): pass - r = self.request_class('http://example.com') - self.assertIsNone(r.callback) - r = self.request_class('http://example.com', a_function) - self.assertIs(r.callback, a_function) - with self.assertRaises(TypeError): - self.request_class('http://example.com', 'a_function') - def test_errback_is_callable(self): - def a_function(): - pass - r = self.request_class('http://example.com') - self.assertIsNone(r.errback) - r = self.request_class('http://example.com', a_function, errback=a_function) - self.assertIs(r.errback, a_function) + r1 = self.request_class('http://example.com') + self.assertIsNone(r1.callback) + self.assertIsNone(r1.errback) + + r2 = self.request_class('http://example.com', callback=a_function) + self.assertIs(r2.callback, a_function) + self.assertIsNone(r2.errback) + + r3 = self.request_class('http://example.com', errback=a_function) + self.assertIsNone(r3.callback) + self.assertIs(r3.errback, a_function) + + r4 = self.request_class( + url='http://example.com', + callback=a_function, + errback=a_function, + ) + self.assertIs(r4.callback, a_function) + self.assertIs(r4.errback, a_function) + + def test_callback_and_errback_type(self): with self.assertRaises(TypeError): - self.request_class('http://example.com', a_function, errback='a_function') + self.request_class('http://example.com', callback='a_function') + with self.assertRaises(TypeError): + self.request_class('http://example.com', errback='a_function') + with self.assertRaises(TypeError): + self.request_class( + url='http://example.com', + callback='a_function', + errback='a_function', + ) def test_from_curl(self): # Note: more curated tests regarding curl conversion are in From 5d8d4bb7d7d998ae2324c4995bafaafdef752572 Mon Sep 17 00:00:00 2001 From: Grammy Jiang <719388+grammy-jiang@users.noreply.github.com> Date: Thu, 5 Dec 2019 00:22:10 +1100 Subject: [PATCH 305/387] Re-arrange the imports in the httpproxy module (#4210) This commit re-arranges the imports in the httpproxy module to follow pep8 --- scrapy/downloadermiddlewares/httpproxy.py | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/scrapy/downloadermiddlewares/httpproxy.py b/scrapy/downloadermiddlewares/httpproxy.py index 2212d9688..5e4542b6c 100644 --- a/scrapy/downloadermiddlewares/httpproxy.py +++ b/scrapy/downloadermiddlewares/httpproxy.py @@ -1,7 +1,8 @@ import base64 +from urllib.request import _parse_proxy + from six.moves.urllib.parse import unquote, urlunparse from six.moves.urllib.request import getproxies, proxy_bypass -from urllib.request import _parse_proxy from scrapy.exceptions import NotConfigured from scrapy.utils.httpobj import urlparse_cached From 702333478d072c3c043c64ec7ad3997befb87943 Mon Sep 17 00:00:00 2001 From: Grammy Jiang <719388+grammy-jiang@users.noreply.github.com> Date: Thu, 5 Dec 2019 00:23:28 +1100 Subject: [PATCH 306/387] Re-arrange the imports in httpcache module (#4209) This commit re-arrange the imports in httpcache module to follow pep8 --- scrapy/downloadermiddlewares/httpcache.py | 16 ++++++++++++---- 1 file changed, 12 insertions(+), 4 deletions(-) diff --git a/scrapy/downloadermiddlewares/httpcache.py b/scrapy/downloadermiddlewares/httpcache.py index 495b103d1..4e06f8236 100644 --- a/scrapy/downloadermiddlewares/httpcache.py +++ b/scrapy/downloadermiddlewares/httpcache.py @@ -1,11 +1,19 @@ from email.utils import formatdate + from twisted.internet import defer -from twisted.internet.error import TimeoutError, DNSLookupError, \ - ConnectionRefusedError, ConnectionDone, ConnectError, \ - ConnectionLost, TCPTimedOutError +from twisted.internet.error import ( + ConnectError, + ConnectionDone, + ConnectionLost, + ConnectionRefusedError, + DNSLookupError, + TCPTimedOutError, + TimeoutError, +) from twisted.web.client import ResponseFailed + from scrapy import signals -from scrapy.exceptions import NotConfigured, IgnoreRequest +from scrapy.exceptions import IgnoreRequest, NotConfigured from scrapy.utils.misc import load_object From 74627033c4a1701f3e197216f9f2801d497f5535 Mon Sep 17 00:00:00 2001 From: Grammy Jiang <719388+grammy-jiang@users.noreply.github.com> Date: Thu, 5 Dec 2019 00:24:14 +1100 Subject: [PATCH 307/387] Remove the used import and re-arrange the imports (#4208) This commit removes unused import and re-arrange the imports in cookies module --- scrapy/downloadermiddlewares/cookies.py | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/scrapy/downloadermiddlewares/cookies.py b/scrapy/downloadermiddlewares/cookies.py index 0d2b9900c..aeb7578b8 100644 --- a/scrapy/downloadermiddlewares/cookies.py +++ b/scrapy/downloadermiddlewares/cookies.py @@ -1,8 +1,8 @@ -import os -import six import logging from collections import defaultdict +import six + from scrapy.exceptions import NotConfigured from scrapy.http import Response from scrapy.http.cookies import CookieJar From 1b35260625c3ffec9885265d9ac92771ade67ad9 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 25 Jul 2019 18:18:34 +0500 Subject: [PATCH 308/387] Add a test for downloader middlewares using Deferreds. --- tests/test_downloadermiddleware.py | 29 +++++++++++++++++++++++++++++ 1 file changed, 29 insertions(+) diff --git a/tests/test_downloadermiddleware.py b/tests/test_downloadermiddleware.py index 6b9a5bee8..1b81ea949 100644 --- a/tests/test_downloadermiddleware.py +++ b/tests/test_downloadermiddleware.py @@ -1,5 +1,6 @@ from unittest import mock +from twisted.internet.defer import Deferred from twisted.trial.unittest import TestCase from twisted.python.failure import Failure @@ -177,3 +178,31 @@ class ProcessExceptionInvalidOutput(ManagerTestCase): dfd.addBoth(results.append) self.assertIsInstance(results[0], Failure) self.assertIsInstance(results[0].value, _InvalidOutput) + + +class MiddlewareUsingDeferreds(ManagerTestCase): + """Middlewares using Deferreds should work""" + + def test_deferred(self): + resp = Response('http://example.com/index.html') + + class DeferredMiddleware: + def cb(self, result): + return result + + def process_request(self, request, spider): + d = Deferred() + d.addCallback(self.cb) + d.callback(resp) + return d + + self.mwman._add_middleware(DeferredMiddleware()) + req = Request('http://example.com/index.html') + download_func = mock.MagicMock() + dfd = self.mwman.download(download_func, req, self.spider) + results = [] + dfd.addBoth(results.append) + self._wait(dfd) + + self.assertIs(results[0], resp) + self.assertFalse(download_func.called) From 1b437bbe9fa0eb9736f35e510d486805706c783e Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Tue, 30 Jul 2019 19:02:16 +0500 Subject: [PATCH 309/387] Install the asyncio reactor on "import scrapy". --- scrapy/__init__.py | 22 ++++++++++++++++++++++ 1 file changed, 22 insertions(+) diff --git a/scrapy/__init__.py b/scrapy/__init__.py index 230e5cee3..41eaee959 100644 --- a/scrapy/__init__.py +++ b/scrapy/__init__.py @@ -23,6 +23,28 @@ import warnings warnings.filterwarnings('ignore', category=DeprecationWarning, module='twisted') del warnings +# Install twisted asyncio loop +def _install_asyncio_reactor(): + global asyncio_supported + try: + import asyncio + from twisted.internet import asyncioreactor + except ImportError: + pass + else: + from twisted.internet.error import ReactorAlreadyInstalledError + try: + asyncioreactor.install(asyncio.get_event_loop()) + asyncio_supported = True + except ReactorAlreadyInstalledError: + import twisted.internet.reactor + if isinstance(twisted.internet.reactor, + asyncioreactor.AsyncioSelectorReactor): + asyncio_supported = True +asyncio_supported = False +_install_asyncio_reactor() +del _install_asyncio_reactor + # Apply monkey patches to fix issues in external libraries from . import _monkeypatches del _monkeypatches From 9777639533373951652f3865a6321a4dff73246a Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Tue, 30 Jul 2019 19:02:59 +0500 Subject: [PATCH 310/387] Run tests using the asyncio reactor. --- pytest.ini | 1 + tests/mockserver.py | 3 +++ tests/requirements-py3.txt | 3 ++- 3 files changed, 6 insertions(+), 1 deletion(-) diff --git a/pytest.ini b/pytest.ini index 33c34b8e8..6c4c21baf 100644 --- a/pytest.ini +++ b/pytest.ini @@ -5,6 +5,7 @@ python_classes= addopts = --assert=plain --doctest-modules + --reactor=asyncio --ignore=docs/_ext --ignore=docs/conf.py --ignore=docs/news.rst diff --git a/tests/mockserver.py b/tests/mockserver.py index 7ebb8bb62..b6aee009a 100644 --- a/tests/mockserver.py +++ b/tests/mockserver.py @@ -6,6 +6,9 @@ from subprocess import Popen, PIPE from OpenSSL import SSL from six.moves.urllib.parse import urlencode + +import scrapy # needed before importing twisted.internet.reactor + from twisted.web.server import Site, NOT_DONE_YET from twisted.web.resource import Resource from twisted.web.static import File diff --git a/tests/requirements-py3.txt b/tests/requirements-py3.txt index c4bc1f278..26ab08b04 100644 --- a/tests/requirements-py3.txt +++ b/tests/requirements-py3.txt @@ -4,7 +4,8 @@ mitmproxy; python_version >= '3.6' mitmproxy==3.0.4; python_version < '3.6' pytest pytest-cov -pytest-twisted +#pytest-twisted +-e git+https://github.com/pytest-dev/pytest-twisted@81b91f17#egg=pytest-twisted pytest-xdist sybil testfixtures From 63c3c62305a8c9c52d02ff5524301b7d45eb5724 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Tue, 30 Jul 2019 19:45:56 +0500 Subject: [PATCH 311/387] Add utils.deferred_from_coro. --- scrapy/utils/defer.py | 22 ++++++++++++++++++++++ 1 file changed, 22 insertions(+) diff --git a/scrapy/utils/defer.py b/scrapy/utils/defer.py index c5916c21c..1f6a2584c 100644 --- a/scrapy/utils/defer.py +++ b/scrapy/utils/defer.py @@ -1,10 +1,14 @@ """ Helper functions for dealing with Twisted deferreds """ +import asyncio +import asyncio.futures +import inspect from twisted.internet import defer, reactor, task from twisted.python import failure +from scrapy import asyncio_supported from scrapy.exceptions import IgnoreRequest @@ -113,3 +117,21 @@ def iter_errback(iterable, errback, *a, **kw): break except Exception: errback(failure.Failure(), *a, **kw) + + +def isfuture(o): + # workaround for Python before 3.5.3 not having asyncio.isfuture + if hasattr(asyncio, 'isfuture'): + return asyncio.isfuture(o) + return isinstance(o, asyncio.futures.Future) + + +def deferred_from_coro(o): + """Converts a coroutine into a Deferred, or returns the object as is if it isn't a coroutine""" + if isinstance(o, defer.Deferred): + return o + if asyncio.iscoroutine(o) or isfuture(o) or inspect.isawaitable(o): + if not asyncio_supported: + raise TypeError('Using coroutines requires installing AsyncioSelectorReactor') + return defer.Deferred.fromFuture(asyncio.ensure_future(o)) + return o From 8d8fbddbde133a94bb8741e48fabec05437a3df9 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Wed, 21 Aug 2019 00:07:08 +0500 Subject: [PATCH 312/387] Switch to the released version of pytest-twisted. --- tests/requirements-py3.txt | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/tests/requirements-py3.txt b/tests/requirements-py3.txt index 26ab08b04..2ac434f41 100644 --- a/tests/requirements-py3.txt +++ b/tests/requirements-py3.txt @@ -4,8 +4,7 @@ mitmproxy; python_version >= '3.6' mitmproxy==3.0.4; python_version < '3.6' pytest pytest-cov -#pytest-twisted --e git+https://github.com/pytest-dev/pytest-twisted@81b91f17#egg=pytest-twisted +pytest-twisted >= 1.11 pytest-xdist sybil testfixtures From b04b541372b219a02bb5fda2cc15cfd9fa1aac66 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Wed, 21 Aug 2019 17:14:46 +0500 Subject: [PATCH 313/387] Install the asyncio reactor only in scrapy.cmdline. --- scrapy/__init__.py | 22 ---------------------- scrapy/cmdline.py | 8 +++++++- scrapy/utils/asyncio.py | 26 ++++++++++++++++++++++++++ scrapy/utils/defer.py | 5 +++-- tests/mockserver.py | 2 -- 5 files changed, 36 insertions(+), 27 deletions(-) create mode 100644 scrapy/utils/asyncio.py diff --git a/scrapy/__init__.py b/scrapy/__init__.py index 41eaee959..230e5cee3 100644 --- a/scrapy/__init__.py +++ b/scrapy/__init__.py @@ -23,28 +23,6 @@ import warnings warnings.filterwarnings('ignore', category=DeprecationWarning, module='twisted') del warnings -# Install twisted asyncio loop -def _install_asyncio_reactor(): - global asyncio_supported - try: - import asyncio - from twisted.internet import asyncioreactor - except ImportError: - pass - else: - from twisted.internet.error import ReactorAlreadyInstalledError - try: - asyncioreactor.install(asyncio.get_event_loop()) - asyncio_supported = True - except ReactorAlreadyInstalledError: - import twisted.internet.reactor - if isinstance(twisted.internet.reactor, - asyncioreactor.AsyncioSelectorReactor): - asyncio_supported = True -asyncio_supported = False -_install_asyncio_reactor() -del _install_asyncio_reactor - # Apply monkey patches to fix issues in external libraries from . import _monkeypatches del _monkeypatches diff --git a/scrapy/cmdline.py b/scrapy/cmdline.py index 418dc1ac9..d66f0cc2d 100644 --- a/scrapy/cmdline.py +++ b/scrapy/cmdline.py @@ -7,9 +7,9 @@ import inspect import pkg_resources import scrapy -from scrapy.crawler import CrawlerProcess from scrapy.commands import ScrapyCommand from scrapy.exceptions import UsageError +from scrapy.utils.asyncio import install_asyncio_reactor, is_asyncio_supported from scrapy.utils.misc import walk_modules from scrapy.utils.project import inside_project, get_project_settings from scrapy.utils.python import garbage_collect @@ -121,6 +121,10 @@ def execute(argv=None, settings=None): settings['EDITOR'] = editor check_deprecated_settings(settings) + # needs to be before _get_commands_dict() as that imports the command modules + # which may import twisted.internet.reactor + install_asyncio_reactor() + inproject = inside_project() cmds = _get_commands_dict(settings, inproject) cmdname = _pop_command_name(argv) @@ -142,6 +146,8 @@ def execute(argv=None, settings=None): opts, args = parser.parse_args(args=argv[1:]) _run_print_help(parser, cmd.process_options, args, opts) + # needs to be after install_asyncio_reactor() as it imports twisted.internet.reactor + from scrapy.crawler import CrawlerProcess cmd.crawler_process = CrawlerProcess(settings) _run_print_help(parser, _run_command, cmd, args, opts) sys.exit(cmd.exitcode) diff --git a/scrapy/utils/asyncio.py b/scrapy/utils/asyncio.py new file mode 100644 index 000000000..e9e3bdd88 --- /dev/null +++ b/scrapy/utils/asyncio.py @@ -0,0 +1,26 @@ +#coding: utf-8 + + +def install_asyncio_reactor(): + """ Tries to install AsyncioSelectorReactor + """ + try: + import asyncio + from twisted.internet import asyncioreactor + except ImportError: + pass + else: + from twisted.internet.error import ReactorAlreadyInstalledError + try: + asyncioreactor.install(asyncio.get_event_loop()) + except ReactorAlreadyInstalledError: + pass + + +def is_asyncio_supported(): + try: + import twisted.internet.reactor + from twisted.internet import asyncioreactor + return isinstance(twisted.internet.reactor, asyncioreactor.AsyncioSelectorReactor) + except ImportError: + return False diff --git a/scrapy/utils/defer.py b/scrapy/utils/defer.py index 1f6a2584c..955fc820a 100644 --- a/scrapy/utils/defer.py +++ b/scrapy/utils/defer.py @@ -8,8 +8,9 @@ import inspect from twisted.internet import defer, reactor, task from twisted.python import failure -from scrapy import asyncio_supported from scrapy.exceptions import IgnoreRequest +from scrapy.utils.asyncio import is_asyncio_supported + def defer_fail(_failure): @@ -131,7 +132,7 @@ def deferred_from_coro(o): if isinstance(o, defer.Deferred): return o if asyncio.iscoroutine(o) or isfuture(o) or inspect.isawaitable(o): - if not asyncio_supported: + if not is_asyncio_supported(): raise TypeError('Using coroutines requires installing AsyncioSelectorReactor') return defer.Deferred.fromFuture(asyncio.ensure_future(o)) return o diff --git a/tests/mockserver.py b/tests/mockserver.py index b6aee009a..d09fbc171 100644 --- a/tests/mockserver.py +++ b/tests/mockserver.py @@ -7,8 +7,6 @@ from subprocess import Popen, PIPE from OpenSSL import SSL from six.moves.urllib.parse import urlencode -import scrapy # needed before importing twisted.internet.reactor - from twisted.web.server import Site, NOT_DONE_YET from twisted.web.resource import Resource from twisted.web.static import File From 2fbe7d49dc084b7770cc4dc6bbe65eb380b5f498 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Wed, 21 Aug 2019 17:16:33 +0500 Subject: [PATCH 314/387] Log asyncio support on spider start. --- scrapy/utils/log.py | 3 +++ 1 file changed, 3 insertions(+) diff --git a/scrapy/utils/log.py b/scrapy/utils/log.py index e07fb8698..b74b7a4af 100644 --- a/scrapy/utils/log.py +++ b/scrapy/utils/log.py @@ -11,6 +11,7 @@ from twisted.python import log as twisted_log import scrapy from scrapy.settings import Settings from scrapy.exceptions import ScrapyDeprecationWarning +from scrapy.utils.asyncio import is_asyncio_supported from scrapy.utils.versions import scrapy_components_versions @@ -148,6 +149,8 @@ def log_scrapy_info(settings): {'versions': ", ".join("%s %s" % (name, version) for name, version in scrapy_components_versions() if name != "Scrapy")}) + if is_asyncio_supported(): + logger.debug("Asyncio support enabled") class StreamLogger(object): From cc19ab5439f20ba6995528542cc064ddab86273c Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 22 Aug 2019 18:15:02 +0500 Subject: [PATCH 315/387] Add tests that check asyncio support. --- conftest.py | 4 ++++ tests/test_commands.py | 7 +++++++ tests/test_crawler.py | 20 +++++++++++++++++++- tests/test_utils_asyncio.py | 17 +++++++++++++++++ tox.ini | 6 ++++++ 5 files changed, 53 insertions(+), 1 deletion(-) create mode 100644 tests/test_utils_asyncio.py diff --git a/conftest.py b/conftest.py index d54ce155c..24e31f130 100644 --- a/conftest.py +++ b/conftest.py @@ -27,3 +27,7 @@ def pytest_collection_modifyitems(session, config, items): items[:] = [item for item in items if isinstance(item, Flake8Item)] except ImportError: pass + +@pytest.fixture() +def reactor_pytest(request): + request.cls.reactor_pytest = request.config.getoption("--reactor") diff --git a/tests/test_commands.py b/tests/test_commands.py index 536379170..8aa7ee109 100644 --- a/tests/test_commands.py +++ b/tests/test_commands.py @@ -9,6 +9,7 @@ from tempfile import mkdtemp from contextlib import contextmanager from threading import Timer +from pytest import mark from twisted.trial import unittest import scrapy @@ -178,6 +179,7 @@ class MiscCommandsTest(CommandTest): self.assertEqual(0, self.call('list')) +@mark.usefixtures('reactor_pytest') class RunSpiderCommandTest(CommandTest): debug_log_spider = """ @@ -295,6 +297,11 @@ class BadSpider(scrapy.Spider): self.assertIn("start_requests", log) self.assertIn("badspider.py", log) + def test_asyncio_supported(self): + if self.reactor_pytest == 'asyncio': + log = self.get_log(self.debug_log_spider) + self.assertIn("DEBUG: Asyncio support enabled", log) + class BenchCommandTest(CommandTest): diff --git a/tests/test_crawler.py b/tests/test_crawler.py index 8eb2389e2..151acb459 100644 --- a/tests/test_crawler.py +++ b/tests/test_crawler.py @@ -1,14 +1,16 @@ import logging import warnings +from pytest import raises, mark +from testfixtures import LogCapture from twisted.internet import defer from twisted.trial import unittest -from pytest import raises import scrapy from scrapy.crawler import Crawler, CrawlerRunner, CrawlerProcess from scrapy.settings import Settings, default_settings from scrapy.spiderloader import SpiderLoader +from scrapy.utils.asyncio import is_asyncio_supported from scrapy.utils.log import configure_logging, get_scrapy_root_handler from scrapy.utils.spider import DefaultSpider from scrapy.utils.misc import load_object @@ -203,6 +205,15 @@ class NoRequestsSpider(scrapy.Spider): return [] +class AsyncioSpider(scrapy.Spider): + name = 'asyncio' + + def start_requests(self): + self.logger.info('Asyncio support: %s', is_asyncio_supported()) + return [] + + +@mark.usefixtures('reactor_pytest') class CrawlerRunnerHasSpider(unittest.TestCase): @defer.inlineCallbacks @@ -245,3 +256,10 @@ class CrawlerRunnerHasSpider(unittest.TestCase): yield runner.crawl(NoRequestsSpider) self.assertEqual(runner.bootstrap_failed, True) + + @defer.inlineCallbacks + def test_asyncio_supported(self): + runner = CrawlerRunner() + with LogCapture() as log: + yield runner.crawl(AsyncioSpider) + log.check_present(('asyncio', 'INFO', 'Asyncio support: %s' % (self.reactor_pytest == 'asyncio'))) diff --git a/tests/test_utils_asyncio.py b/tests/test_utils_asyncio.py new file mode 100644 index 000000000..e34d3002a --- /dev/null +++ b/tests/test_utils_asyncio.py @@ -0,0 +1,17 @@ +from unittest import TestCase + +from pytest import mark + +from scrapy.utils.asyncio import is_asyncio_supported, install_asyncio_reactor + + +@mark.usefixtures('reactor_pytest') +class AsyncioTest(TestCase): + + def test_is_asyncio_supported(self): + # the result should depend only on the pytest --reactor argument + self.assertEquals(is_asyncio_supported(), self.reactor_pytest == 'asyncio') + + def test_install_asyncio_reactor(self): + # this should do nothing + install_asyncio_reactor() diff --git a/tox.ini b/tox.ini index fd75d18e2..844956e5f 100644 --- a/tox.ini +++ b/tox.ini @@ -106,3 +106,9 @@ deps = {[testenv]deps} reppy robotexclusionrulesparser + +[testenv:py38-no-asyncio] +basepython = python3.8 +deps = {[testenv]deps} +commands = + py.test --cov=scrapy --cov-report= --reactor=default {posargs:scrapy tests} From f41c2f3874d2f9deac365d633a71f032e1339e3c Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 22 Aug 2019 21:24:30 +0500 Subject: [PATCH 316/387] Add py38-no-asyncio to Travis. --- .travis.yml | 2 ++ 1 file changed, 2 insertions(+) diff --git a/.travis.yml b/.travis.yml index 9f477e860..fdf40fdf1 100644 --- a/.travis.yml +++ b/.travis.yml @@ -25,6 +25,8 @@ matrix: python: 3.8 - env: TOXENV=py38-extra-deps python: 3.8 + - env: TOXENV=py38-no-asyncio + python: 3.8 - env: TOXENV=docs python: 3.6 install: From 3ba25ccbd3b456024b1d350645407557c81d73c7 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Fri, 8 Nov 2019 00:09:28 +0500 Subject: [PATCH 317/387] Don't use asyncio.iscoroutine, as it is True for generators. --- scrapy/utils/defer.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/utils/defer.py b/scrapy/utils/defer.py index 955fc820a..30163d2fb 100644 --- a/scrapy/utils/defer.py +++ b/scrapy/utils/defer.py @@ -131,7 +131,7 @@ def deferred_from_coro(o): """Converts a coroutine into a Deferred, or returns the object as is if it isn't a coroutine""" if isinstance(o, defer.Deferred): return o - if asyncio.iscoroutine(o) or isfuture(o) or inspect.isawaitable(o): + if isfuture(o) or inspect.isawaitable(o): if not is_asyncio_supported(): raise TypeError('Using coroutines requires installing AsyncioSelectorReactor') return defer.Deferred.fromFuture(asyncio.ensure_future(o)) From 794cf71806a94ba68238a93e75c396f355159ab5 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 14 Nov 2019 13:27:21 +0500 Subject: [PATCH 318/387] Fix or ignore flake8 problems. --- pytest.ini | 2 ++ scrapy/cmdline.py | 2 +- scrapy/utils/asyncio.py | 3 --- 3 files changed, 3 insertions(+), 4 deletions(-) diff --git a/pytest.ini b/pytest.ini index 6c4c21baf..8b97237c3 100644 --- a/pytest.ini +++ b/pytest.ini @@ -116,6 +116,7 @@ flake8-ignore = scrapy/spiders/feed.py E501 E261 scrapy/spiders/sitemap.py E501 # scrapy/utils + scrapy/utils/asyncio.py E501 scrapy/utils/benchserver.py E501 scrapy/utils/conf.py E402 E502 E501 scrapy/utils/console.py E261 E306 E305 @@ -227,6 +228,7 @@ flake8-ignore = tests/test_spidermiddleware_output_chain.py E501 W293 E226 tests/test_spidermiddleware_referer.py E501 F841 E125 E201 E261 E124 E501 E241 E121 tests/test_squeues.py E501 E701 E741 + tests/test_utils_asyncio.py E501 tests/test_utils_conf.py E501 E303 E128 tests/test_utils_curl.py E501 tests/test_utils_datatypes.py E402 E501 E305 diff --git a/scrapy/cmdline.py b/scrapy/cmdline.py index d66f0cc2d..213e99bc0 100644 --- a/scrapy/cmdline.py +++ b/scrapy/cmdline.py @@ -9,7 +9,7 @@ import pkg_resources import scrapy from scrapy.commands import ScrapyCommand from scrapy.exceptions import UsageError -from scrapy.utils.asyncio import install_asyncio_reactor, is_asyncio_supported +from scrapy.utils.asyncio import install_asyncio_reactor from scrapy.utils.misc import walk_modules from scrapy.utils.project import inside_project, get_project_settings from scrapy.utils.python import garbage_collect diff --git a/scrapy/utils/asyncio.py b/scrapy/utils/asyncio.py index e9e3bdd88..f732774f1 100644 --- a/scrapy/utils/asyncio.py +++ b/scrapy/utils/asyncio.py @@ -1,6 +1,3 @@ -#coding: utf-8 - - def install_asyncio_reactor(): """ Tries to install AsyncioSelectorReactor """ From c079d5002bae90dbd85bee1f61fdc359b9f39d29 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Thu, 21 Nov 2019 23:40:16 +0500 Subject: [PATCH 319/387] Run tests without asyncio support by default, add py35-asyncio and py38-asyncio envs. --- pytest.ini | 1 - tox.ini | 10 ++++++++-- 2 files changed, 8 insertions(+), 3 deletions(-) diff --git a/pytest.ini b/pytest.ini index 8b97237c3..336ef041d 100644 --- a/pytest.ini +++ b/pytest.ini @@ -5,7 +5,6 @@ python_classes= addopts = --assert=plain --doctest-modules - --reactor=asyncio --ignore=docs/_ext --ignore=docs/conf.py --ignore=docs/news.rst diff --git a/tox.ini b/tox.ini index 844956e5f..a4edae439 100644 --- a/tox.ini +++ b/tox.ini @@ -107,8 +107,14 @@ deps = reppy robotexclusionrulesparser -[testenv:py38-no-asyncio] +[testenv:py35-asyncio] +basepython = python3.5 +deps = {[testenv]deps} +commands = + py.test --cov=scrapy --cov-report= --reactor=asyncio {posargs:scrapy tests} + +[testenv:py38-asyncio] basepython = python3.8 deps = {[testenv]deps} commands = - py.test --cov=scrapy --cov-report= --reactor=default {posargs:scrapy tests} + py.test --cov=scrapy --cov-report= --reactor=asyncio {posargs:scrapy tests} From ed34ce14c0c06d4539d4fdeb0ad014f4e6fb5b94 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Wed, 4 Dec 2019 21:32:16 +0500 Subject: [PATCH 320/387] Add the ASYNCIO_SUPPORT setting, reshuffle other logic accordingly. --- docs/topics/settings.rst | 25 +++++++++++++++++++++++++ scrapy/cmdline.py | 8 ++------ scrapy/commands/crawl.py | 3 +++ scrapy/commands/runspider.py | 3 +++ scrapy/settings/default_settings.py | 2 ++ scrapy/utils/asyncio.py | 2 +- scrapy/utils/defer.py | 17 +++++++++++------ scrapy/utils/log.py | 11 ++++++++--- tests/test_commands.py | 13 +++++++------ tests/test_crawler.py | 24 +++++++++++++++++++++--- tests/test_utils_asyncio.py | 6 +++--- 11 files changed, 86 insertions(+), 28 deletions(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index a1d15a760..43f59f7cc 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -160,6 +160,31 @@ to any particular component. In that case the module of that component will be shown, typically an extension, middleware or pipeline. It also means that the component must be enabled in order for the setting to have any effect. +.. setting:: ASYNCIO_SUPPORT + +ASYNCIO_SUPPORT +--------------- + +Default: ``False`` + +Whether to support ``async def`` methods and callbacks which use code that +requires an asyncio loop. + +If an ``async def`` coroutine doesn't require the asyncio loop, it will work +even if this is set to ``False``. Coroutines that require the asyncio loop may +silently fail to run or raise errors unless this is set to ``True``. + +When this option is set to ``True``, Scrapy will require +:class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor`. It will +install this reactor if no reactor is installed yet, such as when using the +``scrapy`` script or :class:`~scrapy.crawler.CrawlerProcess`. If you are using +:class:`~scrapy.crawler.CrawlerRunner`, you need to install the correct reactor +manually. + +The default value for this option is currently ``False`` to maintain backward +compatibility and avoid possible problems caused by using a different Twisted +reactor. + .. setting:: AWS_ACCESS_KEY_ID AWS_ACCESS_KEY_ID diff --git a/scrapy/cmdline.py b/scrapy/cmdline.py index 213e99bc0..ce030cf75 100644 --- a/scrapy/cmdline.py +++ b/scrapy/cmdline.py @@ -9,7 +9,6 @@ import pkg_resources import scrapy from scrapy.commands import ScrapyCommand from scrapy.exceptions import UsageError -from scrapy.utils.asyncio import install_asyncio_reactor from scrapy.utils.misc import walk_modules from scrapy.utils.project import inside_project, get_project_settings from scrapy.utils.python import garbage_collect @@ -121,10 +120,6 @@ def execute(argv=None, settings=None): settings['EDITOR'] = editor check_deprecated_settings(settings) - # needs to be before _get_commands_dict() as that imports the command modules - # which may import twisted.internet.reactor - install_asyncio_reactor() - inproject = inside_project() cmds = _get_commands_dict(settings, inproject) cmdname = _pop_command_name(argv) @@ -146,7 +141,8 @@ def execute(argv=None, settings=None): opts, args = parser.parse_args(args=argv[1:]) _run_print_help(parser, cmd.process_options, args, opts) - # needs to be after install_asyncio_reactor() as it imports twisted.internet.reactor + # needs to be after cmd.process_options() as it imports twisted.internet.reactor + # while commands may want to install the asyncio reactor from scrapy.crawler import CrawlerProcess cmd.crawler_process = CrawlerProcess(settings) _run_print_help(parser, _run_command, cmd, args, opts) diff --git a/scrapy/commands/crawl.py b/scrapy/commands/crawl.py index 8093fd402..e2e69be49 100644 --- a/scrapy/commands/crawl.py +++ b/scrapy/commands/crawl.py @@ -1,5 +1,6 @@ import os from scrapy.commands import ScrapyCommand +from scrapy.utils.asyncio import install_asyncio_reactor from scrapy.utils.conf import arglist_to_dict from scrapy.utils.python import without_none_values from scrapy.exceptions import UsageError @@ -26,6 +27,8 @@ class Command(ScrapyCommand): def process_options(self, args, opts): ScrapyCommand.process_options(self, args, opts) + if self.settings.getbool('ASYNCIO_SUPPORT'): + install_asyncio_reactor() try: opts.spargs = arglist_to_dict(opts.spargs) except ValueError: diff --git a/scrapy/commands/runspider.py b/scrapy/commands/runspider.py index 57d8471ca..ebd4eb620 100644 --- a/scrapy/commands/runspider.py +++ b/scrapy/commands/runspider.py @@ -2,6 +2,7 @@ import sys import os from importlib import import_module +from scrapy.utils.asyncio import install_asyncio_reactor from scrapy.utils.spider import iter_spider_classes from scrapy.commands import ScrapyCommand from scrapy.exceptions import UsageError @@ -50,6 +51,8 @@ class Command(ScrapyCommand): def process_options(self, args, opts): ScrapyCommand.process_options(self, args, opts) + if self.settings.getbool('ASYNCIO_SUPPORT'): + install_asyncio_reactor() try: opts.spargs = arglist_to_dict(opts.spargs) except ValueError: diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 5c9678c01..c9097bd1f 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -19,6 +19,8 @@ from os.path import join, abspath, dirname AJAXCRAWL_ENABLED = False +ASYNCIO_SUPPORT = False + AUTOTHROTTLE_ENABLED = False AUTOTHROTTLE_DEBUG = False AUTOTHROTTLE_MAX_DELAY = 60.0 diff --git a/scrapy/utils/asyncio.py b/scrapy/utils/asyncio.py index f732774f1..b5d5f92d9 100644 --- a/scrapy/utils/asyncio.py +++ b/scrapy/utils/asyncio.py @@ -14,7 +14,7 @@ def install_asyncio_reactor(): pass -def is_asyncio_supported(): +def is_asyncio_reactor_installed(): try: import twisted.internet.reactor from twisted.internet import asyncioreactor diff --git a/scrapy/utils/defer.py b/scrapy/utils/defer.py index 30163d2fb..3b7ef75ab 100644 --- a/scrapy/utils/defer.py +++ b/scrapy/utils/defer.py @@ -9,8 +9,7 @@ from twisted.internet import defer, reactor, task from twisted.python import failure from scrapy.exceptions import IgnoreRequest -from scrapy.utils.asyncio import is_asyncio_supported - +from scrapy.utils.asyncio import is_asyncio_reactor_installed def defer_fail(_failure): @@ -127,12 +126,18 @@ def isfuture(o): return isinstance(o, asyncio.futures.Future) -def deferred_from_coro(o): +def deferred_from_coro(o, asyncio_enabled=False): """Converts a coroutine into a Deferred, or returns the object as is if it isn't a coroutine""" if isinstance(o, defer.Deferred): return o if isfuture(o) or inspect.isawaitable(o): - if not is_asyncio_supported(): - raise TypeError('Using coroutines requires installing AsyncioSelectorReactor') - return defer.Deferred.fromFuture(asyncio.ensure_future(o)) + if not asyncio_enabled: + # wrapping the coroutine directly into a Deferred, this doesn't work correctly with coroutines + # that use asyncio, e.g. "await asyncio.sleep(1)" + return defer.ensureDeferred(o) + else: + # wrapping the coroutine into a Future and then into a Deferred, this requires AsyncioSelectorReactor + if not is_asyncio_reactor_installed(): + raise TypeError('Using coroutines requires installing AsyncioSelectorReactor') + return defer.Deferred.fromFuture(asyncio.ensure_future(o)) return o diff --git a/scrapy/utils/log.py b/scrapy/utils/log.py index b74b7a4af..8c56cfa42 100644 --- a/scrapy/utils/log.py +++ b/scrapy/utils/log.py @@ -11,7 +11,7 @@ from twisted.python import log as twisted_log import scrapy from scrapy.settings import Settings from scrapy.exceptions import ScrapyDeprecationWarning -from scrapy.utils.asyncio import is_asyncio_supported +from scrapy.utils.asyncio import is_asyncio_reactor_installed from scrapy.utils.versions import scrapy_components_versions @@ -149,8 +149,13 @@ def log_scrapy_info(settings): {'versions': ", ".join("%s %s" % (name, version) for name, version in scrapy_components_versions() if name != "Scrapy")}) - if is_asyncio_supported(): - logger.debug("Asyncio support enabled") + if settings.getbool('ASYNCIO_SUPPORT'): + if is_asyncio_reactor_installed(): + logger.debug("Asyncio support enabled") + else: + logger.error("ASYNCIO_SUPPORT is on but the Twisted asyncio " + "reactor is not installed, this is not supported " + "and asyncio coroutines will not work.") class StreamLogger(object): diff --git a/tests/test_commands.py b/tests/test_commands.py index 8aa7ee109..3b64bfa23 100644 --- a/tests/test_commands.py +++ b/tests/test_commands.py @@ -9,7 +9,6 @@ from tempfile import mkdtemp from contextlib import contextmanager from threading import Timer -from pytest import mark from twisted.trial import unittest import scrapy @@ -179,7 +178,6 @@ class MiscCommandsTest(CommandTest): self.assertEqual(0, self.call('list')) -@mark.usefixtures('reactor_pytest') class RunSpiderCommandTest(CommandTest): debug_log_spider = """ @@ -297,10 +295,13 @@ class BadSpider(scrapy.Spider): self.assertIn("start_requests", log) self.assertIn("badspider.py", log) - def test_asyncio_supported(self): - if self.reactor_pytest == 'asyncio': - log = self.get_log(self.debug_log_spider) - self.assertIn("DEBUG: Asyncio support enabled", log) + def test_asyncio_support_true(self): + log = self.get_log(self.debug_log_spider, args=['-s', 'ASYNCIO_SUPPORT=True']) + self.assertIn("DEBUG: Asyncio support enabled", log) + + def test_asyncio_support_false(self): + log = self.get_log(self.debug_log_spider, args=['-s', 'ASYNCIO_SUPPORT=False']) + self.assertNotIn("DEBUG: Asyncio support enabled", log) class BenchCommandTest(CommandTest): diff --git a/tests/test_crawler.py b/tests/test_crawler.py index 151acb459..3ac45ca1d 100644 --- a/tests/test_crawler.py +++ b/tests/test_crawler.py @@ -10,7 +10,7 @@ import scrapy from scrapy.crawler import Crawler, CrawlerRunner, CrawlerProcess from scrapy.settings import Settings, default_settings from scrapy.spiderloader import SpiderLoader -from scrapy.utils.asyncio import is_asyncio_supported +from scrapy.utils.asyncio import is_asyncio_reactor_installed from scrapy.utils.log import configure_logging, get_scrapy_root_handler from scrapy.utils.spider import DefaultSpider from scrapy.utils.misc import load_object @@ -209,7 +209,7 @@ class AsyncioSpider(scrapy.Spider): name = 'asyncio' def start_requests(self): - self.logger.info('Asyncio support: %s', is_asyncio_supported()) + self.logger.info('Asyncio support: %s', is_asyncio_reactor_installed()) return [] @@ -258,7 +258,25 @@ class CrawlerRunnerHasSpider(unittest.TestCase): self.assertEqual(runner.bootstrap_failed, True) @defer.inlineCallbacks - def test_asyncio_supported(self): + def test_crawler_process_asyncio_supported_true(self): + with LogCapture(level=logging.DEBUG) as log: + runner = CrawlerProcess(settings={'ASYNCIO_SUPPORT': True}) + yield runner.crawl(NoRequestsSpider) + if self.reactor_pytest == 'asyncio': + self.assertIn("Asyncio support enabled", str(log)) + else: + self.assertNotIn("Asyncio support enabled", str(log)) + self.assertIn("ASYNCIO_SUPPORT is on but the Twisted asyncio reactor is not installed", str(log)) + + @defer.inlineCallbacks + def test_crawler_process_asyncio_supported_false(self): + runner = CrawlerProcess(settings={'ASYNCIO_SUPPORT': False}) + with LogCapture(level=logging.DEBUG) as log: + yield runner.crawl(NoRequestsSpider) + self.assertNotIn("Asyncio support enabled", str(log)) + + @defer.inlineCallbacks + def test_crawler_runner_asyncio_supported(self): runner = CrawlerRunner() with LogCapture() as log: yield runner.crawl(AsyncioSpider) diff --git a/tests/test_utils_asyncio.py b/tests/test_utils_asyncio.py index e34d3002a..a6ba24876 100644 --- a/tests/test_utils_asyncio.py +++ b/tests/test_utils_asyncio.py @@ -2,15 +2,15 @@ from unittest import TestCase from pytest import mark -from scrapy.utils.asyncio import is_asyncio_supported, install_asyncio_reactor +from scrapy.utils.asyncio import is_asyncio_reactor_installed, install_asyncio_reactor @mark.usefixtures('reactor_pytest') class AsyncioTest(TestCase): - def test_is_asyncio_supported(self): + def test_is_asyncio_reactor_installed(self): # the result should depend only on the pytest --reactor argument - self.assertEquals(is_asyncio_supported(), self.reactor_pytest == 'asyncio') + self.assertEquals(is_asyncio_reactor_installed(), self.reactor_pytest == 'asyncio') def test_install_asyncio_reactor(self): # this should do nothing From 97fb61cec846641eb1c8e224ae24e55558746f4f Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Wed, 4 Dec 2019 21:53:07 +0500 Subject: [PATCH 321/387] Move an import to postpone another "import twisted.internet.reactor". --- scrapy/commands/shell.py | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/scrapy/commands/shell.py b/scrapy/commands/shell.py index e05084272..7516e2aba 100644 --- a/scrapy/commands/shell.py +++ b/scrapy/commands/shell.py @@ -6,7 +6,6 @@ See documentation in docs/topics/shell.rst from threading import Thread from scrapy.commands import ScrapyCommand -from scrapy.shell import Shell from scrapy.http import Request from scrapy.utils.spider import spidercls_for_request, DefaultSpider from scrapy.utils.url import guess_scheme @@ -70,6 +69,8 @@ class Command(ScrapyCommand): self._start_crawler_thread() + # moved from the top-level because it imports twisted.internet.reactor + from scrapy.shell import Shell shell = Shell(crawler, update_vars=self.update_vars, code=opts.code) shell.start(url=url, redirect=not opts.no_redirect) From 0b9f29215ff7f81203efbbc25b8e9cf3e9719920 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin Date: Wed, 4 Dec 2019 22:06:35 +0500 Subject: [PATCH 322/387] Update .travis.yml. --- .travis.yml | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/.travis.yml b/.travis.yml index fdf40fdf1..98dab01f2 100644 --- a/.travis.yml +++ b/.travis.yml @@ -17,6 +17,8 @@ matrix: python: 3.5 - env: TOXENV=py35-pinned python: 3.5 + - env: TOXENV=py35-asyncio + python: 3.5 - env: TOXENV=py36 python: 3.6 - env: TOXENV=py37 @@ -25,7 +27,7 @@ matrix: python: 3.8 - env: TOXENV=py38-extra-deps python: 3.8 - - env: TOXENV=py38-no-asyncio + - env: TOXENV=py38-asyncio python: 3.8 - env: TOXENV=docs python: 3.6 From 9b4b43f8ac3a39d5dd6137223699d3d9e7dacec2 Mon Sep 17 00:00:00 2001 From: Grammy Jiang Date: Thu, 5 Dec 2019 11:25:19 +1100 Subject: [PATCH 323/387] Convert the relative imports to absolute imports This commits converts the relative imports to absolute imports in the entire package --- scrapy/__init__.py | 2 +- scrapy/contracts/default.py | 2 +- scrapy/core/downloader/__init__.py | 4 ++-- scrapy/core/downloader/handlers/http.py | 6 ++++-- scrapy/core/downloader/handlers/s3.py | 2 +- scrapy/linkextractors/__init__.py | 2 +- scrapy/linkextractors/regex.py | 2 +- 7 files changed, 11 insertions(+), 9 deletions(-) diff --git a/scrapy/__init__.py b/scrapy/__init__.py index 230e5cee3..fb8357f3c 100644 --- a/scrapy/__init__.py +++ b/scrapy/__init__.py @@ -24,7 +24,7 @@ warnings.filterwarnings('ignore', category=DeprecationWarning, module='twisted') del warnings # Apply monkey patches to fix issues in external libraries -from . import _monkeypatches +from scrapy import _monkeypatches del _monkeypatches from twisted import version as _txv diff --git a/scrapy/contracts/default.py b/scrapy/contracts/default.py index 24f6c2e77..e0d425874 100644 --- a/scrapy/contracts/default.py +++ b/scrapy/contracts/default.py @@ -4,7 +4,7 @@ from scrapy.item import BaseItem from scrapy.http import Request from scrapy.exceptions import ContractFail -from . import Contract +from scrapy.contracts import Contract # contracts diff --git a/scrapy/core/downloader/__init__.py b/scrapy/core/downloader/__init__.py index 213268741..c5474a57f 100644 --- a/scrapy/core/downloader/__init__.py +++ b/scrapy/core/downloader/__init__.py @@ -10,8 +10,8 @@ from scrapy.utils.defer import mustbe_deferred from scrapy.utils.httpobj import urlparse_cached from scrapy.resolver import dnscache from scrapy import signals -from .middleware import DownloaderMiddlewareManager -from .handlers import DownloadHandlers +from scrapy.core.downloader.middleware import DownloaderMiddlewareManager +from scrapy.core.downloader.handlers import DownloadHandlers class Slot(object): diff --git a/scrapy/core/downloader/handlers/http.py b/scrapy/core/downloader/handlers/http.py index 6111e132a..52535bd8b 100644 --- a/scrapy/core/downloader/handlers/http.py +++ b/scrapy/core/downloader/handlers/http.py @@ -1,2 +1,4 @@ -from .http10 import HTTP10DownloadHandler -from .http11 import HTTP11DownloadHandler as HTTPDownloadHandler +from scrapy.core.downloader.handlers.http10 import HTTP10DownloadHandler +from scrapy.core.downloader.handlers.http11 import ( + HTTP11DownloadHandler as HTTPDownloadHandler, +) diff --git a/scrapy/core/downloader/handlers/s3.py b/scrapy/core/downloader/handlers/s3.py index 808d1bf21..f4a42ce12 100644 --- a/scrapy/core/downloader/handlers/s3.py +++ b/scrapy/core/downloader/handlers/s3.py @@ -3,7 +3,7 @@ from six.moves.urllib.parse import unquote from scrapy.exceptions import NotConfigured from scrapy.utils.httpobj import urlparse_cached from scrapy.utils.boto import is_botocore -from .http import HTTPDownloadHandler +from scrapy.core.downloader.handlers.http import HTTPDownloadHandler def _get_boto_connection(): diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index 8c3693f04..8411c4d59 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -118,4 +118,4 @@ class FilteringLinkExtractor(object): # Top-level imports -from .lxmlhtml import LxmlLinkExtractor as LinkExtractor # noqa: F401 +from scrapy.linkextractors.lxmlhtml import LxmlLinkExtractor as LinkExtractor # noqa: F401 diff --git a/scrapy/linkextractors/regex.py b/scrapy/linkextractors/regex.py index e689b4727..49aa2be46 100644 --- a/scrapy/linkextractors/regex.py +++ b/scrapy/linkextractors/regex.py @@ -4,7 +4,7 @@ from six.moves.urllib.parse import urljoin from w3lib.html import remove_tags, replace_entities, replace_escape_chars, get_base_url from scrapy.link import Link -from .sgml import SgmlLinkExtractor +from scrapy.linkextractors.sgml import SgmlLinkExtractor linkre = re.compile( "|\s.*?>)(.*?)<[/ ]?a>", From af624ef414ab10d833925b4d6f9048468be01273 Mon Sep 17 00:00:00 2001 From: Wang Qin <37098874+dqwerter@users.noreply.github.com> Date: Thu, 5 Dec 2019 09:29:12 +0800 Subject: [PATCH 324/387] Update overview.rst | Fix an inconsistency There exists an inconsistency between the code (line 37 - 38) and the output 'quotes.json' (line 56 - 68). Note that even though according to line 53 - 54 'quotes.json' is "reformatted here for better readability", it cannot explain why the "author" field precedes the "text" field. Intended output for the code BEFORE change: [{ "text": "\u201cThe person, be it gentleman or lady, who has not pleasure in a good novel, must be intolerably stupid.\u201d", "author": "Jane Austen" }, { "text": "\u201cOutside of a dog, a book is man's best friend. Inside of a dog it's too dark to read.\u201d", "author": "Groucho Marx" }, { "text": "\u201cA day without sunshine is like, you know, night.\u201d", "author": "Steve Martin" }, ...] Intended output for the code After change (the inconsistency is fixed): [{ "author": "Jane Austen", "text": "\u201cThe person, be it gentleman or lady, who has not pleasure in a good novel, must be intolerably stupid.\u201d" }, { "author": "Groucho Marx", "text": "\u201cOutside of a dog, a book is man's best friend. Inside of a dog it's too dark to read.\u201d" }, { "author": "Steve Martin", "text": "\u201cA day without sunshine is like, you know, night.\u201d" }, ...] --- docs/intro/overview.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/intro/overview.rst b/docs/intro/overview.rst index 8b2fef065..01986b594 100644 --- a/docs/intro/overview.rst +++ b/docs/intro/overview.rst @@ -34,8 +34,8 @@ http://quotes.toscrape.com, following the pagination:: def parse(self, response): for quote in response.css('div.quote'): yield { - 'text': quote.css('span.text::text').get(), 'author': quote.xpath('span/small/text()').get(), + 'text': quote.css('span.text::text').get(), } next_page = response.css('li.next a::attr("href")').get() From 83b8046fdcae660276634a03b99d51c629fa8b7c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= Date: Fri, 8 Nov 2019 16:26:46 +0100 Subject: [PATCH 325/387] Do not indent doctests from the documentation unnecessarily --- docs/intro/tutorial.rst | 120 +++---- docs/news.rst | 12 +- docs/topics/developer-tools.rst | 12 +- docs/topics/dynamic-content.rst | 28 +- docs/topics/items.rst | 106 +++--- docs/topics/leaks.rst | 140 ++++---- docs/topics/loaders.rst | 114 ++++--- docs/topics/selectors.rst | 565 ++++++++++++++++---------------- docs/topics/shell.rst | 77 +++-- docs/topics/stats.rst | 12 +- 10 files changed, 589 insertions(+), 597 deletions(-) diff --git a/docs/intro/tutorial.rst b/docs/intro/tutorial.rst index 6b15a5fbd..33b1a969b 100644 --- a/docs/intro/tutorial.rst +++ b/docs/intro/tutorial.rst @@ -252,30 +252,30 @@ The result of running ``response.css('title')`` is a list-like object called and allow you to run further queries to fine-grain the selection or extract the data. -To extract the text from the title above, you can do:: +To extract the text from the title above, you can do: - >>> response.css('title::text').getall() - ['Quotes to Scrape'] +>>> response.css('title::text').getall() +['Quotes to Scrape'] There are two things to note here: one is that we've added ``::text`` to the CSS query, to mean we want to select only the text elements directly inside ```` element. If we don't specify ``::text``, we'd get the full title -element, including its tags:: +element, including its tags: - >>> response.css('title').getall() - ['<title>Quotes to Scrape'] +>>> response.css('title').getall() +['Quotes to Scrape'] The other thing is that the result of calling ``.getall()`` is a list: it is possible that a selector returns more than one result, so we extract them all. -When you know you just want the first result, as in this case, you can do:: +When you know you just want the first result, as in this case, you can do: - >>> response.css('title::text').get() - 'Quotes to Scrape' +>>> response.css('title::text').get() +'Quotes to Scrape' -As an alternative, you could've written:: +As an alternative, you could've written: - >>> response.css('title::text')[0].get() - 'Quotes to Scrape' +>>> response.css('title::text')[0].get() +'Quotes to Scrape' However, using ``.get()`` directly on a :class:`~scrapy.selector.SelectorList` instance avoids an ``IndexError`` and returns ``None`` when it doesn't @@ -288,14 +288,14 @@ to be scraped, you can at least get **some** data. Besides the :meth:`~scrapy.selector.SelectorList.getall` and :meth:`~scrapy.selector.SelectorList.get` methods, you can also use the :meth:`~scrapy.selector.SelectorList.re` method to extract using `regular -expressions`_:: +expressions`_: - >>> response.css('title::text').re(r'Quotes.*') - ['Quotes to Scrape'] - >>> response.css('title::text').re(r'Q\w+') - ['Quotes'] - >>> response.css('title::text').re(r'(\w+) to (\w+)') - ['Quotes', 'Scrape'] +>>> response.css('title::text').re(r'Quotes.*') +['Quotes to Scrape'] +>>> response.css('title::text').re(r'Q\w+') +['Quotes'] +>>> response.css('title::text').re(r'(\w+) to (\w+)') +['Quotes', 'Scrape'] In order to find the proper CSS selectors to use, you might find useful opening the response page from the shell in your web browser using ``view(response)``. @@ -312,12 +312,12 @@ visually selected elements, which works in many browsers. XPath: a brief intro ^^^^^^^^^^^^^^^^^^^^ -Besides `CSS`_, Scrapy selectors also support using `XPath`_ expressions:: +Besides `CSS`_, Scrapy selectors also support using `XPath`_ expressions: - >>> response.xpath('//title') - [] - >>> response.xpath('//title/text()').get() - 'Quotes to Scrape' +>>> response.xpath('//title') +[] +>>> response.xpath('//title/text()').get() +'Quotes to Scrape' XPath expressions are very powerful, and are the foundation of Scrapy Selectors. In fact, CSS selectors are converted to XPath under-the-hood. You @@ -372,35 +372,35 @@ we want:: $ scrapy shell 'http://quotes.toscrape.com' -We get a list of selectors for the quote HTML elements with:: +We get a list of selectors for the quote HTML elements with: - >>> response.css("div.quote") - [, - , - ...] +>>> response.css("div.quote") +[, + , + ...] Each of the selectors returned by the query above allows us to run further queries over their sub-elements. Let's assign the first selector to a -variable, so that we can run our CSS selectors directly on a particular quote:: +variable, so that we can run our CSS selectors directly on a particular quote: - >>> quote = response.css("div.quote")[0] +>>> quote = response.css("div.quote")[0] Now, let's extract ``text``, ``author`` and the ``tags`` from that quote -using the ``quote`` object we just created:: +using the ``quote`` object we just created: - >>> text = quote.css("span.text::text").get() - >>> text - '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”' - >>> author = quote.css("small.author::text").get() - >>> author - 'Albert Einstein' +>>> text = quote.css("span.text::text").get() +>>> text +'“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”' +>>> author = quote.css("small.author::text").get() +>>> author +'Albert Einstein' Given that the tags are a list of strings, we can use the ``.getall()`` method -to get all of them:: +to get all of them: - >>> tags = quote.css("div.tags a.tag::text").getall() - >>> tags - ['change', 'deep-thoughts', 'thinking', 'world'] +>>> tags = quote.css("div.tags a.tag::text").getall() +>>> tags +['change', 'deep-thoughts', 'thinking', 'world'] .. invisible-code-block: python @@ -409,16 +409,16 @@ to get all of them:: .. skip: next if(version_info < (3, 6), reason="Only Python 3.6+ dictionaries match the output") Having figured out how to extract each bit, we can now iterate over all the -quotes elements and put them together into a Python dictionary:: +quotes elements and put them together into a Python dictionary: - >>> for quote in response.css("div.quote"): - ... text = quote.css("span.text::text").get() - ... author = quote.css("small.author::text").get() - ... tags = quote.css("div.tags a.tag::text").getall() - ... print(dict(text=text, author=author, tags=tags)) - {'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein', 'tags': ['change', 'deep-thoughts', 'thinking', 'world']} - {'text': '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', 'author': 'J.K. Rowling', 'tags': ['abilities', 'choices']} - ... +>>> for quote in response.css("div.quote"): +... text = quote.css("span.text::text").get() +... author = quote.css("small.author::text").get() +... tags = quote.css("div.tags a.tag::text").getall() +... print(dict(text=text, author=author, tags=tags)) +{'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein', 'tags': ['change', 'deep-thoughts', 'thinking', 'world']} +{'text': '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', 'author': 'J.K. Rowling', 'tags': ['abilities', 'choices']} +... Extracting data in our spider ----------------------------- @@ -516,23 +516,23 @@ markup: -We can try extracting it in the shell:: +We can try extracting it in the shell: - >>> response.css('li.next a').get() - 'Next ' +>>> response.css('li.next a').get() +'Next ' This gets the anchor element, but we want the attribute ``href``. For that, Scrapy supports a CSS extension that lets you select the attribute contents, -like this:: +like this: - >>> response.css('li.next a::attr(href)').get() - '/page/2/' +>>> response.css('li.next a::attr(href)').get() +'/page/2/' There is also an ``attrib`` property available -(see :ref:`selecting-attributes` for more):: +(see :ref:`selecting-attributes` for more): - >>> response.css('li.next a').attrib['href'] - '/page/2/' +>>> response.css('li.next a').attrib['href'] +'/page/2/' Let's see now our spider modified to recursively follow the link to the next page, extracting data from it:: diff --git a/docs/news.rst b/docs/news.rst index 9dfd28508..217382c57 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -47,13 +47,13 @@ Backward-incompatible changes (:issue:`2111`, :issue:`3392`, :issue:`3442`, :issue:`3450`) * :class:`~scrapy.loader.ItemLoader` now turns the values of its input item - into lists:: + into lists: - >>> item = MyItem() - >>> item['field'] = 'value1' - >>> loader = ItemLoader(item=item) - >>> item['field'] - ['value1'] + >>> item = MyItem() + >>> item['field'] = 'value1' + >>> loader = ItemLoader(item=item) + >>> item['field'] + ['value1'] This is needed to allow adding values to existing fields (``loader.add_value('field', 'value2')``). diff --git a/docs/topics/developer-tools.rst b/docs/topics/developer-tools.rst index bf14643be..1c9315cd8 100644 --- a/docs/topics/developer-tools.rst +++ b/docs/topics/developer-tools.rst @@ -112,13 +112,13 @@ see each quote: With this knowledge we can refine our XPath: Instead of a path to follow, we'll simply select all ``span`` tags with the ``class="text"`` by using -the `has-class-extension`_:: +the `has-class-extension`_: - >>> response.xpath('//span[has-class("text")]/text()').getall() - ['"The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”, - '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', - '“There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.”', - (...)] +>>> response.xpath('//span[has-class("text")]/text()').getall() +['"The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', + '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', + '“There are only two ways to live your life. One is as though nothing is a miracle. The other is as though everything is a miracle.”', + ...] And with one simple, cleverer XPath we are able to extract all quotes from the page. We could have constructed a loop over our first XPath to increase diff --git a/docs/topics/dynamic-content.rst b/docs/topics/dynamic-content.rst index 8334ddcec..1c3607860 100644 --- a/docs/topics/dynamic-content.rst +++ b/docs/topics/dynamic-content.rst @@ -172,27 +172,27 @@ data from it: data in JSON format, which you can then parse with `json.loads`_. For example, if the JavaScript code contains a separate line like - ``var data = {"field": "value"};`` you can extract that data as follows:: + ``var data = {"field": "value"};`` you can extract that data as follows: - >>> pattern = r'\bvar\s+data\s*=\s*(\{.*?\})\s*;\s*\n' - >>> json_data = response.css('script::text').re_first(pattern) - >>> json.loads(json_data) - {'field': 'value'} + >>> pattern = r'\bvar\s+data\s*=\s*(\{.*?\})\s*;\s*\n' + >>> json_data = response.css('script::text').re_first(pattern) + >>> json.loads(json_data) + {'field': 'value'} - Otherwise, use js2xml_ to convert the JavaScript code into an XML document that you can parse using :ref:`selectors `. For example, if the JavaScript code contains - ``var data = {field: "value"};`` you can extract that data as follows:: + ``var data = {field: "value"};`` you can extract that data as follows: - >>> import js2xml - >>> import lxml.etree - >>> from parsel import Selector - >>> javascript = response.css('script::text').get() - >>> xml = lxml.etree.tostring(js2xml.parse(javascript), encoding='unicode') - >>> selector = Selector(text=xml) - >>> selector.css('var[name="data"]').get() - 'value' + >>> import js2xml + >>> import lxml.etree + >>> from parsel import Selector + >>> javascript = response.css('script::text').get() + >>> xml = lxml.etree.tostring(js2xml.parse(javascript), encoding='unicode') + >>> selector = Selector(text=xml) + >>> selector.css('var[name="data"]').get() + 'value' .. _topics-javascript-rendering: diff --git a/docs/topics/items.rst b/docs/topics/items.rst index cdf60208e..15313775b 100644 --- a/docs/topics/items.rst +++ b/docs/topics/items.rst @@ -84,77 +84,74 @@ notice the API is very similar to the `dict API`_. Creating items -------------- -:: +>>> product = Product(name='Desktop PC', price=1000) +>>> print(product) +Product(name='Desktop PC', price=1000) - >>> product = Product(name='Desktop PC', price=1000) - >>> print(product) - Product(name='Desktop PC', price=1000) Getting field values -------------------- -:: +>>> product['name'] +Desktop PC +>>> product.get('name') +Desktop PC - >>> product['name'] - Desktop PC - >>> product.get('name') - Desktop PC +>>> product['price'] +1000 - >>> product['price'] - 1000 +>>> product['last_updated'] +Traceback (most recent call last): + ... +KeyError: 'last_updated' - >>> product['last_updated'] - Traceback (most recent call last): - ... - KeyError: 'last_updated' +>>> product.get('last_updated', 'not set') +not set - >>> product.get('last_updated', 'not set') - not set +>>> product['lala'] # getting unknown field +Traceback (most recent call last): + ... +KeyError: 'lala' - >>> product['lala'] # getting unknown field - Traceback (most recent call last): - ... - KeyError: 'lala' +>>> product.get('lala', 'unknown field') +'unknown field' - >>> product.get('lala', 'unknown field') - 'unknown field' +>>> 'name' in product # is name field populated? +True - >>> 'name' in product # is name field populated? - True +>>> 'last_updated' in product # is last_updated populated? +False - >>> 'last_updated' in product # is last_updated populated? - False +>>> 'last_updated' in product.fields # is last_updated a declared field? +True - >>> 'last_updated' in product.fields # is last_updated a declared field? - True +>>> 'lala' in product.fields # is lala a declared field? +False - >>> 'lala' in product.fields # is lala a declared field? - False Setting field values -------------------- -:: +>>> product['last_updated'] = 'today' +>>> product['last_updated'] +today - >>> product['last_updated'] = 'today' - >>> product['last_updated'] - today +>>> product['lala'] = 'test' # setting unknown field +Traceback (most recent call last): + ... +KeyError: 'Product does not support field: lala' - >>> product['lala'] = 'test' # setting unknown field - Traceback (most recent call last): - ... - KeyError: 'Product does not support field: lala' Accessing all populated values ------------------------------ -To access all populated values, just use the typical `dict API`_:: +To access all populated values, just use the typical `dict API`_: - >>> product.keys() - ['price', 'name'] +>>> product.keys() +['price', 'name'] - >>> product.items() - [('price', 1000), ('name', 'Desktop PC')] +>>> product.items() +[('price', 1000), ('name', 'Desktop PC')] .. _copying-items: @@ -194,20 +191,21 @@ To create a deep copy, call :meth:`~scrapy.item.Item.deepcopy` instead Other common tasks ------------------ -Creating dicts from items:: +Creating dicts from items: - >>> dict(product) # create a dict from all populated values - {'price': 1000, 'name': 'Desktop PC'} +>>> dict(product) # create a dict from all populated values +{'price': 1000, 'name': 'Desktop PC'} -Creating items from dicts:: +Creating items from dicts: - >>> Product({'name': 'Laptop PC', 'price': 1500}) - Product(price=1500, name='Laptop PC') +>>> Product({'name': 'Laptop PC', 'price': 1500}) +Product(price=1500, name='Laptop PC') + +>>> Product({'name': 'Laptop PC', 'lala': 1500}) # warning: unknown field in dict +Traceback (most recent call last): + ... +KeyError: 'Product does not support field: lala' - >>> Product({'name': 'Laptop PC', 'lala': 1500}) # warning: unknown field in dict - Traceback (most recent call last): - ... - KeyError: 'Product does not support field: lala' Extending Items =============== diff --git a/docs/topics/leaks.rst b/docs/topics/leaks.rst index 793636f59..87d9d262f 100644 --- a/docs/topics/leaks.rst +++ b/docs/topics/leaks.rst @@ -132,21 +132,21 @@ and check the code of the spider to discover the nasty line that is generating the leaks (passing response references inside requests). Sometimes extra information about live objects can be helpful. -Let's check the oldest response:: +Let's check the oldest response: - >>> from scrapy.utils.trackref import get_oldest - >>> r = get_oldest('HtmlResponse') - >>> r.url - 'http://www.somenastyspider.com/product.php?pid=123' +>>> from scrapy.utils.trackref import get_oldest +>>> r = get_oldest('HtmlResponse') +>>> r.url +'http://www.somenastyspider.com/product.php?pid=123' If you want to iterate over all objects, instead of getting the oldest one, you -can use the :func:`scrapy.utils.trackref.iter_all` function:: +can use the :func:`scrapy.utils.trackref.iter_all` function: - >>> from scrapy.utils.trackref import iter_all - >>> [r.url for r in iter_all('HtmlResponse')] - ['http://www.somenastyspider.com/product.php?pid=123', - 'http://www.somenastyspider.com/product.php?pid=584', - ... +>>> from scrapy.utils.trackref import iter_all +>>> [r.url for r in iter_all('HtmlResponse')] +['http://www.somenastyspider.com/product.php?pid=123', + 'http://www.somenastyspider.com/product.php?pid=584', +...] Too many spiders? ----------------- @@ -155,10 +155,10 @@ If your project has too many spiders executed in parallel, the output of :func:`prefs()` can be difficult to read. For this reason, that function has a ``ignore`` argument which can be used to ignore a particular class (and all its subclases). For -example, this won't show any live references to spiders:: +example, this won't show any live references to spiders: - >>> from scrapy.spiders import Spider - >>> prefs(ignore=Spider) +>>> from scrapy.spiders import Spider +>>> prefs(ignore=Spider) .. module:: scrapy.utils.trackref :synopsis: Track references of live objects @@ -214,41 +214,41 @@ If you use ``pip``, you can install Guppy with the following command:: The telnet console also comes with a built-in shortcut (``hpy``) for accessing Guppy heap objects. Here's an example to view all Python objects available in -the heap using Guppy:: +the heap using Guppy: - >>> x = hpy.heap() - >>> x.bytype - Partition of a set of 297033 objects. Total size = 52587824 bytes. - Index Count % Size % Cumulative % Type - 0 22307 8 16423880 31 16423880 31 dict - 1 122285 41 12441544 24 28865424 55 str - 2 68346 23 5966696 11 34832120 66 tuple - 3 227 0 5836528 11 40668648 77 unicode - 4 2461 1 2222272 4 42890920 82 type - 5 16870 6 2024400 4 44915320 85 function - 6 13949 5 1673880 3 46589200 89 types.CodeType - 7 13422 5 1653104 3 48242304 92 list - 8 3735 1 1173680 2 49415984 94 _sre.SRE_Pattern - 9 1209 0 456936 1 49872920 95 scrapy.http.headers.Headers - <1676 more rows. Type e.g. '_.more' to view.> +>>> x = hpy.heap() +>>> x.bytype +Partition of a set of 297033 objects. Total size = 52587824 bytes. + Index Count % Size % Cumulative % Type + 0 22307 8 16423880 31 16423880 31 dict + 1 122285 41 12441544 24 28865424 55 str + 2 68346 23 5966696 11 34832120 66 tuple + 3 227 0 5836528 11 40668648 77 unicode + 4 2461 1 2222272 4 42890920 82 type + 5 16870 6 2024400 4 44915320 85 function + 6 13949 5 1673880 3 46589200 89 types.CodeType + 7 13422 5 1653104 3 48242304 92 list + 8 3735 1 1173680 2 49415984 94 _sre.SRE_Pattern + 9 1209 0 456936 1 49872920 95 scrapy.http.headers.Headers +<1676 more rows. Type e.g. '_.more' to view.> You can see that most space is used by dicts. Then, if you want to see from -which attribute those dicts are referenced, you could do:: +which attribute those dicts are referenced, you could do: - >>> x.bytype[0].byvia - Partition of a set of 22307 objects. Total size = 16423880 bytes. - Index Count % Size % Cumulative % Referred Via: - 0 10982 49 9416336 57 9416336 57 '.__dict__' - 1 1820 8 2681504 16 12097840 74 '.__dict__', '.func_globals' - 2 3097 14 1122904 7 13220744 80 - 3 990 4 277200 2 13497944 82 "['cookies']" - 4 987 4 276360 2 13774304 84 "['cache']" - 5 985 4 275800 2 14050104 86 "['meta']" - 6 897 4 251160 2 14301264 87 '[2]' - 7 1 0 196888 1 14498152 88 "['moduleDict']", "['modules']" - 8 672 3 188160 1 14686312 89 "['cb_kwargs']" - 9 27 0 155016 1 14841328 90 '[1]' - <333 more rows. Type e.g. '_.more' to view.> +>>> x.bytype[0].byvia +Partition of a set of 22307 objects. Total size = 16423880 bytes. + Index Count % Size % Cumulative % Referred Via: + 0 10982 49 9416336 57 9416336 57 '.__dict__' + 1 1820 8 2681504 16 12097840 74 '.__dict__', '.func_globals' + 2 3097 14 1122904 7 13220744 80 + 3 990 4 277200 2 13497944 82 "['cookies']" + 4 987 4 276360 2 13774304 84 "['cache']" + 5 985 4 275800 2 14050104 86 "['meta']" + 6 897 4 251160 2 14301264 87 '[2]' + 7 1 0 196888 1 14498152 88 "['moduleDict']", "['modules']" + 8 672 3 188160 1 14686312 89 "['cb_kwargs']" + 9 27 0 155016 1 14841328 90 '[1]' +<333 more rows. Type e.g. '_.more' to view.> As you can see, the Guppy module is very powerful but also requires some deep knowledge about Python internals. For more info about Guppy, refer to the @@ -269,32 +269,32 @@ If you use ``pip``, you can install muppy with the following command:: pip install Pympler Here's an example to view all Python objects available in -the heap using muppy:: +the heap using muppy: - >>> from pympler import muppy - >>> all_objects = muppy.get_objects() - >>> len(all_objects) - 28667 - >>> from pympler import summary - >>> suml = summary.summarize(all_objects) - >>> summary.print_(suml) - types | # objects | total size - ==================================== | =========== | ============ - >> from pympler import muppy +>>> all_objects = muppy.get_objects() +>>> len(all_objects) +28667 +>>> from pympler import summary +>>> suml = summary.summarize(all_objects) +>>> summary.print_(suml) + types | # objects | total size +==================================== | =========== | ============ + >> from scrapy.loader import ItemLoader - >>> il = ItemLoader(item=Product()) - >>> il.add_value('name', [u'Welcome to my', u'website']) - >>> il.add_value('price', [u'€', u'1000']) - >>> il.load_item() - {'name': u'Welcome to my website', 'price': u'1000'} +>>> from scrapy.loader import ItemLoader +>>> il = ItemLoader(item=Product()) +>>> il.add_value('name', [u'Welcome to my', u'website']) +>>> il.add_value('price', [u'€', u'1000']) +>>> il.load_item() +{'name': u'Welcome to my website', 'price': u'1000'} The precedence order, for both input and output processors, is as follows: @@ -314,11 +312,11 @@ ItemLoader objects applied before processors :type re: str or compiled regex - Examples:: + Examples: - >>> from scrapy.loader.processors import TakeFirst - >>> loader.get_value(u'name: foo', TakeFirst(), unicode.upper, re='name: (.+)') - 'FOO` + >>> from scrapy.loader.processors import TakeFirst + >>> loader.get_value(u'name: foo', TakeFirst(), unicode.upper, re='name: (.+)') + 'FOO` .. method:: add_value(field_name, value, \*processors, \**kwargs) @@ -639,12 +637,12 @@ Here is a list of all built-in processors: values unchanged. It doesn't receive any ``__init__`` method arguments, nor does it accept Loader contexts. - Example:: + Example: - >>> from scrapy.loader.processors import Identity - >>> proc = Identity() - >>> proc(['one', 'two', 'three']) - ['one', 'two', 'three'] + >>> from scrapy.loader.processors import Identity + >>> proc = Identity() + >>> proc(['one', 'two', 'three']) + ['one', 'two', 'three'] .. class:: TakeFirst @@ -652,12 +650,12 @@ Here is a list of all built-in processors: so it's typically used as an output processor to single-valued fields. It doesn't receive any ``__init__`` method arguments, nor does it accept Loader contexts. - Example:: + Example: - >>> from scrapy.loader.processors import TakeFirst - >>> proc = TakeFirst() - >>> proc(['', 'one', 'two', 'three']) - 'one' + >>> from scrapy.loader.processors import TakeFirst + >>> proc = TakeFirst() + >>> proc(['', 'one', 'two', 'three']) + 'one' .. class:: Join(separator=u' ') @@ -667,15 +665,15 @@ Here is a list of all built-in processors: When using the default separator, this processor is equivalent to the function: ``u' '.join`` - Examples:: + Examples: - >>> from scrapy.loader.processors import Join - >>> proc = Join() - >>> proc(['one', 'two', 'three']) - 'one two three' - >>> proc = Join('
') - >>> proc(['one', 'two', 'three']) - 'one
two
three' + >>> from scrapy.loader.processors import Join + >>> proc = Join() + >>> proc(['one', 'two', 'three']) + 'one two three' + >>> proc = Join('
') + >>> proc(['one', 'two', 'three']) + 'one
two
three' .. class:: Compose(\*functions, \**default_loader_context) @@ -688,12 +686,12 @@ Here is a list of all built-in processors: By default, stop process on ``None`` value. This behaviour can be changed by passing keyword argument ``stop_on_none=False``. - Example:: + Example: - >>> from scrapy.loader.processors import Compose - >>> proc = Compose(lambda v: v[0], str.upper) - >>> proc(['hello', 'world']) - 'HELLO' + >>> from scrapy.loader.processors import Compose + >>> proc = Compose(lambda v: v[0], str.upper) + >>> proc(['hello', 'world']) + 'HELLO' Each function can optionally receive a ``loader_context`` parameter. For those which do, this processor will pass the currently active :ref:`Loader @@ -732,15 +730,15 @@ Here is a list of all built-in processors: :meth:`~scrapy.selector.Selector.extract` method of :ref:`selectors `, which returns a list of unicode strings. - The example below should clarify how it works:: + The example below should clarify how it works: - >>> def filter_world(x): - ... return None if x == 'world' else x - ... - >>> from scrapy.loader.processors import MapCompose - >>> proc = MapCompose(filter_world, str.upper) - >>> proc(['hello', 'world', 'this', 'is', 'scrapy']) - ['HELLO, 'THIS', 'IS', 'SCRAPY'] + >>> def filter_world(x): + ... return None if x == 'world' else x + ... + >>> from scrapy.loader.processors import MapCompose + >>> proc = MapCompose(filter_world, str.upper) + >>> proc(['hello', 'world', 'this', 'is', 'scrapy']) + ['HELLO, 'THIS', 'IS', 'SCRAPY'] As with the Compose processor, functions can receive Loader contexts, and ``__init__`` method keyword arguments are used as default context values. See @@ -752,21 +750,21 @@ Here is a list of all built-in processors: Requires jmespath (https://github.com/jmespath/jmespath.py) to run. This processor takes only one input at a time. - Example:: + Example: - >>> from scrapy.loader.processors import SelectJmes, Compose, MapCompose - >>> proc = SelectJmes("foo") #for direct use on lists and dictionaries - >>> proc({'foo': 'bar'}) - 'bar' - >>> proc({'foo': {'bar': 'baz'}}) - {'bar': 'baz'} + >>> from scrapy.loader.processors import SelectJmes, Compose, MapCompose + >>> proc = SelectJmes("foo") #for direct use on lists and dictionaries + >>> proc({'foo': 'bar'}) + 'bar' + >>> proc({'foo': {'bar': 'baz'}}) + {'bar': 'baz'} - Working with Json:: + Working with Json: - >>> import json - >>> proc_single_json_str = Compose(json.loads, SelectJmes("foo")) - >>> proc_single_json_str('{"foo": "bar"}') - 'bar' - >>> proc_json_list = Compose(json.loads, MapCompose(SelectJmes('foo'))) - >>> proc_json_list('[{"foo":"bar"}, {"baz":"tar"}]') - ['bar'] + >>> import json + >>> proc_single_json_str = Compose(json.loads, SelectJmes("foo")) + >>> proc_single_json_str('{"foo": "bar"}') + 'bar' + >>> proc_json_list = Compose(json.loads, MapCompose(SelectJmes('foo'))) + >>> proc_json_list('[{"foo":"bar"}, {"baz":"tar"}]') + ['bar'] diff --git a/docs/topics/selectors.rst b/docs/topics/selectors.rst index 282a585d4..8ec758b0e 100644 --- a/docs/topics/selectors.rst +++ b/docs/topics/selectors.rst @@ -51,18 +51,18 @@ Constructing selectors .. highlight:: python Response objects expose a :class:`~scrapy.selector.Selector` instance -on ``.selector`` attribute:: +on ``.selector`` attribute: - >>> response.selector.xpath('//span/text()').get() - 'good' +>>> response.selector.xpath('//span/text()').get() +'good' Querying responses using XPath and CSS is so common that responses include two -more shortcuts: ``response.xpath()`` and ``response.css()``:: +more shortcuts: ``response.xpath()`` and ``response.css()``: - >>> response.xpath('//span/text()').get() - 'good' - >>> response.css('span::text').get() - 'good' +>>> response.xpath('//span/text()').get() +'good' +>>> response.css('span::text').get() +'good' Scrapy selectors are instances of :class:`~scrapy.selector.Selector` class constructed by passing either :class:`~scrapy.http.TextResponse` object or @@ -74,21 +74,21 @@ shortcuts. By using ``response.selector`` or one of these shortcuts you can also ensure the response body is parsed only once. But if required, it is possible to use ``Selector`` directly. -Constructing from text:: +Constructing from text: - >>> from scrapy.selector import Selector - >>> body = 'good' - >>> Selector(text=body).xpath('//span/text()').get() - 'good' +>>> from scrapy.selector import Selector +>>> body = 'good' +>>> Selector(text=body).xpath('//span/text()').get() +'good' Constructing from response - :class:`~scrapy.http.HtmlResponse` is one of -:class:`~scrapy.http.TextResponse` subclasses:: +:class:`~scrapy.http.TextResponse` subclasses: - >>> from scrapy.selector import Selector - >>> from scrapy.http import HtmlResponse - >>> response = HtmlResponse(url='http://example.com', body=body) - >>> Selector(response=response).xpath('//span/text()').get() - 'good' +>>> from scrapy.selector import Selector +>>> from scrapy.http import HtmlResponse +>>> response = HtmlResponse(url='http://example.com', body=body) +>>> Selector(response=response).xpath('//span/text()').get() +'good' ``Selector`` automatically chooses the best parsing rules (XML vs HTML) based on input type. @@ -123,118 +123,118 @@ Since we're dealing with HTML, the selector will automatically use an HTML parse .. highlight:: python So, by looking at the :ref:`HTML code ` of that -page, let's construct an XPath for selecting the text inside the title tag:: +page, let's construct an XPath for selecting the text inside the title tag: - >>> response.xpath('//title/text()') - [] +>>> response.xpath('//title/text()') +[] To actually extract the textual data, you must call the selector ``.get()`` -or ``.getall()`` methods, as follows:: +or ``.getall()`` methods, as follows: - >>> response.xpath('//title/text()').getall() - ['Example website'] - >>> response.xpath('//title/text()').get() - 'Example website' +>>> response.xpath('//title/text()').getall() +['Example website'] +>>> response.xpath('//title/text()').get() +'Example website' ``.get()`` always returns a single result; if there are several matches, content of a first match is returned; if there are no matches, None is returned. ``.getall()`` returns a list with all results. Notice that CSS selectors can select text or attribute nodes using CSS3 -pseudo-elements:: +pseudo-elements: - >>> response.css('title::text').get() - 'Example website' +>>> response.css('title::text').get() +'Example website' As you can see, ``.xpath()`` and ``.css()`` methods return a :class:`~scrapy.selector.SelectorList` instance, which is a list of new -selectors. This API can be used for quickly selecting nested data:: +selectors. This API can be used for quickly selecting nested data: - >>> response.css('img').xpath('@src').getall() - ['image1_thumb.jpg', - 'image2_thumb.jpg', - 'image3_thumb.jpg', - 'image4_thumb.jpg', - 'image5_thumb.jpg'] +>>> response.css('img').xpath('@src').getall() +['image1_thumb.jpg', + 'image2_thumb.jpg', + 'image3_thumb.jpg', + 'image4_thumb.jpg', + 'image5_thumb.jpg'] If you want to extract only the first matched element, you can call the selector ``.get()`` (or its alias ``.extract_first()`` commonly used in -previous Scrapy versions):: +previous Scrapy versions): - >>> response.xpath('//div[@id="images"]/a/text()').get() - 'Name: My image 1 ' +>>> response.xpath('//div[@id="images"]/a/text()').get() +'Name: My image 1 ' -It returns ``None`` if no element was found:: +It returns ``None`` if no element was found: - >>> response.xpath('//div[@id="not-exists"]/text()').get() is None - True +>>> response.xpath('//div[@id="not-exists"]/text()').get() is None +True A default return value can be provided as an argument, to be used instead of ``None``: - >>> response.xpath('//div[@id="not-exists"]/text()').get(default='not-found') - 'not-found' +>>> response.xpath('//div[@id="not-exists"]/text()').get(default='not-found') +'not-found' Instead of using e.g. ``'@src'`` XPath it is possible to query for attributes -using ``.attrib`` property of a :class:`~scrapy.selector.Selector`:: +using ``.attrib`` property of a :class:`~scrapy.selector.Selector`: - >>> [img.attrib['src'] for img in response.css('img')] - ['image1_thumb.jpg', - 'image2_thumb.jpg', - 'image3_thumb.jpg', - 'image4_thumb.jpg', - 'image5_thumb.jpg'] +>>> [img.attrib['src'] for img in response.css('img')] +['image1_thumb.jpg', + 'image2_thumb.jpg', + 'image3_thumb.jpg', + 'image4_thumb.jpg', + 'image5_thumb.jpg'] As a shortcut, ``.attrib`` is also available on SelectorList directly; -it returns attributes for the first matching element:: +it returns attributes for the first matching element: - >>> response.css('img').attrib['src'] - 'image1_thumb.jpg' +>>> response.css('img').attrib['src'] +'image1_thumb.jpg' This is most useful when only a single result is expected, e.g. when selecting -by id, or selecting unique elements on a web page:: +by id, or selecting unique elements on a web page: - >>> response.css('base').attrib['href'] - 'http://example.com/' +>>> response.css('base').attrib['href'] +'http://example.com/' -Now we're going to get the base URL and some image links:: +Now we're going to get the base URL and some image links: - >>> response.xpath('//base/@href').get() - 'http://example.com/' +>>> response.xpath('//base/@href').get() +'http://example.com/' - >>> response.css('base::attr(href)').get() - 'http://example.com/' +>>> response.css('base::attr(href)').get() +'http://example.com/' - >>> response.css('base').attrib['href'] - 'http://example.com/' +>>> response.css('base').attrib['href'] +'http://example.com/' - >>> response.xpath('//a[contains(@href, "image")]/@href').getall() - ['image1.html', - 'image2.html', - 'image3.html', - 'image4.html', - 'image5.html'] +>>> response.xpath('//a[contains(@href, "image")]/@href').getall() +['image1.html', + 'image2.html', + 'image3.html', + 'image4.html', + 'image5.html'] - >>> response.css('a[href*=image]::attr(href)').getall() - ['image1.html', - 'image2.html', - 'image3.html', - 'image4.html', - 'image5.html'] +>>> response.css('a[href*=image]::attr(href)').getall() +['image1.html', + 'image2.html', + 'image3.html', + 'image4.html', + 'image5.html'] - >>> response.xpath('//a[contains(@href, "image")]/img/@src').getall() - ['image1_thumb.jpg', - 'image2_thumb.jpg', - 'image3_thumb.jpg', - 'image4_thumb.jpg', - 'image5_thumb.jpg'] +>>> response.xpath('//a[contains(@href, "image")]/img/@src').getall() +['image1_thumb.jpg', + 'image2_thumb.jpg', + 'image3_thumb.jpg', + 'image4_thumb.jpg', + 'image5_thumb.jpg'] - >>> response.css('a[href*=image] img::attr(src)').getall() - ['image1_thumb.jpg', - 'image2_thumb.jpg', - 'image3_thumb.jpg', - 'image4_thumb.jpg', - 'image5_thumb.jpg'] +>>> response.css('a[href*=image] img::attr(src)').getall() +['image1_thumb.jpg', + 'image2_thumb.jpg', + 'image3_thumb.jpg', + 'image4_thumb.jpg', + 'image5_thumb.jpg'] .. _topics-selectors-css-extensions: @@ -259,47 +259,47 @@ that Scrapy (parsel) implements a couple of **non-standard pseudo-elements**: Examples: -* ``title::text`` selects children text nodes of a descendant ```` element:: +* ``title::text`` selects children text nodes of a descendant ``<title>`` element: - >>> response.css('title::text').get() - 'Example website' +>>> response.css('title::text').get() +'Example website' -* ``*::text`` selects all descendant text nodes of the current selector context:: +* ``*::text`` selects all descendant text nodes of the current selector context: - >>> response.css('#images *::text').getall() - ['\n ', - 'Name: My image 1 ', - '\n ', - 'Name: My image 2 ', - '\n ', - 'Name: My image 3 ', - '\n ', - 'Name: My image 4 ', - '\n ', - 'Name: My image 5 ', - '\n '] +>>> response.css('#images *::text').getall() +['\n ', + 'Name: My image 1 ', + '\n ', + 'Name: My image 2 ', + '\n ', + 'Name: My image 3 ', + '\n ', + 'Name: My image 4 ', + '\n ', + 'Name: My image 5 ', + '\n '] * ``foo::text`` returns no results if ``foo`` element exists, but contains - no text (i.e. text is empty):: + no text (i.e. text is empty): - >>> response.css('img::text').getall() - [] +>>> response.css('img::text').getall() +[] This means ``.css('foo::text').get()`` could return None even if an element - exists. Use ``default=''`` if you always want a string:: + exists. Use ``default=''`` if you always want a string: - >>> response.css('img::text').get() - >>> response.css('img::text').get(default='') - '' +>>> response.css('img::text').get() +>>> response.css('img::text').get(default='') +'' -* ``a::attr(href)`` selects the *href* attribute value of descendant links:: +* ``a::attr(href)`` selects the *href* attribute value of descendant links: - >>> response.css('a::attr(href)').getall() - ['image1.html', - 'image2.html', - 'image3.html', - 'image4.html', - 'image5.html'] +>>> response.css('a::attr(href)').getall() +['image1.html', + 'image2.html', + 'image3.html', + 'image4.html', + 'image5.html'] .. note:: See also: :ref:`selecting-attributes`. @@ -318,25 +318,24 @@ Nesting selectors The selection methods (``.xpath()`` or ``.css()``) return a list of selectors of the same type, so you can call the selection methods for those selectors -too. Here's an example:: +too. Here's an example: - >>> links = response.xpath('//a[contains(@href, "image")]') - >>> links.getall() - ['<a href="image1.html">Name: My image 1 <br><img src="image1_thumb.jpg"></a>', - '<a href="image2.html">Name: My image 2 <br><img src="image2_thumb.jpg"></a>', - '<a href="image3.html">Name: My image 3 <br><img src="image3_thumb.jpg"></a>', - '<a href="image4.html">Name: My image 4 <br><img src="image4_thumb.jpg"></a>', - '<a href="image5.html">Name: My image 5 <br><img src="image5_thumb.jpg"></a>'] +>>> links = response.xpath('//a[contains(@href, "image")]') +>>> links.getall() +['<a href="image1.html">Name: My image 1 <br><img src="image1_thumb.jpg"></a>', + '<a href="image2.html">Name: My image 2 <br><img src="image2_thumb.jpg"></a>', + '<a href="image3.html">Name: My image 3 <br><img src="image3_thumb.jpg"></a>', + '<a href="image4.html">Name: My image 4 <br><img src="image4_thumb.jpg"></a>', + '<a href="image5.html">Name: My image 5 <br><img src="image5_thumb.jpg"></a>'] - >>> for index, link in enumerate(links): - ... args = (index, link.xpath('@href').get(), link.xpath('img/@src').get()) - ... print('Link number %d points to url %r and image %r' % args) - - Link number 0 points to url 'image1.html' and image 'image1_thumb.jpg' - Link number 1 points to url 'image2.html' and image 'image2_thumb.jpg' - Link number 2 points to url 'image3.html' and image 'image3_thumb.jpg' - Link number 3 points to url 'image4.html' and image 'image4_thumb.jpg' - Link number 4 points to url 'image5.html' and image 'image5_thumb.jpg' +>>> for index, link in enumerate(links): +... args = (index, link.xpath('@href').get(), link.xpath('img/@src').get()) +... print('Link number %d points to url %r and image %r' % args) +Link number 0 points to url 'image1.html' and image 'image1_thumb.jpg' +Link number 1 points to url 'image2.html' and image 'image2_thumb.jpg' +Link number 2 points to url 'image3.html' and image 'image3_thumb.jpg' +Link number 3 points to url 'image4.html' and image 'image4_thumb.jpg' +Link number 4 points to url 'image5.html' and image 'image5_thumb.jpg' .. _selecting-attributes: @@ -344,42 +343,42 @@ Selecting element attributes ---------------------------- There are several ways to get a value of an attribute. First, one can use -XPath syntax:: +XPath syntax: - >>> response.xpath("//a/@href").getall() - ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] +>>> response.xpath("//a/@href").getall() +['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] XPath syntax has a few advantages: it is a standard XPath feature, and ``@attributes`` can be used in other parts of an XPath expression - e.g. it is possible to filter by attribute value. Scrapy also provides an extension to CSS selectors (``::attr(...)``) -which allows to get attribute values:: +which allows to get attribute values: - >>> response.css('a::attr(href)').getall() - ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] +>>> response.css('a::attr(href)').getall() +['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] In addition to that, there is a ``.attrib`` property of Selector. You can use it if you prefer to lookup attributes in Python -code, without using XPaths or CSS extensions:: +code, without using XPaths or CSS extensions: - >>> [a.attrib['href'] for a in response.css('a')] - ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] +>>> [a.attrib['href'] for a in response.css('a')] +['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] This property is also available on SelectorList; it returns a dictionary with attributes of a first matching element. It is convenient to use when a selector is expected to give a single result (e.g. when selecting by element -ID, or when selecting an unique element on a page):: +ID, or when selecting an unique element on a page): - >>> response.css('base').attrib - {'href': 'http://example.com/'} - >>> response.css('base').attrib['href'] - 'http://example.com/' +>>> response.css('base').attrib +{'href': 'http://example.com/'} +>>> response.css('base').attrib['href'] +'http://example.com/' -``.attrib`` property of an empty SelectorList is empty:: +``.attrib`` property of an empty SelectorList is empty: - >>> response.css('foo').attrib - {} +>>> response.css('foo').attrib +{} Using selectors with regular expressions ---------------------------------------- @@ -390,21 +389,21 @@ data using regular expressions. However, unlike using ``.xpath()`` or can't construct nested ``.re()`` calls. Here's an example used to extract image names from the :ref:`HTML code -<topics-selectors-htmlcode>` above:: +<topics-selectors-htmlcode>` above: - >>> response.xpath('//a[contains(@href, "image")]/text()').re(r'Name:\s*(.*)') - ['My image 1', - 'My image 2', - 'My image 3', - 'My image 4', - 'My image 5'] +>>> response.xpath('//a[contains(@href, "image")]/text()').re(r'Name:\s*(.*)') +['My image 1', + 'My image 2', + 'My image 3', + 'My image 4', + 'My image 5'] There's an additional helper reciprocating ``.get()`` (and its alias ``.extract_first()``) for ``.re()``, named ``.re_first()``. -Use it to extract just the first matching string:: +Use it to extract just the first matching string: - >>> response.xpath('//a[contains(@href, "image")]/text()').re_first(r'Name:\s*(.*)') - 'My image 1' +>>> response.xpath('//a[contains(@href, "image")]/text()').re_first(r'Name:\s*(.*)') +'My image 1' .. _old-extraction-api: @@ -422,28 +421,28 @@ and readable code. The following examples show how these methods map to each other. -1. ``SelectorList.get()`` is the same as ``SelectorList.extract_first()``:: +1. ``SelectorList.get()`` is the same as ``SelectorList.extract_first()``: - >>> response.css('a::attr(href)').get() - 'image1.html' - >>> response.css('a::attr(href)').extract_first() - 'image1.html' + >>> response.css('a::attr(href)').get() + 'image1.html' + >>> response.css('a::attr(href)').extract_first() + 'image1.html' -2. ``SelectorList.getall()`` is the same as ``SelectorList.extract()``:: +2. ``SelectorList.getall()`` is the same as ``SelectorList.extract()``: - >>> response.css('a::attr(href)').getall() - ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] - >>> response.css('a::attr(href)').extract() - ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] + >>> response.css('a::attr(href)').getall() + ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] + >>> response.css('a::attr(href)').extract() + ['image1.html', 'image2.html', 'image3.html', 'image4.html', 'image5.html'] -3. ``Selector.get()`` is the same as ``Selector.extract()``:: +3. ``Selector.get()`` is the same as ``Selector.extract()``: - >>> response.css('a::attr(href)')[0].get() - 'image1.html' - >>> response.css('a::attr(href)')[0].extract() - 'image1.html' + >>> response.css('a::attr(href)')[0].get() + 'image1.html' + >>> response.css('a::attr(href)')[0].extract() + 'image1.html' -4. For consistency, there is also ``Selector.getall()``, which returns a list:: +4. For consistency, there is also ``Selector.getall()``, which returns a list: >>> response.css('a::attr(href)')[0].getall() ['image1.html'] @@ -481,26 +480,26 @@ with ``/``, that XPath will be absolute to the document and not relative to the ``Selector`` you're calling it from. For example, suppose you want to extract all ``<p>`` elements inside ``<div>`` -elements. First, you would get all ``<div>`` elements:: +elements. First, you would get all ``<div>`` elements: - >>> divs = response.xpath('//div') +>>> divs = response.xpath('//div') At first, you may be tempted to use the following approach, which is wrong, as it actually extracts all ``<p>`` elements from the document, not only those -inside ``<div>`` elements:: +inside ``<div>`` elements: - >>> for p in divs.xpath('//p'): # this is wrong - gets all <p> from the whole document - ... print(p.get()) +>>> for p in divs.xpath('//p'): # this is wrong - gets all <p> from the whole document +... print(p.get()) -This is the proper way to do it (note the dot prefixing the ``.//p`` XPath):: +This is the proper way to do it (note the dot prefixing the ``.//p`` XPath): - >>> for p in divs.xpath('.//p'): # extracts all <p> inside - ... print(p.get()) +>>> for p in divs.xpath('.//p'): # extracts all <p> inside +... print(p.get()) -Another common case would be to extract all direct ``<p>`` children:: +Another common case would be to extract all direct ``<p>`` children: - >>> for p in divs.xpath('p'): - ... print(p.get()) +>>> for p in divs.xpath('p'): +... print(p.get()) For more details about relative XPaths see the `Location Paths`_ section in the XPath specification. @@ -521,12 +520,12 @@ for that you may end up with more elements that you want, if they have a differe class name that shares the string ``someclass``. As it turns out, Scrapy selectors allow you to chain selectors, so most of the time -you can just select by class using CSS and then switch to XPath when needed:: +you can just select by class using CSS and then switch to XPath when needed: - >>> from scrapy import Selector - >>> sel = Selector(text='<div class="hero shout"><time datetime="2014-07-23 19:00">Special date</time></div>') - >>> sel.css('.shout').xpath('./time/@datetime').getall() - ['2014-07-23 19:00'] +>>> from scrapy import Selector +>>> sel = Selector(text='<div class="hero shout"><time datetime="2014-07-23 19:00">Special date</time></div>') +>>> sel.css('.shout').xpath('./time/@datetime').getall() +['2014-07-23 19:00'] This is cleaner than using the verbose XPath trick shown above. Just remember to use the ``.`` in the XPath expressions that will follow. @@ -538,41 +537,41 @@ Beware of the difference between //node[1] and (//node)[1] ``(//node)[1]`` selects all the nodes in the document, and then gets only the first of them. -Example:: +Example: - >>> from scrapy import Selector - >>> sel = Selector(text=""" - ....: <ul class="list"> - ....: <li>1</li> - ....: <li>2</li> - ....: <li>3</li> - ....: </ul> - ....: <ul class="list"> - ....: <li>4</li> - ....: <li>5</li> - ....: <li>6</li> - ....: </ul>""") - >>> xp = lambda x: sel.xpath(x).getall() +>>> from scrapy import Selector +>>> sel = Selector(text=""" +....: <ul class="list"> +....: <li>1</li> +....: <li>2</li> +....: <li>3</li> +....: </ul> +....: <ul class="list"> +....: <li>4</li> +....: <li>5</li> +....: <li>6</li> +....: </ul>""") +>>> xp = lambda x: sel.xpath(x).getall() -This gets all first ``<li>`` elements under whatever it is its parent:: +This gets all first ``<li>`` elements under whatever it is its parent: - >>> xp("//li[1]") - ['<li>1</li>', '<li>4</li>'] +>>> xp("//li[1]") +['<li>1</li>', '<li>4</li>'] -And this gets the first ``<li>`` element in the whole document:: +And this gets the first ``<li>`` element in the whole document: - >>> xp("(//li)[1]") - ['<li>1</li>'] +>>> xp("(//li)[1]") +['<li>1</li>'] -This gets all first ``<li>`` elements under an ``<ul>`` parent:: +This gets all first ``<li>`` elements under an ``<ul>`` parent: - >>> xp("//ul/li[1]") - ['<li>1</li>', '<li>4</li>'] +>>> xp("//ul/li[1]") +['<li>1</li>', '<li>4</li>'] -And this gets the first ``<li>`` element under an ``<ul>`` parent in the whole document:: +And this gets the first ``<li>`` element under an ``<ul>`` parent in the whole document: - >>> xp("(//ul/li)[1]") - ['<li>1</li>'] +>>> xp("(//ul/li)[1]") +['<li>1</li>'] Using text nodes in a condition ------------------------------- @@ -584,34 +583,34 @@ This is because the expression ``.//text()`` yields a collection of text element And when a node-set is converted to a string, which happens when it is passed as argument to a string function like ``contains()`` or ``starts-with()``, it results in the text for the first element only. -Example:: +Example: - >>> from scrapy import Selector - >>> sel = Selector(text='<a href="#">Click here to go to the <strong>Next Page</strong></a>') +>>> from scrapy import Selector +>>> sel = Selector(text='<a href="#">Click here to go to the <strong>Next Page</strong></a>') -Converting a *node-set* to string:: +Converting a *node-set* to string: - >>> sel.xpath('//a//text()').getall() # take a peek at the node-set - ['Click here to go to the ', 'Next Page'] - >>> sel.xpath("string(//a[1]//text())").getall() # convert it to string - ['Click here to go to the '] +>>> sel.xpath('//a//text()').getall() # take a peek at the node-set +['Click here to go to the ', 'Next Page'] +>>> sel.xpath("string(//a[1]//text())").getall() # convert it to string +['Click here to go to the '] -A *node* converted to a string, however, puts together the text of itself plus of all its descendants:: +A *node* converted to a string, however, puts together the text of itself plus of all its descendants: - >>> sel.xpath("//a[1]").getall() # select the first node - ['<a href="#">Click here to go to the <strong>Next Page</strong></a>'] - >>> sel.xpath("string(//a[1])").getall() # convert it to string - ['Click here to go to the Next Page'] +>>> sel.xpath("//a[1]").getall() # select the first node +['<a href="#">Click here to go to the <strong>Next Page</strong></a>'] +>>> sel.xpath("string(//a[1])").getall() # convert it to string +['Click here to go to the Next Page'] -So, using the ``.//text()`` node-set won't select anything in this case:: +So, using the ``.//text()`` node-set won't select anything in this case: - >>> sel.xpath("//a[contains(.//text(), 'Next Page')]").getall() - [] +>>> sel.xpath("//a[contains(.//text(), 'Next Page')]").getall() +[] -But using the ``.`` to mean the node, works:: +But using the ``.`` to mean the node, works: - >>> sel.xpath("//a[contains(., 'Next Page')]").getall() - ['<a href="#">Click here to go to the <strong>Next Page</strong></a>'] +>>> sel.xpath("//a[contains(., 'Next Page')]").getall() +['<a href="#">Click here to go to the <strong>Next Page</strong></a>'] .. _`XPath string function`: https://www.w3.org/TR/xpath/#section-String-Functions @@ -627,17 +626,17 @@ some arguments in your queries with placeholders like ``?``, which are then substituted with values passed with the query. Here's an example to match an element based on its "id" attribute value, -without hard-coding it (that was shown previously):: +without hard-coding it (that was shown previously): - >>> # `$val` used in the expression, a `val` argument needs to be passed - >>> response.xpath('//div[@id=$val]/a/text()', val='images').get() - 'Name: My image 1 ' +>>> # `$val` used in the expression, a `val` argument needs to be passed +>>> response.xpath('//div[@id=$val]/a/text()', val='images').get() +'Name: My image 1 ' Here's another example, to find the "id" attribute of a ``<div>`` tag containing -five ``<a>`` children (here we pass the value ``5`` as an integer):: +five ``<a>`` children (here we pass the value ``5`` as an integer): - >>> response.xpath('//div[count(a)=$cnt]/@id', cnt=5).get() - 'images' +>>> response.xpath('//div[count(a)=$cnt]/@id', cnt=5).get() +'images' All variable references must have a binding value when calling ``.xpath()`` (otherwise you'll get a ``ValueError: XPath error:`` exception). @@ -687,19 +686,19 @@ You can see several namespace declarations including a default .. highlight:: python Once in the shell we can try selecting all ``<link>`` objects and see that it -doesn't work (because the Atom XML namespace is obfuscating those nodes):: +doesn't work (because the Atom XML namespace is obfuscating those nodes): - >>> response.xpath("//link") - [] +>>> response.xpath("//link") +[] But once we call the :meth:`Selector.remove_namespaces` method, all -nodes can be accessed directly by their names:: +nodes can be accessed directly by their names: - >>> response.selector.remove_namespaces() - >>> response.xpath("//link") - [<Selector xpath='//link' data='<link rel="alternate" type="text/html" h'>, - <Selector xpath='//link' data='<link rel="next" type="application/atom+'>, - ... +>>> response.selector.remove_namespaces() +>>> response.xpath("//link") +[<Selector xpath='//link' data='<link rel="alternate" type="text/html" h'>, + <Selector xpath='//link' data='<link rel="next" type="application/atom+'>, + ... If you wonder why the namespace removal procedure isn't always called by default instead of having to call it manually, this is because of two reasons, which, in order @@ -734,26 +733,25 @@ Regular expressions The ``test()`` function, for example, can prove quite useful when XPath's ``starts-with()`` or ``contains()`` are not sufficient. -Example selecting links in list item with a "class" attribute ending with a digit:: +Example selecting links in list item with a "class" attribute ending with a digit: - >>> from scrapy import Selector - >>> doc = u""" - ... <div> - ... <ul> - ... <li class="item-0"><a href="link1.html">first item</a></li> - ... <li class="item-1"><a href="link2.html">second item</a></li> - ... <li class="item-inactive"><a href="link3.html">third item</a></li> - ... <li class="item-1"><a href="link4.html">fourth item</a></li> - ... <li class="item-0"><a href="link5.html">fifth item</a></li> - ... </ul> - ... </div> - ... """ - >>> sel = Selector(text=doc, type="html") - >>> sel.xpath('//li//@href').getall() - ['link1.html', 'link2.html', 'link3.html', 'link4.html', 'link5.html'] - >>> sel.xpath('//li[re:test(@class, "item-\d$")]//@href').getall() - ['link1.html', 'link2.html', 'link4.html', 'link5.html'] - >>> +>>> from scrapy import Selector +>>> doc = u""" +... <div> +... <ul> +... <li class="item-0"><a href="link1.html">first item</a></li> +... <li class="item-1"><a href="link2.html">second item</a></li> +... <li class="item-inactive"><a href="link3.html">third item</a></li> +... <li class="item-1"><a href="link4.html">fourth item</a></li> +... <li class="item-0"><a href="link5.html">fifth item</a></li> +... </ul> +... </div> +... """ +>>> sel = Selector(text=doc, type="html") +>>> sel.xpath('//li//@href').getall() +['link1.html', 'link2.html', 'link3.html', 'link4.html', 'link5.html'] +>>> sel.xpath('//li[re:test(@class, "item-\d$")]//@href').getall() +['link1.html', 'link2.html', 'link4.html', 'link5.html'] .. warning:: C library ``libxslt`` doesn't natively support EXSLT regular expressions so `lxml`_'s implementation uses hooks to Python's ``re`` module. @@ -849,7 +847,6 @@ with groups of itemscopes and corresponding itemprops:: current scope: ['http://schema.org/Rating'] properties: ['worstRating', 'ratingValue', 'bestRating'] - >>> Here we first iterate over ``itemscope`` elements, and for each one, we look for all ``itemprops`` elements and exclude those that are themselves @@ -877,15 +874,15 @@ For the following HTML:: .. highlight:: python -You can use it like this:: +You can use it like this: - >>> response.xpath('//p[has-class("foo")]') - [<Selector xpath='//p[has-class("foo")]' data='<p class="foo bar-baz">First</p>'>, - <Selector xpath='//p[has-class("foo")]' data='<p class="foo">Second</p>'>] - >>> response.xpath('//p[has-class("foo", "bar-baz")]') - [<Selector xpath='//p[has-class("foo", "bar-baz")]' data='<p class="foo bar-baz">First</p>'>] - >>> response.xpath('//p[has-class("foo", "bar")]') - [] +>>> response.xpath('//p[has-class("foo")]') +[<Selector xpath='//p[has-class("foo")]' data='<p class="foo bar-baz">First</p>'>, + <Selector xpath='//p[has-class("foo")]' data='<p class="foo">Second</p>'>] +>>> response.xpath('//p[has-class("foo", "bar-baz")]') +[<Selector xpath='//p[has-class("foo", "bar-baz")]' data='<p class="foo bar-baz">First</p>'>] +>>> response.xpath('//p[has-class("foo", "bar")]') +[] So XPath ``//p[has-class("foo", "bar-baz")]`` is roughly equivalent to CSS ``p.foo.bar-baz``. Please note, that it is slower in most of the cases, diff --git a/docs/topics/shell.rst b/docs/topics/shell.rst index 68a0b19b5..c1fdfd221 100644 --- a/docs/topics/shell.rst +++ b/docs/topics/shell.rst @@ -177,47 +177,46 @@ all start with the ``[s]`` prefix):: >>> -After that, we can start playing with the objects:: +After that, we can start playing with the objects: - >>> response.xpath('//title/text()').get() - 'Scrapy | A Fast and Powerful Scraping and Web Crawling Framework' +>>> response.xpath('//title/text()').get() +'Scrapy | A Fast and Powerful Scraping and Web Crawling Framework' - >>> fetch("https://reddit.com") +>>> fetch("https://reddit.com") - >>> response.xpath('//title/text()').get() - 'reddit: the front page of the internet' +>>> response.xpath('//title/text()').get() +'reddit: the front page of the internet' - >>> request = request.replace(method="POST") +>>> request = request.replace(method="POST") - >>> fetch(request) +>>> fetch(request) - >>> response.status - 404 +>>> response.status +404 - >>> from pprint import pprint +>>> from pprint import pprint - >>> pprint(response.headers) - {'Accept-Ranges': ['bytes'], - 'Cache-Control': ['max-age=0, must-revalidate'], - 'Content-Type': ['text/html; charset=UTF-8'], - 'Date': ['Thu, 08 Dec 2016 16:21:19 GMT'], - 'Server': ['snooserv'], - 'Set-Cookie': ['loid=KqNLou0V9SKMX4qb4n; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', - 'loidcreated=2016-12-08T16%3A21%3A19.445Z; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', - 'loid=vi0ZVe4NkxNWdlH7r7; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', - 'loidcreated=2016-12-08T16%3A21%3A19.459Z; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure'], - 'Vary': ['accept-encoding'], - 'Via': ['1.1 varnish'], - 'X-Cache': ['MISS'], - 'X-Cache-Hits': ['0'], - 'X-Content-Type-Options': ['nosniff'], - 'X-Frame-Options': ['SAMEORIGIN'], - 'X-Moose': ['majestic'], - 'X-Served-By': ['cache-cdg8730-CDG'], - 'X-Timer': ['S1481214079.394283,VS0,VE159'], - 'X-Ua-Compatible': ['IE=edge'], - 'X-Xss-Protection': ['1; mode=block']} - >>> +>>> pprint(response.headers) +{'Accept-Ranges': ['bytes'], + 'Cache-Control': ['max-age=0, must-revalidate'], + 'Content-Type': ['text/html; charset=UTF-8'], + 'Date': ['Thu, 08 Dec 2016 16:21:19 GMT'], + 'Server': ['snooserv'], + 'Set-Cookie': ['loid=KqNLou0V9SKMX4qb4n; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', + 'loidcreated=2016-12-08T16%3A21%3A19.445Z; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', + 'loid=vi0ZVe4NkxNWdlH7r7; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure', + 'loidcreated=2016-12-08T16%3A21%3A19.459Z; Domain=reddit.com; Max-Age=63071999; Path=/; expires=Sat, 08-Dec-2018 16:21:19 GMT; secure'], + 'Vary': ['accept-encoding'], + 'Via': ['1.1 varnish'], + 'X-Cache': ['MISS'], + 'X-Cache-Hits': ['0'], + 'X-Content-Type-Options': ['nosniff'], + 'X-Frame-Options': ['SAMEORIGIN'], + 'X-Moose': ['majestic'], + 'X-Served-By': ['cache-cdg8730-CDG'], + 'X-Timer': ['S1481214079.394283,VS0,VE159'], + 'X-Ua-Compatible': ['IE=edge'], + 'X-Xss-Protection': ['1; mode=block']} .. _topics-shell-inspect-response: @@ -263,16 +262,16 @@ When you run the spider, you will get something similar to this:: >>> response.url 'http://example.org' -Then, you can check if the extraction code is working:: +Then, you can check if the extraction code is working: - >>> response.xpath('//h1[@class="fn"]') - [] +>>> response.xpath('//h1[@class="fn"]') +[] Nope, it doesn't. So you can open the response in your web browser and see if -it's the response you were expecting:: +it's the response you were expecting: - >>> view(response) - True +>>> view(response) +True Finally you hit Ctrl-D (or Ctrl-Z in Windows) to exit the shell and resume the crawling:: diff --git a/docs/topics/stats.rst b/docs/topics/stats.rst index 38648ec55..3dd829ebe 100644 --- a/docs/topics/stats.rst +++ b/docs/topics/stats.rst @@ -57,15 +57,15 @@ Set stat value only if lower than previous:: stats.min_value('min_free_memory_percent', value) -Get stat value:: +Get stat value: - >>> stats.get_value('custom_count') - 1 +>>> stats.get_value('custom_count') +1 -Get all stats:: +Get all stats: - >>> stats.get_stats() - {'custom_count': 1, 'start_time': datetime.datetime(2009, 7, 14, 21, 47, 28, 977139)} +>>> stats.get_stats() +{'custom_count': 1, 'start_time': datetime.datetime(2009, 7, 14, 21, 47, 28, 977139)} Available Stats Collectors ========================== From 07b8cd28aa84fb322a072467b49e2557d4aa3881 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 5 Dec 2019 14:48:31 +0100 Subject: [PATCH 326/387] =?UTF-8?q?Mark=20bandit=E2=80=99s=20402=20check?= =?UTF-8?q?=20as=20addressed=20by=20#4180=20(#4181)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .bandit.yml | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/.bandit.yml b/.bandit.yml index cc7db3a66..243379b0b 100644 --- a/.bandit.yml +++ b/.bandit.yml @@ -8,7 +8,7 @@ skips: - B311 - B320 - B321 -- B402 +- B402 # https://github.com/scrapy/scrapy/issues/4180 - B403 - B404 - B406 From 02cdc53fb82e3cc5e51771180f1f79186b52670a Mon Sep 17 00:00:00 2001 From: Andrey Rahmatullin <wrar@wrar.name> Date: Fri, 13 Dec 2019 18:04:05 +0500 Subject: [PATCH 327/387] Add a test for a CrawlerProcess script. (#4218) * Add a test for a CrawlerProcess script. * Add tests/CrawlerProcess to collect_ignore. * Remove an extra line. * Fix/improve conftest.py. --- conftest.py | 9 ++++++++- tests/CrawlerProcess/simple.py | 15 +++++++++++++++ tests/test_crawler.py | 20 ++++++++++++++++++++ 3 files changed, 43 insertions(+), 1 deletion(-) create mode 100644 tests/CrawlerProcess/simple.py diff --git a/conftest.py b/conftest.py index d54ce155c..d37c22436 100644 --- a/conftest.py +++ b/conftest.py @@ -1,12 +1,19 @@ +from pathlib import Path + import pytest +def _py_files(folder): + return (str(p) for p in Path(folder).rglob('*.py')) + + collect_ignore = [ # not a test, but looks like a test "scrapy/utils/testsite.py", + # contains scripts to be run by tests/test_crawler.py::CrawlerProcessSubprocess + *_py_files("tests/CrawlerProcess") ] - for line in open('tests/ignores.txt'): file_path = line.strip() if file_path and file_path[0] != '#': diff --git a/tests/CrawlerProcess/simple.py b/tests/CrawlerProcess/simple.py new file mode 100644 index 000000000..5f6f1ae30 --- /dev/null +++ b/tests/CrawlerProcess/simple.py @@ -0,0 +1,15 @@ +import scrapy +from scrapy.crawler import CrawlerProcess + + +class NoRequestsSpider(scrapy.Spider): + name = 'no_request' + + def start_requests(self): + return [] + + +process = CrawlerProcess(settings={}) + +process.crawl(NoRequestsSpider) +process.start() diff --git a/tests/test_crawler.py b/tests/test_crawler.py index 8eb2389e2..e37a2ff0e 100644 --- a/tests/test_crawler.py +++ b/tests/test_crawler.py @@ -1,4 +1,7 @@ import logging +import os +import subprocess +import sys import warnings from twisted.internet import defer @@ -14,6 +17,7 @@ from scrapy.utils.spider import DefaultSpider from scrapy.utils.misc import load_object from scrapy.extensions.throttle import AutoThrottle from scrapy.extensions import telnet +from scrapy.utils.test import get_testenv class BaseCrawlerTest(unittest.TestCase): @@ -245,3 +249,19 @@ class CrawlerRunnerHasSpider(unittest.TestCase): yield runner.crawl(NoRequestsSpider) self.assertEqual(runner.bootstrap_failed, True) + + +class CrawlerProcessSubprocess(unittest.TestCase): + script_dir = os.path.join(os.path.abspath(os.path.dirname(__file__)), 'CrawlerProcess') + + def run_script(self, script_name): + script_path = os.path.join(self.script_dir, script_name) + args = (sys.executable, script_path) + p = subprocess.Popen(args, env=get_testenv(), + stdout=subprocess.PIPE, stderr=subprocess.PIPE) + stdout, stderr = p.communicate() + return stderr.decode('utf-8') + + def test_simple(self): + log = self.run_script('simple.py') + self.assertIn('Spider closed (finished)', log) From 3560123090c1660fcfad7c5e7da08c9af503940f Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 5 Dec 2019 19:06:51 +0500 Subject: [PATCH 328/387] Rename ASYNCIO_SUPPORT to ASYNCIO_ENABLED. --- docs/topics/settings.rst | 4 ++-- scrapy/commands/crawl.py | 2 +- scrapy/commands/runspider.py | 2 +- scrapy/settings/default_settings.py | 2 +- scrapy/utils/log.py | 4 ++-- tests/test_commands.py | 8 ++++---- tests/test_crawler.py | 10 +++++----- 7 files changed, 16 insertions(+), 16 deletions(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 43f59f7cc..5cbf7450e 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -160,9 +160,9 @@ to any particular component. In that case the module of that component will be shown, typically an extension, middleware or pipeline. It also means that the component must be enabled in order for the setting to have any effect. -.. setting:: ASYNCIO_SUPPORT +.. setting:: ASYNCIO_ENABLED -ASYNCIO_SUPPORT +ASYNCIO_ENABLED --------------- Default: ``False`` diff --git a/scrapy/commands/crawl.py b/scrapy/commands/crawl.py index e2e69be49..b50761e4a 100644 --- a/scrapy/commands/crawl.py +++ b/scrapy/commands/crawl.py @@ -27,7 +27,7 @@ class Command(ScrapyCommand): def process_options(self, args, opts): ScrapyCommand.process_options(self, args, opts) - if self.settings.getbool('ASYNCIO_SUPPORT'): + if self.settings.getbool('ASYNCIO_ENABLED'): install_asyncio_reactor() try: opts.spargs = arglist_to_dict(opts.spargs) diff --git a/scrapy/commands/runspider.py b/scrapy/commands/runspider.py index ebd4eb620..bfe844eb5 100644 --- a/scrapy/commands/runspider.py +++ b/scrapy/commands/runspider.py @@ -51,7 +51,7 @@ class Command(ScrapyCommand): def process_options(self, args, opts): ScrapyCommand.process_options(self, args, opts) - if self.settings.getbool('ASYNCIO_SUPPORT'): + if self.settings.getbool('ASYNCIO_ENABLED'): install_asyncio_reactor() try: opts.spargs = arglist_to_dict(opts.spargs) diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index c9097bd1f..153b8037a 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -19,7 +19,7 @@ from os.path import join, abspath, dirname AJAXCRAWL_ENABLED = False -ASYNCIO_SUPPORT = False +ASYNCIO_ENABLED = False AUTOTHROTTLE_ENABLED = False AUTOTHROTTLE_DEBUG = False diff --git a/scrapy/utils/log.py b/scrapy/utils/log.py index 8c56cfa42..0fe3d1549 100644 --- a/scrapy/utils/log.py +++ b/scrapy/utils/log.py @@ -149,11 +149,11 @@ def log_scrapy_info(settings): {'versions': ", ".join("%s %s" % (name, version) for name, version in scrapy_components_versions() if name != "Scrapy")}) - if settings.getbool('ASYNCIO_SUPPORT'): + if settings.getbool('ASYNCIO_ENABLED'): if is_asyncio_reactor_installed(): logger.debug("Asyncio support enabled") else: - logger.error("ASYNCIO_SUPPORT is on but the Twisted asyncio " + logger.error("ASYNCIO_ENABLED is on but the Twisted asyncio " "reactor is not installed, this is not supported " "and asyncio coroutines will not work.") diff --git a/tests/test_commands.py b/tests/test_commands.py index 3b64bfa23..197d80217 100644 --- a/tests/test_commands.py +++ b/tests/test_commands.py @@ -295,12 +295,12 @@ class BadSpider(scrapy.Spider): self.assertIn("start_requests", log) self.assertIn("badspider.py", log) - def test_asyncio_support_true(self): - log = self.get_log(self.debug_log_spider, args=['-s', 'ASYNCIO_SUPPORT=True']) + def test_asyncio_enabled_true(self): + log = self.get_log(self.debug_log_spider, args=['-s', 'ASYNCIO_ENABLED=True']) self.assertIn("DEBUG: Asyncio support enabled", log) - def test_asyncio_support_false(self): - log = self.get_log(self.debug_log_spider, args=['-s', 'ASYNCIO_SUPPORT=False']) + def test_asyncio_enabled_false(self): + log = self.get_log(self.debug_log_spider, args=['-s', 'ASYNCIO_ENABLED=False']) self.assertNotIn("DEBUG: Asyncio support enabled", log) diff --git a/tests/test_crawler.py b/tests/test_crawler.py index 9410b0e7a..05909d995 100644 --- a/tests/test_crawler.py +++ b/tests/test_crawler.py @@ -262,19 +262,19 @@ class CrawlerRunnerHasSpider(unittest.TestCase): self.assertEqual(runner.bootstrap_failed, True) @defer.inlineCallbacks - def test_crawler_process_asyncio_supported_true(self): + def test_crawler_process_asyncio_enabled_true(self): with LogCapture(level=logging.DEBUG) as log: - runner = CrawlerProcess(settings={'ASYNCIO_SUPPORT': True}) + runner = CrawlerProcess(settings={'ASYNCIO_ENABLED': True}) yield runner.crawl(NoRequestsSpider) if self.reactor_pytest == 'asyncio': self.assertIn("Asyncio support enabled", str(log)) else: self.assertNotIn("Asyncio support enabled", str(log)) - self.assertIn("ASYNCIO_SUPPORT is on but the Twisted asyncio reactor is not installed", str(log)) + self.assertIn("ASYNCIO_ENABLED is on but the Twisted asyncio reactor is not installed", str(log)) @defer.inlineCallbacks - def test_crawler_process_asyncio_supported_false(self): - runner = CrawlerProcess(settings={'ASYNCIO_SUPPORT': False}) + def test_crawler_process_asyncio_enabled_false(self): + runner = CrawlerProcess(settings={'ASYNCIO_ENABLED': False}) with LogCapture(level=logging.DEBUG) as log: yield runner.crawl(NoRequestsSpider) self.assertNotIn("Asyncio support enabled", str(log)) From 69cd2e247efe1823ab188eacb241b9e5596a879c Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Sat, 7 Dec 2019 00:09:53 +0500 Subject: [PATCH 329/387] Move a bunch of "from twisted.internet import reactor" inside functions. --- scrapy/cmdline.py | 4 +--- scrapy/commands/shell.py | 3 +-- scrapy/crawler.py | 7 ++++++- scrapy/shell.py | 3 ++- scrapy/utils/defer.py | 4 +++- scrapy/utils/ossignal.py | 3 +-- scrapy/utils/reactor.py | 4 +++- 7 files changed, 17 insertions(+), 11 deletions(-) diff --git a/scrapy/cmdline.py b/scrapy/cmdline.py index 3c2efe58f..69e917004 100644 --- a/scrapy/cmdline.py +++ b/scrapy/cmdline.py @@ -6,6 +6,7 @@ import inspect import pkg_resources import scrapy +from scrapy.crawler import CrawlerProcess from scrapy.commands import ScrapyCommand from scrapy.exceptions import UsageError from scrapy.utils.misc import walk_modules @@ -140,9 +141,6 @@ def execute(argv=None, settings=None): opts, args = parser.parse_args(args=argv[1:]) _run_print_help(parser, cmd.process_options, args, opts) - # needs to be after cmd.process_options() as it imports twisted.internet.reactor - # while commands may want to install the asyncio reactor - from scrapy.crawler import CrawlerProcess cmd.crawler_process = CrawlerProcess(settings) _run_print_help(parser, _run_command, cmd, args, opts) sys.exit(cmd.exitcode) diff --git a/scrapy/commands/shell.py b/scrapy/commands/shell.py index 7516e2aba..d44a32d5f 100644 --- a/scrapy/commands/shell.py +++ b/scrapy/commands/shell.py @@ -7,6 +7,7 @@ from threading import Thread from scrapy.commands import ScrapyCommand from scrapy.http import Request +from scrapy.shell import Shell from scrapy.utils.spider import spidercls_for_request, DefaultSpider from scrapy.utils.url import guess_scheme @@ -69,8 +70,6 @@ class Command(ScrapyCommand): self._start_crawler_thread() - # moved from the top-level because it imports twisted.internet.reactor - from scrapy.shell import Shell shell = Shell(crawler, update_vars=self.update_vars, code=opts.code) shell.start(url=url, redirect=not opts.no_redirect) diff --git a/scrapy/crawler.py b/scrapy/crawler.py index 6c7eb737b..450260004 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -3,7 +3,7 @@ import pprint import signal import warnings -from twisted.internet import reactor, defer +from twisted.internet import defer from zope.interface.verify import verifyClass, DoesNotImplement from scrapy import Spider @@ -261,6 +261,7 @@ class CrawlerProcess(CrawlerRunner): log_scrapy_info(self.settings) def _signal_shutdown(self, signum, _): + from twisted.internet import reactor install_shutdown_handlers(self._signal_kill) signame = signal_names[signum] logger.info("Received %(signame)s, shutting down gracefully. Send again to force ", @@ -268,6 +269,7 @@ class CrawlerProcess(CrawlerRunner): reactor.callFromThread(self._graceful_stop_reactor) def _signal_kill(self, signum, _): + from twisted.internet import reactor install_shutdown_handlers(signal.SIG_IGN) signame = signal_names[signum] logger.info('Received %(signame)s twice, forcing unclean shutdown', @@ -286,6 +288,7 @@ class CrawlerProcess(CrawlerRunner): :param boolean stop_after_crawl: stop or not the reactor when all crawlers have finished """ + from twisted.internet import reactor if stop_after_crawl: d = self.join() # Don't start the reactor if the deferreds are already fired @@ -300,6 +303,7 @@ class CrawlerProcess(CrawlerRunner): reactor.run(installSignalHandlers=False) # blocking call def _get_dns_resolver(self): + from twisted.internet import reactor if self.settings.getbool('DNSCACHE_ENABLED'): cache_size = self.settings.getint('DNSCACHE_SIZE') else: @@ -316,6 +320,7 @@ class CrawlerProcess(CrawlerRunner): return d def _stop_reactor(self, _=None): + from twisted.internet import reactor try: reactor.stop() except RuntimeError: # raised if already stopped or in shutdown stage diff --git a/scrapy/shell.py b/scrapy/shell.py index a649d555f..a23b04df9 100644 --- a/scrapy/shell.py +++ b/scrapy/shell.py @@ -7,7 +7,7 @@ import os import signal import warnings -from twisted.internet import reactor, threads, defer +from twisted.internet import threads, defer from twisted.python import threadable from w3lib.url import any_to_uri @@ -98,6 +98,7 @@ class Shell(object): return spider def fetch(self, request_or_url, spider=None, redirect=True, **kwargs): + from twisted.internet import reactor if isinstance(request_or_url, Request): request = request_or_url else: diff --git a/scrapy/utils/defer.py b/scrapy/utils/defer.py index 3b7ef75ab..6a91776c7 100644 --- a/scrapy/utils/defer.py +++ b/scrapy/utils/defer.py @@ -5,7 +5,7 @@ import asyncio import asyncio.futures import inspect -from twisted.internet import defer, reactor, task +from twisted.internet import defer, task from twisted.python import failure from scrapy.exceptions import IgnoreRequest @@ -19,6 +19,7 @@ def defer_fail(_failure): It delays by 100ms so reactor has a chance to go through readers and writers before attending pending delayed calls, so do not set delay to zero. """ + from twisted.internet import reactor d = defer.Deferred() reactor.callLater(0.1, d.errback, _failure) return d @@ -31,6 +32,7 @@ def defer_succeed(result): It delays by 100ms so reactor has a chance to go trough readers and writers before attending pending delayed calls, so do not set delay to zero. """ + from twisted.internet import reactor d = defer.Deferred() reactor.callLater(0.1, d.callback, result) return d diff --git a/scrapy/utils/ossignal.py b/scrapy/utils/ossignal.py index 7a7aec9be..45c9cef0c 100644 --- a/scrapy/utils/ossignal.py +++ b/scrapy/utils/ossignal.py @@ -1,7 +1,5 @@ import signal -from twisted.internet import reactor - signal_names = {} for signame in dir(signal): @@ -17,6 +15,7 @@ def install_shutdown_handlers(function, override_sigint=True): SIGINT handler won't be install if there is already a handler in place (e.g. Pdb) """ + from twisted.internet import reactor reactor._handleSignals() signal.signal(signal.SIGTERM, function) if signal.getsignal(signal.SIGINT) == signal.default_int_handler or \ diff --git a/scrapy/utils/reactor.py b/scrapy/utils/reactor.py index 493d26d4c..b98fff6ec 100644 --- a/scrapy/utils/reactor.py +++ b/scrapy/utils/reactor.py @@ -1,8 +1,9 @@ -from twisted.internet import reactor, error +from twisted.internet import error def listen_tcp(portrange, host, factory): """Like reactor.listenTCP but tries different ports in a range.""" + from twisted.internet import reactor assert len(portrange) <= 2, "invalid portrange: %s" % portrange if not portrange: return reactor.listenTCP(0, factory, interface=host) @@ -30,6 +31,7 @@ class CallLaterOnce(object): self._call = None def schedule(self, delay=0): + from twisted.internet import reactor if self._call is None: self._call = reactor.callLater(delay, self) From 855bbebc8bb862aa02e48f65fd861b1ddf78b57a Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 13 Dec 2019 18:11:49 +0500 Subject: [PATCH 330/387] Move install_asyncio_reactor() from commands to CrawlerProcess. --- scrapy/commands/crawl.py | 3 --- scrapy/commands/runspider.py | 3 --- scrapy/crawler.py | 3 +++ 3 files changed, 3 insertions(+), 6 deletions(-) diff --git a/scrapy/commands/crawl.py b/scrapy/commands/crawl.py index b50761e4a..8093fd402 100644 --- a/scrapy/commands/crawl.py +++ b/scrapy/commands/crawl.py @@ -1,6 +1,5 @@ import os from scrapy.commands import ScrapyCommand -from scrapy.utils.asyncio import install_asyncio_reactor from scrapy.utils.conf import arglist_to_dict from scrapy.utils.python import without_none_values from scrapy.exceptions import UsageError @@ -27,8 +26,6 @@ class Command(ScrapyCommand): def process_options(self, args, opts): ScrapyCommand.process_options(self, args, opts) - if self.settings.getbool('ASYNCIO_ENABLED'): - install_asyncio_reactor() try: opts.spargs = arglist_to_dict(opts.spargs) except ValueError: diff --git a/scrapy/commands/runspider.py b/scrapy/commands/runspider.py index bfe844eb5..57d8471ca 100644 --- a/scrapy/commands/runspider.py +++ b/scrapy/commands/runspider.py @@ -2,7 +2,6 @@ import sys import os from importlib import import_module -from scrapy.utils.asyncio import install_asyncio_reactor from scrapy.utils.spider import iter_spider_classes from scrapy.commands import ScrapyCommand from scrapy.exceptions import UsageError @@ -51,8 +50,6 @@ class Command(ScrapyCommand): def process_options(self, args, opts): ScrapyCommand.process_options(self, args, opts) - if self.settings.getbool('ASYNCIO_ENABLED'): - install_asyncio_reactor() try: opts.spargs = arglist_to_dict(opts.spargs) except ValueError: diff --git a/scrapy/crawler.py b/scrapy/crawler.py index 450260004..706c8a59d 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -14,6 +14,7 @@ from scrapy.extension import ExtensionManager from scrapy.settings import overridden_settings, Settings from scrapy.signalmanager import SignalManager from scrapy.exceptions import ScrapyDeprecationWarning +from scrapy.utils.asyncio import install_asyncio_reactor from scrapy.utils.ossignal import install_shutdown_handlers, signal_names from scrapy.utils.misc import load_object from scrapy.utils.log import ( @@ -256,6 +257,8 @@ class CrawlerProcess(CrawlerRunner): def __init__(self, settings=None, install_root_handler=True): super(CrawlerProcess, self).__init__(settings) + if self.settings.getbool('ASYNCIO_ENABLED'): + install_asyncio_reactor() install_shutdown_handlers(self._signal_shutdown) configure_logging(self.settings, install_root_handler) log_scrapy_info(self.settings) From bfb78b8dea44a5db3f4a3bca83ab58c7ca0e3ef3 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 13 Dec 2019 18:12:07 +0500 Subject: [PATCH 331/387] Add CrawlerProcess tests for ASYNCIO_ENABLED. --- .../asyncio_enabled_no_reactor.py | 17 ++++++++++++++ .../CrawlerProcess/asyncio_enabled_reactor.py | 22 +++++++++++++++++++ tests/test_crawler.py | 11 ++++++++++ 3 files changed, 50 insertions(+) create mode 100644 tests/CrawlerProcess/asyncio_enabled_no_reactor.py create mode 100644 tests/CrawlerProcess/asyncio_enabled_reactor.py diff --git a/tests/CrawlerProcess/asyncio_enabled_no_reactor.py b/tests/CrawlerProcess/asyncio_enabled_no_reactor.py new file mode 100644 index 000000000..dfe028ef4 --- /dev/null +++ b/tests/CrawlerProcess/asyncio_enabled_no_reactor.py @@ -0,0 +1,17 @@ +import scrapy +from scrapy.crawler import CrawlerProcess + + +class NoRequestsSpider(scrapy.Spider): + name = 'no_request' + + def start_requests(self): + return [] + + +process = CrawlerProcess(settings={ + 'ASYNCIO_ENABLED': True, +}) + +process.crawl(NoRequestsSpider) +process.start() diff --git a/tests/CrawlerProcess/asyncio_enabled_reactor.py b/tests/CrawlerProcess/asyncio_enabled_reactor.py new file mode 100644 index 000000000..7a172ea28 --- /dev/null +++ b/tests/CrawlerProcess/asyncio_enabled_reactor.py @@ -0,0 +1,22 @@ +import asyncio + +from twisted.internet import asyncioreactor +asyncioreactor.install(asyncio.get_event_loop()) + +import scrapy +from scrapy.crawler import CrawlerProcess + + +class NoRequestsSpider(scrapy.Spider): + name = 'no_request' + + def start_requests(self): + return [] + + +process = CrawlerProcess(settings={ + 'ASYNCIO_ENABLED': True, +}) + +process.crawl(NoRequestsSpider) +process.start() diff --git a/tests/test_crawler.py b/tests/test_crawler.py index 05909d995..0b2645280 100644 --- a/tests/test_crawler.py +++ b/tests/test_crawler.py @@ -301,3 +301,14 @@ class CrawlerProcessSubprocess(unittest.TestCase): def test_simple(self): log = self.run_script('simple.py') self.assertIn('Spider closed (finished)', log) + self.assertNotIn("DEBUG: Asyncio support enabled", log) + + def test_asyncio_enabled_no_reactor(self): + log = self.run_script('asyncio_enabled_no_reactor.py') + self.assertIn('Spider closed (finished)', log) + self.assertIn("DEBUG: Asyncio support enabled", log) + + def test_asyncio_enabled_reactor(self): + log = self.run_script('asyncio_enabled_reactor.py') + self.assertIn('Spider closed (finished)', log) + self.assertIn("DEBUG: Asyncio support enabled", log) From b5c4c2cae89714479d6fbb5a497fe43d48d886bc Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Fri, 13 Dec 2019 14:20:48 +0100 Subject: [PATCH 332/387] Keep 2 spaces between code and inline comments (#4195) --- pytest.ini | 35 +++++++++++++------------- scrapy/core/downloader/webclient.py | 2 +- scrapy/core/engine.py | 2 +- scrapy/core/scraper.py | 2 +- scrapy/exporters.py | 2 +- scrapy/settings/default_settings.py | 4 +-- scrapy/spiders/feed.py | 4 +-- scrapy/utils/console.py | 8 +++--- tests/test_command_fetch.py | 2 +- tests/test_downloader_handlers.py | 2 +- tests/test_engine.py | 1 - tests/test_http_request.py | 2 +- tests/test_pipeline_media.py | 6 ++--- tests/test_spidermiddleware_referer.py | 20 +++++++-------- tests/test_utils_defer.py | 8 +++--- tests/test_utils_spider.py | 13 ++++------ 16 files changed, 54 insertions(+), 59 deletions(-) diff --git a/pytest.ini b/pytest.ini index 33c34b8e8..1b23595c0 100644 --- a/pytest.ini +++ b/pytest.ini @@ -26,8 +26,8 @@ flake8-ignore = scrapy/http/__init__.py F401 # Issues pending a review: # extras - extras/qps-bench-server.py E261 E501 - extras/qpsclient.py E501 E261 E501 + extras/qps-bench-server.py E501 + extras/qpsclient.py E501 E501 # scrapy/commands scrapy/commands/__init__.py E128 E501 scrapy/commands/check.py E501 @@ -45,15 +45,15 @@ flake8-ignore = scrapy/contracts/__init__.py E501 W504 scrapy/contracts/default.py E502 E128 # scrapy/core - scrapy/core/engine.py E261 E501 E128 E127 E306 E502 + scrapy/core/engine.py E501 E128 E127 E306 E502 scrapy/core/scheduler.py E501 - scrapy/core/scraper.py E501 E306 E261 E128 W504 + scrapy/core/scraper.py E501 E306 E128 W504 scrapy/core/spidermw.py E501 E731 E502 E126 E226 scrapy/core/downloader/__init__.py E501 scrapy/core/downloader/contextfactory.py E501 E128 E126 scrapy/core/downloader/middleware.py E501 E502 scrapy/core/downloader/tls.py E501 E305 E241 - scrapy/core/downloader/webclient.py E731 E501 E261 E502 E128 E126 E226 + scrapy/core/downloader/webclient.py E731 E501 E502 E128 E126 E226 scrapy/core/downloader/handlers/__init__.py E501 scrapy/core/downloader/handlers/ftp.py E501 E305 E128 E127 scrapy/core/downloader/handlers/http10.py E501 @@ -102,7 +102,7 @@ flake8-ignore = scrapy/selector/unified.py E501 E111 # scrapy/settings scrapy/settings/__init__.py E501 - scrapy/settings/default_settings.py E501 E261 E114 E116 E226 + scrapy/settings/default_settings.py E501 E114 E116 E226 scrapy/settings/deprecated.py E501 # scrapy/spidermiddlewares scrapy/spidermiddlewares/httperror.py E501 @@ -112,17 +112,16 @@ flake8-ignore = # scrapy/spiders scrapy/spiders/__init__.py E501 E402 scrapy/spiders/crawl.py E501 - scrapy/spiders/feed.py E501 E261 + scrapy/spiders/feed.py E501 scrapy/spiders/sitemap.py E501 # scrapy/utils scrapy/utils/benchserver.py E501 scrapy/utils/conf.py E402 E502 E501 - scrapy/utils/console.py E261 E306 E305 + scrapy/utils/console.py E306 E305 scrapy/utils/datatypes.py E501 E226 scrapy/utils/decorators.py E501 scrapy/utils/defer.py E501 E128 scrapy/utils/deprecate.py E128 E501 E127 E502 - scrapy/utils/engine.py E261 scrapy/utils/gz.py E305 E501 W504 scrapy/utils/http.py F403 E226 scrapy/utils/httpobj.py E501 @@ -150,7 +149,7 @@ flake8-ignore = scrapy/crawler.py E501 scrapy/dupefilters.py E501 E202 scrapy/exceptions.py E501 - scrapy/exporters.py E501 E261 E226 + scrapy/exporters.py E501 E226 scrapy/interfaces.py E501 scrapy/item.py E501 E128 scrapy/link.py E501 @@ -171,7 +170,7 @@ flake8-ignore = tests/pipelines.py F841 E226 tests/spiders.py E501 E127 tests/test_closespider.py E501 E127 - tests/test_command_fetch.py E501 E261 + tests/test_command_fetch.py E501 tests/test_command_parse.py E501 E128 E303 E226 tests/test_command_shell.py E501 E128 tests/test_commands.py E128 E501 @@ -179,7 +178,7 @@ flake8-ignore = tests/test_crawl.py E501 E741 E265 tests/test_crawler.py F841 E306 E501 tests/test_dependencies.py F841 E501 E305 - tests/test_downloader_handlers.py E124 E127 E128 E225 E261 E265 E501 E502 E701 E126 E226 E123 + tests/test_downloader_handlers.py E124 E127 E128 E225 E265 E501 E502 E701 E126 E226 E123 tests/test_downloadermiddleware.py E501 tests/test_downloadermiddleware_ajaxcrawlable.py E501 tests/test_downloadermiddleware_cookies.py E731 E741 E501 E128 E303 E265 E126 @@ -194,13 +193,13 @@ flake8-ignore = tests/test_downloadermiddleware_robotstxt.py E501 tests/test_downloadermiddleware_stats.py E501 tests/test_dupefilters.py E221 E501 E741 W293 W291 E128 E124 - tests/test_engine.py E401 E501 E502 E128 E261 + tests/test_engine.py E401 E501 E502 E128 tests/test_exporters.py E501 E731 E306 E128 E124 tests/test_extension_telnet.py F841 tests/test_feedexport.py E501 F841 E241 tests/test_http_cookies.py E501 tests/test_http_headers.py E501 - tests/test_http_request.py E402 E501 E261 E127 E128 W293 E502 E128 E502 E126 E123 + tests/test_http_request.py E402 E501 E127 E128 W293 E502 E128 E502 E126 E123 tests/test_http_response.py E501 E301 E502 E128 E265 tests/test_item.py E701 E128 F841 E306 tests/test_link.py E501 @@ -212,7 +211,7 @@ flake8-ignore = tests/test_pipeline_crawl.py E131 E501 E128 E126 tests/test_pipeline_files.py E501 W293 E303 E272 E226 tests/test_pipeline_images.py F841 E501 E303 - tests/test_pipeline_media.py E501 E741 E731 E128 E261 E306 E502 + tests/test_pipeline_media.py E501 E741 E731 E128 E306 E502 tests/test_proxy_connect.py E501 E741 tests/test_request_cb_kwargs.py E501 tests/test_responsetypes.py E501 E305 @@ -224,12 +223,12 @@ flake8-ignore = tests/test_spidermiddleware_httperror.py E128 E501 E127 E121 tests/test_spidermiddleware_offsite.py E501 E128 E111 W293 tests/test_spidermiddleware_output_chain.py E501 W293 E226 - tests/test_spidermiddleware_referer.py E501 F841 E125 E201 E261 E124 E501 E241 E121 + tests/test_spidermiddleware_referer.py E501 F841 E125 E201 E124 E501 E241 E121 tests/test_squeues.py E501 E701 E741 tests/test_utils_conf.py E501 E303 E128 tests/test_utils_curl.py E501 tests/test_utils_datatypes.py E402 E501 E305 - tests/test_utils_defer.py E306 E261 E501 F841 E226 + tests/test_utils_defer.py E306 E501 F841 E226 tests/test_utils_deprecate.py F841 E306 E501 tests/test_utils_http.py E501 E502 E128 W504 tests/test_utils_iterators.py E501 E128 E129 E303 E241 @@ -240,7 +239,7 @@ flake8-ignore = tests/test_utils_response.py E501 tests/test_utils_signal.py E741 F841 E731 E226 tests/test_utils_sitemap.py E128 E501 E124 - tests/test_utils_spider.py E261 E305 + tests/test_utils_spider.py E305 tests/test_utils_template.py E305 tests/test_utils_url.py E501 E127 E305 E211 E125 E501 E226 E241 E126 E123 tests/test_webclient.py E501 E128 E122 E303 E402 E306 E226 E241 E123 E126 diff --git a/scrapy/core/downloader/webclient.py b/scrapy/core/downloader/webclient.py index 798346f19..f368c3bae 100644 --- a/scrapy/core/downloader/webclient.py +++ b/scrapy/core/downloader/webclient.py @@ -42,7 +42,7 @@ class ScrapyHTTPPageGetter(HTTPClient): delimiter = b'\n' def connectionMade(self): - self.headers = Headers() # bucket for response headers + self.headers = Headers() # bucket for response headers # Method command self.sendCommand(self.factory.method, self.factory.path) diff --git a/scrapy/core/engine.py b/scrapy/core/engine.py index fa913e528..829e69993 100644 --- a/scrapy/core/engine.py +++ b/scrapy/core/engine.py @@ -25,7 +25,7 @@ class Slot(object): def __init__(self, start_requests, close_if_idle, nextcall, scheduler): self.closing = False - self.inprogress = set() # requests in progress + self.inprogress = set() # requests in progress self.start_requests = iter(start_requests) self.close_if_idle = close_if_idle self.nextcall = nextcall diff --git a/scrapy/core/scraper.py b/scrapy/core/scraper.py index db463f989..b3d585cce 100644 --- a/scrapy/core/scraper.py +++ b/scrapy/core/scraper.py @@ -123,7 +123,7 @@ class Scraper(object): callback/errback""" assert isinstance(response, (Response, Failure)) - dfd = self._scrape2(response, request, spider) # returns spiders processed output + dfd = self._scrape2(response, request, spider) # returns spiders processed output dfd.addErrback(self.handle_spider_error, request, response, spider) dfd.addCallback(self.handle_spider_output, request, response, spider) return dfd diff --git a/scrapy/exporters.py b/scrapy/exporters.py index fcb55da67..5bf131312 100644 --- a/scrapy/exporters.py +++ b/scrapy/exporters.py @@ -199,7 +199,7 @@ class CsvItemExporter(BaseItemExporter): line_buffering=False, write_through=True, encoding=self.encoding, - newline='' # Windows needs this https://github.com/scrapy/scrapy/issues/3034 + newline='' # Windows needs this https://github.com/scrapy/scrapy/issues/3034 ) self.csv_writer = csv.writer(self.stream, **kwargs) self._headers_not_written = True diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index 5c9678c01..1e163e1fb 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -84,8 +84,8 @@ DOWNLOADER = 'scrapy.core.downloader.Downloader' DOWNLOADER_HTTPCLIENTFACTORY = 'scrapy.core.downloader.webclient.ScrapyHTTPClientFactory' DOWNLOADER_CLIENTCONTEXTFACTORY = 'scrapy.core.downloader.contextfactory.ScrapyClientContextFactory' DOWNLOADER_CLIENT_TLS_CIPHERS = 'DEFAULT' -DOWNLOADER_CLIENT_TLS_METHOD = 'TLS' # Use highest TLS/SSL protocol version supported by the platform, - # also allowing negotiation +# Use highest TLS/SSL protocol version supported by the platform, also allowing negotiation: +DOWNLOADER_CLIENT_TLS_METHOD = 'TLS' DOWNLOADER_CLIENT_TLS_VERBOSE_LOGGING = False DOWNLOADER_MIDDLEWARES = {} diff --git a/scrapy/spiders/feed.py b/scrapy/spiders/feed.py index 197812a26..c566f0236 100644 --- a/scrapy/spiders/feed.py +++ b/scrapy/spiders/feed.py @@ -100,8 +100,8 @@ class CSVFeedSpider(Spider): and the file's headers. """ - delimiter = None # When this is None, python's csv module's default delimiter is used - quotechar = None # When this is None, python's csv module's default quotechar is used + delimiter = None # When this is None, python's csv module's default delimiter is used + quotechar = None # When this is None, python's csv module's default quotechar is used headers = None def process_results(self, response, results): diff --git a/scrapy/utils/console.py b/scrapy/utils/console.py index 688e28c34..7eb40f0ce 100644 --- a/scrapy/utils/console.py +++ b/scrapy/utils/console.py @@ -47,7 +47,7 @@ def _embed_ptpython_shell(namespace={}, banner=''): def _embed_standard_shell(namespace={}, banner=''): """Start a standard python shell""" import code - try: # readline module is only available on unix systems + try: # readline module is only available on unix systems import readline except ImportError: pass @@ -72,9 +72,9 @@ def get_shell_embed_func(shells=None, known_shells=None): """Return the first acceptable shell-embed function from a given list of shell names. """ - if shells is None: # list, preference order of shells + if shells is None: # list, preference order of shells shells = DEFAULT_PYTHON_SHELLS.keys() - if known_shells is None: # available embeddable shells + if known_shells is None: # available embeddable shells known_shells = DEFAULT_PYTHON_SHELLS.copy() for shell in shells: if shell in known_shells: @@ -97,5 +97,5 @@ def start_python_console(namespace=None, banner='', shells=None): shell = get_shell_embed_func(shells) if shell is not None: shell(namespace=namespace, banner=banner) - except SystemExit: # raised when using exit() in python code.interact + except SystemExit: # raised when using exit() in python code.interact pass diff --git a/tests/test_command_fetch.py b/tests/test_command_fetch.py index 3fa3ed930..9d3c8fe73 100644 --- a/tests/test_command_fetch.py +++ b/tests/test_command_fetch.py @@ -29,6 +29,6 @@ class FetchTest(ProcessTest, SiteTest, unittest.TestCase): @defer.inlineCallbacks def test_headers(self): _, out, _ = yield self.execute([self.url('/text'), '--headers']) - out = out.replace(b'\r', b'') # required on win32 + out = out.replace(b'\r', b'') # required on win32 assert b'Server: TwistedWeb' in out, out assert b'Content-Type: text/plain' in out diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 60124b93f..2db2417e8 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -349,7 +349,7 @@ class HttpTestCase(unittest.TestCase): return self.download_request(request, Spider('foo')).addCallback(_test) def test_payload(self): - body = b'1'*100 # PayloadResource requires body length to be 100 + body = b'1'*100 # PayloadResource requires body length to be 100 request = Request(self.getURL('payload'), method='POST', body=body) d = self.download_request(request, Spider('foo')) d.addCallback(lambda r: r.body) diff --git a/tests/test_engine.py b/tests/test_engine.py index 002c4e6bb..537df8d91 100644 --- a/tests/test_engine.py +++ b/tests/test_engine.py @@ -271,7 +271,6 @@ class EngineTest(unittest.TestCase): self.run.signals_catched[signals.spider_opened]) self.assertEqual({'spider': self.run.spider}, self.run.signals_catched[signals.spider_idle]) - self.run.signals_catched[signals.spider_closed].pop('spider_stats', None) # XXX: remove for scrapy 0.17 self.assertEqual({'spider': self.run.spider, 'reason': 'finished'}, self.run.signals_catched[signals.spider_closed]) diff --git a/tests/test_http_request.py b/tests/test_http_request.py index a98aa1e6f..57e7b457d 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -149,7 +149,7 @@ class RequestTest(unittest.TestCase): r2 = self.request_class(url="http://www.example.com/", body=b"") assert isinstance(r2.body, bytes) - self.assertEqual(r2.encoding, 'utf-8') # default encoding + self.assertEqual(r2.encoding, 'utf-8') # default encoding r3 = self.request_class(url="http://www.example.com/", body=u"Price: \xa3100", encoding='utf-8') assert isinstance(r3.body, bytes) diff --git a/tests/test_pipeline_media.py b/tests/test_pipeline_media.py index 0d23f51cc..1fcc5799e 100644 --- a/tests/test_pipeline_media.py +++ b/tests/test_pipeline_media.py @@ -238,10 +238,10 @@ class MediaPipelineTestCase(BaseMediaPipelineTestCase): self.assertEqual(new_item['results'], [(True, rsp1), (False, fail)]) m = self.pipe._mockcalled # only once - self.assertEqual(m[0], 'get_media_requests') # first hook called + self.assertEqual(m[0], 'get_media_requests') # first hook called self.assertEqual(m.count('get_media_requests'), 1) self.assertEqual(m.count('item_completed'), 1) - self.assertEqual(m[-1], 'item_completed') # last hook called + self.assertEqual(m[-1], 'item_completed') # last hook called # twice, one per request self.assertEqual(m.count('media_to_download'), 2) # one to handle success and other for failure @@ -252,7 +252,7 @@ class MediaPipelineTestCase(BaseMediaPipelineTestCase): def test_get_media_requests(self): # returns single Request (without callback) req = Request('http://url') - item = dict(requests=req) # pass a single item + item = dict(requests=req) # pass a single item new_item = yield self.pipe.process_item(item, self.spider) assert new_item is item assert request_fingerprint(req) in self.info.downloaded diff --git a/tests/test_spidermiddleware_referer.py b/tests/test_spidermiddleware_referer.py index ecec6135d..7cc17600c 100644 --- a/tests/test_spidermiddleware_referer.py +++ b/tests/test_spidermiddleware_referer.py @@ -548,8 +548,8 @@ class TestReferrerOnRedirect(TestRefererMiddleware): (301, 'http://scrapytest.org/3'), (301, 'http://scrapytest.org/4'), ), - b'http://scrapytest.org/1', # expected initial referer - b'http://scrapytest.org/1', # expected referer for the redirection request + b'http://scrapytest.org/1', # expected initial referer + b'http://scrapytest.org/1', # expected referer for the redirection request ), ( 'https://scrapytest.org/1', 'https://scrapytest.org/2', @@ -609,8 +609,8 @@ class TestReferrerOnRedirectNoReferrer(TestReferrerOnRedirect): (301, 'http://scrapytest.org/3'), (301, 'http://scrapytest.org/4'), ), - None, # expected initial "Referer" - None, # expected "Referer" for the redirection request + None, # expected initial "Referer" + None, # expected "Referer" for the redirection request ), ( 'https://scrapytest.org/1', 'https://scrapytest.org/2', @@ -648,8 +648,8 @@ class TestReferrerOnRedirectSameOrigin(TestReferrerOnRedirect): (301, 'http://scrapytest.org/103'), (301, 'http://scrapytest.org/104'), ), - b'http://scrapytest.org/101', # expected initial "Referer" - b'http://scrapytest.org/101', # expected referer for the redirection request + b'http://scrapytest.org/101', # expected initial "Referer" + b'http://scrapytest.org/101', # expected referer for the redirection request ), ( 'https://scrapytest.org/201', 'https://scrapytest.org/202', @@ -757,8 +757,8 @@ class TestReferrerOnRedirectOriginWhenCrossOrigin(TestReferrerOnRedirect): (301, 'http://scrapytest.org/103'), (301, 'http://scrapytest.org/104'), ), - b'http://scrapytest.org/101', # expected initial referer - b'http://scrapytest.org/101', # expected referer for the redirection request + b'http://scrapytest.org/101', # expected initial referer + b'http://scrapytest.org/101', # expected referer for the redirection request ), ( 'https://scrapytest.org/201', 'https://scrapytest.org/202', @@ -827,8 +827,8 @@ class TestReferrerOnRedirectStrictOriginWhenCrossOrigin(TestReferrerOnRedirect): (301, 'http://scrapytest.org/103'), (301, 'http://scrapytest.org/104'), ), - b'http://scrapytest.org/101', # expected initial referer - b'http://scrapytest.org/101', # expected referer for the redirection request + b'http://scrapytest.org/101', # expected initial referer + b'http://scrapytest.org/101', # expected referer for the redirection request ), ( 'https://scrapytest.org/201', 'https://scrapytest.org/202', diff --git a/tests/test_utils_defer.py b/tests/test_utils_defer.py index 49c2befb5..dfbe71ae2 100644 --- a/tests/test_utils_defer.py +++ b/tests/test_utils_defer.py @@ -14,8 +14,8 @@ class MustbeDeferredTest(unittest.TestCase): return steps dfd = mustbe_deferred(_append, 1) - dfd.addCallback(self.assertEqual, [1, 2]) # it is [1] with maybeDeferred - steps.append(2) # add another value, that should be catched by assertEqual + dfd.addCallback(self.assertEqual, [1, 2]) # it is [1] with maybeDeferred + steps.append(2) # add another value, that should be catched by assertEqual return dfd def test_unfired_deferred(self): @@ -27,8 +27,8 @@ class MustbeDeferredTest(unittest.TestCase): return dfd dfd = mustbe_deferred(_append, 1) - dfd.addCallback(self.assertEqual, [1, 2]) # it is [1] with maybeDeferred - steps.append(2) # add another value, that should be catched by assertEqual + dfd.addCallback(self.assertEqual, [1, 2]) # it is [1] with maybeDeferred + steps.append(2) # add another value, that should be catched by assertEqual return dfd diff --git a/tests/test_utils_spider.py b/tests/test_utils_spider.py index edeeacc80..ee7d17062 100644 --- a/tests/test_utils_spider.py +++ b/tests/test_utils_spider.py @@ -1,20 +1,16 @@ import unittest + +from scrapy import Spider from scrapy.http import Request from scrapy.item import BaseItem from scrapy.utils.spider import iterate_spider_output, iter_spider_classes -from scrapy.spiders import CrawlSpider - -class MyBaseSpider(CrawlSpider): - pass # abstract spider - - -class MySpider1(MyBaseSpider): +class MySpider1(Spider): name = 'myspider1' -class MySpider2(MyBaseSpider): +class MySpider2(Spider): name = 'myspider2' @@ -35,5 +31,6 @@ class UtilsSpidersTestCase(unittest.TestCase): it = iter_spider_classes(tests.test_utils_spider) self.assertEqual(set(it), {MySpider1, MySpider2}) + if __name__ == "__main__": unittest.main() From a4ef9750f9058fbe1041bad0fb28f39c693e5659 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Fri, 13 Dec 2019 14:32:06 +0100 Subject: [PATCH 333/387] Fix Flake8-reported issues --- pytest.ini | 2 +- tests/test_linkextractors.py | 12 +++++++----- 2 files changed, 8 insertions(+), 6 deletions(-) diff --git a/pytest.ini b/pytest.ini index 33c34b8e8..02014d1a5 100644 --- a/pytest.ini +++ b/pytest.ini @@ -88,7 +88,7 @@ flake8-ignore = scrapy/http/response/__init__.py E501 E128 W293 W291 scrapy/http/response/text.py E501 W293 E128 E124 # scrapy/linkextractors - scrapy/linkextractors/__init__.py E731 E502 E501 E402 + scrapy/linkextractors/__init__.py E731 E502 E501 E402 W504 scrapy/linkextractors/lxmlhtml.py E501 E731 E226 # scrapy/loader scrapy/loader/__init__.py E501 E502 E128 diff --git a/tests/test_linkextractors.py b/tests/test_linkextractors.py index ebe497913..0ffeaecc3 100644 --- a/tests/test_linkextractors.py +++ b/tests/test_linkextractors.py @@ -514,25 +514,27 @@ class LxmlLinkExtractorTestCase(Base.LinkExtractorTestCase): """Make sure the FilteringLinkExtractor deprecation warning is not issued for LxmlLinkExtractor""" with catch_warnings(record=True) as warnings: - extractor = LxmlLinkExtractor() + LxmlLinkExtractor() self.assertEqual(len(warnings), 0) + class SubclassedItem(LxmlLinkExtractor): pass - subclassed_extractor = SubclassedItem() + + SubclassedItem() self.assertEqual(len(warnings), 0) class FilteringLinkExtractorTest(unittest.TestCase): def test_deprecation_warning(self): - args = [None]*10 + args = [None] * 10 with catch_warnings(record=True) as warnings: - extractor = FilteringLinkExtractor(*args) + FilteringLinkExtractor(*args) self.assertEqual(len(warnings), 1) self.assertEqual(warnings[0].category, ScrapyDeprecationWarning) with catch_warnings(record=True) as warnings: class SubclassedFilteringLinkExtractor(FilteringLinkExtractor): pass - subclassed_extractor = SubclassedFilteringLinkExtractor(*args) + SubclassedFilteringLinkExtractor(*args) self.assertEqual(len(warnings), 1) self.assertEqual(warnings[0].category, ScrapyDeprecationWarning) From afc886e57865e82e63f3f8f3326f481908919086 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 13 Dec 2019 19:34:47 +0500 Subject: [PATCH 334/387] Simplify tox.ini asyncio entries. --- tox.ini | 10 ++++++---- 1 file changed, 6 insertions(+), 4 deletions(-) diff --git a/tox.ini b/tox.ini index a4edae439..795c20233 100644 --- a/tox.ini +++ b/tox.ini @@ -107,14 +107,16 @@ deps = reppy robotexclusionrulesparser +[asyncio] +commands = + py.test --cov=scrapy --cov-report= --reactor=asyncio {posargs:scrapy tests} + [testenv:py35-asyncio] basepython = python3.5 deps = {[testenv]deps} -commands = - py.test --cov=scrapy --cov-report= --reactor=asyncio {posargs:scrapy tests} +commands = {[asyncio]commands} [testenv:py38-asyncio] basepython = python3.8 deps = {[testenv]deps} -commands = - py.test --cov=scrapy --cov-report= --reactor=asyncio {posargs:scrapy tests} +commands = {[asyncio]commands} From a1605cade6286dd5f7f1c9e4c9660d44ed15ed19 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 13 Dec 2019 19:35:09 +0500 Subject: [PATCH 335/387] Hide utils.defer.isfuture(). --- scrapy/utils/defer.py | 7 +++---- 1 file changed, 3 insertions(+), 4 deletions(-) diff --git a/scrapy/utils/defer.py b/scrapy/utils/defer.py index 6a91776c7..530bf0e9d 100644 --- a/scrapy/utils/defer.py +++ b/scrapy/utils/defer.py @@ -2,7 +2,6 @@ Helper functions for dealing with Twisted deferreds """ import asyncio -import asyncio.futures import inspect from twisted.internet import defer, task @@ -121,18 +120,18 @@ def iter_errback(iterable, errback, *a, **kw): errback(failure.Failure(), *a, **kw) -def isfuture(o): +def _isfuture(o): # workaround for Python before 3.5.3 not having asyncio.isfuture if hasattr(asyncio, 'isfuture'): return asyncio.isfuture(o) - return isinstance(o, asyncio.futures.Future) + return isinstance(o, asyncio.Future) def deferred_from_coro(o, asyncio_enabled=False): """Converts a coroutine into a Deferred, or returns the object as is if it isn't a coroutine""" if isinstance(o, defer.Deferred): return o - if isfuture(o) or inspect.isawaitable(o): + if _isfuture(o) or inspect.isawaitable(o): if not asyncio_enabled: # wrapping the coroutine directly into a Deferred, this doesn't work correctly with coroutines # that use asyncio, e.g. "await asyncio.sleep(1)" From e3a3ad4aaff6d07252bb4f65be4667e21070aea1 Mon Sep 17 00:00:00 2001 From: marc <Marc> Date: Sat, 14 Dec 2019 10:34:31 +0100 Subject: [PATCH 336/387] remove reference to old (Python 2.7) environment --- docs/contributing.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/contributing.rst b/docs/contributing.rst index 234c4bcee..eaaf86c29 100644 --- a/docs/contributing.rst +++ b/docs/contributing.rst @@ -229,7 +229,7 @@ You can also specify a comma-separated list of environmets, and use :ref:`tox’ parallel mode <tox:parallel_mode>` to run the tests on multiple environments in parallel:: - tox -e py27,py36 -p auto + tox -e py36,py38 -p auto To pass command-line options to :doc:`pytest <pytest:index>`, add them after ``--`` in your call to :doc:`tox <tox:index>`. Using ``--`` overrides the From 1aab20e1cec9e2e9aac4ce9efd37b07811b8f7d6 Mon Sep 17 00:00:00 2001 From: marc <Marc> Date: Sat, 14 Dec 2019 10:37:31 +0100 Subject: [PATCH 337/387] update copyright notice year --- docs/conf.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/conf.py b/docs/conf.py index 914d1d05f..5fdaa4e4c 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -50,7 +50,7 @@ master_doc = 'index' # General information about the project. project = u'Scrapy' -copyright = u'2008–2018, Scrapy developers' +copyright = u'2008–2020, Scrapy developers' # The version info for the project you're documenting, acts as replacement for # |version| and |release|, also used in various other places throughout the From a59bb279d18966e6a47f71d21c690598fc69678e Mon Sep 17 00:00:00 2001 From: marc <Marc> Date: Sun, 15 Dec 2019 17:33:00 +0100 Subject: [PATCH 338/387] add year through code --- docs/conf.py | 9 +++++---- 1 file changed, 5 insertions(+), 4 deletions(-) diff --git a/docs/conf.py b/docs/conf.py index 5fdaa4e4c..ed56c5cd1 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -12,6 +12,7 @@ # serve to show the default. import sys +from datetime import datetime from os import path # If your extensions are in another directory, add it here. If the directory @@ -49,8 +50,8 @@ source_suffix = '.rst' master_doc = 'index' # General information about the project. -project = u'Scrapy' -copyright = u'2008–2020, Scrapy developers' +project = 'Scrapy' +copyright = '2008–{}, Scrapy developers'.format(datetime.now().year) # The version info for the project you're documenting, acts as replacement for # |version| and |release|, also used in various other places throughout the @@ -194,8 +195,8 @@ htmlhelp_basename = 'Scrapydoc' # Grouping the document tree into LaTeX files. List of tuples # (source start file, target name, title, author, document class [howto/manual]). latex_documents = [ - ('index', 'Scrapy.tex', u'Scrapy Documentation', - u'Scrapy developers', 'manual'), + ('index', 'Scrapy.tex', 'Scrapy Documentation', + 'Scrapy developers', 'manual'), ] # The name of an image file (relative to this directory) to place at the top of From 2db7d453788f5c638d0921b0f7f8bab58e2a58bc Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Mon, 16 Dec 2019 19:24:25 +0500 Subject: [PATCH 339/387] Enable skipping tests based on --reactor. --- conftest.py | 8 ++++++++ pytest.ini | 2 ++ 2 files changed, 10 insertions(+) diff --git a/conftest.py b/conftest.py index 64136b48d..56d552953 100644 --- a/conftest.py +++ b/conftest.py @@ -35,6 +35,14 @@ def pytest_collection_modifyitems(session, config, items): except ImportError: pass + @pytest.fixture() def reactor_pytest(request): request.cls.reactor_pytest = request.config.getoption("--reactor") + return request.cls.reactor_pytest + + +@pytest.fixture(autouse=True) +def only_asyncio(request, reactor_pytest): + if request.node.get_closest_marker('only_asyncio') and reactor_pytest != 'asyncio': + pytest.skip('This test is only run with --reactor-asyncio') diff --git a/pytest.ini b/pytest.ini index 336ef041d..7b62a1bd8 100644 --- a/pytest.ini +++ b/pytest.ini @@ -19,6 +19,8 @@ addopts = --ignore=docs/topics/telnetconsole.rst --ignore=docs/utils twisted = 1 +markers = + only_asyncio: marks tests as only enabled when --reactor=asyncio is passed flake8-ignore = # Files that are only meant to provide top-level imports are expected not # to use any of their imports: From 451e7a616e6d23e88528836296a9abdc887834c6 Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Fri, 12 Jul 2019 00:26:01 -0300 Subject: [PATCH 340/387] Scan callbacks/errbacks for return statements with values different than None --- scrapy/core/scraper.py | 19 ++++-- scrapy/utils/datatypes.py | 31 ++++++++- scrapy/utils/misc.py | 46 +++++++++++++ tests/test_utils_datatypes.py | 66 ++++++++++++++++++- ...t_return_with_argument_inside_generator.py | 37 +++++++++++ 5 files changed, 190 insertions(+), 9 deletions(-) create mode 100644 tests/test_utils_misc/test_return_with_argument_inside_generator.py diff --git a/scrapy/core/scraper.py b/scrapy/core/scraper.py index b3d585cce..99114d3bb 100644 --- a/scrapy/core/scraper.py +++ b/scrapy/core/scraper.py @@ -9,7 +9,7 @@ from twisted.internet import defer from scrapy.utils.defer import defer_result, defer_succeed, parallel, iter_errback from scrapy.utils.spider import iterate_spider_output -from scrapy.utils.misc import load_object +from scrapy.utils.misc import load_object, warn_on_generator_with_return_value from scrapy.utils.log import logformatter_adapter, failure_to_exc_info from scrapy.exceptions import CloseSpider, DropItem, IgnoreRequest from scrapy import signals @@ -18,6 +18,7 @@ from scrapy.item import BaseItem from scrapy.core.spidermw import SpiderMiddlewareManager from scrapy.utils.request import referer_str + logger = logging.getLogger(__name__) @@ -99,11 +100,13 @@ class Scraper(object): def enqueue_scrape(self, response, request, spider): slot = self.slot dfd = slot.add_response_request(response, request) + def finish_scraping(_): slot.finish_response(response, request) self._check_if_closing(spider, slot) self._scrape_next(spider, slot) return _ + dfd.addBoth(finish_scraping) dfd.addErrback( lambda f: logger.error('Scraper bug processing %(request)s', @@ -123,7 +126,7 @@ class Scraper(object): callback/errback""" assert isinstance(response, (Response, Failure)) - dfd = self._scrape2(response, request, spider) # returns spiders processed output + dfd = self._scrape2(response, request, spider) # returns spider's processed output dfd.addErrback(self.handle_spider_error, request, response, spider) dfd.addCallback(self.handle_spider_output, request, response, spider) return dfd @@ -142,7 +145,10 @@ class Scraper(object): def call_spider(self, result, request, spider): result.request = request dfd = defer_result(result) - dfd.addCallbacks(callback=request.callback or spider.parse, + callback = request.callback or spider.parse + warn_on_generator_with_return_value(spider, callback) + warn_on_generator_with_return_value(spider, request.errback) + dfd.addCallbacks(callback=callback, errback=request.errback, callbackKeywords=request.cb_kwargs) return dfd.addCallback(iterate_spider_output) @@ -172,8 +178,8 @@ class Scraper(object): if not result: return defer_succeed(None) it = iter_errback(result, self.handle_spider_error, request, response, spider) - dfd = parallel(it, self.concurrent_items, - self._process_spidermw_output, request, response, spider) + dfd = parallel(it, self.concurrent_items, self._process_spidermw_output, + request, response, spider) return dfd def _process_spidermw_output(self, output, request, response, spider): @@ -200,8 +206,7 @@ class Scraper(object): """Log and silence errors that come from the engine (typically download errors that got propagated thru here) """ - if (isinstance(download_failure, Failure) and - not download_failure.check(IgnoreRequest)): + if isinstance(download_failure, Failure) and not download_failure.check(IgnoreRequest): if download_failure.frames: logger.error('Error downloading %(request)s', {'request': request}, diff --git a/scrapy/utils/datatypes.py b/scrapy/utils/datatypes.py index ffd1537c3..a52bbc70e 100644 --- a/scrapy/utils/datatypes.py +++ b/scrapy/utils/datatypes.py @@ -8,6 +8,7 @@ This module must not depend on any module outside the Standard Library. import collections import copy import warnings +import weakref from collections.abc import Mapping from scrapy.exceptions import ScrapyDeprecationWarning @@ -240,7 +241,6 @@ class LocalCache(collections.OrderedDict): """Dictionary with a finite number of keys. Older items expires first. - """ def __init__(self, limit=None): @@ -254,6 +254,35 @@ class LocalCache(collections.OrderedDict): super(LocalCache, self).__setitem__(key, value) +class LocalWeakReferencedCache(weakref.WeakKeyDictionary): + """ + A weakref.WeakKeyDictionary implementation that uses LocalCache as its + underlying data structure, making it ordered and capable of being size-limited. + + Useful for memoization, while avoiding keeping received + arguments in memory only because of the cached references. + + Note: like LocalCache and unlike weakref.WeakKeyDictionary, + it cannot be instantiated with an initial dictionary. + """ + + def __init__(self, limit=None): + super(LocalWeakReferencedCache, self).__init__() + self.data = LocalCache(limit=limit) + + def __setitem__(self, key, value): + try: + super(LocalWeakReferencedCache, self).__setitem__(key, value) + except TypeError: + pass # key is not weak-referenceable, skip caching + + def __getitem__(self, key): + try: + return super(LocalWeakReferencedCache, self).__getitem__(key) + except TypeError: + return None # key is not weak-referenceable, it's not cached + + class SequenceExclude(object): """Object to test if an item is NOT within some sequence.""" diff --git a/scrapy/utils/misc.py b/scrapy/utils/misc.py index 9955fb1e7..cb0ee5af3 100644 --- a/scrapy/utils/misc.py +++ b/scrapy/utils/misc.py @@ -1,13 +1,18 @@ """Helper functions which don't fit anywhere else""" +import ast +import inspect import os import re import hashlib +import warnings from contextlib import contextmanager from importlib import import_module from pkgutil import iter_modules +from textwrap import dedent from w3lib.html import replace_entities +from scrapy.utils.datatypes import LocalWeakReferencedCache from scrapy.utils.python import flatten, to_unicode from scrapy.item import BaseItem @@ -161,3 +166,44 @@ def set_environ(**kwargs): del os.environ[k] else: os.environ[k] = v + + +_generator_callbacks_cache = LocalWeakReferencedCache(limit=128) + + +def is_generator_with_return_value(callable): + """ + Returns True if a callable is a generator function which includes a + 'return' statement with a value different than None, False otherwise + """ + if callable in _generator_callbacks_cache: + return _generator_callbacks_cache[callable] + + def returns_none(return_node): + value = return_node.value + return value is None or isinstance(value, ast.NameConstant) and value.value is None + + if inspect.isgeneratorfunction(callable): + tree = ast.parse(dedent(inspect.getsource(callable))) + for node in ast.walk(tree): + if isinstance(node, ast.Return) and not returns_none(node): + _generator_callbacks_cache[callable] = True + return _generator_callbacks_cache[callable] + + _generator_callbacks_cache[callable] = False + return _generator_callbacks_cache[callable] + + +def warn_on_generator_with_return_value(spider, callable): + """ + Logs a warning if a callable is a generator function and includes + a 'return' statement with a value different than None + """ + if is_generator_with_return_value(callable): + warnings.warn( + 'The "{}.{}" method is a generator and includes a "return" statement with a ' + 'value different than None. This could lead to unexpected behaviour. Please see ' + 'https://docs.python.org/3/reference/simple_stmts.html#the-return-statement ' + 'for details about the semantics of the "return" statement within generators' + .format(spider.__class__.__name__, callable.__name__), stacklevel=2, + ) diff --git a/tests/test_utils_datatypes.py b/tests/test_utils_datatypes.py index 38a25778e..e5aa56eb9 100644 --- a/tests/test_utils_datatypes.py +++ b/tests/test_utils_datatypes.py @@ -2,7 +2,9 @@ import copy import unittest from collections.abc import Mapping, MutableMapping -from scrapy.utils.datatypes import CaselessDict, LocalCache, SequenceExclude +from scrapy.http import Request +from scrapy.utils.datatypes import CaselessDict, LocalCache, LocalWeakReferencedCache, SequenceExclude +from scrapy.utils.python import garbage_collect __doctests__ = ['scrapy.utils.datatypes'] @@ -255,5 +257,67 @@ class LocalCacheTest(unittest.TestCase): self.assertEqual(cache[str(x)], x) +class LocalWeakReferencedCacheTest(unittest.TestCase): + + def test_cache_with_limit(self): + cache = LocalWeakReferencedCache(limit=2) + r1 = Request('https://example.org') + r2 = Request('https://example.com') + r3 = Request('https://example.net') + cache[r1] = 1 + cache[r2] = 2 + cache[r3] = 3 + self.assertEqual(len(cache), 2) + self.assertNotIn(r1, cache) + self.assertIn(r2, cache) + self.assertIn(r3, cache) + self.assertEqual(cache[r2], 2) + self.assertEqual(cache[r3], 3) + del r2 + + # PyPy takes longer to collect dead references + garbage_collect() + + self.assertEqual(len(cache), 1) + + def test_cache_non_weak_referenceable_objects(self): + cache = LocalWeakReferencedCache() + k1 = None + k2 = 1 + k3 = [1, 2, 3] + cache[k1] = 1 + cache[k2] = 2 + cache[k3] = 3 + self.assertNotIn(k1, cache) + self.assertNotIn(k2, cache) + self.assertNotIn(k3, cache) + self.assertEqual(len(cache), 0) + + def test_cache_without_limit(self): + max = 10**4 + cache = LocalWeakReferencedCache() + refs = [] + for x in range(max): + refs.append(Request('https://example.org/{}'.format(x))) + cache[refs[-1]] = x + self.assertEqual(len(cache), max) + for i, r in enumerate(refs): + self.assertIn(r, cache) + self.assertEqual(cache[r], i) + del r # delete reference to the last object in the list + + # delete half of the objects, make sure that is reflected in the cache + for _ in range(max // 2): + refs.pop() + + # PyPy takes longer to collect dead references + garbage_collect() + + self.assertEqual(len(cache), max // 2) + for i, r in enumerate(refs): + self.assertIn(r, cache) + self.assertEqual(cache[r], i) + + if __name__ == "__main__": unittest.main() diff --git a/tests/test_utils_misc/test_return_with_argument_inside_generator.py b/tests/test_utils_misc/test_return_with_argument_inside_generator.py new file mode 100644 index 000000000..bdbec1beb --- /dev/null +++ b/tests/test_utils_misc/test_return_with_argument_inside_generator.py @@ -0,0 +1,37 @@ +import unittest + +from scrapy.utils.misc import is_generator_with_return_value + + +class UtilsMiscPy3TestCase(unittest.TestCase): + + def test_generators_with_return_statements(self): + def f(): + yield 1 + return 2 + + def g(): + yield 1 + return 'asdf' + + def h(): + yield 1 + return None + + def i(): + yield 1 + return + + def j(): + yield 1 + + def k(): + yield 1 + yield from g() + + assert is_generator_with_return_value(f) + assert is_generator_with_return_value(g) + assert not is_generator_with_return_value(h) + assert not is_generator_with_return_value(i) + assert not is_generator_with_return_value(j) + assert not is_generator_with_return_value(k) # not recursive From 039e6fe6919341dbfd864c4a406d35389c8e2992 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Mon, 16 Dec 2019 20:17:41 +0500 Subject: [PATCH 341/387] Refactor install_asyncio_reactor slightly. --- scrapy/utils/asyncio.py | 16 +++++++++------- 1 file changed, 9 insertions(+), 7 deletions(-) diff --git a/scrapy/utils/asyncio.py b/scrapy/utils/asyncio.py index b5d5f92d9..b53c8a8b0 100644 --- a/scrapy/utils/asyncio.py +++ b/scrapy/utils/asyncio.py @@ -1,3 +1,8 @@ +from contextlib import suppress + +from twisted.internet.error import ReactorAlreadyInstalledError + + def install_asyncio_reactor(): """ Tries to install AsyncioSelectorReactor """ @@ -5,13 +10,10 @@ def install_asyncio_reactor(): import asyncio from twisted.internet import asyncioreactor except ImportError: - pass - else: - from twisted.internet.error import ReactorAlreadyInstalledError - try: - asyncioreactor.install(asyncio.get_event_loop()) - except ReactorAlreadyInstalledError: - pass + return + + with suppress(ReactorAlreadyInstalledError): + asyncioreactor.install(asyncio.get_event_loop()) def is_asyncio_reactor_installed(): From 900de7c14607fbe2936fa682d03747916337f075 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Mon, 16 Dec 2019 21:11:58 +0500 Subject: [PATCH 342/387] Fix the reactor_pytest fixture. --- conftest.py | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/conftest.py b/conftest.py index 56d552953..6d9696a3f 100644 --- a/conftest.py +++ b/conftest.py @@ -36,8 +36,11 @@ def pytest_collection_modifyitems(session, config, items): pass -@pytest.fixture() +@pytest.fixture(scope='class') def reactor_pytest(request): + if not request.cls: + # doctests + return request.cls.reactor_pytest = request.config.getoption("--reactor") return request.cls.reactor_pytest From 5980d3bbff67d060a1a1b15372293ced972dbe8b Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 10 Sep 2019 14:23:11 +0500 Subject: [PATCH 343/387] Add simple tests for pipelines. --- tests/test_pipelines.py | 71 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 71 insertions(+) create mode 100644 tests/test_pipelines.py diff --git a/tests/test_pipelines.py b/tests/test_pipelines.py new file mode 100644 index 000000000..bc53f5427 --- /dev/null +++ b/tests/test_pipelines.py @@ -0,0 +1,71 @@ +from twisted.internet import defer +from twisted.internet.defer import Deferred +from twisted.trial import unittest + +from scrapy import Spider, signals, Request +from scrapy.utils.test import get_crawler + +from tests.mockserver import MockServer + + +class SimplePipeline: + def process_item(self, item, spider): + item['pipeline_passed'] = True + return item + + +class DeferredPipeline: + def cb(self, item): + item['pipeline_passed'] = True + return item + + def process_item(self, item, spider): + d = Deferred() + d.addCallback(self.cb) + d.callback(item) + return d + + +class ItemSpider(Spider): + name = 'itemspider' + + def start_requests(self): + yield Request(self.mockserver.url('/status?n=200')) + + def parse(self, response): + return {'field': 42} + + +class PipelineTestCase(unittest.TestCase): + def setUp(self): + self.mockserver = MockServer() + self.mockserver.__enter__() + + def tearDown(self): + self.mockserver.__exit__(None, None, None) + + def _on_item_scraped(self, item): + self.assertIsInstance(item, dict) + self.assertTrue(item.get('pipeline_passed')) + self.items.append(item) + + def _create_crawler(self, pipeline_class): + settings = { + 'ITEM_PIPELINES': {__name__ + '.' + pipeline_class.__name__: 1}, + } + crawler = get_crawler(ItemSpider, settings) + crawler.signals.connect(self._on_item_scraped, signals.item_scraped) + self.items = [] + return crawler + + @defer.inlineCallbacks + def test_simple_pipeline(self): + crawler = self._create_crawler(SimplePipeline) + yield crawler.crawl(mockserver=self.mockserver) + self.assertEqual(len(self.items), 1) + + @defer.inlineCallbacks + def test_deferred_pipeline(self): + crawler = self._create_crawler(DeferredPipeline) + yield crawler.crawl(mockserver=self.mockserver) + self.assertEqual(len(self.items), 1) From 7d0096da6e37e508280e3af7388122bcdf0e3dcf Mon Sep 17 00:00:00 2001 From: apu <1173372284@qq.com> Date: Tue, 17 Dec 2019 09:47:01 +0800 Subject: [PATCH 344/387] Fix mail attachs tcmime *** (#4229) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit When the file name consists of alphanumeric characters, it is normal to receive the attachment name. However,However, problems will occur if the file name is changed to Chinese. This has nothing to do with the file type --- scrapy/mail.py | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/scrapy/mail.py b/scrapy/mail.py index 891bb5e09..9655b8114 100644 --- a/scrapy/mail.py +++ b/scrapy/mail.py @@ -73,8 +73,7 @@ class MailSender(object): part = MIMEBase(*mimetype.split('/')) part.set_payload(f.read()) Encoders.encode_base64(part) - part.add_header('Content-Disposition', 'attachment; filename="%s"' \ - % attach_name) + part.add_header('Content-Disposition', 'attachment', filename=attach_name) msg.attach(part) else: msg.set_payload(body) From 63cf5c75c850e48b6f267574e4a8f7ae3293deac Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Marc=20Hern=C3=A1ndez?= <noviluni@gmail.com> Date: Tue, 17 Dec 2019 13:53:15 +0100 Subject: [PATCH 345/387] Fix E502: backslash is redundant between brackets (#4238) --- pytest.ini | 40 +++++------ scrapy/cmdline.py | 4 +- scrapy/commands/fetch.py | 11 ++- scrapy/commands/startproject.py | 4 +- scrapy/contracts/default.py | 4 +- scrapy/core/downloader/handlers/s3.py | 4 +- scrapy/core/downloader/webclient.py | 6 +- scrapy/core/spidermw.py | 6 +- .../downloadermiddlewares/httpcompression.py | 5 +- scrapy/extensions/closespider.py | 6 +- scrapy/linkextractors/__init__.py | 3 +- scrapy/middleware.py | 4 +- scrapy/utils/conf.py | 2 +- tests/test_cmdline/__init__.py | 10 ++- tests/test_downloader_handlers.py | 72 ++++++++++--------- tests/test_downloadermiddleware_retry.py | 6 +- tests/test_engine.py | 4 +- tests/test_http_request.py | 34 +++++---- tests/test_http_response.py | 4 +- tests/test_utils_http.py | 8 +-- 20 files changed, 121 insertions(+), 116 deletions(-) diff --git a/pytest.ini b/pytest.ini index 1b23595c0..f088e10ef 100644 --- a/pytest.ini +++ b/pytest.ini @@ -33,45 +33,45 @@ flake8-ignore = scrapy/commands/check.py E501 scrapy/commands/crawl.py E501 scrapy/commands/edit.py E501 - scrapy/commands/fetch.py E401 E501 E128 E502 E731 + scrapy/commands/fetch.py E401 E501 E128 E731 scrapy/commands/genspider.py E128 E501 E502 scrapy/commands/parse.py E128 E501 E731 E226 scrapy/commands/runspider.py E501 scrapy/commands/settings.py E128 scrapy/commands/shell.py E128 E501 E502 - scrapy/commands/startproject.py E502 E127 E501 E128 + scrapy/commands/startproject.py E127 E501 E128 scrapy/commands/version.py E501 E128 # scrapy/contracts scrapy/contracts/__init__.py E501 W504 - scrapy/contracts/default.py E502 E128 + scrapy/contracts/default.py E128 # scrapy/core scrapy/core/engine.py E501 E128 E127 E306 E502 scrapy/core/scheduler.py E501 scrapy/core/scraper.py E501 E306 E128 W504 - scrapy/core/spidermw.py E501 E731 E502 E126 E226 + scrapy/core/spidermw.py E501 E731 E126 E226 scrapy/core/downloader/__init__.py E501 scrapy/core/downloader/contextfactory.py E501 E128 E126 scrapy/core/downloader/middleware.py E501 E502 scrapy/core/downloader/tls.py E501 E305 E241 - scrapy/core/downloader/webclient.py E731 E501 E502 E128 E126 E226 + scrapy/core/downloader/webclient.py E731 E501 E128 E126 E226 scrapy/core/downloader/handlers/__init__.py E501 scrapy/core/downloader/handlers/ftp.py E501 E305 E128 E127 scrapy/core/downloader/handlers/http10.py E501 scrapy/core/downloader/handlers/http11.py E501 - scrapy/core/downloader/handlers/s3.py E501 E502 E128 E126 + scrapy/core/downloader/handlers/s3.py E501 E128 E126 # scrapy/downloadermiddlewares scrapy/downloadermiddlewares/ajaxcrawl.py E501 E226 scrapy/downloadermiddlewares/decompression.py E501 scrapy/downloadermiddlewares/defaultheaders.py E501 scrapy/downloadermiddlewares/httpcache.py E501 E126 - scrapy/downloadermiddlewares/httpcompression.py E502 E128 + scrapy/downloadermiddlewares/httpcompression.py E501 E128 scrapy/downloadermiddlewares/httpproxy.py E501 scrapy/downloadermiddlewares/redirect.py E501 W504 scrapy/downloadermiddlewares/retry.py E501 E126 scrapy/downloadermiddlewares/robotstxt.py E501 scrapy/downloadermiddlewares/stats.py E501 # scrapy/extensions - scrapy/extensions/closespider.py E501 E502 E128 E123 + scrapy/extensions/closespider.py E501 E128 E123 scrapy/extensions/corestats.py E501 scrapy/extensions/feedexport.py E128 E501 scrapy/extensions/httpcache.py E128 E501 E303 @@ -88,10 +88,10 @@ flake8-ignore = scrapy/http/response/__init__.py E501 E128 W293 W291 scrapy/http/response/text.py E501 W293 E128 E124 # scrapy/linkextractors - scrapy/linkextractors/__init__.py E731 E502 E501 E402 + scrapy/linkextractors/__init__.py E731 E501 E402 scrapy/linkextractors/lxmlhtml.py E501 E731 E226 # scrapy/loader - scrapy/loader/__init__.py E501 E502 E128 + scrapy/loader/__init__.py E501 E128 scrapy/loader/processors.py E501 # scrapy/pipelines scrapy/pipelines/files.py E116 E501 E266 @@ -116,7 +116,7 @@ flake8-ignore = scrapy/spiders/sitemap.py E501 # scrapy/utils scrapy/utils/benchserver.py E501 - scrapy/utils/conf.py E402 E502 E501 + scrapy/utils/conf.py E402 E501 scrapy/utils/console.py E306 E305 scrapy/utils/datatypes.py E501 E226 scrapy/utils/decorators.py E501 @@ -145,7 +145,7 @@ flake8-ignore = # scrapy scrapy/__init__.py E402 E501 scrapy/_monkeypatches.py W293 - scrapy/cmdline.py E502 E501 + scrapy/cmdline.py E501 scrapy/crawler.py E501 scrapy/dupefilters.py E501 E202 scrapy/exceptions.py E501 @@ -155,7 +155,7 @@ flake8-ignore = scrapy/link.py E501 scrapy/logformatter.py E501 W293 scrapy/mail.py E402 E128 E501 E502 - scrapy/middleware.py E502 E128 E501 + scrapy/middleware.py E128 E501 scrapy/pqueues.py E501 scrapy/responsetypes.py E128 E501 E305 scrapy/robotstxt.py E501 @@ -178,7 +178,7 @@ flake8-ignore = tests/test_crawl.py E501 E741 E265 tests/test_crawler.py F841 E306 E501 tests/test_dependencies.py F841 E501 E305 - tests/test_downloader_handlers.py E124 E127 E128 E225 E265 E501 E502 E701 E126 E226 E123 + tests/test_downloader_handlers.py E124 E127 E128 E225 E265 E501 E701 E126 E226 E123 tests/test_downloadermiddleware.py E501 tests/test_downloadermiddleware_ajaxcrawlable.py E501 tests/test_downloadermiddleware_cookies.py E731 E741 E501 E128 E303 E265 E126 @@ -189,18 +189,18 @@ flake8-ignore = tests/test_downloadermiddleware_httpcompression.py E501 E251 E126 E123 tests/test_downloadermiddleware_httpproxy.py E501 E128 tests/test_downloadermiddleware_redirect.py E501 E303 E128 E306 E127 E305 - tests/test_downloadermiddleware_retry.py E501 E128 W293 E251 E502 E303 E126 + tests/test_downloadermiddleware_retry.py E501 E128 W293 E251 E303 E126 tests/test_downloadermiddleware_robotstxt.py E501 tests/test_downloadermiddleware_stats.py E501 tests/test_dupefilters.py E221 E501 E741 W293 W291 E128 E124 - tests/test_engine.py E401 E501 E502 E128 + tests/test_engine.py E401 E501 E128 tests/test_exporters.py E501 E731 E306 E128 E124 tests/test_extension_telnet.py F841 tests/test_feedexport.py E501 F841 E241 tests/test_http_cookies.py E501 tests/test_http_headers.py E501 - tests/test_http_request.py E402 E501 E127 E128 W293 E502 E128 E502 E126 E123 - tests/test_http_response.py E501 E301 E502 E128 E265 + tests/test_http_request.py E402 E501 E127 E128 W293 E128 E126 E123 + tests/test_http_response.py E501 E301 E128 E265 tests/test_item.py E701 E128 F841 E306 tests/test_link.py E501 tests/test_linkextractors.py E501 E128 E124 @@ -230,7 +230,7 @@ flake8-ignore = tests/test_utils_datatypes.py E402 E501 E305 tests/test_utils_defer.py E306 E501 F841 E226 tests/test_utils_deprecate.py F841 E306 E501 - tests/test_utils_http.py E501 E502 E128 W504 + tests/test_utils_http.py E501 E128 W504 tests/test_utils_iterators.py E501 E128 E129 E303 E241 tests/test_utils_log.py E741 E226 tests/test_utils_python.py E501 E303 E731 E701 E305 @@ -243,7 +243,7 @@ flake8-ignore = tests/test_utils_template.py E305 tests/test_utils_url.py E501 E127 E305 E211 E125 E501 E226 E241 E126 E123 tests/test_webclient.py E501 E128 E122 E303 E402 E306 E226 E241 E123 E126 - tests/test_cmdline/__init__.py E502 E501 + tests/test_cmdline/__init__.py E501 tests/test_settings/__init__.py E501 E128 tests/test_spiderloader/__init__.py E128 E501 tests/test_utils_misc/__init__.py E501 diff --git a/scrapy/cmdline.py b/scrapy/cmdline.py index 69e917004..ec78f7c91 100644 --- a/scrapy/cmdline.py +++ b/scrapy/cmdline.py @@ -67,7 +67,7 @@ def _pop_command_name(argv): def _print_header(settings, inproject): if inproject: - print("Scrapy %s - project: %s\n" % (scrapy.__version__, \ + print("Scrapy %s - project: %s\n" % (scrapy.__version__, settings['BOT_NAME'])) else: print("Scrapy %s - no active project\n" % scrapy.__version__) @@ -123,7 +123,7 @@ def execute(argv=None, settings=None): inproject = inside_project() cmds = _get_commands_dict(settings, inproject) cmdname = _pop_command_name(argv) - parser = optparse.OptionParser(formatter=optparse.TitledHelpFormatter(), \ + parser = optparse.OptionParser(formatter=optparse.TitledHelpFormatter(), conflict_handler='resolve') if not cmdname: _print_commands(settings, inproject) diff --git a/scrapy/commands/fetch.py b/scrapy/commands/fetch.py index 8a22ebabe..0e149941d 100644 --- a/scrapy/commands/fetch.py +++ b/scrapy/commands/fetch.py @@ -24,12 +24,11 @@ class Command(ScrapyCommand): def add_options(self, parser): ScrapyCommand.add_options(self, parser) - parser.add_option("--spider", dest="spider", - help="use this spider") - parser.add_option("--headers", dest="headers", action="store_true", \ - help="print response HTTP headers instead of body") - parser.add_option("--no-redirect", dest="no_redirect", action="store_true", \ - default=False, help="do not handle HTTP 3xx status codes and print response as-is") + parser.add_option("--spider", dest="spider", help="use this spider") + parser.add_option("--headers", dest="headers", action="store_true", + help="print response HTTP headers instead of body") + parser.add_option("--no-redirect", dest="no_redirect", action="store_true", + default=False, help="do not handle HTTP 3xx status codes and print response as-is") def _print_headers(self, headers, prefix): for key, values in headers.items(): diff --git a/scrapy/commands/startproject.py b/scrapy/commands/startproject.py index e65131ae8..b123e5c84 100644 --- a/scrapy/commands/startproject.py +++ b/scrapy/commands/startproject.py @@ -43,8 +43,8 @@ class Command(ScrapyCommand): return False if not re.search(r'^[_a-zA-Z]\w*$', project_name): - print('Error: Project names must begin with a letter and contain'\ - ' only\nletters, numbers and underscores') + print('Error: Project names must begin with a letter and contain' + ' only\nletters, numbers and underscores') elif _module_exists(project_name): print('Error: Module %r already exists' % project_name) else: diff --git a/scrapy/contracts/default.py b/scrapy/contracts/default.py index e0d425874..3002fc702 100644 --- a/scrapy/contracts/default.py +++ b/scrapy/contracts/default.py @@ -86,8 +86,8 @@ class ReturnsContract(Contract): else: expected = '%s..%s' % (self.min_bound, self.max_bound) - raise ContractFail("Returned %s %s, expected %s" % \ - (occurrences, self.obj_name, expected)) + raise ContractFail("Returned %s %s, expected %s" % + (occurrences, self.obj_name, expected)) class ScrapesContract(Contract): diff --git a/scrapy/core/downloader/handlers/s3.py b/scrapy/core/downloader/handlers/s3.py index e2a07bdef..d6fbd54ee 100644 --- a/scrapy/core/downloader/handlers/s3.py +++ b/scrapy/core/downloader/handlers/s3.py @@ -32,8 +32,8 @@ def _get_boto_connection(): class S3DownloadHandler(object): - def __init__(self, settings, aws_access_key_id=None, aws_secret_access_key=None, \ - httpdownloadhandler=HTTPDownloadHandler, **kw): + def __init__(self, settings, aws_access_key_id=None, aws_secret_access_key=None, + httpdownloadhandler=HTTPDownloadHandler, **kw): if not aws_access_key_id: aws_access_key_id = settings['AWS_ACCESS_KEY_ID'] diff --git a/scrapy/core/downloader/webclient.py b/scrapy/core/downloader/webclient.py index f368c3bae..fc796e8bb 100644 --- a/scrapy/core/downloader/webclient.py +++ b/scrapy/core/downloader/webclient.py @@ -88,9 +88,9 @@ class ScrapyHTTPPageGetter(HTTPClient): if self.factory.url.startswith(b'https'): self.transport.stopProducing() - self.factory.noPage(\ - defer.TimeoutError("Getting %s took longer than %s seconds." % \ - (self.factory.url, self.factory.timeout))) + self.factory.noPage( + defer.TimeoutError("Getting %s took longer than %s seconds." % + (self.factory.url, self.factory.timeout))) class ScrapyHTTPClientFactory(HTTPClientFactory): diff --git a/scrapy/core/spidermw.py b/scrapy/core/spidermw.py index e2ade8256..097a374bf 100644 --- a/scrapy/core/spidermw.py +++ b/scrapy/core/spidermw.py @@ -44,7 +44,7 @@ class SpiderMiddlewareManager(MiddlewareManager): try: result = method(response=response, spider=spider) if result is not None: - raise _InvalidOutput('Middleware {} must return None or raise an exception, got {}' \ + raise _InvalidOutput('Middleware {} must return None or raise an exception, got {}' .format(fname(method), type(result))) except _InvalidOutput: raise @@ -69,7 +69,7 @@ class SpiderMiddlewareManager(MiddlewareManager): elif result is None: continue else: - raise _InvalidOutput('Middleware {} must return None or an iterable, got {}' \ + raise _InvalidOutput('Middleware {} must return None or an iterable, got {}' .format(fname(method), type(result))) return _failure @@ -103,7 +103,7 @@ class SpiderMiddlewareManager(MiddlewareManager): if _isiterable(result): result = evaluate_iterable(result, method_index) else: - raise _InvalidOutput('Middleware {} must return an iterable, got {}' \ + raise _InvalidOutput('Middleware {} must return an iterable, got {}' .format(fname(method), type(result))) return chain(result, recovered) diff --git a/scrapy/downloadermiddlewares/httpcompression.py b/scrapy/downloadermiddlewares/httpcompression.py index 203dee42d..65b652953 100644 --- a/scrapy/downloadermiddlewares/httpcompression.py +++ b/scrapy/downloadermiddlewares/httpcompression.py @@ -37,8 +37,9 @@ class HttpCompressionMiddleware(object): if content_encoding: encoding = content_encoding.pop() decoded_body = self._decode(response.body, encoding.lower()) - respcls = responsetypes.from_args(headers=response.headers, \ - url=response.url, body=decoded_body) + respcls = responsetypes.from_args( + headers=response.headers, url=response.url, body=decoded_body + ) kwargs = dict(cls=respcls, body=decoded_body) if issubclass(respcls, TextResponse): # force recalculating the encoding until we make sure the diff --git a/scrapy/extensions/closespider.py b/scrapy/extensions/closespider.py index 9ccf356ec..afb2ed049 100644 --- a/scrapy/extensions/closespider.py +++ b/scrapy/extensions/closespider.py @@ -54,9 +54,9 @@ class CloseSpider(object): self.crawler.engine.close_spider(spider, 'closespider_pagecount') def spider_opened(self, spider): - self.task = reactor.callLater(self.close_on['timeout'], \ - self.crawler.engine.close_spider, spider, \ - reason='closespider_timeout') + self.task = reactor.callLater(self.close_on['timeout'], + self.crawler.engine.close_spider, spider, + reason='closespider_timeout') def item_scraped(self, item, spider): self.counter['itemcount'] += 1 diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index 4a3e74fbe..bc65f41cc 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -44,8 +44,7 @@ IGNORED_EXTENSIONS = [ _re_type = type(re.compile("", 0)) _matches = lambda url, regexs: any(r.search(url) for r in regexs) -_is_valid_url = lambda url: url.split('://', 1)[0] in {'http', 'https', \ - 'file', 'ftp'} +_is_valid_url = lambda url: url.split('://', 1)[0] in {'http', 'https', 'file', 'ftp'} class FilteringLinkExtractor(object): diff --git a/scrapy/middleware.py b/scrapy/middleware.py index 1cfd8a782..53fa435bb 100644 --- a/scrapy/middleware.py +++ b/scrapy/middleware.py @@ -65,8 +65,8 @@ class MiddlewareManager(object): return process_chain(self.methods[methodname], obj, *args) def _process_chain_both(self, cb_methodname, eb_methodname, obj, *args): - return process_chain_both(self.methods[cb_methodname], \ - self.methods[eb_methodname], obj, *args) + return process_chain_both(self.methods[cb_methodname], + self.methods[eb_methodname], obj, *args) def open_spider(self, spider): return self._process_parallel('open_spider', spider) diff --git a/scrapy/utils/conf.py b/scrapy/utils/conf.py index 7a15e77ff..23306ca28 100644 --- a/scrapy/utils/conf.py +++ b/scrapy/utils/conf.py @@ -37,7 +37,7 @@ def build_component_list(compdict, custom=None, convert=update_classpath): """Fail if a value in the components dict is not a real number or None.""" for name, value in compdict.items(): if value is not None and not isinstance(value, numbers.Real): - raise ValueError('Invalid value {} for component {}, please provide ' \ + raise ValueError('Invalid value {} for component {}, please provide ' 'a real number or None instead'.format(value, name)) # BEGIN Backward compatibility for old (base, custom) call signature diff --git a/tests/test_cmdline/__init__.py b/tests/test_cmdline/__init__.py index 909ea90e0..da99a6be8 100644 --- a/tests/test_cmdline/__init__.py +++ b/tests/test_cmdline/__init__.py @@ -25,17 +25,15 @@ class CmdlineTest(unittest.TestCase): return comm.decode(encoding) def test_default_settings(self): - self.assertEqual(self._execute('settings', '--get', 'TEST1'), \ - 'default') + self.assertEqual(self._execute('settings', '--get', 'TEST1'), 'default') def test_override_settings_using_set_arg(self): - self.assertEqual(self._execute('settings', '--get', 'TEST1', '-s', 'TEST1=override'), \ - 'override') + self.assertEqual(self._execute('settings', '--get', 'TEST1', '-s', + 'TEST1=override'), 'override') def test_override_settings_using_envvar(self): self.env['SCRAPY_TEST1'] = 'override' - self.assertEqual(self._execute('settings', '--get', 'TEST1'), \ - 'override') + self.assertEqual(self._execute('settings', '--get', 'TEST1'), 'override') def test_profiling(self): path = tempfile.mkdtemp() diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 2db2417e8..82d7b18d6 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -800,8 +800,8 @@ class S3TestCase(unittest.TestCase): req = Request('s3://johnsmith/photos/puppy.jpg', headers={'Date': date}) with self._mocked_date(date): httpreq = self.download_request(req, self.spider) - self.assertEqual(httpreq.headers['Authorization'], \ - b'AWS 0PN5J17HBGZHT7JJ3X82:xXjDGYUmKxnwqr5KXNPGldn5LbA=') + self.assertEqual(httpreq.headers['Authorization'], + b'AWS 0PN5J17HBGZHT7JJ3X82:xXjDGYUmKxnwqr5KXNPGldn5LbA=') def test_request_signing2(self): # puts an object into the johnsmith bucket. @@ -813,21 +813,22 @@ class S3TestCase(unittest.TestCase): }) with self._mocked_date(date): httpreq = self.download_request(req, self.spider) - self.assertEqual(httpreq.headers['Authorization'], \ - b'AWS 0PN5J17HBGZHT7JJ3X82:hcicpDDvL9SsO6AkvxqmIWkmOuQ=') + self.assertEqual(httpreq.headers['Authorization'], + b'AWS 0PN5J17HBGZHT7JJ3X82:hcicpDDvL9SsO6AkvxqmIWkmOuQ=') def test_request_signing3(self): # lists the content of the johnsmith bucket. date = 'Tue, 27 Mar 2007 19:42:41 +0000' - req = Request('s3://johnsmith/?prefix=photos&max-keys=50&marker=puppy', \ - method='GET', headers={ - 'User-Agent': 'Mozilla/5.0', - 'Date': date, - }) + req = Request( + 's3://johnsmith/?prefix=photos&max-keys=50&marker=puppy', + method='GET', headers={ + 'User-Agent': 'Mozilla/5.0', + 'Date': date, + }) with self._mocked_date(date): httpreq = self.download_request(req, self.spider) - self.assertEqual(httpreq.headers['Authorization'], \ - b'AWS 0PN5J17HBGZHT7JJ3X82:jsRt/rhG+Vtp88HrYL706QhE4w4=') + self.assertEqual(httpreq.headers['Authorization'], + b'AWS 0PN5J17HBGZHT7JJ3X82:jsRt/rhG+Vtp88HrYL706QhE4w4=') def test_request_signing4(self): # fetches the access control policy sub-resource for the 'johnsmith' bucket. @@ -836,8 +837,8 @@ class S3TestCase(unittest.TestCase): method='GET', headers={'Date': date}) with self._mocked_date(date): httpreq = self.download_request(req, self.spider) - self.assertEqual(httpreq.headers['Authorization'], \ - b'AWS 0PN5J17HBGZHT7JJ3X82:thdUi9VAkzhkniLj96JIrOPGi0g=') + self.assertEqual(httpreq.headers['Authorization'], + b'AWS 0PN5J17HBGZHT7JJ3X82:thdUi9VAkzhkniLj96JIrOPGi0g=') def test_request_signing5(self): try: @@ -850,11 +851,11 @@ class S3TestCase(unittest.TestCase): # deletes an object from the 'johnsmith' bucket using the # path-style and Date alternative. date = 'Tue, 27 Mar 2007 21:20:27 +0000' - req = Request('s3://johnsmith/photos/puppy.jpg', \ - method='DELETE', headers={ - 'Date': date, - 'x-amz-date': 'Tue, 27 Mar 2007 21:20:26 +0000', - }) + req = Request( + 's3://johnsmith/photos/puppy.jpg', method='DELETE', headers={ + 'Date': date, + 'x-amz-date': 'Tue, 27 Mar 2007 21:20:26 +0000', + }) with self._mocked_date(date): httpreq = self.download_request(req, self.spider) # botocore does not override Date with x-amz-date @@ -864,25 +865,26 @@ class S3TestCase(unittest.TestCase): def test_request_signing6(self): # uploads an object to a CNAME style virtual hosted bucket with metadata. date = 'Tue, 27 Mar 2007 21:06:08 +0000' - req = Request('s3://static.johnsmith.net:8080/db-backup.dat.gz', \ - method='PUT', headers={ - 'User-Agent': 'curl/7.15.5', - 'Host': 'static.johnsmith.net:8080', - 'Date': date, - 'x-amz-acl': 'public-read', - 'content-type': 'application/x-download', - 'Content-MD5': '4gJE4saaMU4BqNR0kLY+lw==', - 'X-Amz-Meta-ReviewedBy': 'joe@johnsmith.net,jane@johnsmith.net', - 'X-Amz-Meta-FileChecksum': '0x02661779', - 'X-Amz-Meta-ChecksumAlgorithm': 'crc32', - 'Content-Disposition': 'attachment; filename=database.dat', - 'Content-Encoding': 'gzip', - 'Content-Length': '5913339', - }) + req = Request( + 's3://static.johnsmith.net:8080/db-backup.dat.gz', + method='PUT', headers={ + 'User-Agent': 'curl/7.15.5', + 'Host': 'static.johnsmith.net:8080', + 'Date': date, + 'x-amz-acl': 'public-read', + 'content-type': 'application/x-download', + 'Content-MD5': '4gJE4saaMU4BqNR0kLY+lw==', + 'X-Amz-Meta-ReviewedBy': 'joe@johnsmith.net,jane@johnsmith.net', + 'X-Amz-Meta-FileChecksum': '0x02661779', + 'X-Amz-Meta-ChecksumAlgorithm': 'crc32', + 'Content-Disposition': 'attachment; filename=database.dat', + 'Content-Encoding': 'gzip', + 'Content-Length': '5913339', + }) with self._mocked_date(date): httpreq = self.download_request(req, self.spider) - self.assertEqual(httpreq.headers['Authorization'], \ - b'AWS 0PN5J17HBGZHT7JJ3X82:C0FlOtU8Ylb9KDTpZqYkZPX91iI=') + self.assertEqual(httpreq.headers['Authorization'], + b'AWS 0PN5J17HBGZHT7JJ3X82:C0FlOtU8Ylb9KDTpZqYkZPX91iI=') def test_request_signing7(self): # ensure that spaces are quoted properly before signing diff --git a/tests/test_downloadermiddleware_retry.py b/tests/test_downloadermiddleware_retry.py index 51b79b6c3..e09d66086 100644 --- a/tests/test_downloadermiddleware_retry.py +++ b/tests/test_downloadermiddleware_retry.py @@ -165,12 +165,12 @@ class MaxRetryTimesTest(unittest.TestCase): # SETTINGS: meta(max_retry_times) = 4 meta_max_retry_times = 4 - req = Request(self.invalid_url, meta= \ - {'max_retry_times': meta_max_retry_times, 'dont_retry': True}) + req = Request(self.invalid_url, meta={ + 'max_retry_times': meta_max_retry_times, 'dont_retry': True + }) self._test_retry(req, DNSLookupError('foo'), 0) - def _test_retry(self, req, exception, max_retry_times): for i in range(0, max_retry_times): diff --git a/tests/test_engine.py b/tests/test_engine.py index 537df8d91..a48b63025 100644 --- a/tests/test_engine.py +++ b/tests/test_engine.py @@ -91,8 +91,8 @@ def start_test_site(debug=False): port = reactor.listenTCP(0, server.Site(r), interface="127.0.0.1") if debug: - print("Test server running at http://localhost:%d/ - hit Ctrl-C to finish." \ - % port.getHost().port) + print("Test server running at http://localhost:%d/ - hit Ctrl-C to finish." + % port.getHost().port) return port diff --git a/tests/test_http_request.py b/tests/test_http_request.py index 57e7b457d..e30417b30 100644 --- a/tests/test_http_request.py +++ b/tests/test_http_request.py @@ -622,8 +622,9 @@ class FormRequestTest(RequestTest): <input type="hidden" name="two" value="3"> <input type="submit" name="clickable2" value="clicked2"> </form>""") - req = self.request_class.from_response(response, formdata={'two': '2'}, \ - clickdata={'name': 'clickable2'}) + req = self.request_class.from_response( + response, formdata={'two': '2'}, clickdata={'name': 'clickable2'} + ) fs = _qs(req) self.assertEqual(fs[b'clickable2'], [b'clicked2']) self.assertFalse(b'clickable1' in fs, fs) @@ -671,8 +672,9 @@ class FormRequestTest(RequestTest): <input type="hidden" name="one" value="clicked1"> <input type="hidden" name="two" value="clicked2"> </form>""") - req = self.request_class.from_response(response, \ - clickdata={u'name': u'clickable', u'value': u'clicked2'}) + req = self.request_class.from_response( + response, clickdata={u'name': u'clickable', u'value': u'clicked2'} + ) fs = _qs(req) self.assertEqual(fs[b'clickable'], [b'clicked2']) self.assertEqual(fs[b'one'], [b'clicked1']) @@ -686,8 +688,9 @@ class FormRequestTest(RequestTest): <input type="hidden" name="poundsign" value="\u00a3"> <input type="hidden" name="eurosign" value="\u20ac"> </form>""") - req = self.request_class.from_response(response, \ - clickdata={u'name': u'price in \u00a3'}) + req = self.request_class.from_response( + response, clickdata={u'name': u'price in \u00a3'} + ) fs = _qs(req, to_unicode=True) self.assertTrue(fs[u'price in \u00a3']) @@ -700,8 +703,9 @@ class FormRequestTest(RequestTest): <input type="hidden" name="yensign" value="\u00a5"> </form>""", encoding='latin1') - req = self.request_class.from_response(response, \ - clickdata={u'name': u'price in \u00a5'}) + req = self.request_class.from_response( + response, clickdata={u'name': u'price in \u00a5'} + ) fs = _qs(req, to_unicode=True, encoding='latin1') self.assertTrue(fs[u'price in \u00a5']) @@ -716,8 +720,9 @@ class FormRequestTest(RequestTest): <input type="hidden" name="field2" value="value2"> </form> """) - req = self.request_class.from_response(response, formname='form2', \ - clickdata={u'name': u'clickable'}) + req = self.request_class.from_response( + response, formname='form2', clickdata={u'name': u'clickable'} + ) fs = _qs(req) self.assertEqual(fs[b'clickable'], [b'clicked2']) self.assertEqual(fs[b'field2'], [b'value2']) @@ -725,8 +730,9 @@ class FormRequestTest(RequestTest): def test_from_response_override_clickable(self): response = _buildresponse('''<form><input type="submit" name="clickme" value="one"> </form>''') - req = self.request_class.from_response(response, \ - formdata={'clickme': 'two'}, clickdata={'name': 'clickme'}) + req = self.request_class.from_response( + response, formdata={'clickme': 'two'}, clickdata={'name': 'clickme'} + ) fs = _qs(req) self.assertEqual(fs[b'clickme'], [b'two']) @@ -853,7 +859,7 @@ class FormRequestTest(RequestTest): <form name="form2" action="post.php" method="POST"> <input type="hidden" name="two" value="2"> </form>""") - self.assertRaises(IndexError, self.request_class.from_response, \ + self.assertRaises(IndexError, self.request_class.from_response, response, formname="form3", formnumber=2) def test_from_response_formid_exists(self): @@ -907,7 +913,7 @@ class FormRequestTest(RequestTest): <form id="form2" name="form2" action="post.php" method="POST"> <input type="hidden" name="two" value="2"> </form>""") - self.assertRaises(IndexError, self.request_class.from_response, \ + self.assertRaises(IndexError, self.request_class.from_response, response, formid="form3", formnumber=2) def test_from_response_select(self): diff --git a/tests/test_http_response.py b/tests/test_http_response.py index 79bb745cc..960ecea3e 100644 --- a/tests/test_http_response.py +++ b/tests/test_http_response.py @@ -316,8 +316,8 @@ class TextResponseTest(BaseResponseTest): assert u'SUFFIX' in r.text, repr(r.text) # Do not destroy html tags due to encoding bugs - r = self.response_class("http://example.com", encoding='utf-8', \ - body=b'\xf0<span>value</span>') + r = self.response_class("http://example.com", encoding='utf-8', + body=b'\xf0<span>value</span>') assert u'<span>value</span>' in r.text, repr(r.text) # FIXME: This test should pass once we stop using BeautifulSoup's UnicodeDammit in TextResponse diff --git a/tests/test_utils_http.py b/tests/test_utils_http.py index f9af4bf87..2fac3da1f 100644 --- a/tests/test_utils_http.py +++ b/tests/test_utils_http.py @@ -13,7 +13,7 @@ class ChunkedTest(unittest.TestCase): chunked_body += "8\r\n" + "sequence\r\n" chunked_body += "0\r\n\r\n" body = decode_chunked_transfer(chunked_body) - self.assertEqual(body, \ - "This is the data in the first chunk\r\n" + - "and this is the second one\r\n" + - "consequence") + self.assertEqual(body, + "This is the data in the first chunk\r\n" + + "and this is the second one\r\n" + + "consequence") From 2d92a39003a5a0b01b265b87cf722a61f121ca60 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Wed, 18 Dec 2019 12:07:08 +0500 Subject: [PATCH 346/387] Restore test_download_with_proxy_https_noconnect, check for a warning there. --- tests/test_downloader_handlers.py | 14 +++++++++++++- 1 file changed, 13 insertions(+), 1 deletion(-) diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 45d4aa952..87ab4c9e5 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -33,7 +33,7 @@ from scrapy.responsetypes import responsetypes from scrapy.settings import Settings from scrapy.utils.test import get_crawler, skip_if_no_boto from scrapy.utils.python import to_bytes -from scrapy.exceptions import NotConfigured +from scrapy.exceptions import NotConfigured, ScrapyDeprecationWarning from tests.mockserver import MockServer, ssl_context_factory, Echo from tests.spiders import SingleRequestSpider @@ -687,6 +687,18 @@ class HttpProxyTestCase(unittest.TestCase): request = Request('http://example.com', meta={'proxy': http_proxy}) return self.download_request(request, Spider('foo')).addCallback(_test) + def test_download_with_proxy_https_noconnect(self): + def _test(response): + self.assertEqual(response.status, 200) + self.assertEqual(response.url, request.url) + self.assertEqual(response.body, b'https://example.com') + + http_proxy = '%s?noconnect' % self.getURL('') + request = Request('https://example.com', meta={'proxy': http_proxy}) + with self.assertWarnsRegex(ScrapyDeprecationWarning, + r'Using HTTPS proxies in the noconnect mode is deprecated'): + return self.download_request(request, Spider('foo')).addCallback(_test) + def test_download_without_proxy(self): def _test(response): self.assertEqual(response.status, 200) From bb2ff13e4c7c80cfc7925f60eadc97dec0d69026 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Wed, 18 Dec 2019 15:39:08 +0500 Subject: [PATCH 347/387] Skip Http10ProxyTestCase.test_download_with_proxy_https_noconnect --- tests/test_downloader_handlers.py | 2 ++ 1 file changed, 2 insertions(+) diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 87ab4c9e5..412a9c084 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -712,6 +712,8 @@ class HttpProxyTestCase(unittest.TestCase): class Http10ProxyTestCase(HttpProxyTestCase): download_handler_cls = HTTP10DownloadHandler + def test_download_with_proxy_https_noconnect(self): + raise unittest.SkipTest('noconnect is not supported in HTTP10DownloadHandler') class Http11ProxyTestCase(HttpProxyTestCase): download_handler_cls = HTTP11DownloadHandler From ac302c3f615d339ed36000231ca3e1e1c347478a Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Wed, 18 Dec 2019 15:43:05 +0500 Subject: [PATCH 348/387] Fix a flake8 problem. --- tests/test_downloader_handlers.py | 1 + 1 file changed, 1 insertion(+) diff --git a/tests/test_downloader_handlers.py b/tests/test_downloader_handlers.py index 412a9c084..7412d7ebf 100644 --- a/tests/test_downloader_handlers.py +++ b/tests/test_downloader_handlers.py @@ -715,6 +715,7 @@ class Http10ProxyTestCase(HttpProxyTestCase): def test_download_with_proxy_https_noconnect(self): raise unittest.SkipTest('noconnect is not supported in HTTP10DownloadHandler') + class Http11ProxyTestCase(HttpProxyTestCase): download_handler_cls = HTTP11DownloadHandler From 12f9ffeb5d1c1f4d02f616852dcae82ef633e8d9 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita <whalebot.helmsman@gmail.com> Date: Wed, 18 Dec 2019 10:53:27 +0000 Subject: [PATCH 349/387] remove requirements-py3.txt --- requirements-py3.txt | 16 ---------------- tox.ini | 1 - 2 files changed, 17 deletions(-) delete mode 100644 requirements-py3.txt diff --git a/requirements-py3.txt b/requirements-py3.txt deleted file mode 100644 index 28c649e28..000000000 --- a/requirements-py3.txt +++ /dev/null @@ -1,16 +0,0 @@ -parsel>=1.5.0 -PyDispatcher>=2.0.5 -Twisted>=17.9.0 -w3lib>=1.17.0 - -pyOpenSSL>=16.2.0 # Earlier versions fail with "AttributeError: module 'lib' has no attribute 'SSL_ST_INIT'" -queuelib>=1.4.2 # Earlier versions fail with "AttributeError: '...QueueTest' object has no attribute 'qpath'" -cryptography>=2.0 # Earlier versions would fail to install - -# Reference versions taken from -# https://packages.ubuntu.com/xenial/python/ -# https://packages.ubuntu.com/xenial/zope/ -cssselect>=0.9.1 -lxml>=3.5.0 -service_identity>=16.0.0 -zope.interface>=4.1.3 diff --git a/tox.ini b/tox.ini index fd75d18e2..f37c381d0 100644 --- a/tox.ini +++ b/tox.ini @@ -9,7 +9,6 @@ envlist = py35 [testenv] deps = -ctests/constraints.txt - -rrequirements-py3.txt -rtests/requirements-py3.txt # Extras botocore>=1.3.23 From ee9881d2704798c9cd61b6da503bb0694227c58c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Wed, 18 Dec 2019 12:08:34 +0100 Subject: [PATCH 350/387] Improve FilteringLinkExtractor.__new__ --- scrapy/linkextractors/__init__.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index a510fef70..7254bd79c 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -61,7 +61,7 @@ class FilteringLinkExtractor(object): warn('scrapy.linkextractors.FilteringLinkExtractor is deprecated, ' 'please use scrapy.linkextractors.LinkExtractor instead', ScrapyDeprecationWarning, stacklevel=2) - return super(FilteringLinkExtractor, cls).__new__(cls) + return super().__new__(cls, *args, **kwargs) def __init__(self, link_extractor, allow, deny, allow_domains, deny_domains, restrict_xpaths, canonicalize, deny_extensions, restrict_css, restrict_text): From 174769a3f08fcd84eaec8a88217a05f8ebc3f2cb Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Wed, 18 Dec 2019 12:09:03 +0100 Subject: [PATCH 351/387] Use a better name for the LxmlLinkExtractor subclassing test --- tests/test_linkextractors.py | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/tests/test_linkextractors.py b/tests/test_linkextractors.py index 0ffeaecc3..cfd4c6b85 100644 --- a/tests/test_linkextractors.py +++ b/tests/test_linkextractors.py @@ -517,10 +517,10 @@ class LxmlLinkExtractorTestCase(Base.LinkExtractorTestCase): LxmlLinkExtractor() self.assertEqual(len(warnings), 0) - class SubclassedItem(LxmlLinkExtractor): + class SubclassedLxmlLinkExtractor(LxmlLinkExtractor): pass - SubclassedItem() + SubclassedLxmlLinkExtractor() self.assertEqual(len(warnings), 0) From 012533924a799c7bf1d3f2d4a489d8ca1ac462f8 Mon Sep 17 00:00:00 2001 From: Vostretsov Nikita <whalebot.helmsman@gmail.com> Date: Wed, 18 Dec 2019 11:13:36 +0000 Subject: [PATCH 352/387] remove requirements from here too --- docs/requirements.txt | 1 - 1 file changed, 1 deletion(-) diff --git a/docs/requirements.txt b/docs/requirements.txt index 0ed11c4dc..773b92cea 100644 --- a/docs/requirements.txt +++ b/docs/requirements.txt @@ -1,4 +1,3 @@ --r ../requirements-py3.txt Sphinx>=2.1 sphinx-hoverxref sphinx-notfound-page From 7ccb169a27ccbedf3145b824e1bd655cb6902f32 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Wed, 18 Dec 2019 19:41:16 +0500 Subject: [PATCH 353/387] Split a long test in test_engine.py into three. --- tests/test_engine.py | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/tests/test_engine.py b/tests/test_engine.py index a48b63025..25dee7c1f 100644 --- a/tests/test_engine.py +++ b/tests/test_engine.py @@ -179,7 +179,6 @@ class EngineTest(unittest.TestCase): @defer.inlineCallbacks def test_crawler(self): - for spider in TestSpider, DictItemsSpider: self.run = CrawlerRun(spider) yield self.run.run() @@ -189,11 +188,15 @@ class EngineTest(unittest.TestCase): self._assert_scraped_items() self._assert_signals_catched() + @defer.inlineCallbacks + def test_crawler_dupefilter(self): self.run = CrawlerRun(TestDupeFilterSpider) yield self.run.run() self._assert_scheduled_requests(urls_to_visit=7) self._assert_dropped_requests() + @defer.inlineCallbacks + def test_crawler_itemerror(self): self.run = CrawlerRun(ItemZeroDivisionErrorSpider) yield self.run.run() self._assert_items_error() From 916382e109dc06a57f4631448555953d1fa540b5 Mon Sep 17 00:00:00 2001 From: elacuesta <elacuesta@users.noreply.github.com> Date: Wed, 18 Dec 2019 12:05:33 -0300 Subject: [PATCH 354/387] Add errback parameter to scrapy.spiders.crawl.Rule (#4000) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * Add errback parameter to scrapy.spiders.crawl.Rule * CrawlSpider: optimize by reducing iterations * [test] Rule.errback * [doc] Rule.errback * [doc] Use autoclass in docs/topics/spiders.rst Co-Authored-By: Adrián Chaves <adrian@chaves.io> * Rule.process_links takes a list * Fix aesthetic issue reported by Flake8 --- docs/topics/spiders.rst | 6 +++++ scrapy/spiders/crawl.py | 60 ++++++++++++++++++++++++++--------------- tests/mockserver.py | 7 +++++ tests/spiders.py | 36 +++++++++++++++++++++++-- tests/test_crawl.py | 18 ++++++++++--- 5 files changed, 101 insertions(+), 26 deletions(-) diff --git a/docs/topics/spiders.rst b/docs/topics/spiders.rst index d65a43afd..b0fb14e24 100644 --- a/docs/topics/spiders.rst +++ b/docs/topics/spiders.rst @@ -414,6 +414,12 @@ Crawling rules from which the request originated as second argument. It must return a ``Request`` object or ``None`` (to filter out the request). + ``errback`` is a callable or a string (in which case a method from the spider + object with that name will be used) to be called if any exception is + raised while processing a request generated by the rule. + It receives a :class:`Twisted Failure <twisted.python.failure.Failure>` + instance as first parameter. + CrawlSpider example ~~~~~~~~~~~~~~~~~~~ diff --git a/scrapy/spiders/crawl.py b/scrapy/spiders/crawl.py index a5eb1a518..a2c364c0e 100644 --- a/scrapy/spiders/crawl.py +++ b/scrapy/spiders/crawl.py @@ -16,7 +16,11 @@ from scrapy.utils.python import get_func_args from scrapy.utils.spider import iterate_spider_output -def _identity(request, response): +def _identity(x): + return x + + +def _identity_process_request(request, response): return request @@ -32,17 +36,20 @@ _default_link_extractor = LinkExtractor() class Rule(object): - def __init__(self, link_extractor=None, callback=None, cb_kwargs=None, follow=None, process_links=None, process_request=None): + def __init__(self, link_extractor=None, callback=None, cb_kwargs=None, follow=None, + process_links=None, process_request=None, errback=None): self.link_extractor = link_extractor or _default_link_extractor self.callback = callback + self.errback = errback self.cb_kwargs = cb_kwargs or {} - self.process_links = process_links - self.process_request = process_request or _identity + self.process_links = process_links or _identity + self.process_request = process_request or _identity_process_request self.process_request_argcount = None self.follow = follow if follow is not None else not callback def _compile(self, spider): self.callback = _get_method(self.callback, spider) + self.errback = _get_method(self.errback, spider) self.process_links = _get_method(self.process_links, spider) self.process_request = _get_method(self.process_request, spider) self.process_request_argcount = len(get_func_args(self.process_request)) @@ -76,48 +83,59 @@ class CrawlSpider(Spider): def process_results(self, response, results): return results - def _build_request(self, rule, link): - r = Request(url=link.url, callback=self._response_downloaded) - r.meta.update(rule=rule, link_text=link.text) - return r + def _build_request(self, rule_index, link): + return Request( + url=link.url, + callback=self._callback, + errback=self._errback, + meta=dict(rule=rule_index, link_text=link.text), + ) def _requests_to_follow(self, response): if not isinstance(response, HtmlResponse): return seen = set() - for n, rule in enumerate(self._rules): + for rule_index, rule in enumerate(self._rules): links = [lnk for lnk in rule.link_extractor.extract_links(response) if lnk not in seen] - if links and rule.process_links: - links = rule.process_links(links) - for link in links: + for link in rule.process_links(links): seen.add(link) - request = self._build_request(n, link) + request = self._build_request(rule_index, link) yield rule._process_request(request, response) - def _response_downloaded(self, response): + def _callback(self, response): rule = self._rules[response.meta['rule']] return self._parse_response(response, rule.callback, rule.cb_kwargs, rule.follow) + def _errback(self, failure): + rule = self._rules[failure.request.meta['rule']] + return self._handle_failure(failure, rule.errback) + def _parse_response(self, response, callback, cb_kwargs, follow=True): if callback: cb_res = callback(response, **cb_kwargs) or () cb_res = self.process_results(response, cb_res) - for requests_or_item in iterate_spider_output(cb_res): - yield requests_or_item + for request_or_item in iterate_spider_output(cb_res): + yield request_or_item if follow and self._follow_links: for request_or_item in self._requests_to_follow(response): yield request_or_item + def _handle_failure(self, failure, errback): + if errback: + results = errback(failure) or () + for request_or_item in iterate_spider_output(results): + yield request_or_item + def _compile_rules(self): - self._rules = [copy.copy(r) for r in self.rules] - for rule in self._rules: - rule._compile(self) + self._rules = [] + for rule in self.rules: + self._rules.append(copy.copy(rule)) + self._rules[-1]._compile(self) @classmethod def from_crawler(cls, crawler, *args, **kwargs): spider = super(CrawlSpider, cls).from_crawler(crawler, *args, **kwargs) - spider._follow_links = crawler.settings.getbool( - 'CRAWLSPIDER_FOLLOW_LINKS', True) + spider._follow_links = crawler.settings.getbool('CRAWLSPIDER_FOLLOW_LINKS', True) return spider diff --git a/tests/mockserver.py b/tests/mockserver.py index fe28176d4..a45277db9 100644 --- a/tests/mockserver.py +++ b/tests/mockserver.py @@ -164,6 +164,12 @@ class Drop(Partial): request.finish() +class ArbitraryLengthPayloadResource(LeafResource): + + def render(self, request): + return request.content.read() + + class Root(Resource): def __init__(self): @@ -177,6 +183,7 @@ class Root(Resource): self.putChild(b"echo", Echo()) self.putChild(b"payload", PayloadResource()) self.putChild(b"xpayload", EncodingResourceWrapper(PayloadResource(), [GzipEncoderFactory()])) + self.putChild(b"alpayload", ArbitraryLengthPayloadResource()) try: from tests import tests_datadir self.putChild(b"files", File(os.path.join(tests_datadir, 'test_site/files/'))) diff --git a/tests/spiders.py b/tests/spiders.py index 981bd2eb8..39c8da0b6 100644 --- a/tests/spiders.py +++ b/tests/spiders.py @@ -1,14 +1,14 @@ """ Some spiders used for testing and benchmarking """ - import time from urllib.parse import urlencode -from scrapy.spiders import Spider from scrapy.http import Request from scrapy.item import Item from scrapy.linkextractors import LinkExtractor +from scrapy.spiders import Spider +from scrapy.spiders.crawl import CrawlSpider, Rule class MockServerSpider(Spider): @@ -184,3 +184,35 @@ class DuplicateStartRequestsSpider(MockServerSpider): def parse(self, response): self.visited += 1 + + +class CrawlSpiderWithErrback(MockServerSpider, CrawlSpider): + name = 'crawl_spider_with_errback' + custom_settings = { + 'RETRY_HTTP_CODES': [], # no need to retry + } + rules = ( + Rule(LinkExtractor(), callback='callback', errback='errback', follow=True), + ) + + def start_requests(self): + test_body = b""" + <html> + <head><title>Page title<title></head> + <body> + <p><a href="/status?n=200">Item 200</a></p> <!-- callback --> + <p><a href="/status?n=201">Item 201</a></p> <!-- callback --> + <p><a href="/status?n=404">Item 404</a></p> <!-- errback --> + <p><a href="/status?n=500">Item 500</a></p> <!-- errback --> + <p><a href="/status?n=501">Item 501</a></p> <!-- errback --> + </body> + </html> + """ + url = self.mockserver.url("/alpayload") + yield Request(url, method="POST", body=test_body) + + def callback(self, response): + self.logger.info('[callback] status %i', response.status) + + def errback(self, failure): + self.logger.info('[errback] status %i', failure.value.response.status) diff --git a/tests/test_crawl.py b/tests/test_crawl.py index 3307899b7..76f87458b 100644 --- a/tests/test_crawl.py +++ b/tests/test_crawl.py @@ -5,12 +5,12 @@ from testfixtures import LogCapture from twisted.internet import defer from twisted.trial.unittest import TestCase -from scrapy.http import Request from scrapy.crawler import CrawlerRunner +from scrapy.http import Request from scrapy.utils.python import to_unicode -from tests.spiders import FollowAllSpider, DelaySpider, SimpleSpider, \ - BrokenStartRequestsSpider, SingleRequestSpider, DuplicateStartRequestsSpider from tests.mockserver import MockServer +from tests.spiders import (FollowAllSpider, DelaySpider, SimpleSpider, BrokenStartRequestsSpider, + SingleRequestSpider, DuplicateStartRequestsSpider, CrawlSpiderWithErrback) class CrawlTestCase(TestCase): @@ -277,3 +277,15 @@ with multiples lines self._assert_retried(log) self.assertIn("Got response 200", str(log)) + + @defer.inlineCallbacks + def test_crawlspider_with_errback(self): + self.runner.crawl(CrawlSpiderWithErrback, mockserver=self.mockserver) + + with LogCapture() as log: + yield self.runner.join() + + self.assertIn("[callback] status 200", str(log)) + self.assertIn("[callback] status 201", str(log)) + self.assertIn("[errback] status 404", str(log)) + self.assertIn("[errback] status 500", str(log)) From a5de2c64e6f1233a98b1b180606fca0ae4dd0871 Mon Sep 17 00:00:00 2001 From: Marc Hernandez Cabot <noviluni@gmail.com> Date: Wed, 18 Dec 2019 16:24:48 +0100 Subject: [PATCH 355/387] fix W291, W292, W293 (whitespaces) --- pytest.ini | 27 ++++++++++----------- scrapy/http/response/__init__.py | 4 +-- scrapy/http/response/text.py | 4 +-- scrapy/logformatter.py | 4 +-- scrapy/utils/markup.py | 2 +- scrapy/utils/multipart.py | 2 +- tests/test_contracts.py | 2 +- tests/test_downloadermiddleware_retry.py | 8 +++--- tests/test_dupefilters.py | 6 ++--- tests/test_pipeline_files.py | 1 - tests/test_robotstxt_interface.py | 2 +- tests/test_spidermiddleware_offsite.py | 2 +- tests/test_spidermiddleware_output_chain.py | 8 +++--- 13 files changed, 35 insertions(+), 37 deletions(-) diff --git a/pytest.ini b/pytest.ini index f088e10ef..ac3c8cfb5 100644 --- a/pytest.ini +++ b/pytest.ini @@ -85,8 +85,8 @@ flake8-ignore = scrapy/http/request/__init__.py E501 scrapy/http/request/form.py E501 E123 scrapy/http/request/json_request.py E501 - scrapy/http/response/__init__.py E501 E128 W293 W291 - scrapy/http/response/text.py E501 W293 E128 E124 + scrapy/http/response/__init__.py E501 E128 + scrapy/http/response/text.py E501 E128 E124 # scrapy/linkextractors scrapy/linkextractors/__init__.py E731 E501 E402 scrapy/linkextractors/lxmlhtml.py E501 E731 E226 @@ -127,9 +127,9 @@ flake8-ignore = scrapy/utils/httpobj.py E501 scrapy/utils/iterators.py E501 E701 scrapy/utils/log.py E128 W503 - scrapy/utils/markup.py F403 W292 + scrapy/utils/markup.py F403 scrapy/utils/misc.py E501 E226 - scrapy/utils/multipart.py F403 W292 + scrapy/utils/multipart.py F403 scrapy/utils/project.py E501 scrapy/utils/python.py E501 scrapy/utils/reactor.py E226 @@ -144,7 +144,6 @@ flake8-ignore = scrapy/utils/url.py E501 F403 E128 F405 # scrapy scrapy/__init__.py E402 E501 - scrapy/_monkeypatches.py W293 scrapy/cmdline.py E501 scrapy/crawler.py E501 scrapy/dupefilters.py E501 E202 @@ -153,7 +152,7 @@ flake8-ignore = scrapy/interfaces.py E501 scrapy/item.py E501 E128 scrapy/link.py E501 - scrapy/logformatter.py E501 W293 + scrapy/logformatter.py E501 scrapy/mail.py E402 E128 E501 E502 scrapy/middleware.py E128 E501 scrapy/pqueues.py E501 @@ -174,7 +173,7 @@ flake8-ignore = tests/test_command_parse.py E501 E128 E303 E226 tests/test_command_shell.py E501 E128 tests/test_commands.py E128 E501 - tests/test_contracts.py E501 E128 W293 + tests/test_contracts.py E501 E128 tests/test_crawl.py E501 E741 E265 tests/test_crawler.py F841 E306 E501 tests/test_dependencies.py F841 E501 E305 @@ -189,17 +188,17 @@ flake8-ignore = tests/test_downloadermiddleware_httpcompression.py E501 E251 E126 E123 tests/test_downloadermiddleware_httpproxy.py E501 E128 tests/test_downloadermiddleware_redirect.py E501 E303 E128 E306 E127 E305 - tests/test_downloadermiddleware_retry.py E501 E128 W293 E251 E303 E126 + tests/test_downloadermiddleware_retry.py E501 E128 E251 E303 E126 tests/test_downloadermiddleware_robotstxt.py E501 tests/test_downloadermiddleware_stats.py E501 - tests/test_dupefilters.py E221 E501 E741 W293 W291 E128 E124 + tests/test_dupefilters.py E221 E501 E741 E128 E124 tests/test_engine.py E401 E501 E128 tests/test_exporters.py E501 E731 E306 E128 E124 tests/test_extension_telnet.py F841 tests/test_feedexport.py E501 F841 E241 tests/test_http_cookies.py E501 tests/test_http_headers.py E501 - tests/test_http_request.py E402 E501 E127 E128 W293 E128 E126 E123 + tests/test_http_request.py E402 E501 E127 E128 E128 E126 E123 tests/test_http_response.py E501 E301 E128 E265 tests/test_item.py E701 E128 F841 E306 tests/test_link.py E501 @@ -209,20 +208,20 @@ flake8-ignore = tests/test_mail.py E128 E501 E305 tests/test_middleware.py E501 E128 tests/test_pipeline_crawl.py E131 E501 E128 E126 - tests/test_pipeline_files.py E501 W293 E303 E272 E226 + tests/test_pipeline_files.py E501 E303 E272 E226 tests/test_pipeline_images.py F841 E501 E303 tests/test_pipeline_media.py E501 E741 E731 E128 E306 E502 tests/test_proxy_connect.py E501 E741 tests/test_request_cb_kwargs.py E501 tests/test_responsetypes.py E501 E305 - tests/test_robotstxt_interface.py E501 W291 E501 + tests/test_robotstxt_interface.py E501 E501 tests/test_scheduler.py E501 E126 E123 tests/test_selector.py E501 E127 tests/test_spider.py E501 tests/test_spidermiddleware.py E501 E226 tests/test_spidermiddleware_httperror.py E128 E501 E127 E121 - tests/test_spidermiddleware_offsite.py E501 E128 E111 W293 - tests/test_spidermiddleware_output_chain.py E501 W293 E226 + tests/test_spidermiddleware_offsite.py E501 E128 E111 + tests/test_spidermiddleware_output_chain.py E501 E226 tests/test_spidermiddleware_referer.py E501 F841 E125 E201 E124 E501 E241 E121 tests/test_squeues.py E501 E701 E741 tests/test_utils_conf.py E501 E303 E128 diff --git a/scrapy/http/response/__init__.py b/scrapy/http/response/__init__.py index 64e9c6c20..e79ce9acc 100644 --- a/scrapy/http/response/__init__.py +++ b/scrapy/http/response/__init__.py @@ -113,8 +113,8 @@ class Response(object_ref): It accepts the same arguments as ``Request.__init__`` method, but ``url`` can be a relative URL or a ``scrapy.link.Link`` object, not only an absolute URL. - - :class:`~.TextResponse` provides a :meth:`~.TextResponse.follow` + + :class:`~.TextResponse` provides a :meth:`~.TextResponse.follow` method which supports selectors in addition to absolute/relative URLs and Link objects. """ diff --git a/scrapy/http/response/text.py b/scrapy/http/response/text.py index 1079fd6e8..4f9afde87 100644 --- a/scrapy/http/response/text.py +++ b/scrapy/http/response/text.py @@ -125,7 +125,7 @@ class TextResponse(Response): Return a :class:`~.Request` instance to follow a link ``url``. It accepts the same arguments as ``Request.__init__`` method, but ``url`` can be not only an absolute URL, but also - + * a relative URL; * a scrapy.link.Link object (e.g. a link extractor result); * an attribute Selector (not SelectorList) - e.g. @@ -133,7 +133,7 @@ class TextResponse(Response): ``response.xpath('//img/@src')[0]``. * a Selector for ``<a>`` or ``<link>`` element, e.g. ``response.css('a.my_link')[0]``. - + See :ref:`response-follow-example` for usage examples. """ if isinstance(url, parsel.Selector): diff --git a/scrapy/logformatter.py b/scrapy/logformatter.py index 5189d7cfa..4e5963e99 100644 --- a/scrapy/logformatter.py +++ b/scrapy/logformatter.py @@ -13,7 +13,7 @@ ERRORMSG = u"'Error processing %(item)s'" class LogFormatter(object): """Class for generating log messages for different actions. - + All methods must return a dictionary listing the parameters ``level``, ``msg`` and ``args`` which are going to be used for constructing the log message when calling ``logging.log``. @@ -48,7 +48,7 @@ class LogFormatter(object): } } """ - + def crawled(self, request, response, spider): """Logs a message when the crawler finds a webpage.""" request_flags = ' %s' % str(request.flags) if request.flags else '' diff --git a/scrapy/utils/markup.py b/scrapy/utils/markup.py index 2455fcc16..9728c542a 100644 --- a/scrapy/utils/markup.py +++ b/scrapy/utils/markup.py @@ -11,4 +11,4 @@ from w3lib.html import * # noqa: F401 warnings.warn("Module `scrapy.utils.markup` is deprecated. " "Please import from `w3lib.html` instead.", - ScrapyDeprecationWarning, stacklevel=2) \ No newline at end of file + ScrapyDeprecationWarning, stacklevel=2) diff --git a/scrapy/utils/multipart.py b/scrapy/utils/multipart.py index e81f63152..5dcf791b8 100644 --- a/scrapy/utils/multipart.py +++ b/scrapy/utils/multipart.py @@ -12,4 +12,4 @@ from w3lib.form import * # noqa: F401 warnings.warn("Module `scrapy.utils.multipart` is deprecated. " "If you're using `encode_multipart` function, please use " "`urllib3.filepost.encode_multipart_formdata` instead", - ScrapyDeprecationWarning, stacklevel=2) \ No newline at end of file + ScrapyDeprecationWarning, stacklevel=2) diff --git a/tests/test_contracts.py b/tests/test_contracts.py index 582e3d052..11d41c1fe 100644 --- a/tests/test_contracts.py +++ b/tests/test_contracts.py @@ -252,7 +252,7 @@ class ContractsManagerTest(unittest.TestCase): self.assertEqual(len(contracts), 3) self.assertEqual(frozenset(type(x) for x in contracts), frozenset([UrlContract, CallbackKeywordArgumentsContract, ReturnsContract])) - + contracts = self.conman.extract_contracts(spider.returns_item_cb_kwargs) self.assertEqual(len(contracts), 3) self.assertEqual(frozenset(type(x) for x in contracts), diff --git a/tests/test_downloadermiddleware_retry.py b/tests/test_downloadermiddleware_retry.py index e09d66086..9c989977e 100644 --- a/tests/test_downloadermiddleware_retry.py +++ b/tests/test_downloadermiddleware_retry.py @@ -124,7 +124,7 @@ class MaxRetryTimesTest(unittest.TestCase): # SETTINGS: meta(max_retry_times) = 0 meta_max_retry_times = 0 - + req = Request(self.invalid_url, meta={'max_retry_times': meta_max_retry_times}) self._test_retry(req, DNSLookupError('foo'), meta_max_retry_times) @@ -137,7 +137,7 @@ class MaxRetryTimesTest(unittest.TestCase): self._test_retry(req, DNSLookupError('foo'), self.mw.max_retry_times) def test_with_metakey_greater(self): - + # SETINGS: RETRY_TIMES < meta(max_retry_times) self.mw.max_retry_times = 2 meta_max_retry_times = 3 @@ -149,7 +149,7 @@ class MaxRetryTimesTest(unittest.TestCase): self._test_retry(req2, DNSLookupError('foo'), self.mw.max_retry_times) def test_with_metakey_lesser(self): - + # SETINGS: RETRY_TIMES > meta(max_retry_times) self.mw.max_retry_times = 5 meta_max_retry_times = 4 @@ -172,7 +172,7 @@ class MaxRetryTimesTest(unittest.TestCase): self._test_retry(req, DNSLookupError('foo'), 0) def _test_retry(self, req, exception, max_retry_times): - + for i in range(0, max_retry_times): req = self.mw.process_exception(req, exception, self.spider) assert isinstance(req, Request) diff --git a/tests/test_dupefilters.py b/tests/test_dupefilters.py index e4b0bdf83..0546558bc 100644 --- a/tests/test_dupefilters.py +++ b/tests/test_dupefilters.py @@ -142,12 +142,12 @@ class RFPDupeFilterTest(unittest.TestCase): r1 = Request('http://scrapytest.org/index.html') r2 = Request('http://scrapytest.org/index.html') - + dupefilter.log(r1, spider) dupefilter.log(r2, spider) assert crawler.stats.get_value('dupefilter/filtered') == 2 - l.check_present(('scrapy.dupefilters', 'DEBUG', + l.check_present(('scrapy.dupefilters', 'DEBUG', ('Filtered duplicate request: <GET http://scrapytest.org/index.html>' ' - no more duplicates will be shown' ' (see DUPEFILTER_DEBUG to show all duplicates)'))) @@ -169,7 +169,7 @@ class RFPDupeFilterTest(unittest.TestCase): r2 = Request('http://scrapytest.org/index.html', headers={'Referer': 'http://scrapytest.org/INDEX.html'} ) - + dupefilter.log(r1, spider) dupefilter.log(r2, spider) diff --git a/tests/test_pipeline_files.py b/tests/test_pipeline_files.py index 52f2b554e..141141671 100644 --- a/tests/test_pipeline_files.py +++ b/tests/test_pipeline_files.py @@ -58,7 +58,6 @@ class FilesPipelineTestCase(unittest.TestCase): self.assertEqual(file_path(Request("data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAR0AAACxCAMAAADOHZloAAACClBMVEX/\ //+F0tzCwMK76ZKQ21AMqr7oAAC96JvD5aWM2kvZ78J0N7fmAAC46Y4Ap7y")), 'full/178059cbeba2e34120a67f2dc1afc3ecc09b61cb.png') - def test_fs_store(self): assert isinstance(self.pipeline.store, FSFilesStore) diff --git a/tests/test_robotstxt_interface.py b/tests/test_robotstxt_interface.py index 27d79437b..24aaaf7ec 100644 --- a/tests/test_robotstxt_interface.py +++ b/tests/test_robotstxt_interface.py @@ -44,7 +44,7 @@ class BaseRobotParserTest: def test_allowed_wildcards(self): robotstxt_robotstxt_body = """User-agent: first - Disallow: /disallowed/*/end$ + Disallow: /disallowed/*/end$ User-agent: second Allow: /*allowed diff --git a/tests/test_spidermiddleware_offsite.py b/tests/test_spidermiddleware_offsite.py index 992e60be2..7511aa568 100644 --- a/tests/test_spidermiddleware_offsite.py +++ b/tests/test_spidermiddleware_offsite.py @@ -73,7 +73,7 @@ class TestOffsiteMiddleware4(TestOffsiteMiddleware3): class TestOffsiteMiddleware5(TestOffsiteMiddleware4): - + def test_get_host_regex(self): self.spider.allowed_domains = ['http://scrapytest.org', 'scrapy.org', 'scrapy.test.org'] with warnings.catch_warnings(record=True) as w: diff --git a/tests/test_spidermiddleware_output_chain.py b/tests/test_spidermiddleware_output_chain.py index 5b7b5e7aa..739cf1c2d 100644 --- a/tests/test_spidermiddleware_output_chain.py +++ b/tests/test_spidermiddleware_output_chain.py @@ -156,7 +156,7 @@ class GeneratorFailMiddleware: r['processed'].append('{}.process_spider_output'.format(self.__class__.__name__)) yield r raise LookupError() - + def process_spider_exception(self, response, exception, spider): method = '{}.process_spider_exception'.format(self.__class__.__name__) spider.logger.info('%s: %s caught', method, exception.__class__.__name__) @@ -264,7 +264,7 @@ class TestSpiderMiddleware(TestCase): @classmethod def tearDownClass(cls): cls.mockserver.__exit__(None, None, None) - + @defer.inlineCallbacks def crawl_log(self, spider): crawler = get_crawler(spider) @@ -308,7 +308,7 @@ class TestSpiderMiddleware(TestCase): self.assertIn("{'from': 'errback'}", str(log1)) self.assertNotIn("{'from': 'callback'}", str(log1)) self.assertIn("'item_scraped_count': 1", str(log1)) - + @defer.inlineCallbacks def test_generator_callback(self): """ @@ -319,7 +319,7 @@ class TestSpiderMiddleware(TestCase): log2 = yield self.crawl_log(GeneratorCallbackSpider) self.assertIn("Middleware: ImportError exception caught", str(log2)) self.assertIn("'item_scraped_count': 2", str(log2)) - + @defer.inlineCallbacks def test_not_a_generator_callback(self): """ From c0d84f0962a4c269441db562e8cbc10298c53b72 Mon Sep 17 00:00:00 2001 From: Marc Hernandez Cabot <noviluni@gmail.com> Date: Wed, 18 Dec 2019 19:39:21 +0100 Subject: [PATCH 356/387] fix typos --- docs/contributing.rst | 2 +- docs/faq.rst | 2 +- docs/news.rst | 69 +++++++++++++++++----------------- docs/topics/jobs.rst | 2 +- docs/topics/leaks.rst | 2 +- docs/topics/media-pipeline.rst | 2 +- docs/topics/settings.rst | 2 +- docs/topics/telnetconsole.rst | 2 +- sep/sep-001.rst | 4 +- sep/sep-019.rst | 2 +- 10 files changed, 44 insertions(+), 45 deletions(-) diff --git a/docs/contributing.rst b/docs/contributing.rst index 3aebb3d50..b56295027 100644 --- a/docs/contributing.rst +++ b/docs/contributing.rst @@ -217,7 +217,7 @@ the tests with Python 3.6 use:: tox -e py36 -You can also specify a comma-separated list of environmets, and use :ref:`tox’s +You can also specify a comma-separated list of environments, and use :ref:`tox’s parallel mode <tox:parallel_mode>` to run the tests on multiple environments in parallel:: diff --git a/docs/faq.rst b/docs/faq.rst index 080d81981..aae2411e0 100644 --- a/docs/faq.rst +++ b/docs/faq.rst @@ -338,7 +338,7 @@ How to split an item into multiple items in an item pipeline? input item. :ref:`Create a spider middleware <custom-spider-middleware>` instead, and use its :meth:`~scrapy.spidermiddlewares.SpiderMiddleware.process_spider_output` -method for this puspose. For example:: +method for this purpose. For example:: from copy import deepcopy diff --git a/docs/news.rst b/docs/news.rst index 9dfd28508..28db40e3c 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -678,7 +678,7 @@ Usability improvements * a message is added to IgnoreRequest in RobotsTxtMiddleware (:issue:`3113`) * better validation of ``url`` argument in ``Response.follow`` (:issue:`3131`) * non-zero exit code is returned from Scrapy commands when error happens - on spider inititalization (:issue:`3226`) + on spider initialization (:issue:`3226`) * Link extraction improvements: "ftp" is added to scheme list (:issue:`3152`); "flv" is added to common video extensions (:issue:`3165`) * better error message when an exporter is disabled (:issue:`3358`); @@ -1156,7 +1156,7 @@ Bug fixes - Fix :command:`view` command ; it was a regression in v1.3.0 (:issue:`2503`). - Fix tests regarding ``*_EXPIRES settings`` with Files/Images pipelines (:issue:`2460`). - Fix name of generated pipeline class when using basic project template (:issue:`2466`). -- Fix compatiblity with Twisted 17+ (:issue:`2496`, :issue:`2528`). +- Fix compatibility with Twisted 17+ (:issue:`2496`, :issue:`2528`). - Fix ``scrapy.Item`` inheritance on Python 3.6 (:issue:`2511`). - Enforce numeric values for components order in ``SPIDER_MIDDLEWARES``, ``DOWNLOADER_MIDDLEWARES``, ``EXTENIONS`` and ``SPIDER_CONTRACTS`` (:issue:`2420`). @@ -1164,7 +1164,7 @@ Bug fixes Documentation ~~~~~~~~~~~~~ -- Reword Code of Coduct section and upgrade to Contributor Covenant v1.4 +- Reword Code of Conduct section and upgrade to Contributor Covenant v1.4 (:issue:`2469`). - Clarify that passing spider arguments converts them to spider attributes (:issue:`2483`). @@ -1178,7 +1178,7 @@ Documentation Cleanups ~~~~~~~~ -- Remove reduntant check in ``MetaRefreshMiddleware`` (:issue:`2542`). +- Remove redundant check in ``MetaRefreshMiddleware`` (:issue:`2542`). - Faster checks in ``LinkExtractor`` for allow/deny patterns (:issue:`2538`). - Remove dead code supporting old Twisted versions (:issue:`2544`). @@ -1204,7 +1204,7 @@ New Features - ``MailSender`` now accepts single strings as values for ``to`` and ``cc`` arguments (:issue:`2272`) - ``scrapy fetch url``, ``scrapy shell url`` and ``fetch(url)`` inside - scrapy shell now follow HTTP redirections by default (:issue:`2290`); + Scrapy shell now follow HTTP redirections by default (:issue:`2290`); See :command:`fetch` and :command:`shell` for details. - ``HttpErrorMiddleware`` now logs errors with ``INFO`` level instead of ``DEBUG``; this is technically **backward incompatible** so please check your log parsers. @@ -1705,7 +1705,7 @@ Scrapy 1.0.4 (2015-12-30) - fix ValueError: Invalid XPath: //div/[id="not-exists"]/text() on selectors.rst (:commit:`ca8d60f`) - Typos corrections (:commit:`7067117`) - fix typos in downloader-middleware.rst and exceptions.rst, middlware -> middleware (:commit:`32f115c`) -- Add note to ubuntu install section about debian compatibility (:commit:`23fda69`) +- Add note to Ubuntu install section about Debian compatibility (:commit:`23fda69`) - Replace alternative OSX install workaround with virtualenv (:commit:`98b63ee`) - Reference Homebrew's homepage for installation instructions (:commit:`1925db1`) - Add oldest supported tox version to contributing docs (:commit:`5d10d6d`) @@ -1758,7 +1758,7 @@ Scrapy 1.0.1 (2015-07-01) - include tests/ to source distribution in MANIFEST.in (:commit:`eca227e`) - DOC Fix SelectJmes documentation (:commit:`b8567bc`) - DOC Bring Ubuntu and Archlinux outside of Windows subsection (:commit:`392233f`) -- DOC remove version suffix from ubuntu package (:commit:`5303c66`) +- DOC remove version suffix from Ubuntu package (:commit:`5303c66`) - DOC Update release date for 1.0 (:commit:`c89fa29`) .. _release-1.0.0: @@ -2211,7 +2211,7 @@ Scrapy 0.24.2 (2014-07-08) - Use a mutable mapping to proxy deprecated settings.overrides and settings.defaults attribute (:commit:`e5e8133`) - there is not support for python3 yet (:commit:`3cd6146`) -- Update python compatible version set to debian packages (:commit:`fa5d76b`) +- Update python compatible version set to Debian packages (:commit:`fa5d76b`) - DOC fix formatting in release notes (:commit:`c6a9e20`) Scrapy 0.24.1 (2014-06-27) @@ -2229,12 +2229,12 @@ Enhancements - Improve Scrapy top-level namespace (:issue:`494`, :issue:`684`) - Add selector shortcuts to responses (:issue:`554`, :issue:`690`) -- Add new lxml based LinkExtractor to replace unmantained SgmlLinkExtractor +- Add new lxml based LinkExtractor to replace unmaintained SgmlLinkExtractor (:issue:`559`, :issue:`761`, :issue:`763`) - Cleanup settings API - part of per-spider settings **GSoC project** (:issue:`737`) - Add UTF8 encoding header to templates (:issue:`688`, :issue:`762`) - Telnet console now binds to 127.0.0.1 by default (:issue:`699`) -- Update debian/ubuntu install instructions (:issue:`509`, :issue:`549`) +- Update Debian/Ubuntu install instructions (:issue:`509`, :issue:`549`) - Disable smart strings in lxml XPath evaluations (:issue:`535`) - Restore filesystem based cache as default for http cache middleware (:issue:`541`, :issue:`500`, :issue:`571`) @@ -2267,7 +2267,7 @@ Enhancements - Tests and docs for ``request_fingerprint`` function (:issue:`597`) - Update SEP-19 for GSoC project ``per-spider settings`` (:issue:`705`) - Set exit code to non-zero when contracts fails (:issue:`727`) -- Add a setting to control what class is instanciated as Downloader component +- Add a setting to control what class is instantiated as Downloader component (:issue:`738`) - Pass response in ``item_dropped`` signal (:issue:`724`) - Improve ``scrapy check`` contracts command (:issue:`733`, :issue:`752`) @@ -2276,7 +2276,7 @@ Enhancements - Add a note about reporting security issues (:issue:`697`) - Add LevelDB http cache storage backend (:issue:`626`, :issue:`500`) - Sort spider list output of ``scrapy list`` command (:issue:`742`) -- Multiple documentation enhancemens and fixes +- Multiple documentation enhancements and fixes (:issue:`575`, :issue:`587`, :issue:`590`, :issue:`596`, :issue:`610`, :issue:`617`, :issue:`618`, :issue:`627`, :issue:`613`, :issue:`643`, :issue:`654`, :issue:`675`, :issue:`663`, :issue:`711`, :issue:`714`) @@ -2321,19 +2321,19 @@ Scrapy 0.22.1 (released 2014-02-08) - BaseSgmlLinkExtractor: Added unit test of a link with an inner tag (:commit:`c1cb418`) - BaseSgmlLinkExtractor: Fixed unknown_endtag() so that it only set current_link=None when the end tag match the opening tag (:commit:`7e4d627`) - Fix tests for Travis-CI build (:commit:`76c7e20`) -- replace unencodeable codepoints with html entities. fixes #562 and #285 (:commit:`5f87b17`) +- replace unencodable codepoints with html entities. fixes #562 and #285 (:commit:`5f87b17`) - RegexLinkExtractor: encode URL unicode value when creating Links (:commit:`d0ee545`) - Updated the tutorial crawl output with latest output. (:commit:`8da65de`) - Updated shell docs with the crawler reference and fixed the actual shell output. (:commit:`875b9ab`) - PEP8 minor edits. (:commit:`f89efaf`) -- Expose current crawler in the scrapy shell. (:commit:`5349cec`) +- Expose current crawler in the Scrapy shell. (:commit:`5349cec`) - Unused re import and PEP8 minor edits. (:commit:`387f414`) - Ignore None's values when using the ItemLoader. (:commit:`0632546`) - DOC Fixed HTTPCACHE_STORAGE typo in the default value which is now Filesystem instead Dbm. (:commit:`cde9a8c`) -- show ubuntu setup instructions as literal code (:commit:`fb5c9c5`) +- show Ubuntu setup instructions as literal code (:commit:`fb5c9c5`) - Update Ubuntu installation instructions (:commit:`70fb105`) - Merge pull request #550 from stray-leone/patch-1 (:commit:`6f70b6a`) -- modify the version of scrapy ubuntu package (:commit:`725900d`) +- modify the version of Scrapy Ubuntu package (:commit:`725900d`) - fix 0.22.0 release date (:commit:`af0219a`) - fix typos in news.rst and remove (not released yet) header (:commit:`b7f58f4`) @@ -2354,7 +2354,7 @@ Enhancements - Improve test coverage and forthcoming Python 3 support (:issue:`525`) - Promote startup info on settings and middleware to INFO level (:issue:`520`) - Support partials in ``get_func_args`` util (:issue:`506`, issue:`504`) -- Allow running indiviual tests via tox (:issue:`503`) +- Allow running individual tests via tox (:issue:`503`) - Update extensions ignored by link extractors (:issue:`498`) - Add middleware methods to get files/images/thumbs paths (:issue:`490`) - Improve offsite middleware tests (:issue:`478`) @@ -2411,7 +2411,7 @@ Enhancements - scrapy.mail.MailSender now can connect over TLS or upgrade using STARTTLS (:issue:`327`) - New FilesPipeline with functionality factored out from ImagesPipeline (:issue:`370`, :issue:`409`) - Recommend Pillow instead of PIL for image handling (:issue:`317`) -- Added debian packages for Ubuntu quantal and raring (:commit:`86230c0`) +- Added Debian packages for Ubuntu Quantal and raring (:commit:`86230c0`) - Mock server (used for tests) can listen for HTTPS requests (:issue:`410`) - Remove multi spider support from multiple core components (:issue:`422`, :issue:`421`, :issue:`420`, :issue:`419`, :issue:`423`, :issue:`418`) @@ -2430,7 +2430,7 @@ Bugfixes - Fix tests under Django 1.6 (:commit:`b6bed44c`) - Lot of bugfixes to retry middleware under disconnections using HTTP 1.1 download handler - Fix inconsistencies among Twisted releases (:issue:`406`) -- Fix scrapy shell bugs (:issue:`418`, :issue:`407`) +- Fix Scrapy shell bugs (:issue:`418`, :issue:`407`) - Fix invalid variable name in setup.py (:issue:`429`) - Fix tutorial references (:issue:`387`) - Improve request-response docs (:issue:`391`) @@ -2512,15 +2512,15 @@ Scrapy 0.18.1 (released 2013-08-27) - test PotentiaDataLoss errors on unbound responses (:commit:`b15470d`) - Treat responses without content-length or Transfer-Encoding as good responses (:commit:`c4bf324`) - do no include ResponseFailed if http11 handler is not enabled (:commit:`6cbe684`) -- New HTTP client wraps connection losts in ResponseFailed exception. fix #373 (:commit:`1a20bba`) +- New HTTP client wraps connection lost in ResponseFailed exception. fix #373 (:commit:`1a20bba`) - limit travis-ci build matrix (:commit:`3b01bb8`) - Merge pull request #375 from peterarenot/patch-1 (:commit:`fa766d7`) - Fixed so it refers to the correct folder (:commit:`3283809`) -- added quantal & raring to support ubuntu releases (:commit:`1411923`) +- added Quantal & raring to support Ubuntu releases (:commit:`1411923`) - fix retry middleware which didn't retry certain connection errors after the upgrade to http1 client, closes GH-373 (:commit:`bb35ed0`) - fix XmlItemExporter in Python 2.7.4 and 2.7.5 (:commit:`de3e451`) - minor updates to 0.18 release notes (:commit:`c45e5f1`) -- fix contributters list format (:commit:`0b60031`) +- fix contributors list format (:commit:`0b60031`) Scrapy 0.18.0 (released 2013-08-09) ----------------------------------- @@ -2617,7 +2617,7 @@ contributors sorted by number of commits:: Scrapy 0.16.5 (released 2013-05-30) ----------------------------------- -- obey request method when scrapy deploy is redirected to a new endpoint (:commit:`8c4fcee`) +- obey request method when Scrapy deploy is redirected to a new endpoint (:commit:`8c4fcee`) - fix inaccurate downloader middleware documentation. refs #280 (:commit:`40667cb`) - doc: remove links to diveintopython.org, which is no longer available. closes #246 (:commit:`bd58bfa`) - Find form nodes in invalid html5 documents (:commit:`e3d6945`) @@ -2631,8 +2631,8 @@ Scrapy 0.16.4 (released 2013-01-23) - Fixed error message formatting. log.err() doesn't support cool formatting and when error occurred, the message was: "ERROR: Error processing %(item)s" (:commit:`c16150c`) - lint and improve images pipeline error logging (:commit:`56b45fc`) - fixed doc typos (:commit:`243be84`) -- add documentation topics: Broad Crawls & Common Practies (:commit:`1fbb715`) -- fix bug in scrapy parse command when spider is not specified explicitly. closes #209 (:commit:`c72e682`) +- add documentation topics: Broad Crawls & Common Practices (:commit:`1fbb715`) +- fix bug in Scrapy parse command when spider is not specified explicitly. closes #209 (:commit:`c72e682`) - Update docs/topics/commands.rst (:commit:`28eac7a`) Scrapy 0.16.3 (released 2012-12-07) @@ -2651,11 +2651,11 @@ Scrapy 0.16.3 (released 2012-12-07) Scrapy 0.16.2 (released 2012-11-09) ----------------------------------- -- scrapy contracts: python2.6 compat (:commit:`a4a9199`) -- scrapy contracts verbose option (:commit:`ec41673`) -- proper unittest-like output for scrapy contracts (:commit:`86635e4`) +- Scrapy contracts: python2.6 compat (:commit:`a4a9199`) +- Scrapy contracts verbose option (:commit:`ec41673`) +- proper unittest-like output for Scrapy contracts (:commit:`86635e4`) - added open_in_browser to debugging doc (:commit:`c9b690d`) -- removed reference to global scrapy stats from settings doc (:commit:`dd55067`) +- removed reference to global Scrapy stats from settings doc (:commit:`dd55067`) - Fix SpiderState bug in Windows platforms (:commit:`58998f4`) @@ -2665,7 +2665,7 @@ Scrapy 0.16.1 (released 2012-10-26) - fixed LogStats extension, which got broken after a wrong merge before the 0.16 release (:commit:`8c780fd`) - better backward compatibility for scrapy.conf.settings (:commit:`3403089`) - extended documentation on how to access crawler stats from extensions (:commit:`c4da0b5`) -- removed .hgtags (no longer needed now that scrapy uses git) (:commit:`d52c188`) +- removed .hgtags (no longer needed now that Scrapy uses git) (:commit:`d52c188`) - fix dashes under rst headers (:commit:`fa4f7f9`) - set release date for 0.16.0 in news (:commit:`e292246`) @@ -2680,8 +2680,7 @@ Scrapy changes: - documented :doc:`topics/autothrottle` and added to extensions installed by default. You still need to enable it with :setting:`AUTOTHROTTLE_ENABLED` - major Stats Collection refactoring: removed separation of global/per-spider stats, removed stats-related signals (``stats_spider_opened``, etc). Stats are much simpler now, backward compatibility is kept on the Stats Collector API and signals. - added :meth:`~scrapy.spidermiddlewares.SpiderMiddleware.process_start_requests` method to spider middlewares -- dropped Signals singleton. Signals should now be accesed through the Crawler.signals attribute. See the signals documentation for more info. -- dropped Signals singleton. Signals should now be accesed through the Crawler.signals attribute. See the signals documentation for more info. +- dropped Signals singleton. Signals should now be accessed through the Crawler.signals attribute. See the signals documentation for more info. - dropped Stats Collector singleton. Stats can now be accessed through the Crawler.stats attribute. See the stats collection documentation for more info. - documented :ref:`topics-api` - ``lxml`` is now the default selectors backend instead of ``libxml2`` @@ -2715,7 +2714,7 @@ Scrapy changes: Scrapy 0.14.4 ------------- -- added precise to supported ubuntu distros (:commit:`b7e46df`) +- added precise to supported Ubuntu distros (:commit:`b7e46df`) - fixed bug in json-rpc webservice reported in https://groups.google.com/forum/#!topic/scrapy-users/qgVBmFybNAQ/discussion. also removed no longer supported 'run' command from extras/scrapy-ws.py (:commit:`340fbdb`) - meta tag attributes for content-type http equiv can be in any order. #123 (:commit:`0cb68af`) - replace "import Image" by more standard "from PIL import Image". closes #88 (:commit:`4d17048`) @@ -2728,11 +2727,11 @@ Scrapy 0.14.3 - include egg files used by testsuite in source distribution. #118 (:commit:`c897793`) - update docstring in project template to avoid confusion with genspider command, which may be considered as an advanced feature. refs #107 (:commit:`2548dcc`) - added note to docs/topics/firebug.rst about google directory being shut down (:commit:`668e352`) -- dont discard slot when empty, just save in another dict in order to recycle if needed again. (:commit:`8e9f607`) +- don't discard slot when empty, just save in another dict in order to recycle if needed again. (:commit:`8e9f607`) - do not fail handling unicode xpaths in libxml2 backed selectors (:commit:`b830e95`) - fixed minor mistake in Request objects documentation (:commit:`bf3c9ee`) - fixed minor defect in link extractors documentation (:commit:`ba14f38`) -- removed some obsolete remaining code related to sqlite support in scrapy (:commit:`0665175`) +- removed some obsolete remaining code related to sqlite support in Scrapy (:commit:`0665175`) Scrapy 0.14.2 ------------- diff --git a/docs/topics/jobs.rst b/docs/topics/jobs.rst index f5542495b..8816a028c 100644 --- a/docs/topics/jobs.rst +++ b/docs/topics/jobs.rst @@ -74,7 +74,7 @@ Request serialization For persistence to work, :class:`~scrapy.http.Request` objects must be serializable with :mod:`pickle`, except for the ``callback`` and ``errback`` values passed to their ``__init__`` method, which must be methods of the -runnning :class:`~scrapy.spiders.Spider` class. +running :class:`~scrapy.spiders.Spider` class. If you wish to log the requests that couldn't be serialized, you can set the :setting:`SCHEDULER_DEBUG` setting to ``True`` in the project's settings page. diff --git a/docs/topics/leaks.rst b/docs/topics/leaks.rst index 793636f59..83b4d8154 100644 --- a/docs/topics/leaks.rst +++ b/docs/topics/leaks.rst @@ -110,7 +110,7 @@ ties the response lifetime to the requests' one, and that would definitely cause memory leaks. Let's see how we can discover the cause (without knowing it -a-priori, of course) by using the ``trackref`` tool. +a priori, of course) by using the ``trackref`` tool. After the crawler is running for a few minutes and we notice its memory usage has grown a lot, we can enter its telnet console and check the live diff --git a/docs/topics/media-pipeline.rst b/docs/topics/media-pipeline.rst index 206e7cfa5..a40682e5b 100644 --- a/docs/topics/media-pipeline.rst +++ b/docs/topics/media-pipeline.rst @@ -558,7 +558,7 @@ See here the methods that you can override in your custom Images Pipeline: Custom Images pipeline example ============================== -Here is a full example of the Images Pipeline whose methods are examplified +Here is a full example of the Images Pipeline whose methods are exemplified above:: import scrapy diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index a1d15a760..f138b7064 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -884,7 +884,7 @@ LOG_FORMAT Default: ``'%(asctime)s [%(name)s] %(levelname)s: %(message)s'`` -String for formatting log messsages. Refer to the `Python logging documentation`_ for the whole list of available +String for formatting log messages. Refer to the `Python logging documentation`_ for the whole list of available placeholders. .. _Python logging documentation: https://docs.python.org/2/library/logging.html#logrecord-attributes diff --git a/docs/topics/telnetconsole.rst b/docs/topics/telnetconsole.rst index 7db7e4f6b..ecdc5423c 100644 --- a/docs/topics/telnetconsole.rst +++ b/docs/topics/telnetconsole.rst @@ -48,7 +48,7 @@ autogenerated Password can be seen on scrapy logs like the example below:: 2018-10-16 14:35:21 [scrapy.extensions.telnet] INFO: Telnet Password: 16f92501e8a59326 -Default Username and Password can be overriden by the settings +Default Username and Password can be overridden by the settings :setting:`TELNETCONSOLE_USERNAME` and :setting:`TELNETCONSOLE_PASSWORD`. .. warning:: diff --git a/sep/sep-001.rst b/sep/sep-001.rst index 2a66f9802..00226283f 100644 --- a/sep/sep-001.rst +++ b/sep/sep-001.rst @@ -254,8 +254,8 @@ ItemForm #!python class MySiteForm(ItemForm): - witdth = adaptor(ItemForm.witdh, default_unit='cm') - volume = adaptor(ItemForm.witdh, default_unit='lt') + width = adaptor(ItemForm.width, default_unit='cm') + volume = adaptor(ItemForm.width, default_unit='lt') ia['width'] = x.x('//p[@class="width"]') ia['volume'] = x.x('//p[@class="volume"]') diff --git a/sep/sep-019.rst b/sep/sep-019.rst index 9fbf6a223..84f3a96c3 100644 --- a/sep/sep-019.rst +++ b/sep/sep-019.rst @@ -185,7 +185,7 @@ These ideas translate to the following changes on the ``SpiderManager`` class: will return a spider class, not an instance. It's basically a ``__get__`` to ``self._spiders``. -- All remaining functions should be deprecated or remove accordantly, since a +- All remaining functions should be deprecated or remove accordingly, since a crawler reference is no longer needed. - New helper ``get_spider_manager_class_from_scrapycfg`` in From 23a67cec271c849ad9b9c07a99531163dc3789fc Mon Sep 17 00:00:00 2001 From: Marc Hernandez Cabot <noviluni@gmail.com> Date: Thu, 19 Dec 2019 09:57:17 +0100 Subject: [PATCH 357/387] fix first letter capitalization for Raring and Scrapy --- docs/contributing.rst | 2 +- docs/index.rst | 2 +- docs/intro/install.rst | 16 ++++++++-------- docs/news.rst | 20 ++++++++++---------- docs/topics/autothrottle.rst | 2 +- docs/topics/commands.rst | 2 +- docs/topics/contracts.rst | 2 +- docs/topics/debug.rst | 2 +- docs/topics/developer-tools.rst | 4 ++-- docs/topics/logging.rst | 4 ++-- docs/topics/settings.rst | 4 ++-- docs/topics/shell.rst | 2 +- docs/topics/telnetconsole.rst | 2 +- sep/sep-004.rst | 2 +- 14 files changed, 33 insertions(+), 33 deletions(-) diff --git a/docs/contributing.rst b/docs/contributing.rst index b56295027..f40a6bba2 100644 --- a/docs/contributing.rst +++ b/docs/contributing.rst @@ -44,7 +44,7 @@ guidelines when you're going to report a new bug. * check the :ref:`FAQ <faq>` first to see if your issue is addressed in a well-known question -* if you have a general question about scrapy usage, please ask it at +* if you have a general question about Scrapy usage, please ask it at `Stack Overflow <https://stackoverflow.com/questions/tagged/scrapy>`__ (use "scrapy" tag). diff --git a/docs/index.rst b/docs/index.rst index 6d5f9e77d..a4343b7e0 100644 --- a/docs/index.rst +++ b/docs/index.rst @@ -170,7 +170,7 @@ Solving specific problems Get answers to most frequently asked questions. :doc:`topics/debug` - Learn how to debug common problems of your scrapy spider. + Learn how to debug common problems of your Scrapy spider. :doc:`topics/contracts` Learn how to use contracts for testing your spiders. diff --git a/docs/intro/install.rst b/docs/intro/install.rst index e924b5303..0d6171884 100644 --- a/docs/intro/install.rst +++ b/docs/intro/install.rst @@ -78,9 +78,9 @@ TL;DR: We recommend installing Scrapy inside a virtual environment on all platforms. Python packages can be installed either globally (a.k.a system wide), -or in user-space. We do not recommend installing scrapy system wide. +or in user-space. We do not recommend installing Scrapy system wide. -Instead, we recommend that you install scrapy within a so-called +Instead, we recommend that you install Scrapy within a so-called "virtual environment" (`virtualenv`_). Virtualenvs allow you to not conflict with already-installed Python system packages (which could break some of your system tools and scripts), @@ -97,7 +97,7 @@ Check this `user guide`_ on how to create your virtualenv. .. note:: If you use Linux or OS X, `virtualenvwrapper`_ is a handy tool to create virtualenvs. -Once you have created a virtualenv, you can install scrapy inside it with ``pip``, +Once you have created a virtualenv, you can install Scrapy inside it with ``pip``, just like any other Python package. (See :ref:`platform-specific guides <intro-install-platform-notes>` below for non-Python dependencies that you may need to install beforehand). @@ -144,7 +144,7 @@ albeit with potential issues with TLS connections. typically too old and slow to catch up with latest Scrapy. -To install scrapy on Ubuntu (or Ubuntu-based) systems, you need to install +To install Scrapy on Ubuntu (or Ubuntu-based) systems, you need to install these dependencies:: sudo apt-get install python3 python3-dev python3-pip libxml2-dev libxslt1-dev zlib1g-dev libffi-dev libssl-dev @@ -225,17 +225,17 @@ PyPy We recommend using the latest PyPy version. The version tested is 5.9.0. For PyPy3, only Linux installation was tested. -Most scrapy dependencides now have binary wheels for CPython, but not for PyPy. +Most Scrapy dependencides now have binary wheels for CPython, but not for PyPy. This means that these dependecies will be built during installation. On OS X, you are likely to face an issue with building Cryptography dependency, solution to this problem is described `here <https://github.com/pyca/cryptography/issues/2692#issuecomment-272773481>`_, that is to ``brew install openssl`` and then export the flags that this command -recommends (only needed when installing scrapy). Installing on Linux has no special +recommends (only needed when installing Scrapy). Installing on Linux has no special issues besides installing build dependencies. -Installing scrapy with PyPy on Windows is not tested. +Installing Scrapy with PyPy on Windows is not tested. -You can check that scrapy is installed correctly by running ``scrapy bench``. +You can check that Scrapy is installed correctly by running ``scrapy bench``. If this command gives errors such as ``TypeError: ... got 2 unexpected keyword arguments``, this means that setuptools was unable to pick up one PyPy-specific dependency. diff --git a/docs/news.rst b/docs/news.rst index 28db40e3c..406e889f5 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -1360,7 +1360,7 @@ Documentation - Grammar fixes: :issue:`2128`, :issue:`1566`. - Download stats badge removed from README (:issue:`2160`). -- New scrapy :ref:`architecture diagram <topics-architecture>` (:issue:`2165`). +- New Scrapy :ref:`architecture diagram <topics-architecture>` (:issue:`2165`). - Updated ``Response`` parameters documentation (:issue:`2197`). - Reworded misleading :setting:`RANDOMIZE_DOWNLOAD_DELAY` description (:issue:`2190`). - Add StackOverflow as a support channel (:issue:`2257`). @@ -1450,7 +1450,7 @@ Documentation - Use "url" variable in downloader middleware example (:issue:`2015`) - Grammar fixes (:issue:`2054`, :issue:`2120`) - New FAQ entry on using BeautifulSoup in spider callbacks (:issue:`2048`) -- Add notes about scrapy not working on Windows with Python 3 (:issue:`2060`) +- Add notes about Scrapy not working on Windows with Python 3 (:issue:`2060`) - Encourage complete titles in pull requests (:issue:`2026`) Tests @@ -1509,7 +1509,7 @@ This 1.1 release brings a lot of interesting features and bug fixes: You can use :setting:`FILES_STORE_S3_ACL` to change it. - We've reimplemented ``canonicalize_url()`` for more correct output, especially for URLs with non-ASCII characters (:issue:`1947`). - This could change link extractors output compared to previous scrapy versions. + This could change link extractors output compared to previous Scrapy versions. This may also invalidate some cache entries you could still have from pre-1.1 runs. **Warning: backward incompatible!**. @@ -1722,7 +1722,7 @@ Scrapy 1.0.4 (2015-12-30) - Merge pull request #1513 from mgedmin/patch-2 (:commit:`5d4daf8`) - Typo (:commit:`f8d0682`) - Fix list formatting (:commit:`5f83a93`) -- fix scrapy squeue tests after recent changes to queuelib (:commit:`3365c01`) +- fix Scrapy squeue tests after recent changes to queuelib (:commit:`3365c01`) - Merge pull request #1475 from rweindl/patch-1 (:commit:`2d688cd`) - Update tutorial.rst (:commit:`fbc1f25`) - Merge pull request #1449 from rhoekman/patch-1 (:commit:`7d6538c`) @@ -1734,7 +1734,7 @@ Scrapy 1.0.4 (2015-12-30) Scrapy 1.0.3 (2015-08-11) ------------------------- -- add service_identity to scrapy install_requires (:commit:`cbc2501`) +- add service_identity to Scrapy install_requires (:commit:`cbc2501`) - Workaround for travis#296 (:commit:`66af9cd`) .. _release-1.0.2: @@ -2411,7 +2411,7 @@ Enhancements - scrapy.mail.MailSender now can connect over TLS or upgrade using STARTTLS (:issue:`327`) - New FilesPipeline with functionality factored out from ImagesPipeline (:issue:`370`, :issue:`409`) - Recommend Pillow instead of PIL for image handling (:issue:`317`) -- Added Debian packages for Ubuntu Quantal and raring (:commit:`86230c0`) +- Added Debian packages for Ubuntu Quantal and Raring (:commit:`86230c0`) - Mock server (used for tests) can listen for HTTPS requests (:issue:`410`) - Remove multi spider support from multiple core components (:issue:`422`, :issue:`421`, :issue:`420`, :issue:`419`, :issue:`423`, :issue:`418`) @@ -2516,7 +2516,7 @@ Scrapy 0.18.1 (released 2013-08-27) - limit travis-ci build matrix (:commit:`3b01bb8`) - Merge pull request #375 from peterarenot/patch-1 (:commit:`fa766d7`) - Fixed so it refers to the correct folder (:commit:`3283809`) -- added Quantal & raring to support Ubuntu releases (:commit:`1411923`) +- added Quantal & Raring to support Ubuntu releases (:commit:`1411923`) - fix retry middleware which didn't retry certain connection errors after the upgrade to http1 client, closes GH-373 (:commit:`bb35ed0`) - fix XmlItemExporter in Python 2.7.4 and 2.7.5 (:commit:`de3e451`) - minor updates to 0.18 release notes (:commit:`c45e5f1`) @@ -2555,8 +2555,8 @@ Scrapy 0.18.0 (released 2013-08-09) - Collect idle downloader slots (:issue:`297`) - Add ``ftp://`` scheme downloader handler (:issue:`329`) - Added downloader benchmark webserver and spider tools :ref:`benchmarking` -- Moved persistent (on disk) queues to a separate project (queuelib_) which scrapy now depends on -- Add scrapy commands using external libraries (:issue:`260`) +- Moved persistent (on disk) queues to a separate project (queuelib_) which Scrapy now depends on +- Add Scrapy commands using external libraries (:issue:`260`) - Added ``--pdb`` option to ``scrapy`` command line tool - Added :meth:`XPathSelector.remove_namespaces <scrapy.selector.Selector.remove_namespaces>` which allows to remove all namespaces from XML documents for convenience (to work with namespace-less XPaths). Documented in :ref:`topics-selectors`. - Several improvements to spider contracts @@ -2568,7 +2568,7 @@ Scrapy 0.18.0 (released 2013-08-09) - several more cleanups to singletons and multi-spider support (thanks Nicolas Ramirez) - support custom download slots - added --spider option to "shell" command. -- log overridden settings when scrapy starts +- log overridden settings when Scrapy starts Thanks to everyone who contribute to this release. Here is a list of contributors sorted by number of commits:: diff --git a/docs/topics/autothrottle.rst b/docs/topics/autothrottle.rst index c9bece753..4317019fc 100644 --- a/docs/topics/autothrottle.rst +++ b/docs/topics/autothrottle.rst @@ -11,7 +11,7 @@ Design goals ============ 1. be nicer to sites instead of using default download delay of zero -2. automatically adjust scrapy to the optimum crawling speed, so the user +2. automatically adjust Scrapy to the optimum crawling speed, so the user doesn't have to tune the download delays to find the optimum one. The user only needs to specify the maximum concurrent requests it allows, and the extension does the rest. diff --git a/docs/topics/commands.rst b/docs/topics/commands.rst index 5b3cd7e75..a0dcba90d 100644 --- a/docs/topics/commands.rst +++ b/docs/topics/commands.rst @@ -29,7 +29,7 @@ in standard locations: 1. ``/etc/scrapy.cfg`` or ``c:\scrapy\scrapy.cfg`` (system-wide), 2. ``~/.config/scrapy.cfg`` (``$XDG_CONFIG_HOME``) and ``~/.scrapy.cfg`` (``$HOME``) for global (user-wide) settings, and -3. ``scrapy.cfg`` inside a scrapy project's root (see next section). +3. ``scrapy.cfg`` inside a Scrapy project's root (see next section). Settings from these files are merged in the listed order of preference: user-defined values have higher priority than system-wide defaults diff --git a/docs/topics/contracts.rst b/docs/topics/contracts.rst index 371ae62d5..43db8f101 100644 --- a/docs/topics/contracts.rst +++ b/docs/topics/contracts.rst @@ -64,7 +64,7 @@ Use the :command:`check` command to run the contract checks. Custom Contracts ================ -If you find you need more power than the built-in scrapy contracts you can +If you find you need more power than the built-in Scrapy contracts you can create and load your own contracts in the project by using the :setting:`SPIDER_CONTRACTS` setting:: diff --git a/docs/topics/debug.rst b/docs/topics/debug.rst index 4b2588518..d75f17301 100644 --- a/docs/topics/debug.rst +++ b/docs/topics/debug.rst @@ -5,7 +5,7 @@ Debugging Spiders ================= This document explains the most common techniques for debugging spiders. -Consider the following scrapy spider below:: +Consider the following Scrapy spider below:: import scrapy from myproject.items import MyItem diff --git a/docs/topics/developer-tools.rst b/docs/topics/developer-tools.rst index bf14643be..d1d4ebf5d 100644 --- a/docs/topics/developer-tools.rst +++ b/docs/topics/developer-tools.rst @@ -80,7 +80,7 @@ expand and collapse a tag by clicking on the arrow in front of it or by double clicking directly on the tag. If we expand the ``span`` tag with the ``class= "text"`` we will see the quote-text we clicked on. The `Inspector` lets you copy XPaths to selected elements. Let's try it out: Right-click on the ``span`` -tag, select ``Copy > XPath`` and paste it in the scrapy shell like so:: +tag, select ``Copy > XPath`` and paste it in the Scrapy shell like so:: $ scrapy shell "http://quotes.toscrape.com/" (...) @@ -159,7 +159,7 @@ The page is quite similar to the basic `quotes.toscrape.com`_-page, but instead of the above-mentioned ``Next`` button, the page automatically loads new quotes when you scroll to the bottom. We could go ahead and try out different XPaths directly, but instead -we'll check another quite useful command from the scrapy shell:: +we'll check another quite useful command from the Scrapy shell:: $ scrapy shell "quotes.toscrape.com/scroll" (...) diff --git a/docs/topics/logging.rst b/docs/topics/logging.rst index dd09477b8..d4d22d889 100644 --- a/docs/topics/logging.rst +++ b/docs/topics/logging.rst @@ -171,9 +171,9 @@ listed in `logging's logrecord attributes docs <https://docs.python.org/2/library/datetime.html#strftime-and-strptime-behavior>`_ respectively. -If :setting:`LOG_SHORT_NAMES` is set, then the logs will not display the scrapy +If :setting:`LOG_SHORT_NAMES` is set, then the logs will not display the Scrapy component that prints the log. It is unset by default, hence logs contain the -scrapy component responsible for that log output. +Scrapy component responsible for that log output. Command-line options -------------------- diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index f138b7064..afe4fade1 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -1264,7 +1264,7 @@ Default:: 'scrapy.contracts.default.ScrapesContract': 3, } -A dict containing the scrapy contracts enabled by default in Scrapy. You should +A dict containing the Scrapy contracts enabled by default in Scrapy. You should never modify this setting in your project, modify :setting:`SPIDER_CONTRACTS` instead. For more info see :ref:`topics-contracts`. @@ -1295,7 +1295,7 @@ SPIDER_LOADER_WARN_ONLY Default: ``False`` -By default, when scrapy tries to import spider classes from :setting:`SPIDER_MODULES`, +By default, when Scrapy tries to import spider classes from :setting:`SPIDER_MODULES`, it will fail loudly if there is any ``ImportError`` exception. But you can choose to silence this exception and turn it into a simple warning by setting ``SPIDER_LOADER_WARN_ONLY = True``. diff --git a/docs/topics/shell.rst b/docs/topics/shell.rst index 68a0b19b5..4fe0dea06 100644 --- a/docs/topics/shell.rst +++ b/docs/topics/shell.rst @@ -31,7 +31,7 @@ for more info. Scrapy also has support for `bpython`_, and will try to use it where `IPython`_ is unavailable. -Through scrapy's settings you can configure it to use any one of +Through Scrapy's settings you can configure it to use any one of ``ipython``, ``bpython`` or the standard ``python`` shell, regardless of which are installed. This is done by setting the ``SCRAPY_PYTHON_SHELL`` environment variable; or by defining it in your :ref:`scrapy.cfg <topics-config-settings>`:: diff --git a/docs/topics/telnetconsole.rst b/docs/topics/telnetconsole.rst index ecdc5423c..47d8d393c 100644 --- a/docs/topics/telnetconsole.rst +++ b/docs/topics/telnetconsole.rst @@ -44,7 +44,7 @@ the console you need to type:: >>> By default Username is ``scrapy`` and Password is autogenerated. The -autogenerated Password can be seen on scrapy logs like the example below:: +autogenerated Password can be seen on Scrapy logs like the example below:: 2018-10-16 14:35:21 [scrapy.extensions.telnet] INFO: Telnet Password: 16f92501e8a59326 diff --git a/sep/sep-004.rst b/sep/sep-004.rst index 69edfa136..05b0eb99c 100644 --- a/sep/sep-004.rst +++ b/sep/sep-004.rst @@ -53,7 +53,7 @@ Here's a simple proof-of-concept code of such script: # ... do something more interesting with scraped_items ... The behaviour of the Scrapy crawler would be controller by the Scrapy settings, -naturally, just like any typical scrapy project. But the default settings +naturally, just like any typical Scrapy project. But the default settings should be sufficient so as to not require adding any specific setting. But, at the same time, you could do it if you need to, say, for specifying a custom middleware. From f6bc1940a3394f74d8aa088faf2912a43690b260 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 19 Dec 2019 12:06:15 +0100 Subject: [PATCH 358/387] Use Python 3.7 to build the documentation --- .readthedocs.yml | 4 +++- .travis.yml | 2 +- 2 files changed, 4 insertions(+), 2 deletions(-) diff --git a/.readthedocs.yml b/.readthedocs.yml index 3c1c3e8be..563add75f 100644 --- a/.readthedocs.yml +++ b/.readthedocs.yml @@ -2,6 +2,8 @@ version: 2 sphinx: configuration: docs/conf.py python: - version: 3.8 + # For available versions, see: + # https://docs.readthedocs.io/en/stable/config-file/v2.html#build-image + version: 3.7 # Keep in sync with .travis.yml install: - requirements: docs/requirements.txt diff --git a/.travis.yml b/.travis.yml index c870934e1..f9d0dc8be 100644 --- a/.travis.yml +++ b/.travis.yml @@ -25,7 +25,7 @@ matrix: - env: TOXENV=extra-deps python: 3.8 - env: TOXENV=docs - python: 3.8 + python: 3.7 # Keep in sync with .readthedocs.yml install: - | if [ "$TOXENV" = "pypy3" ]; then From e22c0c27d9d33383f4ac18a4142ccffe1d9a0d3b Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Adri=C3=A1n=20Chaves?= <adrian@chaves.io> Date: Thu, 19 Dec 2019 12:15:54 +0100 Subject: [PATCH 359/387] Revert "Improve FilteringLinkExtractor.__new__" This reverts commit ee9881d2704798c9cd61b6da503bb0694227c58c. --- scrapy/linkextractors/__init__.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/scrapy/linkextractors/__init__.py b/scrapy/linkextractors/__init__.py index d0d34035d..bdeab3a75 100644 --- a/scrapy/linkextractors/__init__.py +++ b/scrapy/linkextractors/__init__.py @@ -60,7 +60,7 @@ class FilteringLinkExtractor(object): warn('scrapy.linkextractors.FilteringLinkExtractor is deprecated, ' 'please use scrapy.linkextractors.LinkExtractor instead', ScrapyDeprecationWarning, stacklevel=2) - return super().__new__(cls, *args, **kwargs) + return super(FilteringLinkExtractor, cls).__new__(cls) def __init__(self, link_extractor, allow, deny, allow_domains, deny_domains, restrict_xpaths, canonicalize, deny_extensions, restrict_css, restrict_text): From 40697dcbfa17dccf81adce7d033bc466ba6e98a2 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 20 Dec 2019 19:33:44 +0500 Subject: [PATCH 360/387] Remove deferred_from_coro from this PR. --- scrapy/utils/defer.py | 28 ---------------------------- 1 file changed, 28 deletions(-) diff --git a/scrapy/utils/defer.py b/scrapy/utils/defer.py index 530bf0e9d..20ce59297 100644 --- a/scrapy/utils/defer.py +++ b/scrapy/utils/defer.py @@ -1,14 +1,10 @@ """ Helper functions for dealing with Twisted deferreds """ -import asyncio -import inspect - from twisted.internet import defer, task from twisted.python import failure from scrapy.exceptions import IgnoreRequest -from scrapy.utils.asyncio import is_asyncio_reactor_installed def defer_fail(_failure): @@ -118,27 +114,3 @@ def iter_errback(iterable, errback, *a, **kw): break except Exception: errback(failure.Failure(), *a, **kw) - - -def _isfuture(o): - # workaround for Python before 3.5.3 not having asyncio.isfuture - if hasattr(asyncio, 'isfuture'): - return asyncio.isfuture(o) - return isinstance(o, asyncio.Future) - - -def deferred_from_coro(o, asyncio_enabled=False): - """Converts a coroutine into a Deferred, or returns the object as is if it isn't a coroutine""" - if isinstance(o, defer.Deferred): - return o - if _isfuture(o) or inspect.isawaitable(o): - if not asyncio_enabled: - # wrapping the coroutine directly into a Deferred, this doesn't work correctly with coroutines - # that use asyncio, e.g. "await asyncio.sleep(1)" - return defer.ensureDeferred(o) - else: - # wrapping the coroutine into a Future and then into a Deferred, this requires AsyncioSelectorReactor - if not is_asyncio_reactor_installed(): - raise TypeError('Using coroutines requires installing AsyncioSelectorReactor') - return defer.Deferred.fromFuture(asyncio.ensure_future(o)) - return o From e342de5038e3757660f947fe1fadf35b54cd7113 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 20 Dec 2019 19:37:50 +0500 Subject: [PATCH 361/387] Remove a stray newline. --- tests/mockserver.py | 1 - 1 file changed, 1 deletion(-) diff --git a/tests/mockserver.py b/tests/mockserver.py index d4e0362fb..a45277db9 100644 --- a/tests/mockserver.py +++ b/tests/mockserver.py @@ -6,7 +6,6 @@ from subprocess import Popen, PIPE from urllib.parse import urlencode from OpenSSL import SSL - from twisted.web.server import Site, NOT_DONE_YET from twisted.web.resource import Resource from twisted.web.static import File From 8de80f59db19d739a056a7e58662f90544fece16 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Sat, 21 Dec 2019 13:08:29 +0500 Subject: [PATCH 362/387] Raise an exception if ASYNCIO_ENABLED but the reactor is wrong. --- scrapy/crawler.py | 6 +++++- scrapy/utils/log.py | 8 +------- tests/test_crawler.py | 9 +++++---- 3 files changed, 11 insertions(+), 12 deletions(-) diff --git a/scrapy/crawler.py b/scrapy/crawler.py index 706c8a59d..a9443f7ac 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -14,7 +14,7 @@ from scrapy.extension import ExtensionManager from scrapy.settings import overridden_settings, Settings from scrapy.signalmanager import SignalManager from scrapy.exceptions import ScrapyDeprecationWarning -from scrapy.utils.asyncio import install_asyncio_reactor +from scrapy.utils.asyncio import install_asyncio_reactor, is_asyncio_reactor_installed from scrapy.utils.ossignal import install_shutdown_handlers, signal_names from scrapy.utils.misc import load_object from scrapy.utils.log import ( @@ -259,6 +259,10 @@ class CrawlerProcess(CrawlerRunner): super(CrawlerProcess, self).__init__(settings) if self.settings.getbool('ASYNCIO_ENABLED'): install_asyncio_reactor() + if not is_asyncio_reactor_installed(): + raise Exception("ASYNCIO_ENABLED is on but the Twisted asyncio " + "reactor is not installed, this is not supported.") + install_shutdown_handlers(self._signal_shutdown) configure_logging(self.settings, install_root_handler) log_scrapy_info(self.settings) diff --git a/scrapy/utils/log.py b/scrapy/utils/log.py index 0fe3d1549..6179e1bd1 100644 --- a/scrapy/utils/log.py +++ b/scrapy/utils/log.py @@ -11,7 +11,6 @@ from twisted.python import log as twisted_log import scrapy from scrapy.settings import Settings from scrapy.exceptions import ScrapyDeprecationWarning -from scrapy.utils.asyncio import is_asyncio_reactor_installed from scrapy.utils.versions import scrapy_components_versions @@ -150,12 +149,7 @@ def log_scrapy_info(settings): for name, version in scrapy_components_versions() if name != "Scrapy")}) if settings.getbool('ASYNCIO_ENABLED'): - if is_asyncio_reactor_installed(): - logger.debug("Asyncio support enabled") - else: - logger.error("ASYNCIO_ENABLED is on but the Twisted asyncio " - "reactor is not installed, this is not supported " - "and asyncio coroutines will not work.") + logger.debug("Asyncio support enabled") class StreamLogger(object): diff --git a/tests/test_crawler.py b/tests/test_crawler.py index 0b2645280..a2865fcd1 100644 --- a/tests/test_crawler.py +++ b/tests/test_crawler.py @@ -264,13 +264,14 @@ class CrawlerRunnerHasSpider(unittest.TestCase): @defer.inlineCallbacks def test_crawler_process_asyncio_enabled_true(self): with LogCapture(level=logging.DEBUG) as log: - runner = CrawlerProcess(settings={'ASYNCIO_ENABLED': True}) - yield runner.crawl(NoRequestsSpider) if self.reactor_pytest == 'asyncio': + runner = CrawlerProcess(settings={'ASYNCIO_ENABLED': True}) + yield runner.crawl(NoRequestsSpider) self.assertIn("Asyncio support enabled", str(log)) else: - self.assertNotIn("Asyncio support enabled", str(log)) - self.assertIn("ASYNCIO_ENABLED is on but the Twisted asyncio reactor is not installed", str(log)) + msg = "ASYNCIO_ENABLED is on but the Twisted asyncio reactor is not installed" + with self.assertRaisesRegex(Exception, msg): + runner = CrawlerProcess(settings={'ASYNCIO_ENABLED': True}) @defer.inlineCallbacks def test_crawler_process_asyncio_enabled_false(self): From 87ece066ca320b07acda57c99aee8a62992ec144 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 26 Dec 2019 20:41:06 +0500 Subject: [PATCH 363/387] Remove conditional asyncio imports. --- scrapy/utils/asyncio.py | 16 ++++------------ 1 file changed, 4 insertions(+), 12 deletions(-) diff --git a/scrapy/utils/asyncio.py b/scrapy/utils/asyncio.py index b53c8a8b0..917973de2 100644 --- a/scrapy/utils/asyncio.py +++ b/scrapy/utils/asyncio.py @@ -1,25 +1,17 @@ +import asyncio from contextlib import suppress +from twisted.internet import asyncioreactor from twisted.internet.error import ReactorAlreadyInstalledError def install_asyncio_reactor(): """ Tries to install AsyncioSelectorReactor """ - try: - import asyncio - from twisted.internet import asyncioreactor - except ImportError: - return - with suppress(ReactorAlreadyInstalledError): asyncioreactor.install(asyncio.get_event_loop()) def is_asyncio_reactor_installed(): - try: - import twisted.internet.reactor - from twisted.internet import asyncioreactor - return isinstance(twisted.internet.reactor, asyncioreactor.AsyncioSelectorReactor) - except ImportError: - return False + from twisted.internet import reactor + return isinstance(reactor, asyncioreactor.AsyncioSelectorReactor) From 37ac47ff8074959ea66566fcb8b0e9e62272f963 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 26 Dec 2019 20:46:54 +0500 Subject: [PATCH 364/387] Fix a deprecation warning. --- tests/test_utils_asyncio.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tests/test_utils_asyncio.py b/tests/test_utils_asyncio.py index a6ba24876..44acc24af 100644 --- a/tests/test_utils_asyncio.py +++ b/tests/test_utils_asyncio.py @@ -10,7 +10,7 @@ class AsyncioTest(TestCase): def test_is_asyncio_reactor_installed(self): # the result should depend only on the pytest --reactor argument - self.assertEquals(is_asyncio_reactor_installed(), self.reactor_pytest == 'asyncio') + self.assertEqual(is_asyncio_reactor_installed(), self.reactor_pytest == 'asyncio') def test_install_asyncio_reactor(self): # this should do nothing From 30ebd05a5f627262702ad1e1e488d6176a8c7882 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 27 Dec 2019 00:05:14 +0500 Subject: [PATCH 365/387] Simplify the tox asyncio entries. --- tox.ini | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/tox.ini b/tox.ini index ed0d4c9ab..b62100026 100644 --- a/tox.ini +++ b/tox.ini @@ -99,7 +99,7 @@ commands = [asyncio] commands = - py.test --cov=scrapy --cov-report= --reactor=asyncio {posargs:scrapy tests} + {[testenv]commands} --reactor=asyncio [testenv:py35-asyncio] basepython = python3.5 From f75ccc997aa75fadf96a8d4b836248397ef89802 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 27 Dec 2019 19:48:54 +0500 Subject: [PATCH 366/387] FIx a typo in the only_asyncio fixture. --- conftest.py | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/conftest.py b/conftest.py index 6d9696a3f..c0de09909 100644 --- a/conftest.py +++ b/conftest.py @@ -48,4 +48,4 @@ def reactor_pytest(request): @pytest.fixture(autouse=True) def only_asyncio(request, reactor_pytest): if request.node.get_closest_marker('only_asyncio') and reactor_pytest != 'asyncio': - pytest.skip('This test is only run with --reactor-asyncio') + pytest.skip('This test is only run with --reactor=asyncio') From dc1ee09481c7655a8ebf77a75ccc965a4ba5400d Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 27 Dec 2019 21:55:58 +0500 Subject: [PATCH 367/387] Rename ASYNCIO_ENABLED to ASYNCIO_REACTOR, change the logic accordingly. --- docs/topics/settings.rst | 14 +++---- scrapy/crawler.py | 17 +++++--- scrapy/settings/default_settings.py | 2 +- scrapy/utils/log.py | 5 ++- .../asyncio_enabled_no_reactor.py | 2 +- .../CrawlerProcess/asyncio_enabled_reactor.py | 2 +- tests/test_commands.py | 8 ++-- tests/test_crawler.py | 42 ++++++++----------- 8 files changed, 43 insertions(+), 49 deletions(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index 62b2870b2..c02f877fc 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -160,26 +160,22 @@ to any particular component. In that case the module of that component will be shown, typically an extension, middleware or pipeline. It also means that the component must be enabled in order for the setting to have any effect. -.. setting:: ASYNCIO_ENABLED +.. setting:: ASYNCIO_REACTOR -ASYNCIO_ENABLED +ASYNCIO_REACTOR --------------- Default: ``False`` -Whether to support ``async def`` methods and callbacks which use code that -requires an asyncio loop. - -If an ``async def`` coroutine doesn't require the asyncio loop, it will work -even if this is set to ``False``. Coroutines that require the asyncio loop may -silently fail to run or raise errors unless this is set to ``True``. +Whether to install and require the Twisted reactor that uses the asyncio loop. When this option is set to ``True``, Scrapy will require :class:`~twisted.internet.asyncioreactor.AsyncioSelectorReactor`. It will install this reactor if no reactor is installed yet, such as when using the ``scrapy`` script or :class:`~scrapy.crawler.CrawlerProcess`. If you are using :class:`~scrapy.crawler.CrawlerRunner`, you need to install the correct reactor -manually. +manually. If a different reactor is installed outside Scrapy, it will raise an +exception. The default value for this option is currently ``False`` to maintain backward compatibility and avoid possible problems caused by using a different Twisted diff --git a/scrapy/crawler.py b/scrapy/crawler.py index a9443f7ac..f87e67d93 100644 --- a/scrapy/crawler.py +++ b/scrapy/crawler.py @@ -137,6 +137,7 @@ class CrawlerRunner(object): self._crawlers = set() self._active = set() self.bootstrap_failed = False + self._handle_asyncio_reactor() @property def spiders(self): @@ -230,6 +231,11 @@ class CrawlerRunner(object): while self._active: yield defer.DeferredList(self._active) + def _handle_asyncio_reactor(self): + if self.settings.getbool('ASYNCIO_REACTOR') and not is_asyncio_reactor_installed(): + raise Exception("ASYNCIO_REACTOR is on but the Twisted asyncio " + "reactor is not installed.") + class CrawlerProcess(CrawlerRunner): """ @@ -257,12 +263,6 @@ class CrawlerProcess(CrawlerRunner): def __init__(self, settings=None, install_root_handler=True): super(CrawlerProcess, self).__init__(settings) - if self.settings.getbool('ASYNCIO_ENABLED'): - install_asyncio_reactor() - if not is_asyncio_reactor_installed(): - raise Exception("ASYNCIO_ENABLED is on but the Twisted asyncio " - "reactor is not installed, this is not supported.") - install_shutdown_handlers(self._signal_shutdown) configure_logging(self.settings, install_root_handler) log_scrapy_info(self.settings) @@ -333,6 +333,11 @@ class CrawlerProcess(CrawlerRunner): except RuntimeError: # raised if already stopped or in shutdown stage pass + def _handle_asyncio_reactor(self): + if self.settings.getbool('ASYNCIO_REACTOR'): + install_asyncio_reactor() + super()._handle_asyncio_reactor() + def _get_spider_loader(settings): """ Get SpiderLoader instance from settings """ diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index a7792e248..d03fd37b0 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -19,7 +19,7 @@ from os.path import join, abspath, dirname AJAXCRAWL_ENABLED = False -ASYNCIO_ENABLED = False +ASYNCIO_REACTOR = False AUTOTHROTTLE_ENABLED = False AUTOTHROTTLE_DEBUG = False diff --git a/scrapy/utils/log.py b/scrapy/utils/log.py index 6179e1bd1..e4cf0196b 100644 --- a/scrapy/utils/log.py +++ b/scrapy/utils/log.py @@ -11,6 +11,7 @@ from twisted.python import log as twisted_log import scrapy from scrapy.settings import Settings from scrapy.exceptions import ScrapyDeprecationWarning +from scrapy.utils.asyncio import is_asyncio_reactor_installed from scrapy.utils.versions import scrapy_components_versions @@ -148,8 +149,8 @@ def log_scrapy_info(settings): {'versions': ", ".join("%s %s" % (name, version) for name, version in scrapy_components_versions() if name != "Scrapy")}) - if settings.getbool('ASYNCIO_ENABLED'): - logger.debug("Asyncio support enabled") + if is_asyncio_reactor_installed(): + logger.debug("Asyncio reactor is installed") class StreamLogger(object): diff --git a/tests/CrawlerProcess/asyncio_enabled_no_reactor.py b/tests/CrawlerProcess/asyncio_enabled_no_reactor.py index dfe028ef4..db1b75931 100644 --- a/tests/CrawlerProcess/asyncio_enabled_no_reactor.py +++ b/tests/CrawlerProcess/asyncio_enabled_no_reactor.py @@ -10,7 +10,7 @@ class NoRequestsSpider(scrapy.Spider): process = CrawlerProcess(settings={ - 'ASYNCIO_ENABLED': True, + 'ASYNCIO_REACTOR': True, }) process.crawl(NoRequestsSpider) diff --git a/tests/CrawlerProcess/asyncio_enabled_reactor.py b/tests/CrawlerProcess/asyncio_enabled_reactor.py index 7a172ea28..cec3c9c25 100644 --- a/tests/CrawlerProcess/asyncio_enabled_reactor.py +++ b/tests/CrawlerProcess/asyncio_enabled_reactor.py @@ -15,7 +15,7 @@ class NoRequestsSpider(scrapy.Spider): process = CrawlerProcess(settings={ - 'ASYNCIO_ENABLED': True, + 'ASYNCIO_REACTOR': True, }) process.crawl(NoRequestsSpider) diff --git a/tests/test_commands.py b/tests/test_commands.py index 197d80217..6024af71c 100644 --- a/tests/test_commands.py +++ b/tests/test_commands.py @@ -296,12 +296,12 @@ class BadSpider(scrapy.Spider): self.assertIn("badspider.py", log) def test_asyncio_enabled_true(self): - log = self.get_log(self.debug_log_spider, args=['-s', 'ASYNCIO_ENABLED=True']) - self.assertIn("DEBUG: Asyncio support enabled", log) + log = self.get_log(self.debug_log_spider, args=['-s', 'ASYNCIO_REACTOR=True']) + self.assertIn("DEBUG: Asyncio reactor is installed", log) def test_asyncio_enabled_false(self): - log = self.get_log(self.debug_log_spider, args=['-s', 'ASYNCIO_ENABLED=False']) - self.assertNotIn("DEBUG: Asyncio support enabled", log) + log = self.get_log(self.debug_log_spider, args=['-s', 'ASYNCIO_REACTOR=False']) + self.assertNotIn("DEBUG: Asyncio reactor is installed", log) class BenchCommandTest(CommandTest): diff --git a/tests/test_crawler.py b/tests/test_crawler.py index a2865fcd1..fce60ca37 100644 --- a/tests/test_crawler.py +++ b/tests/test_crawler.py @@ -13,7 +13,6 @@ import scrapy from scrapy.crawler import Crawler, CrawlerRunner, CrawlerProcess from scrapy.settings import Settings, default_settings from scrapy.spiderloader import SpiderLoader -from scrapy.utils.asyncio import is_asyncio_reactor_installed from scrapy.utils.log import configure_logging, get_scrapy_root_handler from scrapy.utils.spider import DefaultSpider from scrapy.utils.misc import load_object @@ -209,14 +208,6 @@ class NoRequestsSpider(scrapy.Spider): return [] -class AsyncioSpider(scrapy.Spider): - name = 'asyncio' - - def start_requests(self): - self.logger.info('Asyncio support: %s', is_asyncio_reactor_installed()) - return [] - - @mark.usefixtures('reactor_pytest') class CrawlerRunnerHasSpider(unittest.TestCase): @@ -261,31 +252,32 @@ class CrawlerRunnerHasSpider(unittest.TestCase): self.assertEqual(runner.bootstrap_failed, True) + def test_crawler_runner_asyncio_enabled_true(self): + if self.reactor_pytest == 'asyncio': + runner = CrawlerRunner(settings={'ASYNCIO_REACTOR': True}) + else: + msg = "ASYNCIO_REACTOR is on but the Twisted asyncio reactor is not installed" + with self.assertRaisesRegex(Exception, msg): + runner = CrawlerRunner(settings={'ASYNCIO_REACTOR': True}) + @defer.inlineCallbacks def test_crawler_process_asyncio_enabled_true(self): with LogCapture(level=logging.DEBUG) as log: if self.reactor_pytest == 'asyncio': - runner = CrawlerProcess(settings={'ASYNCIO_ENABLED': True}) + runner = CrawlerProcess(settings={'ASYNCIO_REACTOR': True}) yield runner.crawl(NoRequestsSpider) - self.assertIn("Asyncio support enabled", str(log)) + self.assertIn("Asyncio reactor is installed", str(log)) else: - msg = "ASYNCIO_ENABLED is on but the Twisted asyncio reactor is not installed" + msg = "ASYNCIO_REACTOR is on but the Twisted asyncio reactor is not installed" with self.assertRaisesRegex(Exception, msg): - runner = CrawlerProcess(settings={'ASYNCIO_ENABLED': True}) + runner = CrawlerProcess(settings={'ASYNCIO_REACTOR': True}) @defer.inlineCallbacks def test_crawler_process_asyncio_enabled_false(self): - runner = CrawlerProcess(settings={'ASYNCIO_ENABLED': False}) + runner = CrawlerProcess(settings={'ASYNCIO_REACTOR': False}) with LogCapture(level=logging.DEBUG) as log: yield runner.crawl(NoRequestsSpider) - self.assertNotIn("Asyncio support enabled", str(log)) - - @defer.inlineCallbacks - def test_crawler_runner_asyncio_supported(self): - runner = CrawlerRunner() - with LogCapture() as log: - yield runner.crawl(AsyncioSpider) - log.check_present(('asyncio', 'INFO', 'Asyncio support: %s' % (self.reactor_pytest == 'asyncio'))) + self.assertNotIn("Asyncio reactor is installed", str(log)) class CrawlerProcessSubprocess(unittest.TestCase): @@ -302,14 +294,14 @@ class CrawlerProcessSubprocess(unittest.TestCase): def test_simple(self): log = self.run_script('simple.py') self.assertIn('Spider closed (finished)', log) - self.assertNotIn("DEBUG: Asyncio support enabled", log) + self.assertNotIn("DEBUG: Asyncio reactor is installed", log) def test_asyncio_enabled_no_reactor(self): log = self.run_script('asyncio_enabled_no_reactor.py') self.assertIn('Spider closed (finished)', log) - self.assertIn("DEBUG: Asyncio support enabled", log) + self.assertIn("DEBUG: Asyncio reactor is installed", log) def test_asyncio_enabled_reactor(self): log = self.run_script('asyncio_enabled_reactor.py') self.assertIn('Spider closed (finished)', log) - self.assertIn("DEBUG: Asyncio support enabled", log) + self.assertIn("DEBUG: Asyncio reactor is installed", log) From 82861c73c8f74c3416f5fa5d77ded33802ca535f Mon Sep 17 00:00:00 2001 From: Atul Gopinathan <41539794+atul-g@users.noreply.github.com> Date: Fri, 27 Dec 2019 22:57:58 +0530 Subject: [PATCH 368/387] Edited the link of the homepage of lxml website The link "https://lxml.de" is redirecting to a completely different and unintended website. I changed the link to the index page of lxml's official website. I thought of changing it to the PyPi page of lxml, but even they are providing the same "https://lxml.de" link which doesn't seem to be working now. --- docs/intro/install.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/intro/install.rst b/docs/intro/install.rst index 51b41b4d7..8ce13d715 100644 --- a/docs/intro/install.rst +++ b/docs/intro/install.rst @@ -278,7 +278,7 @@ For details, see `Issue #2473 <https://github.com/scrapy/scrapy/issues/2473>`_. .. _Python: https://www.python.org/ .. _pip: https://pip.pypa.io/en/latest/installing/ -.. _lxml: http://lxml.de/ +.. _lxml: https://lxml.de/index.html .. _parsel: https://pypi.python.org/pypi/parsel .. _w3lib: https://pypi.python.org/pypi/w3lib .. _twisted: https://twistedmatrix.com/ From 14d4428e705e9ca739a4fd041ce7e3134749363c Mon Sep 17 00:00:00 2001 From: 1um0s <abitha95@gmail.com> Date: Mon, 30 Dec 2019 01:26:22 +0530 Subject: [PATCH 369/387] Rephrasing documentation for image and file pipelines (#4252) * scrapy#4034 Clarify documentation for image and file pipelines * scrapy#4034 Clarify documentation for file pipeline * scrapy#4034 Simplify documentation for pipeline * scrapy#4034 Simplify documentation for pipeline * scrapy#4034 Clarify documentation for image and file pipelines * scrapy#4034 Clarify documentation for file pipeline * scrapy#4034 Simplify documentation for pipeline * scrapy#4034 Simplify documentation for pipeline * scrapy#4034 Revert image, file pipeline docs. Enhance custom media pipeline docs. * scrapy#4034 rebase master * scrapy#4034 Clarify documentation for image and file pipelines * scrapy#4034 Clarify documentation for file pipeline * scrapy#4034 Simplify documentation for pipeline * scrapy#4034 Simplify documentation for pipeline * scrapy#4034 Clarify documentation for image and file pipelines * scrapy#4034 Clarify documentation for file pipeline * scrapy#4034 Simplify documentation for pipeline * scrapy#4034 Simplify documentation for pipeline * scrapy#4034 Revert image, file pipeline docs. Enhance custom media pipeline docs. * scrapy#4034 rebase master * Rebase master * Add class to media pipeline docs Co-Authored-By: elacuesta <elacuesta@users.noreply.github.com> Co-authored-by: elacuesta <elacuesta@users.noreply.github.com> --- docs/topics/media-pipeline.rst | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/docs/topics/media-pipeline.rst b/docs/topics/media-pipeline.rst index a40682e5b..332a14eb7 100644 --- a/docs/topics/media-pipeline.rst +++ b/docs/topics/media-pipeline.rst @@ -97,7 +97,6 @@ For Files Pipeline, use:: ITEM_PIPELINES = {'scrapy.pipelines.files.FilesPipeline': 1} - .. note:: You can also use both the Files and Images Pipeline at the same time. @@ -578,4 +577,12 @@ above:: item['image_paths'] = image_paths return item + +To enable your custom media pipeline component you must add its class import path to the +:setting:`ITEM_PIPELINES` setting, like in the following example:: + + ITEM_PIPELINES = { + 'myproject.pipelines.MyImagesPipeline': 300 + } + .. _MD5 hash: https://en.wikipedia.org/wiki/MD5 From 21f50c795ac6978de9e73b64f31542a9df928ac5 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 30 Jul 2019 19:46:18 +0500 Subject: [PATCH 370/387] Add async def support to downloader middlewares. --- scrapy/core/downloader/middleware.py | 8 ++++---- tests/test_downloadermiddleware.py | 24 ++++++++++++++++++++++++ 2 files changed, 28 insertions(+), 4 deletions(-) diff --git a/scrapy/core/downloader/middleware.py b/scrapy/core/downloader/middleware.py index 38608a429..9c0014206 100644 --- a/scrapy/core/downloader/middleware.py +++ b/scrapy/core/downloader/middleware.py @@ -8,7 +8,7 @@ from twisted.internet import defer from scrapy.exceptions import _InvalidOutput from scrapy.http import Request, Response from scrapy.middleware import MiddlewareManager -from scrapy.utils.defer import mustbe_deferred +from scrapy.utils.defer import mustbe_deferred, deferred_from_coro from scrapy.utils.conf import build_component_list @@ -33,7 +33,7 @@ class DownloaderMiddlewareManager(MiddlewareManager): @defer.inlineCallbacks def process_request(request): for method in self.methods['process_request']: - response = yield method(request=request, spider=spider) + response = yield deferred_from_coro(method(request=request, spider=spider)) if response is not None and not isinstance(response, (Response, Request)): raise _InvalidOutput('Middleware %s.process_request must return None, Response or Request, got %s' % \ (method.__self__.__class__.__name__, response.__class__.__name__)) @@ -48,7 +48,7 @@ class DownloaderMiddlewareManager(MiddlewareManager): defer.returnValue(response) for method in self.methods['process_response']: - response = yield method(request=request, response=response, spider=spider) + response = yield deferred_from_coro(method(request=request, response=response, spider=spider)) if not isinstance(response, (Response, Request)): raise _InvalidOutput('Middleware %s.process_response must return Response or Request, got %s' % \ (method.__self__.__class__.__name__, type(response))) @@ -60,7 +60,7 @@ class DownloaderMiddlewareManager(MiddlewareManager): def process_exception(_failure): exception = _failure.value for method in self.methods['process_exception']: - response = yield method(request=request, exception=exception, spider=spider) + response = yield deferred_from_coro(method(request=request, exception=exception, spider=spider)) if response is not None and not isinstance(response, (Response, Request)): raise _InvalidOutput('Middleware %s.process_exception must return None, Response or Request, got %s' % \ (method.__self__.__class__.__name__, type(response))) diff --git a/tests/test_downloadermiddleware.py b/tests/test_downloadermiddleware.py index 1b81ea949..135321d0a 100644 --- a/tests/test_downloadermiddleware.py +++ b/tests/test_downloadermiddleware.py @@ -1,3 +1,4 @@ +import asyncio from unittest import mock from twisted.internet.defer import Deferred @@ -206,3 +207,26 @@ class MiddlewareUsingDeferreds(ManagerTestCase): self.assertIs(results[0], resp) self.assertFalse(download_func.called) + + +class MiddlewareUsingCoro(ManagerTestCase): + """Middlewares using asyncio coroutines should work""" + + def test_asyncdef(self): + resp = Response('http://example.com/index.html') + + class CoroMiddleware: + async def process_request(self, request, spider): + await asyncio.sleep(0.1) + return resp + + self.mwman._add_middleware(CoroMiddleware()) + req = Request('http://example.com/index.html') + download_func = mock.MagicMock() + dfd = self.mwman.download(download_func, req, self.spider) + results = [] + dfd.addBoth(results.append) + self._wait(dfd) + + self.assertIs(results[0], resp) + self.assertFalse(download_func.called) From 3603644552f8d15d203abca221edb56047119528 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 12 Sep 2019 20:25:29 +0500 Subject: [PATCH 371/387] Add a non-asyncio async def middleware test. --- tests/test_downloadermiddleware.py | 20 ++++++++++++++++++++ 1 file changed, 20 insertions(+) diff --git a/tests/test_downloadermiddleware.py b/tests/test_downloadermiddleware.py index 135321d0a..5b0cf1eb7 100644 --- a/tests/test_downloadermiddleware.py +++ b/tests/test_downloadermiddleware.py @@ -1,6 +1,7 @@ import asyncio from unittest import mock +from twisted.internet import defer from twisted.internet.defer import Deferred from twisted.trial.unittest import TestCase from twisted.python.failure import Failure @@ -215,6 +216,25 @@ class MiddlewareUsingCoro(ManagerTestCase): def test_asyncdef(self): resp = Response('http://example.com/index.html') + class CoroMiddleware: + async def process_request(self, request, spider): + await defer.succeed(42) + return resp + + self.mwman._add_middleware(CoroMiddleware()) + req = Request('http://example.com/index.html') + download_func = mock.MagicMock() + dfd = self.mwman.download(download_func, req, self.spider) + results = [] + dfd.addBoth(results.append) + self._wait(dfd) + + self.assertIs(results[0], resp) + self.assertFalse(download_func.called) + + def test_asyncdef_asyncio(self): + resp = Response('http://example.com/index.html') + class CoroMiddleware: async def process_request(self, request, spider): await asyncio.sleep(0.1) From 5cf1ac0005fa3a174b700d88c0e2536b689f13c7 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Mon, 16 Dec 2019 19:24:44 +0500 Subject: [PATCH 372/387] Move the asyncio downloader mw test to a separate class. --- tests/test_downloadermiddleware.py | 7 +++++++ 1 file changed, 7 insertions(+) diff --git a/tests/test_downloadermiddleware.py b/tests/test_downloadermiddleware.py index 5b0cf1eb7..c5c4d13bd 100644 --- a/tests/test_downloadermiddleware.py +++ b/tests/test_downloadermiddleware.py @@ -1,6 +1,7 @@ import asyncio from unittest import mock +from pytest import mark from twisted.internet import defer from twisted.internet.defer import Deferred from twisted.trial.unittest import TestCase @@ -232,6 +233,12 @@ class MiddlewareUsingCoro(ManagerTestCase): self.assertIs(results[0], resp) self.assertFalse(download_func.called) + +@mark.only_asyncio() +class MiddlewareUsingCoroAsyncio(ManagerTestCase): + + settings_dict = {'ASYNCIO_ENABLED': True} + def test_asyncdef_asyncio(self): resp = Response('http://example.com/index.html') From 50aa6ef22cc9dec3655c9e3aaee225f159e94df1 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Sat, 21 Dec 2019 14:36:11 +0500 Subject: [PATCH 373/387] Add deferred_from_coro. --- scrapy/utils/defer.py | 26 ++++++++++++++++++++++++++ 1 file changed, 26 insertions(+) diff --git a/scrapy/utils/defer.py b/scrapy/utils/defer.py index 20ce59297..bbd5ebe52 100644 --- a/scrapy/utils/defer.py +++ b/scrapy/utils/defer.py @@ -1,10 +1,14 @@ """ Helper functions for dealing with Twisted deferreds """ +import asyncio +import inspect + from twisted.internet import defer, task from twisted.python import failure from scrapy.exceptions import IgnoreRequest +from scrapy.utils.asyncio import is_asyncio_reactor_installed def defer_fail(_failure): @@ -114,3 +118,25 @@ def iter_errback(iterable, errback, *a, **kw): break except Exception: errback(failure.Failure(), *a, **kw) + + +def _isfuture(o): + # workaround for Python before 3.5.3 not having asyncio.isfuture + if hasattr(asyncio, 'isfuture'): + return asyncio.isfuture(o) + return isinstance(o, asyncio.Future) + + +def deferred_from_coro(o): + """Converts a coroutine into a Deferred, or returns the object as is if it isn't a coroutine""" + if isinstance(o, defer.Deferred): + return o + if _isfuture(o) or inspect.isawaitable(o): + if not is_asyncio_reactor_installed(): + # wrapping the coroutine directly into a Deferred, this doesn't work correctly with coroutines + # that use asyncio, e.g. "await asyncio.sleep(1)" + return defer.ensureDeferred(o) + else: + # wrapping the coroutine into a Future and then into a Deferred, this requires AsyncioSelectorReactor + return defer.Deferred.fromFuture(asyncio.ensure_future(o)) + return o From 16787f5bf4475fe1604c1e4cac7b491f1df6a1fb Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Mon, 30 Dec 2019 12:02:19 +0500 Subject: [PATCH 374/387] Merge middleware tests back as we don't need to set the setting anymore. --- tests/test_downloadermiddleware.py | 7 +------ 1 file changed, 1 insertion(+), 6 deletions(-) diff --git a/tests/test_downloadermiddleware.py b/tests/test_downloadermiddleware.py index c5c4d13bd..3943cecf7 100644 --- a/tests/test_downloadermiddleware.py +++ b/tests/test_downloadermiddleware.py @@ -233,12 +233,7 @@ class MiddlewareUsingCoro(ManagerTestCase): self.assertIs(results[0], resp) self.assertFalse(download_func.called) - -@mark.only_asyncio() -class MiddlewareUsingCoroAsyncio(ManagerTestCase): - - settings_dict = {'ASYNCIO_ENABLED': True} - + @mark.only_asyncio() def test_asyncdef_asyncio(self): resp = Response('http://example.com/index.html') From e3b8ba6188ce703e43563d8588b7d709b037a0e0 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 31 Dec 2019 17:54:01 +0500 Subject: [PATCH 375/387] Run py35-asyncio also on 3.5.2 to test Xenial. --- .travis.yml | 2 ++ 1 file changed, 2 insertions(+) diff --git a/.travis.yml b/.travis.yml index 82167e10a..c808b3436 100644 --- a/.travis.yml +++ b/.travis.yml @@ -18,6 +18,8 @@ matrix: python: 3.5 - env: TOXENV=py35-asyncio python: 3.5 + - env: TOXENV=py35-asyncio + python: 3.5.2 - env: TOXENV=py36 python: 3.6 - env: TOXENV=py37 From 2b9254c2bde76995e81c54ad112ae884a1499386 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 31 Dec 2019 17:54:41 +0500 Subject: [PATCH 376/387] Add a test function that uses asyncio.Queue(). --- scrapy/utils/test.py | 9 ++++++++- tests/test_downloadermiddleware.py | 5 +++-- 2 files changed, 11 insertions(+), 3 deletions(-) diff --git a/scrapy/utils/test.py b/scrapy/utils/test.py index 307c25352..0f4cf8091 100644 --- a/scrapy/utils/test.py +++ b/scrapy/utils/test.py @@ -1,7 +1,7 @@ """ This module contains some assorted functions used in tests """ - +import asyncio import os from importlib import import_module @@ -96,3 +96,10 @@ def assert_samelines(testcase, text1, text2, msg=None): line endings between platforms """ testcase.assertEqual(text1.splitlines(), text2.splitlines(), msg) + + +def get_from_asyncio_queue(value): + q = asyncio.Queue() + getter = q.get() + q.put_nowait(value) + return getter diff --git a/tests/test_downloadermiddleware.py b/tests/test_downloadermiddleware.py index 3943cecf7..3dd4f2351 100644 --- a/tests/test_downloadermiddleware.py +++ b/tests/test_downloadermiddleware.py @@ -11,7 +11,7 @@ from scrapy.http import Request, Response from scrapy.spiders import Spider from scrapy.exceptions import _InvalidOutput from scrapy.core.downloader.middleware import DownloaderMiddlewareManager -from scrapy.utils.test import get_crawler +from scrapy.utils.test import get_crawler, get_from_asyncio_queue from scrapy.utils.python import to_bytes @@ -240,7 +240,8 @@ class MiddlewareUsingCoro(ManagerTestCase): class CoroMiddleware: async def process_request(self, request, spider): await asyncio.sleep(0.1) - return resp + result = await get_from_asyncio_queue(resp) + return result self.mwman._add_middleware(CoroMiddleware()) req = Request('http://example.com/index.html') From b2dd379bc2b93b4863dd36a296780481d758c4cb Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Fri, 3 Jan 2020 21:38:05 +0500 Subject: [PATCH 377/387] Remove the py35-asyncio env for 3.5 from Travis. --- .travis.yml | 2 -- 1 file changed, 2 deletions(-) diff --git a/.travis.yml b/.travis.yml index c808b3436..66e1a9617 100644 --- a/.travis.yml +++ b/.travis.yml @@ -16,8 +16,6 @@ matrix: python: 3.5 - env: TOXENV=pinned python: 3.5 - - env: TOXENV=py35-asyncio - python: 3.5 - env: TOXENV=py35-asyncio python: 3.5.2 - env: TOXENV=py36 From 81175669746737f1ab0bd0236a70f787ed1855a4 Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Mon, 16 Dec 2019 22:12:27 +0500 Subject: [PATCH 378/387] Add utils.defer.deferred_f_from_coro_f. --- scrapy/utils/defer.py | 13 +++++++++++++ 1 file changed, 13 insertions(+) diff --git a/scrapy/utils/defer.py b/scrapy/utils/defer.py index bbd5ebe52..62b43a96c 100644 --- a/scrapy/utils/defer.py +++ b/scrapy/utils/defer.py @@ -2,6 +2,7 @@ Helper functions for dealing with Twisted deferreds """ import asyncio +from functools import wraps import inspect from twisted.internet import defer, task @@ -140,3 +141,15 @@ def deferred_from_coro(o): # wrapping the coroutine into a Future and then into a Deferred, this requires AsyncioSelectorReactor return defer.Deferred.fromFuture(asyncio.ensure_future(o)) return o + + +def deferred_f_from_coro_f(coro_f): + """ Converts a coroutine function into a function that returns a Deferred. + + The coroutine function will be called at the time when the wrapper is called. Wrapper args will be passed to it. + This is useful for callback chains, as callback functions are called with the previous callback result. + """ + @wraps(coro_f) + def f(*coro_args, **coro_kwargs): + return deferred_from_coro(coro_f(*coro_args, **coro_kwargs)) + return f From 1f9cef787d3ca0c12099f1b1b4c52efc510e381d Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 10 Sep 2019 14:26:21 +0500 Subject: [PATCH 379/387] Add async def support to pipelines. --- scrapy/pipelines/__init__.py | 4 +++- tests/test_pipelines.py | 13 +++++++++++++ 2 files changed, 16 insertions(+), 1 deletion(-) diff --git a/scrapy/pipelines/__init__.py b/scrapy/pipelines/__init__.py index aa1bfb77f..1a45e00a2 100644 --- a/scrapy/pipelines/__init__.py +++ b/scrapy/pipelines/__init__.py @@ -6,6 +6,8 @@ See documentation in docs/item-pipeline.rst from scrapy.middleware import MiddlewareManager from scrapy.utils.conf import build_component_list +from scrapy.utils.defer import deferred_f_from_coro_f + class ItemPipelineManager(MiddlewareManager): @@ -19,7 +21,7 @@ class ItemPipelineManager(MiddlewareManager): def _add_middleware(self, pipe): super(ItemPipelineManager, self)._add_middleware(pipe) if hasattr(pipe, 'process_item'): - self.methods['process_item'].append(pipe.process_item) + self.methods['process_item'].append(deferred_f_from_coro_f(pipe.process_item)) def process_item(self, item, spider): return self._process_chain('process_item', item, spider) diff --git a/tests/test_pipelines.py b/tests/test_pipelines.py index bc53f5427..cfe4471d7 100644 --- a/tests/test_pipelines.py +++ b/tests/test_pipelines.py @@ -26,6 +26,13 @@ class DeferredPipeline: return d +class AsyncDefPipeline: + async def process_item(self, item, spider): + await defer.succeed(42) + item['pipeline_passed'] = True + return item + + class ItemSpider(Spider): name = 'itemspider' @@ -69,3 +76,9 @@ class PipelineTestCase(unittest.TestCase): crawler = self._create_crawler(DeferredPipeline) yield crawler.crawl(mockserver=self.mockserver) self.assertEqual(len(self.items), 1) + + @defer.inlineCallbacks + def test_asyncdef_pipeline(self): + crawler = self._create_crawler(AsyncDefPipeline) + yield crawler.crawl(mockserver=self.mockserver) + self.assertEqual(len(self.items), 1) From bfdd552a32b3a67dfa2895b483c187d87b63b50e Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Tue, 10 Sep 2019 14:57:07 +0500 Subject: [PATCH 380/387] Add a test for pipelines using asyncio. --- tests/test_pipelines.py | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/tests/test_pipelines.py b/tests/test_pipelines.py index cfe4471d7..aba6d85b7 100644 --- a/tests/test_pipelines.py +++ b/tests/test_pipelines.py @@ -1,3 +1,5 @@ +import asyncio + from twisted.internet import defer from twisted.internet.defer import Deferred from twisted.trial import unittest @@ -33,6 +35,13 @@ class AsyncDefPipeline: return item +class AsyncDefAsyncioPipeline: + async def process_item(self, item, spider): + await asyncio.sleep(0.2) + item['pipeline_passed'] = True + return item + + class ItemSpider(Spider): name = 'itemspider' @@ -82,3 +91,9 @@ class PipelineTestCase(unittest.TestCase): crawler = self._create_crawler(AsyncDefPipeline) yield crawler.crawl(mockserver=self.mockserver) self.assertEqual(len(self.items), 1) + + @defer.inlineCallbacks + def test_asyncdef_asyncio_pipeline(self): + crawler = self._create_crawler(AsyncDefAsyncioPipeline) + yield crawler.crawl(mockserver=self.mockserver) + self.assertEqual(len(self.items), 1) From bdef948aaebced3d28ea7b81525ea9edf7cc650c Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Mon, 16 Dec 2019 23:15:43 +0500 Subject: [PATCH 381/387] Mark the asyncio pipelines test as only_asyncio. --- tests/test_pipelines.py | 2 ++ 1 file changed, 2 insertions(+) diff --git a/tests/test_pipelines.py b/tests/test_pipelines.py index aba6d85b7..6f33282f2 100644 --- a/tests/test_pipelines.py +++ b/tests/test_pipelines.py @@ -1,5 +1,6 @@ import asyncio +from pytest import mark from twisted.internet import defer from twisted.internet.defer import Deferred from twisted.trial import unittest @@ -92,6 +93,7 @@ class PipelineTestCase(unittest.TestCase): yield crawler.crawl(mockserver=self.mockserver) self.assertEqual(len(self.items), 1) + @mark.only_asyncio() @defer.inlineCallbacks def test_asyncdef_asyncio_pipeline(self): crawler = self._create_crawler(AsyncDefAsyncioPipeline) From 9d8c54c0f2a98dc627efe8a2b58cfd311f2105cb Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Mon, 16 Dec 2019 22:43:55 +0500 Subject: [PATCH 382/387] Fix/ignore flake8 problems. --- pytest.ini | 1 + scrapy/pipelines/__init__.py | 1 - 2 files changed, 1 insertion(+), 1 deletion(-) diff --git a/pytest.ini b/pytest.ini index c3f3292bb..1030d5530 100644 --- a/pytest.ini +++ b/pytest.ini @@ -95,6 +95,7 @@ flake8-ignore = scrapy/loader/__init__.py E501 E128 scrapy/loader/processors.py E501 # scrapy/pipelines + scrapy/pipelines/__init__.py E501 scrapy/pipelines/files.py E116 E501 E266 scrapy/pipelines/images.py E265 E501 scrapy/pipelines/media.py E125 E501 E266 diff --git a/scrapy/pipelines/__init__.py b/scrapy/pipelines/__init__.py index 1a45e00a2..b5725a8ee 100644 --- a/scrapy/pipelines/__init__.py +++ b/scrapy/pipelines/__init__.py @@ -9,7 +9,6 @@ from scrapy.utils.conf import build_component_list from scrapy.utils.defer import deferred_f_from_coro_f - class ItemPipelineManager(MiddlewareManager): component_name = 'item pipeline' From 7d859848800ef05760e4f57dc206a4d756460a9d Mon Sep 17 00:00:00 2001 From: Andrey Rakhmatullin <wrar@wrar.name> Date: Thu, 9 Jan 2020 14:48:07 +0500 Subject: [PATCH 383/387] Use get_from_asyncio_queue in the pipeline test. --- tests/test_pipelines.py | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/tests/test_pipelines.py b/tests/test_pipelines.py index 6f33282f2..c72f1a338 100644 --- a/tests/test_pipelines.py +++ b/tests/test_pipelines.py @@ -6,7 +6,7 @@ from twisted.internet.defer import Deferred from twisted.trial import unittest from scrapy import Spider, signals, Request -from scrapy.utils.test import get_crawler +from scrapy.utils.test import get_crawler, get_from_asyncio_queue from tests.mockserver import MockServer @@ -39,7 +39,7 @@ class AsyncDefPipeline: class AsyncDefAsyncioPipeline: async def process_item(self, item, spider): await asyncio.sleep(0.2) - item['pipeline_passed'] = True + item['pipeline_passed'] = await get_from_asyncio_queue(True) return item From eaa8ed02d05d0cc5e7be1f773a22755676c7fe56 Mon Sep 17 00:00:00 2001 From: Juan Pablo Balarini <jpbalarini@gmail.com> Date: Wed, 26 Dec 2018 13:11:23 -0300 Subject: [PATCH 384/387] Add ability to change max_active_size by settings --- docs/topics/settings.rst | 11 +++++++++++ scrapy/core/scraper.py | 2 +- scrapy/settings/default_settings.py | 2 ++ 3 files changed, 14 insertions(+), 1 deletion(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index c02f877fc..d4480c879 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -1262,6 +1262,17 @@ Type of priority queue used by the scheduler. Another available type is domains in parallel. But currently ``scrapy.pqueues.DownloaderAwarePriorityQueue`` does not work together with :setting:`CONCURRENT_REQUESTS_PER_IP`. +.. setting:: SCRAPER_SLOT_MAX_ACTIVE_SIZE + +SCRAPER_SLOT_MAX_ACTIVE_SIZE +---------------------------- +Default: ``5000000`` + +Soft limit (in bytes) for response data being processed. + +While the sum of the sizes of all responses being processed is above this value, +Scrapy does not process new requests. + .. setting:: SPIDER_CONTRACTS SPIDER_CONTRACTS diff --git a/scrapy/core/scraper.py b/scrapy/core/scraper.py index 99114d3bb..facbd8b73 100644 --- a/scrapy/core/scraper.py +++ b/scrapy/core/scraper.py @@ -78,7 +78,7 @@ class Scraper(object): @defer.inlineCallbacks def open_spider(self, spider): """Open the given spider for scraping and allocate resources for it""" - self.slot = Slot() + self.slot = Slot(self.crawler.settings.getint('SCRAPER_SLOT_MAX_ACTIVE_SIZE')) yield self.itemproc.open_spider(spider) def close_spider(self, spider): diff --git a/scrapy/settings/default_settings.py b/scrapy/settings/default_settings.py index d03fd37b0..ddd48d327 100644 --- a/scrapy/settings/default_settings.py +++ b/scrapy/settings/default_settings.py @@ -254,6 +254,8 @@ SCHEDULER_DISK_QUEUE = 'scrapy.squeues.PickleLifoDiskQueue' SCHEDULER_MEMORY_QUEUE = 'scrapy.squeues.LifoMemoryQueue' SCHEDULER_PRIORITY_QUEUE = 'scrapy.pqueues.ScrapyPriorityQueue' +SCRAPER_SLOT_MAX_ACTIVE_SIZE = 5000000 + SPIDER_LOADER_CLASS = 'scrapy.spiderloader.SpiderLoader' SPIDER_LOADER_WARN_ONLY = False From 0f2d871d88b6741aaf0348d1f11a1e020124099f Mon Sep 17 00:00:00 2001 From: JP Balarini <jpbalarini@gmail.com> Date: Wed, 22 May 2019 16:26:27 -0300 Subject: [PATCH 385/387] Use PEP 515 style for SCRAPER_SLOT_MAX_ACTIVE_SIZE documentation --- docs/topics/settings.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/topics/settings.rst b/docs/topics/settings.rst index d4480c879..f4c3494f5 100644 --- a/docs/topics/settings.rst +++ b/docs/topics/settings.rst @@ -1266,7 +1266,7 @@ does not work together with :setting:`CONCURRENT_REQUESTS_PER_IP`. SCRAPER_SLOT_MAX_ACTIVE_SIZE ---------------------------- -Default: ``5000000`` +Default: ``5_000_000`` Soft limit (in bytes) for response data being processed. From c75cf15b7a8293fd83f6e8b27abca2e50b0a04ed Mon Sep 17 00:00:00 2001 From: Eugenio Lacuesta <eugenio.lacuesta@gmail.com> Date: Wed, 22 Jan 2020 10:38:59 -0300 Subject: [PATCH 386/387] Update CSS selectors in tutorial --- docs/intro/tutorial.rst | 15 ++++++++++----- 1 file changed, 10 insertions(+), 5 deletions(-) diff --git a/docs/intro/tutorial.rst b/docs/intro/tutorial.rst index c4d8d717b..c9d00eb74 100644 --- a/docs/intro/tutorial.rst +++ b/docs/intro/tutorial.rst @@ -616,19 +616,24 @@ instance; you still have to yield this Request. You can also pass a selector to ``response.follow`` instead of a string; this selector should extract necessary attributes:: - href = response.css('li.next a::attr(href)')[0] - yield response.follow(href, callback=self.parse) + for href in response.css('ul.pager a::attr(href)'): + yield response.follow(href, callback=self.parse) For ``<a>`` elements there is a shortcut: ``response.follow`` uses their href attribute automatically. So the code can be shortened further:: - a = response.css('li.next a')[0] - yield response.follow(a, callback=self.parse) + for a in response.css('ul.pager a'): + yield response.follow(a, callback=self.parse) To create multiple requests from an iterable, you can use :meth:`response.follow_all <scrapy.http.TextResponse.follow_all>` instead:: - yield from response.follow_all(response.css('a'), callback=self.parse) + anchors = response.css('ul.pager a') + yield from response.follow_all(anchors, callback=self.parse) + +or, shortening it further:: + + yield from response.follow_all(css='ul.pager a', callback=self.parse) More examples and patterns From 7d5cebcf773d54d2f7b5ac0b90886847a4a70c26 Mon Sep 17 00:00:00 2001 From: Peter Vandenabeele <peter@vandenabeele.com> Date: Thu, 23 Jan 2020 09:08:21 +0100 Subject: [PATCH 387/387] fix logical documentation error with PER_DOMAIN or PER_DOMAIN --- docs/faq.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/faq.rst b/docs/faq.rst index aae2411e0..169e1a479 100644 --- a/docs/faq.rst +++ b/docs/faq.rst @@ -140,7 +140,7 @@ setting the following settings:: While pending requests are below the configured values of :setting:`CONCURRENT_REQUESTS`, :setting:`CONCURRENT_REQUESTS_PER_DOMAIN` or -:setting:`CONCURRENT_REQUESTS_PER_DOMAIN`, those requests are sent +:setting:`CONCURRENT_REQUESTS_PER_IP`, those requests are sent concurrently. As a result, the first few requests of a crawl rarely follow the desired order. Lowering those settings to ``1`` enforces the desired order, but it significantly slows down the crawl as a whole.