Merge branch 'master' of https://github.com/scrapy/scrapy into scrapy-4996

This commit is contained in:
deepang17 2021-03-09 17:33:05 +05:30
commit 2d41c77e6a
2 changed files with 37 additions and 0 deletions

View File

@ -145,6 +145,41 @@ How can I make Scrapy consume less memory?
See previous question.
How can I prevent memory errors due to many allowed domains?
------------------------------------------------------------
If you have a spider with a long list of
:attr:`~scrapy.spiders.Spider.allowed_domains` (e.g. 50,000+), consider
replacing the default
:class:`~scrapy.spidermiddlewares.offsite.OffsiteMiddleware` spider middleware
with a :ref:`custom spider middleware <custom-spider-middleware>` that requires
less memory. For example:
- If your domain names are similar enough, use your own regular expression
instead joining the strings in
:attr:`~scrapy.spiders.Spider.allowed_domains` into a complex regular
expression.
- If you can `meet the installation requirements`_, use pyre2_ instead of
Pythons re_ to compile your URL-filtering regular expression. See
:issue:`1908`.
See also other suggestions at `StackOverflow`_.
.. note:: Remember to disable
:class:`scrapy.spidermiddlewares.offsite.OffsiteMiddleware` when you enable
your custom implementation::
SPIDER_MIDDLEWARES = {
'scrapy.spidermiddlewares.offsite.OffsiteMiddleware': None,
'myproject.middlewares.CustomOffsiteMiddleware': 500,
}
.. _meet the installation requirements: https://github.com/andreasvc/pyre2#installation
.. _pyre2: https://github.com/andreasvc/pyre2
.. _re: https://docs.python.org/library/re.html
.. _StackOverflow: https://stackoverflow.com/q/36440681/939364
Can I use Basic HTTP Authentication in my spiders?
--------------------------------------------------

View File

@ -18,6 +18,8 @@ deps =
# Extras
botocore>=1.4.87
Pillow>=4.0.0
# Twisted 21+ causes issues in tests that use skipIf
Twisted<21
passenv =
S3_TEST_FILE_URI
AWS_ACCESS_KEY_ID