Merge remote-tracking branch 'origin/master' into feed-mode

This commit is contained in:
Adrian Chaves 2026-08-10 10:15:38 +02:00
commit 0078f308f0
88 changed files with 1690 additions and 888 deletions

View File

@ -86,6 +86,7 @@ jobs:
- python-version: pypy3.11-7.3.20
env:
TOXENV: pypy3-extra-deps
coverage: true
- python-version: "3.14"
env:
TOXENV: botocore

View File

@ -16,8 +16,11 @@ jobs:
tests:
name: tests
runs-on: ubuntu-latest
timeout-minutes: 30
env:
PYTEST_ADDOPTS: -n auto --no-cov
# A development branch of a dependency can make a test hang forever, so
# tests get a time limit here that they do not need elsewhere.
PYTEST_ADDOPTS: -n auto --no-cov --timeout=120
TOXENV: vcs-deps
UV_PYTHON_PREFERENCE: only-system
steps:

View File

@ -31,9 +31,14 @@ extensions = [
"sphinx_scrapy",
"scrapyfixautodoc", # Must be after "sphinx.ext.autodoc"
"sphinx.ext.coverage",
"sphinx_reredirects",
"sphinx_rtd_dark_mode",
]
redirects = {
"topics/broad-crawls": "optimize.html#broad-crawls",
}
templates_path = ["_templates"]
exclude_patterns = ["build", "Thumbs.db", ".DS_Store"]

View File

@ -292,7 +292,7 @@ Does Scrapy manage cookies automatically?
Yes, Scrapy receives and keeps track of cookies sent by servers, and sends them
back on subsequent requests, like any regular web browser does.
For more info see :ref:`topics-request-response` and :ref:`cookies-mw`.
For more info see :ref:`cookies`.
How can I see the cookies being sent and received from Scrapy?
--------------------------------------------------------------

View File

@ -78,6 +78,7 @@ Basic concepts
topics/item-pipeline
topics/feed-exports
topics/request-response
topics/cookies
topics/link-extractors
topics/settings
topics/exceptions
@ -109,6 +110,9 @@ Basic concepts
:doc:`topics/request-response`
Understand the classes used to represent HTTP requests and responses.
:doc:`topics/cookies`
Send and receive cookies.
:doc:`topics/link-extractors`
Convenient classes to extract links to follow from pages.
@ -152,7 +156,7 @@ Solving specific problems
topics/contracts
topics/practices
topics/security
topics/broad-crawls
topics/optimize
topics/developer-tools
topics/dynamic-content
topics/leaks
@ -180,8 +184,8 @@ Solving specific problems
Understand the security implications of Scrapy defaults and how to harden
them.
:doc:`topics/broad-crawls`
Tune Scrapy for crawling a lot domains in parallel.
:doc:`topics/optimize`
Find the bottleneck of your crawls and learn how to address it.
:doc:`topics/developer-tools`
Learn how to scrape with your browser's developer tools.

View File

@ -111,8 +111,6 @@ The following extras are available:
- Provides
* - ``bpython``
- :ref:`bpython shell <shell-config>`
* - ``brotli``
- :ref:`Brotli response decompression <http-compression>`
* - ``gcs``
- :ref:`Google Cloud Storage <topics-feed-storage-gcs>` for
:ref:`feed exports <topics-feed-exports>` and

View File

@ -3,6 +3,7 @@ pydantic
scrapy-spider-metadata
sphinx
sphinx-notfound-page
sphinx-reredirects
sphinx-rtd-theme
sphinx-rtd-dark-mode
sphinx-scrapy @ git+https://github.com/scrapy/sphinx-scrapy.git@0.8.10

View File

@ -134,6 +134,7 @@ sphinx==9.1.0
# sphinx-llms-txt
# sphinx-markdown-builder
# sphinx-notfound-page
# sphinx-reredirects
# sphinx-rtd-theme
# sphinx-scrapy
# sphinxcontrib-jquery
@ -147,6 +148,8 @@ sphinx-markdown-builder @ git+https://github.com/zytedata/sphinx-markdown-builde
# via sphinx-scrapy
sphinx-notfound-page==1.1.0
# via -r docs/requirements.in
sphinx-reredirects==1.1.0
# via -r docs/requirements.in
sphinx-rtd-dark-mode==1.3.0
# via -r docs/requirements.in
sphinx-rtd-theme==3.1.0

View File

@ -35,6 +35,13 @@ how you :ref:`configure the downloader middlewares
:class:`scrapy.Spider` subclass and a
:class:`scrapy.settings.Settings` object.
The :attr:`engine`, :attr:`extensions`, :attr:`logformatter`,
:attr:`request_fingerprinter` and :attr:`stats` attributes get their value
when the crawl starts, and raise :exc:`RuntimeError` when read before that.
.. versionchanged:: VERSION
Those attributes used to be ``None`` before getting their value.
.. attribute:: request_fingerprinter
The request fingerprint builder of this crawler.

View File

@ -1,195 +0,0 @@
.. _topics-broad-crawls:
============
Broad Crawls
============
Scrapy defaults are optimized for crawling specific sites. These sites are
often handled by a single Scrapy spider, although this is not necessary or
required (for example, there are generic spiders that handle any given site
thrown at them).
In addition to this "focused crawl", there is another common type of crawling
which covers a large (potentially unlimited) number of domains, and is only
limited by time or other arbitrary constraint, rather than stopping when the
domain was crawled to completion or when there are no more requests to perform.
These are called "broad crawls" and is the typical crawlers employed by search
engines.
These are some common properties often found in broad crawls:
* they crawl many domains (often, unbounded) instead of a specific set of sites
* they don't necessarily crawl domains to completion, because it would be
impractical (or impossible) to do so, and instead limit the crawl by time or
number of pages crawled
* they are simpler in logic (as opposed to very complex spiders with many
extraction rules) because data is often post-processed in a separate stage
* they crawl many domains concurrently, which allows them to achieve faster
crawl speeds by not being limited by any particular site constraint (each site
is crawled slowly to respect politeness, but many sites are crawled in
parallel)
As said above, Scrapy default settings are optimized for focused crawls, not
broad crawls. However, due to its asynchronous architecture, Scrapy is very
well suited for performing fast broad crawls. This page summarizes some things
you need to keep in mind when using Scrapy for doing broad crawls, along with
concrete suggestions of Scrapy settings to tune in order to achieve an
efficient broad crawl.
.. _broad-crawls-scheduler-priority-queue:
.. _broad-crawls-concurrency:
Increase concurrency
====================
Concurrency is the number of requests that are processed in parallel. There is
a global limit (:setting:`CONCURRENT_REQUESTS`) and an additional limit that
can be set per domain (:setting:`CONCURRENT_REQUESTS_PER_DOMAIN`).
The default global concurrency limit in Scrapy is not suitable for crawling
many different domains in parallel, so you will want to increase it. How much
to increase it will depend on how much CPU and memory your crawler will have
available.
A good starting point is ``100``:
.. code-block:: python
CONCURRENT_REQUESTS = 100
But the best way to find out is by doing some trials and identifying at what
concurrency your Scrapy process gets CPU bounded. For optimum performance, you
should pick a concurrency where CPU usage is at 80-90%.
Increasing concurrency also increases memory usage. If memory usage is a
concern, you might need to lower your global concurrency limit accordingly.
Increase Twisted IO thread pool maximum size
============================================
Currently Scrapy does DNS resolution in a blocking way with usage of thread
pool. With higher concurrency levels the crawling could be slow or even fail
hitting DNS resolver timeouts. Possible solution to increase the number of
threads handling DNS queries. The DNS queue will be processed faster speeding
up establishing of connection and crawling overall.
To increase maximum thread pool size use:
.. code-block:: python
REACTOR_THREADPOOL_MAXSIZE = 20
Setup your own DNS
==================
If you have multiple crawling processes and single central DNS, it can act
like DoS attack on the DNS server resulting to slow down of entire network or
even blocking your machines. To avoid this setup your own DNS server with
local cache and upstream to some large DNS like OpenDNS or Verizon.
Reduce log level
================
When doing broad crawls you are often only interested in the crawl rates you
get and any errors found. These stats are reported by Scrapy when using the
``INFO`` log level. In order to save CPU (and log storage requirements) you
should not use ``DEBUG`` log level when performing large broad crawls in
production. Using ``DEBUG`` level when developing your (broad) crawler may be
fine though.
To set the log level use:
.. code-block:: python
LOG_LEVEL = "INFO"
Disable cookies
===============
Disable cookies unless you *really* need. Cookies are often not needed when
doing broad crawls (search engine crawlers ignore them), and they improve
performance by saving some CPU cycles and reducing the memory footprint of your
Scrapy crawler.
To disable cookies use:
.. code-block:: python
COOKIES_ENABLED = False
Disable retries
===============
Retrying failed HTTP requests can slow down the crawls substantially, especially
when sites causes are very slow (or fail) to respond, thus causing a timeout
error which gets retried many times, unnecessarily, preventing crawler capacity
to be reused for other domains.
To disable retries use:
.. code-block:: python
RETRY_ENABLED = False
Reduce download timeout
=======================
Unless you are crawling from a very slow connection (which shouldn't be the
case for broad crawls) reduce the download timeout so that stuck requests are
discarded quickly and free up capacity to process the next ones.
To reduce the download timeout use:
.. code-block:: python
DOWNLOAD_TIMEOUT = 15
Disable redirects
=================
Consider disabling redirects, unless you are interested in following them. When
doing broad crawls it's common to save redirects and resolve them when
revisiting the site at a later crawl. This also help to keep the number of
request constant per crawl batch, otherwise redirect loops may cause the
crawler to dedicate too many resources on any specific domain.
To disable redirects use:
.. code-block:: python
REDIRECT_ENABLED = False
.. _broad-crawls-bfo:
Crawl in BFO order
==================
:ref:`Scrapy crawls in DFO order by default <faq-bfo-dfo>`.
In broad crawls, however, page crawling tends to be faster than page
processing. As a result, unprocessed early requests stay in memory until the
final depth is reached, which can significantly increase memory usage.
:ref:`Crawl in BFO order <faq-bfo-dfo>` instead to save memory.
Be mindful of memory leaks
==========================
If your broad crawl shows a high memory usage, in addition to :ref:`crawling in
BFO order <broad-crawls-bfo>`, :ref:`lowering concurrency
<broad-crawls-concurrency>` and :ref:`delaying start request iteration
<start-requests-lazy>` you should :ref:`debug your memory leaks
<topics-leaks>`.
Install a specific Twisted reactor
==================================
If the crawl is exceeding the system's capabilities, you might want to try
installing a specific Twisted reactor, via the :setting:`TWISTED_REACTOR` setting.

138
docs/topics/cookies.rst Normal file
View File

@ -0,0 +1,138 @@
.. _cookies:
.. _cookies-mw:
=======
Cookies
=======
Scrapy keeps track of the cookies that websites set and sends them back on
later requests to those websites, just like a web browser does. That is the job
of :class:`~scrapy.downloadermiddlewares.cookies.CookiesMiddleware`, which is
enabled by default.
Setting cookies on a request
============================
.. invisible-code-block: python
from scrapy import Request
Use the ``cookies`` parameter of :class:`~scrapy.Request` to send cookies of
your own, either as a dict:
.. code-block:: python
request = Request(
url="https://example.com",
cookies={"currency": "USD", "country": "UY"},
)
Or as a list of dicts, which also lets you set cookie attributes:
.. code-block:: python
request = Request(
url="https://example.com",
cookies=[
{
"name": "currency",
"value": "USD",
"domain": "example.com",
"path": "/currency",
"secure": True,
},
],
)
Setting attributes is only useful if the cookies are stored for later requests,
i.e. if :reqmeta:`dont_merge_cookies` is not enabled.
.. caution:: Cookies set through the ``Cookie`` header are not handled by
:class:`~scrapy.downloadermiddlewares.cookies.CookiesMiddleware`, which
drops that header.
.. caution:: When a cookie name or value is a byte sequence that is not UTF-8
encoded, the cookie is dropped and a warning is logged. See
:ref:`topics-logging-advanced-customization` to customize the logging
behavior.
.. reqmeta:: cookiejar
Multiple cookie sessions per spider
===================================
By default all requests share a single cookie jar (session). To use different
ones, pass an identifier in the :reqmeta:`cookiejar` request meta key:
.. skip: next
.. code-block:: python
for i, url in enumerate(urls):
yield Request(url, meta={"cookiejar": i}, callback=self.parse_page)
The :reqmeta:`cookiejar` meta key is not "sticky", so you need to keep passing
it along on subsequent requests:
.. code-block:: python
def parse_page(self, response):
return Request(
"https://example.com/otherpage",
meta={"cookiejar": response.meta["cookiejar"]},
callback=self.parse_other_page,
)
.. reqmeta:: dont_merge_cookies
Skipping the cookie jar for a request
=====================================
Set the :reqmeta:`dont_merge_cookies` request meta key to ``True`` to keep a
request from touching the cookie jar in either direction: no stored cookie is
sent with the request, and no cookie received in the response is stored. The
cookies of the request itself are ignored as well.
.. setting:: COOKIES_ENABLED
COOKIES_ENABLED
===============
Default: ``True``
Whether to enable :class:`~scrapy.downloadermiddlewares.cookies.CookiesMiddleware`.
If disabled, no cookies are sent to web servers.
.. setting:: COOKIES_DEBUG
COOKIES_DEBUG
=============
Default: ``False``
If enabled, Scrapy logs all cookies sent in requests (i.e. the ``Cookie``
header) and all cookies received in responses (i.e. the ``Set-Cookie``
header)::
2011-04-06 14:35:10-0300 [scrapy.core.engine] INFO: Spider opened
2011-04-06 14:35:10-0300 [scrapy.downloadermiddlewares.cookies] DEBUG: Sending cookies to: <GET http://www.diningcity.com/netherlands/index.html>
Cookie: clientlanguage_nl=en_EN
2011-04-06 14:35:14-0300 [scrapy.downloadermiddlewares.cookies] DEBUG: Received cookies from: <200 http://www.diningcity.com/netherlands/index.html>
Set-Cookie: JSESSIONID=B~FA4DC0C496C8762AE4F1A620EAB34F38; Path=/
Set-Cookie: ip_isocode=US
Set-Cookie: clientlanguage_nl=en_EN; Expires=Thu, 07-Apr-2011 21:21:34 GMT; Path=/
2011-04-06 14:49:50-0300 [scrapy.core.engine] DEBUG: Crawled (200) <GET http://www.diningcity.com/netherlands/index.html> (referer: None)
[...]
CookiesMiddleware
=================
.. module:: scrapy.downloadermiddlewares.cookies
:synopsis: Cookies Downloader Middleware
.. autoclass:: CookiesMiddleware

View File

@ -156,6 +156,61 @@ defines one or more of these methods:
:param exception: the raised exception
:type exception: an ``Exception`` object
.. _mw-download:
Downloading a request from a downloader middleware
==================================================
A downloader middleware can download a request of its own while it processes
another one, e.g. to fetch something that the request it is processing needs.
The built-in :ref:`robots.txt middleware <topics-dlmw-robots>` does that: it
holds each request while it downloads the ``robots.txt`` file of its website.
Use :meth:`crawler.engine.download_async()
<scrapy.core.engine.ExecutionEngine.download_async>` for that:
.. code-block:: python
from scrapy import Request
from scrapy.http.request import NO_CALLBACK
class TokenMiddleware:
def __init__(self, crawler):
self.crawler = crawler
self.token = None
@classmethod
def from_crawler(cls, crawler):
return cls(crawler)
async def process_request(self, request):
if request.meta.get("dont_obey_robotstxt"):
return
if self.token is None:
response = await self.crawler.engine.download_async(
Request(
"https://example.com/token",
callback=NO_CALLBACK,
meta={"dont_obey_robotstxt": True},
)
)
self.token = response.text
request.headers["Authorization"] = self.token
Requests that you download this way go through the downloader middleware chain
as well, including your own middleware and the :ref:`robots.txt middleware
<topics-dlmw-robots>`, which holds a request until the ``robots.txt`` file of
its website arrives. Be careful not to introduce deadlocks: a request that you
download must not end up waiting for the request that is waiting for it. Hence
:reqmeta:`dont_obey_robotstxt` above, which makes both middlewares let the token
request through.
While the first token response is in transit, ``process_request`` runs for other
requests as well, and the middleware above downloads a token for each of them.
Cache the task that downloads the token, and not only its result, to download
the token only once.
.. _topics-downloader-middleware-ref:
Built-in downloader middleware reference
@ -169,106 +224,10 @@ middleware, see the :ref:`downloader middleware usage guide
For a list of the components enabled by default (and their orders) see the
:setting:`DOWNLOADER_MIDDLEWARES_BASE` setting.
.. _cookies-mw:
CookiesMiddleware
-----------------
.. module:: scrapy.downloadermiddlewares.cookies
:synopsis: Cookies Downloader Middleware
.. class:: CookiesMiddleware
This middleware enables working with sites that require cookies, such as
those that use sessions. It keeps track of cookies sent by web servers, and
sends them back on subsequent requests (from that spider), just like web
browsers do.
.. caution:: When non-UTF8 encoded byte sequences are passed to a
:class:`~scrapy.Request`, the ``CookiesMiddleware`` will log
a warning. Refer to :ref:`topics-logging-advanced-customization`
to customize the logging behaviour.
.. caution:: Cookies set via the ``Cookie`` header are not considered by the
:ref:`cookies-mw`. If you need to set cookies for a request, use the
:class:`Request.cookies <scrapy.Request>` parameter. This is a known
current limitation that is being worked on.
The following settings can be used to configure the cookie middleware:
* :setting:`COOKIES_ENABLED`
* :setting:`COOKIES_DEBUG`
.. reqmeta:: cookiejar
Multiple cookie sessions per spider
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
There is support for keeping multiple cookie sessions per spider by using the
:reqmeta:`cookiejar` Request meta key. By default it uses a single cookie jar
(session), but you can pass an identifier to use different ones.
For example:
.. skip: next
.. code-block:: python
for i, url in enumerate(urls):
yield scrapy.Request(url, meta={"cookiejar": i}, callback=self.parse_page)
Keep in mind that the :reqmeta:`cookiejar` meta key is not "sticky". You need to keep
passing it along on subsequent requests. For example:
.. code-block:: python
def parse_page(self, response):
# do some processing
return scrapy.Request(
"http://www.example.com/otherpage",
meta={"cookiejar": response.meta["cookiejar"]},
callback=self.parse_other_page,
)
.. setting:: COOKIES_ENABLED
COOKIES_ENABLED
~~~~~~~~~~~~~~~
Default: ``True``
Whether to enable the cookies middleware. If disabled, no cookies will be sent
to web servers.
Notice that despite the value of :setting:`COOKIES_ENABLED` setting if
``Request.``:reqmeta:`meta['dont_merge_cookies'] <dont_merge_cookies>`
evaluates to ``True`` the request cookies will **not** be sent to the
web server and received cookies in :class:`~scrapy.http.Response` will
**not** be merged with the existing cookies.
For more detailed information see the ``cookies`` parameter in
:class:`~scrapy.Request`.
.. setting:: COOKIES_DEBUG
COOKIES_DEBUG
~~~~~~~~~~~~~
Default: ``False``
If enabled, Scrapy will log all cookies sent in requests (i.e. ``Cookie``
header) and all cookies received in responses (i.e. ``Set-Cookie`` header).
Here's an example of a log with :setting:`COOKIES_DEBUG` enabled::
2011-04-06 14:35:10-0300 [scrapy.core.engine] INFO: Spider opened
2011-04-06 14:35:10-0300 [scrapy.downloadermiddlewares.cookies] DEBUG: Sending cookies to: <GET http://www.diningcity.com/netherlands/index.html>
Cookie: clientlanguage_nl=en_EN
2011-04-06 14:35:14-0300 [scrapy.downloadermiddlewares.cookies] DEBUG: Received cookies from: <200 http://www.diningcity.com/netherlands/index.html>
Set-Cookie: JSESSIONID=B~FA4DC0C496C8762AE4F1A620EAB34F38; Path=/
Set-Cookie: ip_isocode=US
Set-Cookie: clientlanguage_nl=en_EN; Expires=Thu, 07-Apr-2011 21:21:34 GMT; Path=/
2011-04-06 14:49:50-0300 [scrapy.core.engine] DEBUG: Crawled (200) <GET http://www.diningcity.com/netherlands/index.html> (referer: None)
[...]
See :ref:`cookies`.
DefaultHeadersMiddleware
@ -741,14 +700,13 @@ HttpCompressionMiddleware
.. class:: HttpCompressionMiddleware
This middleware allows compressed (gzip, deflate) traffic to be
This middleware allows compressed (gzip, deflate, `brotli`_) traffic to be
sent/received from web sites.
This middleware also supports decoding `brotli-compressed`_ responses with
the :ref:`brotli <extras>` extra, and `zstd-compressed`_
responses with the :ref:`zstd <extras>` extra.
This middleware also supports decoding `zstd-compressed`_ responses with
the :ref:`zstd <extras>` extra.
.. _brotli-compressed: https://www.ietf.org/rfc/rfc7932.txt
.. _brotli: https://www.ietf.org/rfc/rfc7932.txt
.. _zstd-compressed: https://www.ietf.org/rfc/rfc8478.txt
@ -841,40 +799,9 @@ OffsiteMiddleware
.. module:: scrapy.downloadermiddlewares.offsite
:synopsis: Offsite Middleware
.. class:: OffsiteMiddleware
.. autoclass:: OffsiteMiddleware
.. versionadded:: 2.11.2
Filters out Requests for URLs outside the domains covered by the spider.
This middleware filters out every request whose host names aren't in the
spider's :attr:`~scrapy.Spider.allowed_domains` attribute.
All subdomains of any domain in the list are also allowed.
E.g. the rule ``www.example.org`` will also allow ``bob.www.example.org``
but not ``www2.example.com`` nor ``example.com``.
When your spider returns a request for a domain not belonging to those
covered by the spider, this middleware will log a debug message similar to
this one::
DEBUG: Filtered offsite request to 'offsite.example': <GET http://offsite.example/some/page.html>
To avoid filling the log with too much noise, it will only print one of
these messages for each new domain filtered. So, for example, if another
request for ``offsite.example`` is filtered, no log message will be
printed. But if a request for ``other.example`` is filtered, a message
will be printed (but only for the first request filtered).
If the spider doesn't define an
:attr:`~scrapy.Spider.allowed_domains` attribute, or the
attribute is empty, the offsite middleware will allow all requests.
.. reqmeta:: allow_offsite
If the request has the :attr:`~scrapy.Request.dont_filter` attribute set to
``True`` or :attr:`Request.meta <scrapy.Request.meta>` has ``allow_offsite``
set to ``True``, then the OffsiteMiddleware will allow the request even if
its domain is not listed in allowed domains.
.. automethod:: should_follow
RedirectMiddleware
------------------

View File

@ -136,6 +136,70 @@ Example:
return f"$ {str(value)}"
return super().serialize_field(field, name, value)
.. _custom-exporters:
Writing your own item exporter
==============================
To write an item exporter, subclass :class:`BaseItemExporter` and implement
:meth:`~BaseItemExporter.export_item`, where
:meth:`~BaseItemExporter.get_serialized_fields` gives you the ``(name, value)``
pairs to export.
To make your exporter available to the :ref:`feed exports
<topics-feed-exports>`, list it in the :setting:`FEED_EXPORTERS` setting. Feed
exports :ref:`build <from-crawler>` it with the output file as the first
positional argument, and with the ``fields``, ``encoding`` and ``indent``
:ref:`feed options <feed-options>` and every key of ``item_export_kwargs`` as
keyword arguments, so your ``__init__`` method must forward unknown keyword
arguments to :class:`BaseItemExporter`.
The file object belongs to whoever opened it, i.e. to the feed storage in the
case of feed exports, which also closes it. If you need a text file, for
example to use :func:`csv.writer` or another Python API that does not accept a
binary file, wrap it with :class:`io.TextIOWrapper` and call
:meth:`~io.TextIOBase.detach` on the wrapper in
:meth:`~BaseItemExporter.finish_exporting`; otherwise the wrapper closes the
underlying file when it is garbage-collected.
For example, the following item exporter writes items as blocks of
``name: value`` lines:
.. code-block:: python
from io import TextIOWrapper
from scrapy.exporters import BaseItemExporter
class TextItemExporter(BaseItemExporter):
def __init__(self, file, item_separator="\n", **kwargs):
super().__init__(**kwargs)
self.item_separator = item_separator
self.stream = TextIOWrapper(
file, encoding=self.encoding or "utf-8", write_through=True
)
def export_item(self, item):
for name, value in self.get_serialized_fields(item):
print(f"{name}: {value}", file=self.stream)
self.stream.write(self.item_separator)
def finish_exporting(self):
self.stream.detach()
To use it as the ``txt`` feed format:
.. code-block:: python
FEED_EXPORTERS = {"txt": "myproject.exporters.TextItemExporter"}
FEEDS = {
"items.txt": {
"format": "txt",
"item_export_kwargs": {"item_separator": "---\n"},
},
}
.. _topics-exporters-reference:
Built-in Item Exporters reference
@ -168,6 +232,8 @@ BaseItemExporter
Exports the given item. This method must be implemented in subclasses.
.. automethod:: BaseItemExporter.get_serialized_fields
.. method:: serialize_field(field, name, value)
Return the serialized value for the given field. You can override this

View File

@ -178,6 +178,37 @@ By overriding ``file_path`` like this:
For more information about the ``file_path`` method, see :ref:`topics-media-pipeline-override`.
.. _file-naming-response:
Naming files after the response
-------------------------------
``file_path`` also receives the ``response``, which allows naming files after
response data. For example, to determine the file extension from the
``Content-Type`` header, for URLs that do not end in a file name:
.. code-block:: python
import mimetypes
from scrapy.pipelines.files import FilesPipeline
class ContentTypeFilesPipeline(FilesPipeline):
def file_path(self, request, response=None, info=None, *, item=None):
path = super().file_path(request, response, info, item=item)
if response is None:
return path
content_type = response.headers["Content-Type"].decode()
return path + (mimetypes.guess_extension(content_type) or "")
This requires setting :setting:`FILES_EXPIRES` to ``0``. To find out whether a
file has already been downloaded, Scrapy calls ``file_path`` before the
download, with ``response`` set to ``None``, and checks the age of the file at
the resulting path. A path that depends on the response can never match that
check, and :setting:`FILES_EXPIRES` set to ``0`` disables it, at the cost of
downloading every file on every run.
.. _topics-supported-storage:
Supported Storage
@ -543,7 +574,7 @@ See here the methods that you can override in your custom Files Pipeline:
return "files/" + PurePosixPath(urlparse_cached(request).path).name
Similarly, you can use the ``item`` to determine the file path based on some item
property.
property, or the ``response``, see :ref:`file-naming-response`.
By default the :meth:`file_path` method returns
``full/<request URL hash>.<extension>``.
@ -693,7 +724,7 @@ See here the methods that you can override in your custom Images Pipeline:
return "files/" + PurePosixPath(urlparse_cached(request).path).name
Similarly, you can use the ``item`` to determine the file path based on some item
property.
property, or the ``response``, see :ref:`file-naming-response`.
By default the :meth:`file_path` method returns
``full/<request URL hash>.<extension>``.

353
docs/topics/optimize.rst Normal file
View File

@ -0,0 +1,353 @@
.. _optimize:
============
Optimization
============
A crawl goes as fast as its slowest part allows. :ref:`Find out which part that
is <optimize-bottleneck>` before changing any setting.
:ref:`Broad crawls <broad-crawls>` have their own set of recommended
adjustments.
.. _optimize-bottleneck:
Finding the bottleneck
======================
The bottleneck depends on the spider: on the same machine, one crawl can be
limited by its own parsing code and another by the target website. So measure
the crawl that you want to optimize.
:class:`~scrapy.extensions.logstats.LogStats` reports crawl speed every
:setting:`LOGSTATS_INTERVAL` seconds:
.. code-block:: text
[scrapy.extensions.logstats] INFO: Crawled 1200 pages (at 60 pages/min), scraped 1150 items (at 58 items/min)
A rate that stays flat as you raise :setting:`CONCURRENT_REQUESTS` means
something else is the limit.
Reading the engine status
-------------------------
The :ref:`telnet console <topics-telnetconsole>` reports, through ``est()``,
what every part of the engine is doing at a given moment:
.. code-block:: text
len(engine.downloader.active) : 16
len(engine._slot.scheduler.mqs) : 92
len(engine.scraper.slot.active) : 0
engine.scraper.slot.active_size : 0
engine.scraper.slot.needs_backout() : False
Take a few readings at different points of the crawl:
- ``len(engine.downloader.active)`` stays at :setting:`CONCURRENT_REQUESTS`:
the downloader is the limit. You are waiting on the network or on the
target website. See :ref:`optimize-concurrency`.
- ``len(engine.downloader.active)`` stays below
:setting:`CONCURRENT_REQUESTS` while the scheduler queues (``mqs``,
``dqs``) hold requests: something throttles those requests before they
reach the downloader, usually :setting:`CONCURRENT_REQUESTS_PER_DOMAIN`,
:setting:`DOWNLOAD_DELAY` or :ref:`AutoThrottle <topics-autothrottle>`.
- Both the downloader and the scheduler queues stay near empty: your spider
is not producing requests fast enough. A crawl that walks pagination one
page at a time cannot use more concurrency than it creates. See
:ref:`optimize-requests`.
- ``needs_backout()`` is ``True``, or ``active_size`` approaches
:setting:`SCRAPER_SLOT_MAX_ACTIVE_SIZE`: responses arrive faster than your
callbacks and :ref:`item pipelines <topics-item-pipeline>` handle them. The
bottleneck is your own code.
- ``len(engine._slot.scheduler.mqs)`` grows without settling: the crawl
discovers requests faster than it downloads them. This is what makes long
crawls run out of memory.
Reading resource usage
----------------------
CPU
Scrapy runs in a single process, and everything except DNS resolution and
code you explicitly move to a thread runs in a single thread. One CPU core
is the ceiling; a process sitting at 100% of a core is CPU-bound no matter
how many cores the machine has.
Use a sampling profiler, such as py-spy_, to find out which code is
spending that CPU. :ref:`Selectors <topics-selectors>` and item pipelines
are the usual answer.
.. _py-spy: https://github.com/benfred/py-spy
Memory
The :ref:`memory usage extension <topics-extensions-ref-memusage>` records
:stat:`memusage/startup` and :stat:`memusage/max`. A :stat:`memusage/max`
far above :stat:`memusage/startup` is expected; what matters is whether it
keeps growing for as long as the crawl runs.
Growth that tracks ``len(engine._slot.scheduler.mqs)`` is a scheduling
problem, covered in :ref:`optimize-memory`. Growth that does not is a
:ref:`memory leak <topics-leaks>`.
Network
Compare :stat:`downloader/response_bytes` over the crawl time against your
available bandwidth. Saturated bandwidth caps concurrency regardless of any
setting.
DNS resolution is separate: it runs on a thread pool of
:setting:`REACTOR_THREADPOOL_MAXSIZE` threads, and results are cached
(:setting:`DNSCACHE_ENABLED`, :setting:`DNSCACHE_SIZE`). It only becomes a
limit of its own when there are many different domains to resolve, as in
:ref:`broad crawls <broad-crawls>`, where it shows up as slow starts and
DNS timeouts.
Disk
:ref:`Feed exports <topics-feed-exports>` write to disk on most crawls,
although item data is usually small enough for that not to matter. The ones
to suspect are
:class:`~scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware` and
the :ref:`media pipelines <topics-media-pipeline>`, which write whole
responses, and :setting:`JOBDIR`, which writes every scheduled request.
.. _optimize-concurrency:
Sending more requests at a time
===============================
:setting:`CONCURRENT_REQUESTS` caps how many requests are being downloaded at
any given moment, :setting:`CONCURRENT_REQUESTS_PER_DOMAIN` caps how many of
those may target the same domain, and :setting:`DOWNLOAD_DELAY` sets a minimum
wait between two consecutive requests to the same domain. A project generated by
:command:`startproject` gets one request per second per domain out of these.
Raise them to crawl a single website faster, and see
:ref:`broad-crawls-concurrency` to spread requests across many websites
instead.
The limit that matters, though, is the one the target website tolerates.
Exceeding it gets you throttled, served errors or banned, all of which make the
crawl slower than a lower concurrency would have been. To find that limit:
- Read the :ref:`robots.txt <topics-dlmw-robots>` file of the website. Scrapy
does not act on its ``Crawl-delay`` and ``Request-rate`` directives, so when
they are present, translate them into :setting:`DOWNLOAD_DELAY` and
concurrency settings yourself.
- Check the traffic that the website already gets, using a service like
`SimilarWeb`_ or `Cloudflare Radar`_. A rate that is a rounding error next
to what the website serves anyway is unlikely to be a problem for it.
.. _SimilarWeb: https://www.similarweb.com/
.. _Cloudflare Radar: https://radar.cloudflare.com/
- Look for a documented way in. An API, a bulk export or a search endpoint is
both faster for you and cheaper for the website than crawling its pages, and
the terms of service may state a rate.
- Crawl when the website is idle, in its own timezone, so that the capacity
you take is capacity nobody else wanted.
- Raise concurrency gradually and watch the website respond.
:stat:`downloader/response_status_count/{status_code}` counts for 429, 503
or the ban page of the website, growing :stat:`retry/count`, or a
:ref:`download latency <download-latency>` that climbs as you push harder,
all mean you have gone past the limit.
.. _optimize-requests:
Producing requests faster
=========================
A spider that discovers its requests one response at a time keeps the
downloader idle no matter how high you set :setting:`CONCURRENT_REQUESTS`. To
put more requests in the scheduler earlier:
- Request every page at once when you can work out how many there are, e.g.
from a page count or from a result count and a page size in the first
response, instead of following a link to the next page on every response.
- Get URLs from a source that lists many of them at once, such as a sitemap
or a search or export endpoint of the target website. For a crawl that
needs nothing else, :class:`~scrapy.spiders.SitemapSpider` reads sitemaps
for you.
- Raise the :attr:`~scrapy.Request.priority` of pagination requests, so that
they are downloaded before the requests that they compete with, and
discover the rest of the crawl sooner.
Each of these trades memory for speed: a request produced before the downloader
can take it waits in the scheduler, or on disk if you set :setting:`JOBDIR`.
Pushed far enough, they turn memory or disk into your new bottleneck, which is
why :ref:`optimize-memory` recommends the reverse of the last point.
.. _optimize-resources:
Lowering resource usage
=======================
.. _optimize-memory:
Lowering memory usage
---------------------
- Lower :setting:`SCRAPER_SLOT_MAX_ACTIVE_SIZE`.
- Lower :setting:`DOWNLOAD_MAXSIZE`, which allows a single response to take up
to 1 GiB of memory by default, multiplied by your concurrency. Set
:setting:`DOWNLOAD_WARNSIZE` first to find out whether the website actually
serves responses that big.
- Lower the number of :ref:`scheduled requests <topics-scheduler>` held in
memory:
- Increase the :attr:`~scrapy.Request.priority` of requests whose
:attr:`~scrapy.Request.callback` cannot yield additional requests.
For example, the following spider uses a higher priority (1) for book
requests than for pagination requests:
.. code-block:: python
from scrapy import Spider
class BooksToScrapeComSpider(Spider):
name = "books_toscrape_com"
start_urls = [
"http://books.toscrape.com/catalogue/category/books/mystery_3/index.html"
]
def parse(self, response):
next_page_links = response.css(".next a")
yield from response.follow_all(next_page_links)
book_links = response.css("article a")
yield from response.follow_all(book_links, callback=self.parse_book, priority=1)
def parse_book(self, response):
yield {
"name": response.css("h1::text").get(),
"price": response.css(".price_color::text").re_first("£(.*)"),
"url": response.url,
}
.. note:: If the number of request-yielding, low-priority requests
scheduled at any given time is lower than concurrency settings
(:setting:`CONCURRENT_REQUESTS_PER_DOMAIN` or
:setting:`CONCURRENT_REQUESTS`), as in the example above, this can
slow down your crawl by turning those requests into a bottleneck.
- If you have many :ref:`start requests <start-requests>`, consider
:ref:`delaying their iteration <start-requests-lazy>`.
- Set :setting:`JOBDIR` to offload all scheduled requests to disk.
- Be on the lookout for :ref:`memory leaks <topics-leaks>`.
Lowering network usage
----------------------
- Install brotli_ and zstandard_ to support brotli-compressed_ and
zstd-compressed_ responses.
.. _brotli-compressed: https://www.ietf.org/rfc/rfc7932.txt
.. _brotli: https://pypi.org/project/Brotli/
.. _zstd-compressed: https://www.ietf.org/rfc/rfc8478.txt
.. _zstandard: https://pypi.org/project/zstandard/
- Enable :class:`~scrapy.downloadermiddlewares.httpcache.HttpCacheMiddleware`
while developing your spider, so that re-runs do not download the same
responses again.
Lowering CPU usage
------------------
- Set :setting:`LOG_LEVEL` to ``"INFO"`` or higher.
- Restrict what you parse. A :ref:`selector <topics-selectors>` over a
smaller part of the response, or a single query whose result you reuse,
beats repeated queries over the whole document.
Other tips
----------
- Try :ref:`using the asyncio reactor <install-asyncio>` with uvloop_ as
:ref:`custom event loop <using-custom-loops>`, i.e. setting
:setting:`ASYNCIO_EVENT_LOOP` to ``"uvloop.Loop"``.
.. _uvloop: https://github.com/MagicStack/uvloop
Alternatively, try :ref:`switching to a non-asyncio reactor
<disable-asyncio>`.
- Disable unused :ref:`components <topics-components>`.
For example, set :setting:`COOKIES_ENABLED` to ``False`` unless you need
cookies.
- Split the crawl across separate processes to use more than one CPU core.
See :ref:`distributed-crawls`.
.. _broad-crawls:
.. _topics-broad-crawls:
Speeding up broad crawls
========================
While Scrapy is well suited for **broad crawls**, i.e. crawls that target many
websites, the default :ref:`settings <topics-settings>` are optimized for
crawls targeting a single website.
For broad crawls, consider these adjustments:
- .. _broad-crawls-concurrency:
Increase the global concurrency:
- Set :setting:`CONCURRENT_REQUESTS` as close to
:setting:`CONCURRENT_REQUESTS_PER_DOMAIN` × [number of target domains]
(e.g. 8 × 10 domains = 80 concurrent requests) as your CPU and memory
allow.
- Increase :setting:`SCRAPER_SLOT_MAX_ACTIVE_SIZE` when increasing
:setting:`CONCURRENT_REQUESTS` stops making a difference.
- .. _broad-crawls-bfo:
If memory is a bottleneck, see if :ref:`crawling in BFO order <bfo>` lowers
memory usage.
- Improve DNS resolution speed:
- Set up your own DNS server, with a local cache and upstream to a `large
DNS server`_, to avoid slowing down your network.
.. _large DNS server: https://en.wikipedia.org/wiki/Public_recursive_name_server#Notable_public_DNS_service_operators
- Increase :setting:`REACTOR_THREADPOOL_MAXSIZE` to the minimum value
that avoids DNS resolution timeouts and makes a noticeable positive
impact in crawl speed.
- Lower the negative impact of some responses:
- Set :setting:`RETRY_ENABLED` to ``False`` or, if you need retries,
consider lowering :setting:`RETRY_TIMES`.
- Lower :setting:`DOWNLOAD_TIMEOUT` to a more reasonable value, to
discard stuck requests more quickly.
- Set :setting:`REDIRECT_ENABLED` to ``False`` unless you want to follow
redirects.

View File

@ -53,65 +53,13 @@ Request objects
``None`` is passed as value, the HTTP header will not be sent at all.
.. caution:: Cookies set via the ``Cookie`` header are not considered by the
:ref:`cookies-mw`. If you need to set cookies for a request, use the
``cookies`` argument. This is a known current limitation that is being
worked on.
:ref:`cookie middleware <cookies>`. If you need to set cookies for a
request, use the ``cookies`` argument.
:type headers: dict
:param cookies: the request cookies. These can be sent in two forms.
.. invisible-code-block: python
from scrapy import Request
1. Using a dict:
.. code-block:: python
request_with_cookies = Request(
url="http://www.example.com",
cookies={"currency": "USD", "country": "UY"},
)
2. Using a list of dicts:
.. code-block:: python
request_with_cookies = Request(
url="https://www.example.com",
cookies=[
{
"name": "currency",
"value": "USD",
"domain": "example.com",
"path": "/currency",
"secure": True,
},
],
)
The latter form allows for customizing the ``domain`` and ``path``
attributes of the cookie. This is only useful if the cookies are saved
for later requests.
.. reqmeta:: dont_merge_cookies
When some site returns cookies (in a response) those are stored in the
cookies for that domain and will be sent again in future requests.
That's the typical behaviour of any regular web browser.
Note that setting the :reqmeta:`dont_merge_cookies` key to ``True`` in
:attr:`request.meta <scrapy.Request.meta>` causes custom cookies to be
ignored.
For more info see :ref:`cookies-mw`.
.. caution:: Cookies set via the ``Cookie`` header are not considered by the
:ref:`cookies-mw`. If you need to set cookies for a request, use the
:class:`scrapy.Request.cookies <scrapy.Request>` parameter. This is a known
current limitation that is being worked on.
:param cookies: the request cookies, as a dict of cookie names and values
or as a list of dicts with a cookie each. See :ref:`cookies`.
:type cookies: dict or list
:param encoding: the encoding of this request (defaults to ``'utf-8'``).

View File

@ -36,6 +36,77 @@ their input in an unsafe way, such as :func:`eval`, :func:`exec`, or
:func:`pickle.loads`, and be careful when writing response data to paths
derived from the response itself.
.. _security-response-size:
Memory use when parsing responses
=================================
Parsing a response with :ref:`selectors <topics-selectors>` builds an in-memory
tree of the whole response body, which takes several times as much memory as
the body itself. Scrapy parses without the size limits that libxml2 applies by
default, so the size of that tree is bound only by the size of the response, as
controlled by :setting:`DOWNLOAD_MAXSIZE` (default: 1 GiB).
XML entities are left unresolved, so the tree stays proportional to the
response body even for input crafted as an `XML bomb
<https://lxml.de/FAQ.html#is-lxml-vulnerable-to-xml-bombs>`_. A server can still
make a crawler allocate a lot of memory by returning a very large response,
though, so if you know the size of the responses you care about, lower the
limit:
.. code-block:: python
DOWNLOAD_MAXSIZE = 32 * 1024 * 1024 # 32 MiB
* **Pro:** a server cannot make the crawler allocate more memory than the limit
allows, whether by returning a large response or by crafting one that is
expensive to parse.
* **Con:** you can no longer scrape sites that legitimately serve responses
above the limit, as those responses are dropped.
.. _security-parser-limits:
Parser limits
-------------
The limits that libxml2 applies by default, such as 256 nesting levels and
10 MB per text node, can be restored by overriding
:attr:`~scrapy.http.TextResponse.selector` in a response subclass and swapping
responses in a :ref:`downloader middleware <topics-downloader-middleware>`:
.. code-block:: python
from functools import cached_property
from scrapy import Selector
from scrapy.http import HtmlResponse
class LimitedHtmlResponse(HtmlResponse):
@cached_property
def selector(self):
return Selector(self, huge_tree=False)
class LimitedParsingMiddleware:
def process_response(self, request, response, spider):
if isinstance(response, HtmlResponse):
return response.replace(cls=LimitedHtmlResponse)
return response
Do the same with :class:`~scrapy.http.XmlResponse` if you also parse XML.
These limits apply per node, so :setting:`DOWNLOAD_MAXSIZE` remains your bound
on total memory: a response made of many small elements is parsed in full and
uses as much memory either way.
* **Pro:** deeply nested responses, and responses with very large individual
nodes, become cheaper to parse.
* **Con:** parsing stops at those limits without raising, so a legitimate page
that exceeds them yields incomplete data and no error.
TLS connections
===============

View File

@ -69,9 +69,10 @@ Example::
precedence and override the project ones.
.. note:: :ref:`Pre-crawler settings <pre-crawler-settings>` cannot be defined
per spider, and :ref:`reactor settings <reactor-settings>` should not have
a different value per spider when :ref:`running multiple spiders in the
same process <run-multiple-spiders>`.
per spider, and :ref:`reactor settings <reactor-settings>` and
:ref:`logging settings <logging-settings>` are subject to restrictions when
:ref:`running multiple spiders in the same process
<run-multiple-spiders>`.
One way to do so is by setting their :attr:`~scrapy.Spider.custom_settings`
attribute:
@ -329,32 +330,41 @@ Reactor settings
**Reactor settings** are settings tied to the :doc:`Twisted reactor
<twisted:core/howto/reactor-basics>`.
These settings can be defined from a spider. However, because only 1 reactor
can be used per process, these settings cannot use a different value per spider
when :ref:`running multiple spiders in the same process
<run-multiple-spiders>`.
Because only 1 reactor can be used per process, these settings cannot use a
different value per spider when :ref:`running multiple spiders in the same
process <run-multiple-spiders>`.
In general, if different spiders define different values, the first defined
value is used. However, if two spiders request a different reactor, an
exception is raised.
These settings are:
These settings are used upon installing the reactor:
- :setting:`ASYNCIO_EVENT_LOOP` (not possible to set per-spider when using
:class:`~scrapy.crawler.AsyncCrawlerProcess`, see below)
- :setting:`TWISTED_REACTOR` (ignored when using
:class:`~scrapy.crawler.AsyncCrawlerProcess`, see below)
They can be :ref:`set from a spider <spider-settings>`, but only the values
from the first spider that runs are used, since that is when the reactor is
installed. If a later spider asks for a different reactor or a different event
loop, an exception is raised. With
:class:`~scrapy.crawler.CrawlerRunner` and
:class:`~scrapy.crawler.AsyncCrawlerRunner` the reactor must be installed
beforehand, so these settings are only used to check that the installed reactor
and event loop match them.
These settings are applied when starting the reactor:
- :setting:`TWISTED_DNS_RESOLVER` and settings used by the corresponding
component, e.g. :setting:`DNSCACHE_ENABLED`, :setting:`DNSCACHE_SIZE`
and :setting:`DNS_TIMEOUT` for the default one.
- :setting:`REACTOR_THREADPOOL_MAXSIZE`
- :setting:`TWISTED_REACTOR` (ignored when using
:class:`~scrapy.crawler.AsyncCrawlerProcess`, see below)
:setting:`ASYNCIO_EVENT_LOOP` and :setting:`TWISTED_REACTOR` are used upon
installing the reactor. The rest of the settings are applied when starting
the reactor.
They are read from the settings of the
:class:`~scrapy.crawler.CrawlerProcess` or
:class:`~scrapy.crawler.AsyncCrawlerProcess` object, so setting them from a
spider or an :ref:`add-on <topics-addons>` has no effect. They are ignored
altogether when using :class:`~scrapy.crawler.CrawlerRunner` or
:class:`~scrapy.crawler.AsyncCrawlerRunner`, which do not start the reactor.
There is an additional restriction for :setting:`TWISTED_REACTOR` and
:setting:`ASYNCIO_EVENT_LOOP` when using
@ -654,9 +664,8 @@ The default headers used for Scrapy HTTP Requests. They're populated in the
:class:`~scrapy.downloadermiddlewares.defaultheaders.DefaultHeadersMiddleware`.
.. caution:: Cookies set via the ``Cookie`` header are not considered by the
:ref:`cookies-mw`. If you need to set cookies for a request, use the
:class:`Request.cookies <scrapy.Request>` parameter. This is a known
current limitation that is being worked on.
:ref:`cookie middleware <cookies>`. If you need to set cookies for a
request, use the :class:`Request.cookies <scrapy.Request>` parameter.
.. caution:: A ``Referer`` header defined here only reaches requests for which
:class:`~scrapy.spidermiddlewares.referer.RefererMiddleware` does not set
@ -1913,6 +1922,7 @@ Type of in-memory queue used by the scheduler. Other available type is:
.. setting:: SCHEDULER_PRIORITY_QUEUE
.. _broad-crawls-scheduler-priority-queue:
SCHEDULER_PRIORITY_QUEUE
------------------------

View File

@ -338,15 +338,22 @@ spider_error
.. signal:: spider_error
.. function:: spider_error(failure, response, spider)
Sent when a spider callback generates an error (i.e. raises an exception).
Sent when a spider callback or the :meth:`~scrapy.Spider.start` method of a
spider generates an error (i.e. raises an exception).
.. versionchanged:: VERSION
Exceptions from :meth:`~scrapy.Spider.start` are also reported, see
:ref:`start-error`.
This signal does not support :ref:`asynchronous handlers <signal-deferred>`.
:param failure: the exception raised
:type failure: twisted.python.failure.Failure
:param response: the response being processed when the exception was raised
:type response: :class:`~scrapy.http.Response` object
:param response: the response being processed when the exception was
raised, or ``None`` if the exception came from
:meth:`~scrapy.Spider.start`.
:type response: :class:`~scrapy.http.Response` | ``None``
:param spider: the spider which raised the exception
:type spider: :class:`~scrapy.Spider` object

View File

@ -228,25 +228,7 @@ DepthMiddleware
.. module:: scrapy.spidermiddlewares.depth
:synopsis: Depth Spider Middleware
.. class:: DepthMiddleware
DepthMiddleware is used for tracking the depth of each Request inside the
site being scraped. It works by setting ``request.meta['depth'] = 0`` whenever
there is no value previously set (usually just the first Request) and
incrementing it by 1 otherwise.
It can be used to limit the maximum depth to scrape, control Request
priority based on their depth, and things like that.
The :class:`DepthMiddleware` can be configured through the following
settings (see the settings documentation for more info):
* :setting:`DEPTH_LIMIT` - The maximum depth that will be allowed to
crawl for any site. If zero, no limit will be imposed.
* :setting:`DEPTH_STATS_VERBOSE` - Whether to collect the number of
requests for each depth.
* :setting:`DEPTH_PRIORITY` - Whether to prioritize the requests based on
their depth.
.. autoclass:: DepthMiddleware
HttpErrorMiddleware
-------------------

View File

@ -411,6 +411,38 @@ scheduled requests:
await self.crawler.signals.wait_for(signals.scheduler_empty)
yield item_or_request
.. _start-error:
Handling start errors
---------------------
An exception raised by :meth:`~scrapy.Spider.start` ends its iteration, so any
remaining start items and requests are never sent. Scrapy logs the exception,
sends the :signal:`spider_error` signal, and, once the already scheduled
requests are done, closes the spider with the ``start_error``
:stat:`finish_reason`.
.. versionchanged:: VERSION
The close reason used to be ``finished``, and neither the
:signal:`spider_error` signal nor the :stat:`spider_exceptions/count` stat
reported the exception.
To keep the iteration going, catch the exception yourself:
.. code-block:: python
async def start(self):
for url in self.start_urls:
try:
request = Request(url)
except ValueError:
self.logger.exception(f"Skipping start URL {url}")
else:
yield request
To stop the crawl instead, and choose your own :stat:`finish_reason`, raise
:exc:`~scrapy.exceptions.CloseSpider`.
.. _builtin-spiders:
Generic Spiders

View File

@ -301,6 +301,10 @@ one per actual value of the placeholder.
- ``shutdown``: the crawl was interrupted, e.g. by a system signal such
as ``SIGINT`` (:kbd:`Ctrl-C`).
- ``start_error``: :meth:`~scrapy.Spider.start` raised an exception, so
some :ref:`start requests <start-requests>` may never have been sent,
see :ref:`start-error`.
Third-party components and your own code may use any other reason, e.g. by
raising :exc:`~scrapy.exceptions.CloseSpider` with it.
@ -721,18 +725,21 @@ one per actual value of the placeholder.
.. stat:: spider_exceptions/count
``spider_exceptions/count``
Number of unhandled exceptions raised by spider callbacks.
Number of unhandled exceptions raised by spider callbacks or by
:meth:`~scrapy.Spider.start`.
Set by the :ref:`scraper <topics-architecture>`.
Set by the :ref:`engine <topics-architecture>` and the :ref:`scraper
<topics-architecture>`.
.. stat:: spider_exceptions/{exception}
``spider_exceptions/{exception}``
Number of unhandled exceptions raised by spider callbacks, per exception,
where ``{exception}`` is the class name of the exception, e.g.
Same as :stat:`spider_exceptions/count`, per exception, where
``{exception}`` is the class name of the exception, e.g.
``spider_exceptions/ValueError``.
Set by the :ref:`scraper <topics-architecture>`.
Set by the :ref:`engine <topics-architecture>` and the :ref:`scraper
<topics-architecture>`.
.. stat:: start_time

View File

@ -26,6 +26,8 @@ dependencies = [
# Platform-specific dependencies
'PyDispatcher>=2.0.5; platform_python_implementation == "CPython"',
'PyPyDispatcher>=2.1.0; platform_python_implementation == "PyPy"',
'brotli>=1.2.0; implementation_name != "pypy"',
'brotlicffi>=1.2.0.0; implementation_name == "pypy"',
]
classifiers = [
"Development Status :: 5 - Production/Stable",
@ -62,10 +64,6 @@ Tracker = "https://github.com/scrapy/scrapy/issues"
[project.optional-dependencies]
bpython = ["bpython>=0.7.1"]
brotli = [
"brotli>=1.2.0; implementation_name != 'pypy'",
"brotlicffi>=1.2.0.0; implementation_name == 'pypy'",
]
gcs = ["google-cloud-storage>=1.29.0"]
httpx = ["httpx2[http2,socks]>=2.0.0"]
images = ["Pillow>=8.3.2"]
@ -125,23 +123,11 @@ module = [
"tests.test_downloaderslotssettings",
"tests.test_dupefilters",
"tests.test_engine_loop",
"tests.test_exporters",
"tests.test_extension_statsmailer",
"tests.test_extension_throttle",
"tests.test_feedexport",
"tests.test_feedexport_postprocess",
"tests.test_feedexport_storages",
"tests.test_feedexport_uri_params",
"tests.test_item",
"tests.test_linkextractors",
"tests.test_loader",
"tests.test_logformatter",
"tests.test_mail",
"tests.test_pipeline_crawl",
"tests.test_pipeline_files",
"tests.test_pipeline_images",
"tests.test_pipeline_media",
"tests.test_pipelines",
"tests.test_pqueues",
"tests.test_scheduler_base",
"tests.test_settings",

View File

@ -281,7 +281,6 @@ class Command(BaseRunSpiderCommand):
) -> list[Any]:
items, requests, opts, depth, spider, callback = args
if opts.pipelines:
assert self.pcrawler.engine
itemproc = self.pcrawler.engine.scraper.itemproc
if hasattr(itemproc, "process_item_async"):
for item in items:

View File

@ -548,7 +548,8 @@ class _ScrapyAgent:
txresponse._transport._producer.abortConnection()
raise DownloadCancelledError(warning_msg)
if warnsize and expected_size > warnsize:
reached_warnsize = bool(warnsize and expected_size > warnsize)
if reached_warnsize:
logger.warning(
get_warnsize_msg(expected_size, warnsize, request, expected=True)
)
@ -561,6 +562,7 @@ class _ScrapyAgent:
request=request,
maxsize=maxsize,
warnsize=warnsize,
reached_warnsize=reached_warnsize,
fail_on_dataloss=fail_on_dataloss,
crawler=self._crawler,
tls_verbose_logging=self._tls_verbose_logging,
@ -625,6 +627,7 @@ class _ResponseReader(Protocol):
fail_on_dataloss: bool,
crawler: Crawler,
*,
reached_warnsize: bool = False,
tls_verbose_logging: bool = False,
):
self._finished: Deferred[_ResultT] = finished
@ -634,7 +637,7 @@ class _ResponseReader(Protocol):
self._maxsize: int = maxsize
self._warnsize: int = warnsize
self._fail_on_dataloss: bool = fail_on_dataloss
self._reached_warnsize: bool = False
self._reached_warnsize: bool = reached_warnsize
self._bytes_received: int = 0
self._certificate: ssl.Certificate | None = None
self._ip_address: ipaddress.IPv4Address | ipaddress.IPv6Address | None = None

View File

@ -112,7 +112,6 @@ class ExecutionEngine:
self.crawler: Crawler = crawler
self.settings: Settings = crawler.settings
self.signals: SignalManager = crawler.signals
assert crawler.logformatter
self.logformatter: LogFormatter = crawler.logformatter
self._slot: _Slot | None = None
self.spider: Spider | None = None
@ -125,6 +124,9 @@ class ExecutionEngine:
] = spider_closed_callback
self.start_time: float | None = None
self._start: AsyncIterator[Any] | None = None
# Whether Spider.start() raised, i.e. some start items or requests may
# never have reached the engine.
self._start_error: bool = False
self._closewait: Deferred[None] | None = None
self._start_request_processing_awaitable: (
asyncio.Future[None] | Deferred[None] | None
@ -277,13 +279,29 @@ class ExecutionEngine:
item_or_request = await anext(self._start)
except StopAsyncIteration:
self._start = None
except CloseSpider as exception:
self._start = None
_schedule_coro(
self.close_spider_async(reason=exception.reason or "cancelled")
)
except Exception as exception:
self._start = None
self._start_error = True
exception_traceback = format_exc()
logger.error(
f"Error while reading start items and requests: {exception}.\n{exception_traceback}",
exc_info=True,
)
self.signals.send_catch_log(
signal=signals.spider_error,
failure=Failure(),
response=None,
spider=self.spider,
)
self.crawler.stats.inc_value("spider_exceptions/count")
self.crawler.stats.inc_value(
f"spider_exceptions/{type(exception).__name__}"
)
else:
if not self.spider:
return # spider already closed
@ -543,17 +561,17 @@ class ExecutionEngine:
if hasattr(scheduler, "open") and (d := scheduler.open(self.crawler.spider)):
await maybe_deferred_to_future(d)
await self.scraper.open_spider_async()
assert self.crawler.stats
if argument_is_required(self.crawler.stats.open_spider, "spider"):
stats = self.crawler.stats
if argument_is_required(stats.open_spider, "spider"):
warnings.warn(
f"The open_spider() method of {global_object_name(type(self.crawler.stats))} requires a spider argument,"
f"The open_spider() method of {global_object_name(type(stats))} requires a spider argument,"
f" this is deprecated and the argument will not be passed in future Scrapy versions.",
ScrapyDeprecationWarning,
stacklevel=2,
)
self.crawler.stats.open_spider(spider=self.crawler.spider)
stats.open_spider(spider=self.crawler.spider)
else:
self.crawler.stats.open_spider()
stats.open_spider()
await self.signals.send_catch_log_async(
signals.spider_opened, spider=self.crawler.spider
)
@ -579,7 +597,8 @@ class ExecutionEngine:
if DontCloseSpider in detected_ex:
return
if self.spider_is_idle():
ex = detected_ex.get(CloseSpider, CloseSpider(reason="finished"))
default_reason = "start_error" if self._start_error else "finished"
ex = detected_ex.get(CloseSpider, CloseSpider(reason=default_reason))
assert isinstance(ex, CloseSpider) # typing
_schedule_coro(self.close_spider_async(reason=ex.reason))
@ -655,20 +674,18 @@ class ExecutionEngine:
extra={"spider": spider},
)
assert self.crawler.stats
try:
if argument_is_required(self.crawler.stats.close_spider, "spider"):
stats = self.crawler.stats
if argument_is_required(stats.close_spider, "spider"):
warnings.warn(
f"The close_spider() method of {global_object_name(type(self.crawler.stats))} requires a spider argument,"
f"The close_spider() method of {global_object_name(type(stats))} requires a spider argument,"
f" this is deprecated and the argument will not be passed in future Scrapy versions.",
ScrapyDeprecationWarning,
stacklevel=2,
)
self.crawler.stats.close_spider(
spider=self.crawler.spider, reason=reason
)
stats.close_spider(spider=self.crawler.spider, reason=reason)
else:
self.crawler.stats.close_spider(reason=reason)
stats.close_spider(reason=reason)
except Exception:
logger.error("Stats close failure")

View File

@ -120,7 +120,6 @@ class Scraper:
self.concurrent_items: int = crawler.settings.getint("CONCURRENT_ITEMS")
self.crawler: Crawler = crawler
self.signals: SignalManager = crawler.signals
assert crawler.logformatter
self.logformatter: LogFormatter = crawler.logformatter
def _check_deprecated_itemproc_method(self, method: str) -> None:
@ -355,7 +354,6 @@ class Scraper:
assert self.crawler.spider
exc = _failure.value
if isinstance(exc, CloseSpider):
assert self.crawler.engine is not None # typing
_schedule_coro(
self.crawler.engine.close_spider_async(reason=exc.reason or "cancelled")
)
@ -374,11 +372,9 @@ class Scraper:
response=response,
spider=self.crawler.spider,
)
assert self.crawler.stats
self.crawler.stats.inc_value("spider_exceptions/count")
self.crawler.stats.inc_value(
f"spider_exceptions/{_failure.value.__class__.__name__}"
)
stats = self.crawler.stats
stats.inc_value("spider_exceptions/count")
stats.inc_value(f"spider_exceptions/{_failure.value.__class__.__name__}")
def handle_spider_output(
self,
@ -456,7 +452,6 @@ class Scraper:
Items are sent to the item pipelines, requests are scheduled.
"""
if isinstance(output, Request):
assert self.crawler.engine is not None # typing
self.crawler.engine.crawl(request=output)
return
if output is not None:

View File

@ -8,7 +8,7 @@ import signal
import warnings
from abc import ABC, abstractmethod
from functools import partial
from typing import TYPE_CHECKING, Any, TypeVar
from typing import TYPE_CHECKING, Any, Generic, TypeVar, overload
from twisted.internet.defer import Deferred, DeferredList, inlineCallbacks
@ -58,7 +58,55 @@ logger = logging.getLogger(__name__)
_T = TypeVar("_T")
class _LateAttribute(Generic[_T]):
"""Descriptor for a :class:`Crawler` attribute that only gets a value once
the crawl starts.
The value is kept in an attribute of the same name prefixed with an
underscore, and reading it before it is set raises :exc:`RuntimeError`.
This way the public attribute can be annotated as always set, and its
users, both in Scrapy and in third-party code, do not need to narrow its
type on every use. Code that runs before the crawl starts reads the
underscore-prefixed attribute instead.
"""
def __set_name__(self, owner: type[Crawler], name: str) -> None:
self._name = name
self._private_name = f"_{name}"
@overload
def __get__(self, instance: None, owner: type[Crawler]) -> _LateAttribute[_T]: ...
@overload
def __get__(self, instance: Crawler, owner: type[Crawler]) -> _T: ...
def __get__(
self, instance: Crawler | None, owner: type[Crawler]
) -> _LateAttribute[_T] | _T:
if instance is None:
return self
value: _T | None = getattr(instance, self._private_name)
if value is None:
raise RuntimeError(
f"Crawler.{self._name} is not set yet. It is set when the "
"crawl starts, so it can only be used from then on, e.g. "
"from the spider_opened signal handler onwards."
)
return value
def __set__(self, instance: Crawler, value: _T) -> None:
setattr(instance, self._private_name, value)
class Crawler:
engine: _LateAttribute[ExecutionEngine] = _LateAttribute()
extensions: _LateAttribute[ExtensionManager] = _LateAttribute()
logformatter: _LateAttribute[LogFormatter] = _LateAttribute()
request_fingerprinter: _LateAttribute[RequestFingerprinterProtocol] = (
_LateAttribute()
)
stats: _LateAttribute[StatsCollector] = _LateAttribute()
def __init__(
self,
spidercls: type[Spider],
@ -83,12 +131,13 @@ class Crawler:
self.crawling: bool = False
self._started: bool = False
self.extensions: ExtensionManager | None = None
self.stats: StatsCollector | None = None
self.logformatter: LogFormatter | None = None
self.request_fingerprinter: RequestFingerprinterProtocol | None = None
self.spider: Spider | None = None
self.engine: ExecutionEngine | None = None
self._engine: ExecutionEngine | None = None
self._extensions: ExtensionManager | None = None
self._logformatter: LogFormatter | None = None
self._request_fingerprinter: RequestFingerprinterProtocol | None = None
self._stats: StatsCollector | None = None
def _update_root_log_handler(self) -> None:
if get_scrapy_root_handler() is not None:
@ -225,8 +274,8 @@ class Crawler:
yield deferred_from_coro(self.engine.start_async())
except Exception:
self.crawling = False
if self.engine is not None:
yield deferred_from_coro(self.engine.close_async())
if self._engine is not None:
yield deferred_from_coro(self._engine.close_async())
raise
async def crawl_async(self, *args: Any, **kwargs: Any) -> None:
@ -255,8 +304,8 @@ class Crawler:
await self.engine.start_async()
except Exception:
self.crawling = False
if self.engine is not None:
await self.engine.close_async()
if self._engine is not None:
await self._engine.close_async()
raise
def _create_spider(self, *args: Any, **kwargs: Any) -> Spider:
@ -282,7 +331,6 @@ class Crawler:
"""
if self.crawling:
self.crawling = False
assert self.engine
if self.engine.running:
await self.engine.stop_async()
@ -313,7 +361,7 @@ class Crawler:
This method can only be called after the crawl engine has been created,
e.g. at signals :signal:`engine_started` or :signal:`spider_opened`.
"""
if not self.engine:
if self._engine is None:
raise RuntimeError(
"Crawler.get_downloader_middleware() can only be called after "
"the crawl engine has been created."
@ -331,7 +379,7 @@ class Crawler:
created, e.g. at signals :signal:`engine_started` or
:signal:`spider_opened`.
"""
if not self.extensions:
if self._extensions is None:
raise RuntimeError(
"Crawler.get_extension() can only be called after the "
"extension manager has been created."
@ -348,7 +396,7 @@ class Crawler:
This method can only be called after the crawl engine has been created,
e.g. at signals :signal:`engine_started` or :signal:`spider_opened`.
"""
if not self.engine:
if self._engine is None:
raise RuntimeError(
"Crawler.get_item_pipeline() can only be called after the "
"crawl engine has been created."
@ -365,7 +413,7 @@ class Crawler:
This method can only be called after the crawl engine has been created,
e.g. at signals :signal:`engine_started` or :signal:`spider_opened`.
"""
if not self.engine:
if self._engine is None:
raise RuntimeError(
"Crawler.get_spider_middleware() can only be called after the "
"crawl engine has been created."

View File

@ -55,7 +55,6 @@ class HttpCacheMiddleware:
@classmethod
def from_crawler(cls, crawler: Crawler) -> Self:
assert crawler.stats
o = cls(crawler.settings, crawler.stats)
crawler.signals.connect(o.spider_opened, signal=signals.spider_opened)
crawler.signals.connect(o.spider_closed, signal=signals.spider_closed)

View File

@ -30,27 +30,7 @@ if TYPE_CHECKING:
logger = getLogger(__name__)
ACCEPTED_ENCODINGS: list[bytes] = [b"gzip", b"deflate"]
try:
try:
import brotli
except ImportError:
import brotlicffi as brotli
except ImportError:
pass
else:
try:
brotli.Decompressor.can_accept_more_data # noqa: B018
except AttributeError: # pragma: no cover
warnings.warn(
"You have brotli installed. But 'br' encoding support now requires "
"brotli's or brotlicffi's version >= 1.2.0. Please upgrade "
"brotli/brotlicffi to make Scrapy decode 'br' encoded responses.",
stacklevel=2,
)
else:
ACCEPTED_ENCODINGS.append(b"br")
ACCEPTED_ENCODINGS: list[bytes] = [b"gzip", b"deflate", b"br"]
if find_spec("zstandard") is not None:
ACCEPTED_ENCODINGS.append(b"zstd")
@ -205,8 +185,6 @@ class HttpCompressionMiddleware:
f"{self.__class__.__name__} cannot decode the response for {response.url} "
f"from unsupported encoding(s) '{encodings_str}'."
)
if b"br" in encodings:
msg += " You need to install brotli or brotlicffi >= 1.2.0 to decode 'br'."
if b"zstd" in encodings:
msg += " You need to install zstandard to decode 'zstd'."
logger.warning(msg)

View File

@ -21,6 +21,36 @@ logger = logging.getLogger(__name__)
class OffsiteMiddleware:
"""Filter out requests for URLs outside the domains covered by the spider.
.. versionadded:: 2.11.2
A request is allowed if its host name is in the
:attr:`~scrapy.Spider.allowed_domains` attribute of the spider, or is a
subdomain of one of those domains. E.g. ``www.example.org`` also allows
``bob.www.example.org``, but neither ``www2.example.org`` nor
``example.org``. See :meth:`should_follow` to use a different policy.
If the spider does not define :attr:`~scrapy.Spider.allowed_domains`, or
the attribute is empty, every request is allowed.
Filtered requests are logged as follows::
DEBUG: Filtered offsite request to 'offsite.example': <GET http://offsite.example/some/page.html>
Only the first request filtered for a given domain is logged, to keep the
log readable.
.. reqmeta:: allow_offsite
allow_offsite
-------------
Requests with the ``allow_offsite`` :attr:`~scrapy.Request.meta` key set to
``True``, or with :attr:`~scrapy.Request.dont_filter` set to ``True``, are
allowed regardless of their host name.
"""
crawler: Crawler
host_regex: re.Pattern[str]
@ -31,7 +61,6 @@ class OffsiteMiddleware:
@classmethod
def from_crawler(cls, crawler: Crawler) -> Self:
assert crawler.stats
o = cls(crawler.stats)
crawler.signals.connect(o.spider_opened, signal=signals.spider_opened)
crawler.signals.connect(o.request_scheduled, signal=signals.request_scheduled)
@ -72,6 +101,23 @@ class OffsiteMiddleware:
raise IgnoreRequest(f"Filtered offsite request to {domain!r}")
def should_follow(self, request: Request, spider: Spider) -> bool:
"""Return ``True`` if *request* is on site, ``False`` if it must be
filtered out.
Override this method to implement a different offsite policy. For
example, to allow the domains in
:attr:`~scrapy.Spider.allowed_domains` but none of their subdomains:
.. code-block:: python
from scrapy.downloadermiddlewares.offsite import OffsiteMiddleware
from scrapy.utils.httpobj import urlparse_cached
class RootOnlyOffsiteMiddleware(OffsiteMiddleware):
def should_follow(self, request, spider):
return urlparse_cached(request).hostname in spider.allowed_domains
"""
self._update_host_regex(spider)
regex = self.host_regex
# hostname can be None for wrong urls (like javascript links)
@ -79,7 +125,6 @@ class OffsiteMiddleware:
return bool(regex.search(host))
def get_host_regex(self, spider: Spider) -> re.Pattern[str]:
"""Override this method to implement a different offsite policy"""
allowed_domains = getattr(spider, "allowed_domains", None)
if not allowed_domains:
return re.compile("") # allow all by default

View File

@ -94,7 +94,6 @@ def get_retry_request(
retry-related job stats
"""
settings = spider.crawler.settings
assert spider.crawler.stats
stats = spider.crawler.stats
retry_times = request.meta.get("retry_times", 0) + 1
if max_retry_times is None:

View File

@ -27,6 +27,7 @@ if TYPE_CHECKING:
from scrapy import Spider
from scrapy.crawler import Crawler
from scrapy.robotstxt import RobotParser
from scrapy.statscollectors import StatsCollector
logger = logging.getLogger(__name__)
@ -41,6 +42,7 @@ class RobotsTxtMiddleware:
self._default_useragent: str = crawler.settings["USER_AGENT"]
self._robotstxt_useragent: str | None = crawler.settings["ROBOTSTXT_USER_AGENT"]
self.crawler: Crawler = crawler
self._stats: StatsCollector = crawler.stats
self._parsers: dict[str, RobotParser | Deferred[RobotParser | None] | None] = {}
self._parserimpl: RobotParser = load_object(
crawler.settings.get("ROBOTSTXT_PARSER")
@ -78,8 +80,7 @@ class RobotsTxtMiddleware:
{"request": request},
extra={"spider": self.crawler.spider},
)
assert self.crawler.stats
self.crawler.stats.inc_value("robotstxt/forbidden")
self._stats.inc_value("robotstxt/forbidden")
raise IgnoreRequest("Forbidden by robots.txt")
async def robot_parser(self, request: Request) -> RobotParser | None:
@ -95,8 +96,6 @@ class RobotsTxtMiddleware:
meta={"dont_obey_robotstxt": True},
callback=NO_CALLBACK,
)
assert self.crawler.engine
assert self.crawler.stats
try:
resp = await self.crawler.engine.download_async(robotsreq)
await self._parse_robots(resp, netloc, request)
@ -109,7 +108,7 @@ class RobotsTxtMiddleware:
extra={"spider": self.crawler.spider},
)
self._robots_error(e, netloc)
self.crawler.stats.inc_value("robotstxt/request_count")
self._stats.inc_value("robotstxt/request_count")
parser = self._parsers[netloc]
if isinstance(parser, Deferred):
@ -119,11 +118,8 @@ class RobotsTxtMiddleware:
async def _parse_robots(
self, response: Response, netloc: str, request: Request
) -> None:
assert self.crawler.stats
self.crawler.stats.inc_value("robotstxt/response_count")
self.crawler.stats.inc_value(
f"robotstxt/response_status_count/{response.status}"
)
self._stats.inc_value("robotstxt/response_count")
self._stats.inc_value(f"robotstxt/response_status_count/{response.status}")
rp = self._parserimpl.from_crawler(self.crawler, response.body)
await self.crawler.signals.send_catch_log_async(
signal=signals.robots_parsed,
@ -138,8 +134,7 @@ class RobotsTxtMiddleware:
def _robots_error(self, exc: Exception, netloc: str) -> None:
if not isinstance(exc, IgnoreRequest):
key = f"robotstxt/exception_count/{type(exc)}"
assert self.crawler.stats
self.crawler.stats.inc_value(key)
self._stats.inc_value(key)
rp_dfd = self._parsers[netloc]
assert isinstance(rp_dfd, Deferred)
self._parsers[netloc] = None

View File

@ -43,7 +43,6 @@ class DownloaderStats:
def from_crawler(cls, crawler: Crawler) -> Self:
if not crawler.settings.getbool("DOWNLOADER_STATS"):
raise NotConfigured
assert crawler.stats
return cls(crawler.stats)
@_warn_spider_arg

View File

@ -95,7 +95,6 @@ class RFPDupeFilter(BaseDupeFilter):
@classmethod
def from_crawler(cls, crawler: Crawler) -> Self:
assert crawler.request_fingerprinter
debug = crawler.settings.getbool("DUPEFILTER_DEBUG")
return cls(
job_dir(crawler.settings),
@ -134,5 +133,4 @@ class RFPDupeFilter(BaseDupeFilter):
self.logger.debug(msg, {"request": request}, extra={"spider": spider})
self.logdupes = False
assert spider.crawler.stats
spider.crawler.stats.inc_value("dupefilter/filtered")

View File

@ -56,8 +56,12 @@ class DontCloseSpider(Exception):
class CloseSpider(Exception):
"""Raised from a :ref:`spider callback <topics-spiders>` to request the
spider to be closed/stopped.
"""Raised from a :ref:`spider callback <topics-spiders>` or from
:meth:`~scrapy.Spider.start` to request the spider to be closed/stopped.
.. versionchanged:: VERSION
Raising it from :meth:`~scrapy.Spider.start` closes the spider, instead
of being reported as a start error.
*reason* is a string with the reason for closing.

View File

@ -85,11 +85,16 @@ class BaseItemExporter(ABC):
declared = (name for name in adapter.field_names() if name in populated)
return dict.fromkeys([*declared, *adapter.keys()])
def _get_serialized_fields(
def get_serialized_fields(
self, item: Any, default_value: Any = None, include_empty: bool | None = None
) -> Iterable[tuple[str, Any]]:
"""Return the fields to export as an iterable of tuples
(name, serialized_value)
"""Return the fields of *item* to export, as an iterable of
``(name, serialized_value)`` tuples, taking :attr:`fields_to_export`
into account and applying :meth:`serialize_field` to every value.
Fields missing from *item* are exported with *default_value*.
*include_empty* overrides :attr:`export_empty_fields`.
"""
item = ItemAdapter(item)
@ -136,7 +141,7 @@ class JsonLinesItemExporter(BaseItemExporter):
self.encoder: JSONEncoder = ScrapyJSONEncoder(**self._kwargs)
def export_item(self, item: Any) -> None:
itemdict = dict(self._get_serialized_fields(item))
itemdict = dict(self.get_serialized_fields(item))
data = self.encoder.encode(itemdict) + "\n"
self.file.write(to_bytes(data, self.encoding))
@ -176,7 +181,7 @@ class JsonItemExporter(BaseItemExporter):
self.file.write(b"]")
def export_item(self, item: Any) -> None:
itemdict = dict(self._get_serialized_fields(item))
itemdict = dict(self.get_serialized_fields(item))
data = to_bytes(self.encoder.encode(itemdict), self.encoding)
self._add_comma_after_first()
self.file.write(data)
@ -216,7 +221,7 @@ class XmlItemExporter(BaseItemExporter):
self._beautify_indent(depth=1)
self.xg.startElement(self.item_element, AttributesImpl({}))
self._beautify_newline()
for name, value in self._get_serialized_fields(item, default_value=""):
for name, value in self.get_serialized_fields(item, default_value=""):
self._export_xml_field(name, value, depth=2)
self._beautify_indent(depth=1)
self.xg.endElement(self.item_element)
@ -310,7 +315,7 @@ class CsvItemExporter(BaseItemExporter):
f"See: https://docs.scrapy.org/en/latest/topics/feed-exports.html#feed-export-fields",
)
self._data_loss_warned = True
fields = self._get_serialized_fields(item, default_value="", include_empty=True)
fields = self.get_serialized_fields(item, default_value="", include_empty=True)
values = list(self._build_row(x for _, x in fields))
self.csv_writer.writerow(values)
@ -347,7 +352,7 @@ class PickleItemExporter(BaseItemExporter):
self.protocol: int = protocol
def export_item(self, item: Any) -> None:
d = dict(self._get_serialized_fields(item))
d = dict(self.get_serialized_fields(item))
pickle.dump(d, self.file, self.protocol)
@ -365,7 +370,7 @@ class MarshalItemExporter(BaseItemExporter):
self.file: BytesIO = file
def export_item(self, item: Any) -> None:
marshal.dump(dict(self._get_serialized_fields(item)), self.file)
marshal.dump(dict(self.get_serialized_fields(item)), self.file)
class PprintItemExporter(BaseItemExporter):
@ -374,7 +379,7 @@ class PprintItemExporter(BaseItemExporter):
self.file: BytesIO = file
def export_item(self, item: Any) -> None:
itemdict = dict(self._get_serialized_fields(item))
itemdict = dict(self.get_serialized_fields(item))
self.file.write(to_bytes(pprint.pformat(itemdict) + "\n"))
@ -417,5 +422,5 @@ class PythonItemExporter(BaseItemExporter):
yield key, self._serialize_value(value)
def export_item(self, item: Any) -> dict[str | bytes, Any]: # type: ignore[override]
result: dict[str | bytes, Any] = dict(self._get_serialized_fields(item))
result: dict[str | bytes, Any] = dict(self.get_serialized_fields(item))
return result

View File

@ -102,7 +102,6 @@ class CloseSpider:
self._close_spider("closespider_pagecount_no_item")
def spider_opened(self, spider: Spider) -> None:
assert self.crawler.engine
self.task = call_later(
self.close_on["timeout"], self._close_spider, "closespider_timeout"
)
@ -146,5 +145,4 @@ class CloseSpider:
self._close_spider("closespider_timeout_no_item")
def _close_spider(self, reason: str) -> None:
assert self.crawler.engine
_schedule_coro(self.crawler.engine.close_spider_async(reason=reason))

View File

@ -26,7 +26,6 @@ class CoreStats:
@classmethod
def from_crawler(cls, crawler: Crawler) -> Self:
assert crawler.stats
o = cls(crawler.stats)
crawler.signals.connect(o.spider_opened, signal=signals.spider_opened)
crawler.signals.connect(o.spider_closed, signal=signals.spider_closed)

View File

@ -45,7 +45,6 @@ class StackTraceDump:
return cls(crawler)
def dump_stacktrace(self, signum: int, frame: FrameType | None) -> None:
assert self.crawler.engine
log_args = {
"stackdumps": self._thread_stacks(),
"enginestatus": format_engine_status(self.crawler.engine),

View File

@ -181,7 +181,7 @@ class BlockingFeedStorage(ABC):
return NamedTemporaryFile(prefix="feed-", dir=path)
def store(self, file: IO[bytes]) -> Deferred[None] | None:
def store(self, file: IO[bytes]) -> Deferred[None]:
return deferred_from_coro(run_in_thread(self._store_in_thread, file))
@abstractmethod
@ -721,7 +721,6 @@ class FeedExporter:
logmsg = f"{slot.format} feed ({slot.itemcount} items) in: {slot.uri}"
slot_type = type(slot.storage).__name__
assert self.crawler.stats
try:
await ensure_awaitable(slot.storage.store(self._get_file(slot)))
except FileExistsError:

View File

@ -261,7 +261,6 @@ class DbmCacheStorage:
extra={"spider": spider},
)
assert spider.crawler.request_fingerprinter
self._fingerprinter: RequestFingerprinterProtocol = (
spider.crawler.request_fingerprinter
)
@ -326,7 +325,6 @@ class FilesystemCacheStorage:
extra={"spider": spider},
)
assert spider.crawler.request_fingerprinter
self._fingerprinter = spider.crawler.request_fingerprinter
def close_spider(self, spider: Spider) -> None:

View File

@ -37,7 +37,6 @@ class LogStats:
interval: float = crawler.settings.getfloat("LOGSTATS_INTERVAL")
if not interval:
raise NotConfigured
assert crawler.stats
o = cls(crawler.stats, interval)
crawler.signals.connect(o.spider_opened, signal=signals.spider_opened)
crawler.signals.connect(o.spider_closed, signal=signals.spider_closed)

View File

@ -29,7 +29,6 @@ class MemoryDebugger:
def from_crawler(cls, crawler: Crawler) -> Self:
if not crawler.settings.getbool("MEMDEBUG_ENABLED"):
raise NotConfigured
assert crawler.stats
o = cls(crawler.stats)
crawler.signals.connect(o.spider_closed, signal=signals.spider_closed)
return o

View File

@ -27,6 +27,7 @@ if TYPE_CHECKING:
from typing_extensions import Self
from scrapy.crawler import Crawler
from scrapy.statscollectors import StatsCollector
logger = logging.getLogger(__name__)
@ -43,6 +44,7 @@ class MemoryUsage:
raise NotConfigured from exc
self.crawler: Crawler = crawler
self._stats: StatsCollector = crawler.stats
self.warned: bool = False
self.notify_mails: list[str] = crawler.settings.getlist("MEMUSAGE_NOTIFY_MAIL")
if self.notify_mails: # pragma: no cover
@ -77,8 +79,7 @@ class MemoryUsage:
return size
def engine_started(self) -> None:
assert self.crawler.stats
self.crawler.stats.set_value("memusage/startup", self.get_virtual_size())
self._stats.set_value("memusage/startup", self.get_virtual_size())
self.tasks: list[AsyncioLoopingCall | LoopingCall] = []
tsk = create_looping_call(self.update)
self.tasks.append(tsk)
@ -98,15 +99,12 @@ class MemoryUsage:
tsk.stop()
def update(self) -> None:
assert self.crawler.stats
self.crawler.stats.max_value("memusage/max", self.get_virtual_size())
self._stats.max_value("memusage/max", self.get_virtual_size())
def _check_limit(self) -> None:
assert self.crawler.engine
assert self.crawler.stats
peak_mem_usage = self.get_virtual_size()
if peak_mem_usage > self.limit:
self.crawler.stats.set_value("memusage/limit_reached", 1)
self._stats.set_value("memusage/limit_reached", 1)
mem = self.limit / 1024 / 1024
logger.error(
"Memory usage exceeded %(memusage)dMiB. Shutting down Scrapy...",
@ -119,7 +117,7 @@ class MemoryUsage:
f"memory usage exceeded {mem}MiB at {socket.gethostname()}"
)
self._send_report(self.notify_mails, subj)
self.crawler.stats.set_value("memusage/limit_notified", 1)
self._stats.set_value("memusage/limit_notified", 1)
if self.crawler.engine.spider is not None:
_schedule_coro(
@ -136,9 +134,8 @@ class MemoryUsage:
def _check_warning(self) -> None:
if self.warned: # warn only once
return
assert self.crawler.stats
if self.get_virtual_size() > self.warning:
self.crawler.stats.set_value("memusage/warning_reached", 1)
self._stats.set_value("memusage/warning_reached", 1)
self.crawler.signals.send_catch_log(signal=signals.memusage_warning_reached)
mem = self.warning / 1024 / 1024
logger.warning(
@ -152,16 +149,13 @@ class MemoryUsage:
f"memory usage reached {mem}MiB at {socket.gethostname()}"
)
self._send_report(self.notify_mails, subj)
self.crawler.stats.set_value("memusage/warning_notified", 1)
self._stats.set_value("memusage/warning_notified", 1)
self.warned = True
def _send_report(self, rcpts: list[str], subject: str) -> None: # pragma: no cover
"""send notification mail with some additional useful info"""
assert self.crawler.engine
assert self.crawler.stats
stats = self.crawler.stats
s = f"Memory usage at engine startup : {stats.get_value('memusage/startup') / 1024 / 1024}M\r\n"
s += f"Maximum memory usage : {stats.get_value('memusage/max') / 1024 / 1024}M\r\n"
s = f"Memory usage at engine startup : {self._stats.get_value('memusage/startup') / 1024 / 1024}M\r\n"
s += f"Maximum memory usage : {self._stats.get_value('memusage/max') / 1024 / 1024}M\r\n"
s += f"Current memory usage : {self.get_virtual_size() / 1024 / 1024}M\r\n"
s += (

View File

@ -87,7 +87,6 @@ class PeriodicLog:
)
if not (ext_stats or ext_delta or ext_timing_enabled):
raise NotConfigured
assert crawler.stats
assert ext_stats is not None
assert ext_delta is not None
o = cls(

View File

@ -42,7 +42,6 @@ class StatsMailer:
if not recipients:
raise NotConfigured
mail: MailSender = MailSender.from_crawler(crawler)
assert crawler.stats
o = cls(crawler.stats, recipients, mail)
crawler.signals.connect(o.spider_closed, signal=signals.spider_closed)
return o

View File

@ -108,7 +108,6 @@ class TelnetConsole(protocol.ServerFactory):
def _get_telnet_vars(self) -> dict[str, Any]:
# Note: if you add entries here also update topics/telnetconsole.rst
assert self.crawler.engine
telnet_vars: dict[str, Any] = {
"engine": self.crawler.engine,
"spider": self.crawler.engine.spider,

View File

@ -45,7 +45,6 @@ class AutoThrottle:
def _spider_opened(self, spider: Spider) -> None:
self.mindelay = self._min_delay()
self.maxdelay = self._max_delay()
assert self.crawler.engine
self.crawler.engine.downloader._delay = self._start_delay()
def _min_delay(self) -> float:
@ -98,7 +97,6 @@ class AutoThrottle:
key: str | None = request.meta.get("download_slot")
if key is None:
return None, None
assert self.crawler.engine
return key, self.crawler.engine.downloader.slots.get(key)
def _adjust_delay(self, slot: Slot, latency: float, response: Response) -> None:

View File

@ -305,6 +305,11 @@ class Response(object_ref):
:class:`~.TextResponse` provides a :meth:`~.TextResponse.follow_all`
method which supports selectors in addition to absolute/relative URLs
and Link objects.
.. caution:: Every returned request gets its own *meta* and
*cb_kwargs* dictionaries, but the values within them are shared.
Mutating one of those values, e.g. appending to a list, affects
all the returned requests.
"""
if not hasattr(urls, "__iter__"):
raise TypeError("'urls' argument must be an iterable")

View File

@ -276,6 +276,9 @@ class TextResponse(Response):
using the ``css`` or ``xpath`` parameters, this method will not produce requests for
selectors from which links cannot be obtained (for instance, anchor tags without an
``href`` attribute)
.. seealso:: :meth:`.Response.follow_all`, for a caution about mutable
*meta* and *cb_kwargs* values.
"""
arguments = [x for x in (urls, css, xpath) if x is not None]
if len(arguments) != 1:

View File

@ -332,7 +332,7 @@ class LxmlLinkExtractor:
unique=unique,
process=process_value,
strip=strip,
canonicalized=not canonicalize,
canonicalized=True,
)
self.allow_res: list[re.Pattern[str]] = self._compile_regexes(allow)
self.deny_res: list[re.Pattern[str]] = self._compile_regexes(deny)

View File

@ -697,9 +697,9 @@ class FilesPipeline(MediaPipeline):
}
def inc_stats(self, status: str) -> None:
assert self.crawler.stats
self.crawler.stats.inc_value("file_count")
self.crawler.stats.inc_value(f"file_status_count/{status}")
stats = self.crawler.stats
stats.inc_value("file_count")
stats.inc_value(f"file_status_count/{status}")
async def _file_downloaded(
self,

View File

@ -29,7 +29,7 @@ from scrapy.utils.defer import ensure_awaitable
from scrapy.utils.python import to_bytes
if TYPE_CHECKING:
from collections.abc import Iterable
from collections.abc import Iterator
from os import PathLike
from PIL import Image
@ -180,7 +180,7 @@ class ImagesPipeline(FilesPipeline):
info: MediaPipeline.SpiderInfo,
*,
item: Any = None,
) -> Iterable[tuple[str, Image.Image, BytesIO]]:
) -> Iterator[tuple[str, Image.Image, BytesIO]]:
path = self.file_path(request, response=response, info=info, item=item)
orig_image = self._Image.open(BytesIO(response.body))
transposed_image = self._ImageOps.exif_transpose(orig_image)

View File

@ -100,7 +100,6 @@ class MediaPipeline(ABC):
stacklevel=2,
)
self.crawler: Crawler = crawler
assert crawler.request_fingerprinter
self._fingerprinter: RequestFingerprinterProtocol = (
crawler.request_fingerprinter
)
@ -228,7 +227,6 @@ class MediaPipeline(ABC):
) -> FileInfo:
try:
self._modify_media_request(request)
assert self.crawler.engine
response = await self.crawler.engine.download_async(request)
return await ensure_awaitable(
self.media_downloaded(response, request, info, item=item)

View File

@ -265,7 +265,6 @@ class ScrapyPriorityQueue:
class DownloaderInterface:
def __init__(self, crawler: Crawler):
assert crawler.engine
self.downloader: Downloader = crawler.engine.downloader
def stats(self, possible_slots: Iterable[str]) -> list[tuple[int, str]]:

View File

@ -193,7 +193,6 @@ class Shell:
"""
if not self.spider:
await self._open_spider(spider)
assert self.crawler.engine is not None
# send the request to the engine
self.crawler.engine.crawl(request)
# this will fire when the request callback runs (via the callback hijacking in _request_deferred())
@ -204,7 +203,6 @@ class Shell:
spider = self.crawler.spider or self.crawler._create_spider()
self.crawler.spider = spider
assert self.crawler.engine
await self.crawler.engine.open_spider_async(close_if_idle=False)
self.spider = spider

View File

@ -28,6 +28,27 @@ logger = logging.getLogger(__name__)
class DepthMiddleware(BaseSpiderMiddleware):
"""Track the depth of each request within the site being scraped, setting
``request.meta["depth"]`` to 0 when there is no value previously set
(usually just the first request) and incrementing it by 1 otherwise.
It can be used to limit the maximum depth to scrape, control request
priority based on their depth, and things like that, through the
:setting:`DEPTH_LIMIT`, :setting:`DEPTH_STATS_VERBOSE` and
:setting:`DEPTH_PRIORITY` settings.
.. reqmeta:: depth_reset
depth_reset
-----------
.. versionadded:: VERSION
:attr:`~scrapy.Request.meta` key that, set to ``True``, gives a request
depth 0 instead of the depth of its source response plus 1, e.g. to keep
:setting:`DEPTH_LIMIT` from applying across a domain change.
"""
crawler: Crawler
def __init__( # pylint: disable=super-init-not-called
@ -49,7 +70,6 @@ class DepthMiddleware(BaseSpiderMiddleware):
maxdepth = settings.getint("DEPTH_LIMIT")
verbose = settings.getbool("DEPTH_STATS_VERBOSE")
prio = settings.getint("DEPTH_PRIORITY")
assert crawler.stats
o = cls(maxdepth, crawler.stats, verbose, prio)
o.crawler = crawler
return o
@ -87,10 +107,13 @@ class DepthMiddleware(BaseSpiderMiddleware):
def get_processed_request(
self, request: Request, response: Response | None
) -> Request | None:
# Consumed here so that it cannot reach response.meta and, from there,
# spread to further requests through a meta copy.
depth_reset = request.meta.pop("depth_reset", False)
if response is None:
# start requests
return request
depth = response.meta["depth"] + 1
depth = 0 if depth_reset else response.meta["depth"] + 1
request.meta["depth"] = depth
if self.prio:
request.priority -= depth * self.prio

View File

@ -78,9 +78,9 @@ class HttpErrorMiddleware:
self, response: Response, exception: Exception, spider: Spider | None = None
) -> Iterable[Any] | None:
if isinstance(exception, HttpError):
assert self.crawler.stats
self.crawler.stats.inc_value("httperror/response_ignored_count")
self.crawler.stats.inc_value(
stats = self.crawler.stats
stats.inc_value("httperror/response_ignored_count")
stats.inc_value(
f"httperror/response_ignored_status_count/{response.status}"
)
logger.info(

View File

@ -15,8 +15,7 @@ logger = logging.getLogger(__name__)
class MetaCopyDetectionMiddleware(BaseSpiderMiddleware):
"""Warn when a spider yields a request with internal meta keys that should
not be copied from response.meta, or when two requests share the same meta
dict object.
not be copied from response.meta.
Each warning is emitted at most once per crawl.
"""

View File

@ -48,6 +48,5 @@ class UrlLengthMiddleware(BaseSpiderMiddleware):
{"maxlength": self.maxlength, "url": request.url},
extra={"spider": self.crawler.spider},
)
assert self.crawler.stats
self.crawler.stats.inc_value("urllength/request_ignored_count")
return None

View File

@ -2,11 +2,10 @@ import contextlib
import zlib
from io import BytesIO
with contextlib.suppress(ImportError):
try:
import brotli
except ImportError:
import brotlicffi as brotli
try:
import brotli
except ImportError:
import brotlicffi as brotli
with contextlib.suppress(ImportError):
import zstandard

View File

@ -239,13 +239,12 @@ class LogCounterHandler(logging.Handler):
def emit(self, record: logging.LogRecord) -> None:
sname = f"log_count/{record.levelname}"
assert self.crawler.stats
self.crawler.stats.inc_value(sname)
def logformatter_adapter(
logkws: LogFormatterResult,
) -> tuple[int, str, dict[str, Any] | tuple[Any, ...]]:
) -> tuple[Any, ...]:
"""
Helper that takes the dictionary output from the methods in LogFormatter
and adapts it into a tuple of positional arguments for logger.log calls.
@ -253,10 +252,14 @@ def logformatter_adapter(
level = logkws.get("level", logging.INFO)
message = logkws.get("msg") or ""
# NOTE: This also handles 'args' being an empty dict, that case doesn't
# play well in logger.log calls
args = cast("dict[str, Any]", logkws) if not logkws.get("args") else logkws["args"]
args = logkws.get("args")
# logging interpolates the message whenever it receives any positional
# argument, so empty args are left out. Tuple args become one positional
# argument each, while a dict is a single positional argument.
if not args:
return (level, message)
if isinstance(args, tuple):
return (level, message, *args)
return (level, message, args)

View File

@ -90,6 +90,10 @@ def open_in_browser(
def parse_details(self, response):
if "item name" not in response.text:
open_in_browser(response)
On the Windows Subsystem for Linux, set the ``BROWSER`` environment
variable to `wslview <https://github.com/wslutilities/wslu>`_ to open the
response in a Windows browser, which cannot read Linux paths otherwise.
"""
# circular imports
from scrapy.http import HtmlResponse, TextResponse # noqa: PLC0415

View File

@ -27,11 +27,12 @@ class MockFTPServer:
(anonymous) and a temporary root path that you can read from the
:attr:`path` attribute."""
proc: Popen[str]
port: int
path: Path
def __init__(self) -> None:
self.proc: Popen[str] | None = None
self.host: str = "127.0.0.1"
self.port: int | None = None
self.path: Path | None = None
def __enter__(self) -> Self:
self.path = Path(mkdtemp())
@ -63,7 +64,6 @@ class MockFTPServer:
traceback: TracebackType | None,
) -> None:
rmtree(str(self.path))
assert self.proc is not None
self.proc.kill()
self.proc.communicate()

View File

@ -201,7 +201,7 @@ class TestInteractiveShell:
env = os.environ.copy()
env["SCRAPY_PYTHON_SHELL"] = "python"
logfile = BytesIO()
p = PopenSpawn(args, env=env, timeout=5)
p = PopenSpawn(args, env=env, timeout=60)
p.logfile_read = logfile
p.expect_exact("Available Scrapy objects")
p.sendline(f"fetch('{mockserver.url('/')}')")
@ -235,7 +235,7 @@ class TestInteractiveShell:
def _run_interactive_shell(self, env: dict[str, str]) -> str:
args = (sys.executable, "-m", "scrapy.cmdline", "shell")
logfile = BytesIO()
p = PopenSpawn(args, env=env, timeout=5)
p = PopenSpawn(args, env=env, timeout=60)
p.logfile_read = logfile
p.expect_exact("Available Scrapy objects")
p.sendeof()
@ -256,7 +256,7 @@ class TestInteractiveShell:
self._isolate_config(env, config_home)
args = (sys.executable, "-m", "scrapy.cmdline", "shell")
logfile = BytesIO()
p = PopenSpawn(args, env=env, timeout=10)
p = PopenSpawn(args, env=env, timeout=60)
p.logfile_read = logfile
p.expect_exact("Available Scrapy objects")
# The standard Python shell never imports IPython, whereas the IPython

View File

@ -74,6 +74,31 @@ class TestCrawler:
assert not settings.frozen
assert crawler.settings.frozen
@pytest.mark.parametrize(
"attr",
["extensions", "logformatter", "request_fingerprinter", "stats"],
)
def test_late_attr_before_apply_settings(self, attr: str) -> None:
crawler = get_raw_crawler(DefaultSpider)
with pytest.raises(RuntimeError, match=rf"Crawler\.{attr} is not set yet"):
getattr(crawler, attr)
crawler._apply_settings()
assert getattr(crawler, attr) is not None
@pytest.mark.parametrize(
"attr",
["engine", "extensions", "logformatter", "request_fingerprinter", "stats"],
)
def test_late_attr_on_class(self, attr: str) -> None:
# Introspection tools such as help() read these off the class.
assert getattr(Crawler, attr) is getattr(Crawler, attr)
def test_late_attr_engine_before_crawl(self) -> None:
crawler = get_raw_crawler(DefaultSpider)
crawler._apply_settings()
with pytest.raises(RuntimeError, match=r"Crawler\.engine is not set yet"):
_ = crawler.engine
@pytest.mark.parametrize(
("attr", "setting"),
[

View File

@ -400,7 +400,7 @@ class TestAsyncCrawlerProcessSubprocess(TestCrawlerProcessSubprocessBase):
def test_reactorless_import_hook(self) -> None:
log = self.run_script("reactorless_import_hook.py")
assert "Not using a Twisted reactor" in log
assert "Spider closed (finished)" in log
assert "Spider closed (start_error)" in log
assert "ImportError: Import of twisted.internet.reactor is forbidden" in log
def test_reactorless_import_hook_uninstall(self) -> None:

View File

@ -52,20 +52,6 @@ FORMAT = {
}
def _skip_if_no_br() -> None:
try:
try:
import brotli # noqa: PLC0415
brotli.Decompressor.can_accept_more_data
except (ImportError, AttributeError):
import brotlicffi # noqa: PLC0415
brotlicffi.Decompressor.can_accept_more_data
except (ImportError, AttributeError):
pytest.skip("no brotli support")
def _skip_if_no_zstd() -> None:
pytest.importorskip("zstandard")
@ -161,8 +147,6 @@ class TestHttpCompression:
self.assertStatsEqual("httpcompression/response_bytes", 74837)
def test_process_response_br(self):
_skip_if_no_br()
response = self._getresponse("br")
assert response.request
request = response.request
@ -174,32 +158,6 @@ class TestHttpCompression:
self.assertStatsEqual("httpcompression/response_count", 1)
self.assertStatsEqual("httpcompression/response_bytes", 74837)
def test_process_response_br_unsupported(self, caplog: pytest.LogCaptureFixture):
if find_spec("brotli") is not None or find_spec("brotlicffi") is not None:
pytest.skip("Requires not having brotli support")
response = self._getresponse("br")
assert response.request
request = response.request
assert response.headers["Content-Encoding"] == b"br"
caplog.clear()
with caplog.at_level(
WARNING, logger="scrapy.downloadermiddlewares.httpcompression"
):
newresponse = self.mw.process_response(request, response)
assert caplog.record_tuples == [
(
"scrapy.downloadermiddlewares.httpcompression",
WARNING,
(
"HttpCompressionMiddleware cannot decode the response for "
"http://scrapytest.org/ from unsupported encoding(s) 'br'. "
"You need to install brotli or brotlicffi >= 1.2.0 to decode 'br'."
),
),
]
assert newresponse is not response
assert newresponse.headers.getlist("Content-Encoding") == [b"br"]
def test_process_response_zstd(self):
_skip_if_no_zstd()
@ -550,8 +508,6 @@ class TestHttpCompression:
assert cause.decompressed_size < 1_100_000
def test_compression_bomb_setting_br(self):
_skip_if_no_br()
self._test_compression_bomb_setting("br")
def test_compression_bomb_setting_deflate(self):
@ -609,8 +565,6 @@ class TestHttpCompression:
@pytest.mark.filterwarnings("ignore::scrapy.exceptions.ScrapyDeprecationWarning")
def test_compression_bomb_spider_attr_br(self):
_skip_if_no_br()
self._test_compression_bomb_spider_attr("br")
@pytest.mark.filterwarnings("ignore::scrapy.exceptions.ScrapyDeprecationWarning")
@ -643,8 +597,6 @@ class TestHttpCompression:
assert cause.decompressed_size < 1_100_000
def test_compression_bomb_request_meta_br(self):
_skip_if_no_br()
self._test_compression_bomb_request_meta("br")
def test_compression_bomb_request_meta_deflate(self):
@ -689,8 +641,6 @@ class TestHttpCompression:
def test_download_warnsize_setting_br(
self, caplog: pytest.LogCaptureFixture
) -> None:
_skip_if_no_br()
self._test_download_warnsize_setting(caplog, "br")
def test_download_warnsize_setting_deflate(
@ -744,8 +694,6 @@ class TestHttpCompression:
def test_download_warnsize_spider_attr_br(
self, caplog: pytest.LogCaptureFixture
) -> None:
_skip_if_no_br()
self._test_download_warnsize_spider_attr(caplog, "br")
@pytest.mark.filterwarnings("ignore::scrapy.exceptions.ScrapyDeprecationWarning")
@ -799,8 +747,6 @@ class TestHttpCompression:
def test_download_warnsize_request_meta_br(
self, caplog: pytest.LogCaptureFixture
) -> None:
_skip_if_no_br()
self._test_download_warnsize_request_meta(caplog, "br")
def test_download_warnsize_request_meta_deflate(
@ -834,7 +780,6 @@ class TestHttpCompression:
return new_response
def test_process_truncated_response_br(self):
_skip_if_no_br()
resp = self._get_truncated_response("br")
assert resp.body.startswith(b"<!DOCTYPE")

View File

@ -6,6 +6,7 @@ import pytest
from scrapy import Request, Spider
from scrapy.downloadermiddlewares.offsite import OffsiteMiddleware
from scrapy.exceptions import IgnoreRequest
from scrapy.utils.httpobj import urlparse_cached
from scrapy.utils.test import get_crawler
UNSET = object()
@ -237,6 +238,21 @@ def test_repeated_offsite_domain():
assert crawler.stats.get_value("offsite/filtered") == 2
def test_should_follow_override():
class RootOnlyOffsiteMiddleware(OffsiteMiddleware):
def should_follow(self, request: Request, spider: Spider) -> bool:
allowed_domains: list[str] = getattr(spider, "allowed_domains", [])
return urlparse_cached(request).hostname in allowed_domains
crawler = get_crawler(Spider)
crawler.spider = crawler._create_spider(name="a", allowed_domains=["example.com"])
mw = RootOnlyOffsiteMiddleware.from_crawler(crawler)
mw.spider_opened(crawler.spider)
assert mw.process_request(Request("https://example.com/1")) is None
with pytest.raises(IgnoreRequest):
mw.process_request(Request("https://www.example.com/1"))
def test_ignore_request_reason():
crawler = get_crawler(Spider)
crawler.spider = crawler._create_spider(name="a", allowed_domains=["example.com"])

View File

@ -6,6 +6,7 @@ from typing import TYPE_CHECKING, Any
from scrapy import Request, Spider, signals
from scrapy.core.scheduler import BaseScheduler
from scrapy.exceptions import CloseSpider
from scrapy.utils.asyncio import call_later, sleep
from scrapy.utils.test import get_crawler
from tests.mockserver.http import MockServer
@ -140,6 +141,74 @@ class TestMain:
assert crawler.stats.get_value("finish_reason") == "shutdown"
assert not actual_urls
@coroutine_test
async def test_start_error(self, caplog: pytest.LogCaptureFixture) -> None:
class TestSpider(Spider):
name = "test"
async def start(self):
yield Request("data:,a")
raise ValueError
def parse(self, response):
pass
actual_urls = []
errors = []
def track_url(request, spider):
actual_urls.append(request.url)
def track_error(failure, response, spider):
errors.append((failure, response))
settings = {"SCHEDULER": MemoryScheduler}
crawler = get_crawler(TestSpider, settings_dict=settings)
crawler.signals.connect(track_url, signals.request_reached_downloader)
crawler.signals.connect(track_error, signals.spider_error)
caplog.clear()
with caplog.at_level(ERROR):
await crawler.crawl_async()
# The requests yielded before the exception are still crawled.
assert actual_urls == ["data:,a"]
assert len(caplog.records) == 1
assert len(errors) == 1
failure, response = errors[0]
assert isinstance(failure.value, ValueError)
assert response is None
assert crawler.stats
assert crawler.stats.get_value("finish_reason") == "start_error"
assert crawler.stats.get_value("spider_exceptions/count") == 1
assert crawler.stats.get_value("spider_exceptions/ValueError") == 1
@coroutine_test
async def test_close_spider_from_start(
self, caplog: pytest.LogCaptureFixture
) -> None:
class TestSpider(Spider):
name = "test"
async def start(self):
yield Request("data:,a")
raise CloseSpider("my_reason")
def parse(self, response):
pass
settings = {"SCHEDULER": MemoryScheduler}
crawler = get_crawler(TestSpider, settings_dict=settings)
caplog.clear()
with caplog.at_level(ERROR):
await crawler.crawl_async()
assert not caplog.records
assert crawler.stats
assert crawler.stats.get_value("finish_reason") == "my_reason"
assert crawler.stats.get_value("spider_exceptions/count") is None
class TestRequestSendOrder:
seconds = 0.1 # increase if flaky

View File

@ -4,6 +4,7 @@ import marshal
import pickle
import re
from abc import ABC, abstractmethod
from collections.abc import Mapping
from datetime import datetime
from io import BytesIO
from typing import Any
@ -63,18 +64,18 @@ class TestBaseItemExporter(ABC):
self.ie = self._get_exporter()
@abstractmethod
def _get_exporter(self, **kwargs) -> BaseItemExporter:
def _get_exporter(self, **kwargs: Any) -> BaseItemExporter:
raise NotImplementedError
def _check_output(self): # noqa: B027
def _check_output(self) -> None: # noqa: B027
pass
def _assert_expected_item(self, exported_dict):
def _assert_expected_item(self, exported_dict: dict[str, Any]) -> None:
for k, v in exported_dict.items():
exported_dict[k] = to_unicode(v)
assert self.i == self.item_class(**exported_dict)
def _get_nonstring_types_item(self):
def _get_nonstring_types_item(self) -> dict[str, Any]:
return {
"boolean": False,
"number": 22,
@ -82,7 +83,7 @@ class TestBaseItemExporter(ABC):
"float": 3.14,
}
def assertItemExportWorks(self, item):
def assertItemExportWorks(self, item: Any) -> None:
self.ie.start_exporting()
self.ie.export_item(item)
self.ie.finish_exporting()
@ -92,7 +93,7 @@ class TestBaseItemExporter(ABC):
del self.ie
self._check_output()
def test_export_item(self):
def test_export_item(self) -> None:
self.assertItemExportWorks(self.i)
def test_export_dict_item(self):
@ -108,26 +109,26 @@ class TestBaseItemExporter(ABC):
def test_fields_to_export(self):
ie = self._get_exporter(fields_to_export=["name"])
assert list(ie._get_serialized_fields(self.i)) == [("name", "John\xa3")]
assert list(ie.get_serialized_fields(self.i)) == [("name", "John\xa3")]
ie = self._get_exporter(fields_to_export=["name"], encoding="latin-1")
_, name = next(iter(ie._get_serialized_fields(self.i)))
_, name = next(iter(ie.get_serialized_fields(self.i)))
assert isinstance(name, str)
assert name == "John\xa3"
ie = self._get_exporter(fields_to_export={"name": "名稱"})
assert list(ie._get_serialized_fields(self.i)) == [("名稱", "John\xa3")]
assert list(ie.get_serialized_fields(self.i)) == [("名稱", "John\xa3")]
def test_field_order(self):
item = self.item_class(age="22", name="John\xa3")
ie = self._get_exporter()
assert [name for name, _ in ie._get_serialized_fields(item)] == ["name", "age"]
assert [name for name, _ in ie.get_serialized_fields(item)] == ["name", "age"]
def test_field_order_dict_item(self):
ie = self._get_exporter()
assert [name for name, _ in ie._get_serialized_fields({"age": "22"})] == ["age"]
assert [name for name, _ in ie.get_serialized_fields({"age": "22"})] == ["age"]
assert [
name for name, _ in ie._get_serialized_fields({"age": "22", "name": "John"})
name for name, _ in ie.get_serialized_fields({"age": "22", "name": "John"})
] == ["age", "name"]
def test_field_custom_serializer(self):
@ -142,7 +143,7 @@ class TestBaseItemExporter(ABC):
class TestPythonItemExporter(TestBaseItemExporter):
def _get_exporter(self, **kwargs):
def _get_exporter(self, **kwargs: Any) -> BaseItemExporter:
return PythonItemExporter(**kwargs)
def test_invalid_option(self):
@ -173,6 +174,7 @@ class TestPythonItemExporter(TestBaseItemExporter):
"age": [{"age": [{"age": "22", "name": "Joseph"}], "name": "Maria"}],
"name": "Jesus",
}
assert exported is not None
assert isinstance(exported["age"][0], dict)
assert isinstance(exported["age"][0]["age"][0], dict)
@ -186,6 +188,7 @@ class TestPythonItemExporter(TestBaseItemExporter):
"age": [{"age": [{"age": "22", "name": "Joseph"}], "name": "Maria"}],
"name": "Jesus",
}
assert exported is not None
assert isinstance(exported["age"][0], dict)
assert isinstance(exported["age"][0]["age"][0], dict)
@ -202,10 +205,10 @@ class TestPythonItemExporterDataclass(TestPythonItemExporter):
class TestPprintItemExporter(TestBaseItemExporter):
def _get_exporter(self, **kwargs):
def _get_exporter(self, **kwargs: Any) -> BaseItemExporter:
return PprintItemExporter(self.output, **kwargs)
def _check_output(self):
def _check_output(self) -> None:
self._assert_expected_item(eval(self.output.getvalue()))
@ -215,10 +218,10 @@ class TestPprintItemExporterDataclass(TestPprintItemExporter):
class TestPickleItemExporter(TestBaseItemExporter):
def _get_exporter(self, **kwargs):
def _get_exporter(self, **kwargs: Any) -> BaseItemExporter:
return PickleItemExporter(self.output, **kwargs)
def _check_output(self):
def _check_output(self) -> None:
self._assert_expected_item(pickle.loads(self.output.getvalue()))
def test_export_multiple_items(self):
@ -252,10 +255,10 @@ class TestPickleItemExporterDataclass(TestPickleItemExporter):
class TestMarshalItemExporter(TestBaseItemExporter):
def _get_exporter(self, **kwargs):
def _get_exporter(self, **kwargs: Any) -> BaseItemExporter:
return MarshalItemExporter(self.output, **kwargs)
def _check_output(self):
def _check_output(self) -> None:
self.output.seek(0)
self._assert_expected_item(marshal.load(self.output))
@ -279,7 +282,7 @@ class TestMarshalItemExporterDataclass(TestMarshalItemExporter):
class TestCsvItemExporter(TestBaseItemExporter):
def _get_exporter(self, **kwargs):
def _get_exporter(self, **kwargs: Any) -> BaseItemExporter:
# We need a fresh instance for each exporter, because
# CsvItemExporter.stream.__del__() closes the underlying file
# (CsvItemExporter.finish_exporting() calls detach() but not all tests
@ -287,8 +290,10 @@ class TestCsvItemExporter(TestBaseItemExporter):
self.output = BytesIO()
return CsvItemExporter(self.output, **kwargs)
def assertCsvEqual(self, first, second, msg=None):
def split_csv(csv):
def assertCsvEqual(
self, first: bytes | str, second: bytes | str, msg: str | None = None
) -> None:
def split_csv(csv: bytes | str) -> list[list[str]]:
return [
sorted(re.split(r"(,|\s+)", line))
for line in to_unicode(csv).splitlines(True)
@ -296,13 +301,15 @@ class TestCsvItemExporter(TestBaseItemExporter):
assert split_csv(first) == split_csv(second), msg
def _check_output(self):
def _check_output(self) -> None:
self.output.seek(0)
self.assertCsvEqual(
to_unicode(self.output.read()), "age,name\r\n22,John\xa3\r\n"
)
def assertExportResult(self, item, expected, **kwargs):
def assertExportResult(
self, item: Any, expected: bytes | str = b"", **kwargs: Any
) -> None:
fp = BytesIO()
ie = CsvItemExporter(fp, **kwargs)
ie.start_exporting()
@ -383,7 +390,6 @@ class TestCsvItemExporter(TestBaseItemExporter):
with pytest.raises(UnicodeEncodeError):
self.assertExportResult(
item={"text": "W\u0275\u200brd"},
expected=None,
encoding="windows-1251",
)
@ -417,7 +423,7 @@ class TestCsvItemExporterDataclass(TestCsvItemExporter):
class TestXmlItemExporter(TestBaseItemExporter):
def _get_exporter(self, **kwargs):
def _get_exporter(self, **kwargs: Any) -> BaseItemExporter:
# We need a fresh instance for each exporter, because
# XmlItemExporter.stream.__del__() closes the underlying file
# (XmlItemExporter.finish_exporting() calls detach() but not all tests
@ -425,20 +431,22 @@ class TestXmlItemExporter(TestBaseItemExporter):
self.output = BytesIO()
return XmlItemExporter(self.output, **kwargs)
def assertXmlEquivalent(self, first, second, msg=None):
def xmltuple(elem):
def assertXmlEquivalent(
self, first: bytes, second: bytes, msg: str | None = None
) -> None:
def xmltuple(elem: Any) -> list[Any]:
children = list(elem.iterchildren())
if children:
return [(child.tag, sorted(xmltuple(child))) for child in children]
return [(elem.tag, [(elem.text, ())])]
def xmlsplit(xmlcontent):
def xmlsplit(xmlcontent: bytes) -> list[Any]:
doc = lxml.etree.fromstring(xmlcontent)
return xmltuple(doc)
assert xmlsplit(first) == xmlsplit(second), msg
def assertExportResult(self, item, expected_value):
def assertExportResult(self, item: Any, expected_value: bytes) -> None:
fp = BytesIO()
ie = XmlItemExporter(fp)
ie.start_exporting()
@ -447,7 +455,7 @@ class TestXmlItemExporter(TestBaseItemExporter):
del ie # See the first “del self.ie” in this file for context.
self.assertXmlEquivalent(fp.getvalue(), expected_value)
def _check_output(self):
def _check_output(self) -> None:
expected_value = (
b'<?xml version="1.0" encoding="utf-8"?>\n'
b"<items><item><age>22</age><name>John\xc2\xa3</name></item></items>"
@ -538,10 +546,10 @@ class TestJsonLinesItemExporter(TestBaseItemExporter):
"age": {"name": "Maria", "age": {"name": "Joseph", "age": "22"}},
}
def _get_exporter(self, **kwargs):
def _get_exporter(self, **kwargs: Any) -> BaseItemExporter:
return JsonLinesItemExporter(self.output, **kwargs)
def _check_output(self):
def _check_output(self) -> None:
exported = json.loads(to_unicode(self.output.getvalue().strip()))
assert exported == ItemAdapter(self.i).asdict()
@ -582,14 +590,14 @@ class TestJsonLinesItemExporterDataclass(TestJsonLinesItemExporter):
class TestJsonItemExporter(TestJsonLinesItemExporter):
_expected_nested = [TestJsonLinesItemExporter._expected_nested]
def _get_exporter(self, **kwargs):
def _get_exporter(self, **kwargs: Any) -> BaseItemExporter:
return JsonItemExporter(self.output, **kwargs)
def _check_output(self):
def _check_output(self) -> None:
exported = json.loads(to_unicode(self.output.getvalue().strip()))
assert exported == [ItemAdapter(self.i).asdict()]
def assertTwoItemsExported(self, item):
def assertTwoItemsExported(self, item: Any) -> None:
self.ie.start_exporting()
self.ie.export_item(item)
self.ie.export_item(item)
@ -658,7 +666,7 @@ class TestJsonItemExporter(TestJsonLinesItemExporter):
class TestJsonItemExporterToBytes(TestBaseItemExporter):
def _get_exporter(self, **kwargs):
def _get_exporter(self, **kwargs: Any) -> BaseItemExporter:
kwargs["encoding"] = "latin"
return JsonItemExporter(self.output, **kwargs)
@ -690,7 +698,9 @@ class TestCustomExporterItem:
def test_exporter_custom_serializer(self):
class CustomItemExporter(BaseItemExporter):
def serialize_field(self, field, name, value):
def serialize_field(
self, field: Mapping[str, Any] | Field, name: str, value: Any
) -> Any:
if name == "age":
return str(int(value) + 1)
return super().serialize_field(field, name, value)

View File

@ -134,6 +134,8 @@ class InstrumentedFeedSlot(FeedSlot):
"""Instrumented FeedSlot subclass for keeping track of calls to
start_exporting and finish_exporting."""
update_listener: Callable[[str], None]
def start_exporting(self):
self.update_listener("start")
super().start_exporting()
@ -143,7 +145,7 @@ class InstrumentedFeedSlot(FeedSlot):
super().finish_exporting()
@classmethod
def subscribe__listener(cls, listener):
def subscribe__listener(cls, listener: IsExportingListener) -> None:
cls.update_listener = listener.update
@ -153,7 +155,7 @@ class IsExportingListener:
finish_exporting and when a call to finish_exporting has been made
before a call to start_exporting."""
def __init__(self):
def __init__(self) -> None:
self.start_without_finish = False
self.finish_without_start = False
@ -341,6 +343,7 @@ class TestFeedExport(TestFeedExportBase):
}
crawler = get_crawler(ItemSpider, settings)
yield crawler.crawl(mockserver=self.mockserver)
assert crawler.stats is not None
assert "feedexport/success_count/FileFeedStorage" in crawler.stats.get_stats()
assert crawler.stats.get_value("feedexport/success_count/FileFeedStorage") == 1
@ -364,6 +367,7 @@ class TestFeedExport(TestFeedExportBase):
side_effect=store,
):
yield crawler.crawl(mockserver=self.mockserver)
assert crawler.stats is not None
assert "feedexport/failed_count/FileFeedStorage" in crawler.stats.get_stats()
assert crawler.stats.get_value("feedexport/failed_count/FileFeedStorage") == 1
@ -381,6 +385,7 @@ class TestFeedExport(TestFeedExportBase):
}
crawler = get_crawler(ItemSpider, settings)
yield crawler.crawl(mockserver=self.mockserver)
assert crawler.stats is not None
assert "feedexport/success_count/FileFeedStorage" in crawler.stats.get_stats()
assert "feedexport/success_count/StdoutFeedStorage" in crawler.stats.get_stats()
assert crawler.stats.get_value("feedexport/success_count/FileFeedStorage") == 1
@ -521,7 +526,7 @@ class TestFeedExport(TestFeedExportBase):
@coroutine_test
async def test_start_finish_exporting_no_items(self):
items = []
items: list[Any] = []
settings = {
"FEEDS": {
self._random_temp_filename(): {"format": "json"},
@ -560,7 +565,7 @@ class TestFeedExport(TestFeedExportBase):
@coroutine_test
async def test_start_finish_exporting_no_items_exception(self):
items = []
items: list[Any] = []
settings = {
"FEEDS": {
self._random_temp_filename(): {"format": "json"},
@ -645,7 +650,7 @@ class TestFeedExport(TestFeedExportBase):
items = [{"foo": "bar"}]
header = ["foo"]
rows = [{"foo": "bar"}]
settings = {"FEED_EXPORT_FIELDS": []}
settings: dict[str, Any] = {"FEED_EXPORT_FIELDS": []}
await self.assertExportedCsv(items, header, rows)
await self.assertExportedJsonLines(items, rows, settings)
@ -761,14 +766,14 @@ class TestFeedExport(TestFeedExportBase):
def accepts(self, item):
return isinstance(item, MyItem)
class CustomFilter2(scrapy.extensions.feedexport.ItemFilter):
class CustomFilter2(ItemFilter):
def accepts(self, item):
return "foo" in item.fields
class CustomFilter3(scrapy.extensions.feedexport.ItemFilter):
class CustomFilter3(ItemFilter):
def accepts(self, item):
return (
isinstance(item, tuple(self.item_classes)) and item["foo"] == "bar1"
isinstance(item, tuple(self.item_classes)) and item["foo"] == "bar1" # type: ignore[index]
)
formats = {
@ -868,7 +873,7 @@ class TestFeedExport(TestFeedExportBase):
}
for fmt, expected in formats.items():
settings = {
settings: dict[str, Any] = {
"FEEDS": {
self._random_temp_filename(): {"format": fmt},
},
@ -945,7 +950,7 @@ class TestFeedExport(TestFeedExportBase):
{"key": "value"},
]
test_cases = [
test_cases: list[dict[str, Any]] = [
# JSON
{
"format": "json",
@ -1166,7 +1171,7 @@ class TestFeedExport(TestFeedExportBase):
expected_with_title_csv = b"foo,bar\r\nFOO,BAR\r\n"
expected_without_title_csv = b"FOO,BAR\r\n"
test_cases = [
test_cases: list[dict[str, Any]] = [
# with title
{
"options": {
@ -1200,6 +1205,9 @@ class TestFeedExport(TestFeedExportBase):
@coroutine_test
async def test_storage_file_no_postprocessing(self):
class Storage:
open_file: IO[bytes]
store_file: IO[bytes]
def __init__(self, uri, *, feed_options=None):
pass
@ -1221,6 +1229,10 @@ class TestFeedExport(TestFeedExportBase):
@coroutine_test
async def test_storage_file_postprocessing(self):
class Storage:
open_file: IO[bytes]
store_file: IO[bytes]
file_was_closed: bool
def __init__(self, uri, *, feed_options=None):
pass
@ -1616,7 +1628,7 @@ class TestItemFilter:
class TestFeedExportInit:
def test_unsupported_storage(self):
settings = {
settings: dict[str, Any] = {
"FEEDS": {
"unsupported://uri": {},
},

View File

@ -74,7 +74,13 @@ class TestFeedPostProcessedExports(TestFeedExportBase):
return content
def get_gzip_compressed(self, data, compresslevel=9, mtime=0, filename=""):
def get_gzip_compressed(
self,
data: bytes,
compresslevel: int = 9,
mtime: int = 0,
filename: str = "",
) -> bytes:
data_stream = BytesIO()
gzipf = gzip.GzipFile(
fileobj=data_stream,
@ -539,11 +545,13 @@ class TestFeedPostProcessedExports(TestFeedExportBase):
data = await self.exported_data(self.items, settings)
for filename, result in data.items():
for filename, data_bytes in data.items():
expected: Any
result: Any
if "pickle" in filename:
expected, result = self.items[0], pickle.loads(result)
expected, result = self.items[0], pickle.loads(data_bytes)
elif "marshal" in filename:
expected, result = self.items[0], marshal.loads(result)
expected, result = self.items[0], marshal.loads(data_bytes)
else:
expected = filename_to_expected[filename]
expected, result = filename_to_expected[filename], data_bytes
assert result == expected

View File

@ -128,27 +128,34 @@ class TestFileFeedStorage:
assert storage.path == path
def get_test_spider(settings: dict[str, Any] | None = None) -> scrapy.Spider:
class TestSpider(scrapy.Spider):
name = "test_spider"
crawler = get_crawler(settings_dict=settings)
return TestSpider.from_crawler(crawler)
class TestFTPFeedStorage:
def get_test_spider(self, settings=None):
class TestSpider(scrapy.Spider):
name = "test_spider"
crawler = get_crawler(settings_dict=settings)
return TestSpider.from_crawler(crawler)
async def _store(self, uri, content, feed_options=None, settings=None):
async def _store(
self,
uri: str,
content: bytes,
feed_options: dict[str, Any] | None = None,
settings: dict[str, Any] | None = None,
) -> None:
crawler = get_crawler(settings_dict=settings or {})
storage = FTPFeedStorage.from_crawler(
crawler,
uri,
feed_options=feed_options,
)
spider = self.get_test_spider()
spider = get_test_spider()
file = storage.open(spider)
file.write(content)
await maybe_deferred_to_future(storage.store(file))
def _assert_stored(self, path: Path, content, unlink: bool = True):
def _assert_stored(self, path: Path, content: bytes, unlink: bool = True) -> None:
assert path.exists()
try:
assert path.read_bytes() == content
@ -218,7 +225,7 @@ class TestFTPFeedStorage:
def test_uri_auth_quote(self):
# RFC3986: 3.2.1. User Information
pw_quoted = quote(string.punctuation, safe="")
st = FTPFeedStorage(f"ftp://foo:{pw_quoted}@example.com/some_path", {})
st = FTPFeedStorage(f"ftp://foo:{pw_quoted}@example.com/some_path")
assert st.password == string.punctuation
def test_uri_without_hostname(self):
@ -234,24 +241,17 @@ class MyBlockingFeedStorage(BlockingFeedStorage):
class TestBlockingFeedStorage:
def get_test_spider(self, settings=None):
class TestSpider(scrapy.Spider):
name = "test_spider"
crawler = get_crawler(settings_dict=settings)
return TestSpider.from_crawler(crawler)
def test_default_temp_dir(self):
b = MyBlockingFeedStorage()
storage_file = b.open(self.get_test_spider())
storage_file = b.open(get_test_spider())
storage_dir = Path(storage_file.name).parent
assert str(storage_dir) == tempfile.gettempdir()
def test_temp_file(self, tmp_path):
b = MyBlockingFeedStorage()
spider = self.get_test_spider({"FEED_TEMPDIR": str(tmp_path)})
spider = get_test_spider({"FEED_TEMPDIR": str(tmp_path)})
storage_file = b.open(spider)
storage_dir = Path(storage_file.name).parent
assert storage_dir == tmp_path
@ -260,7 +260,7 @@ class TestBlockingFeedStorage:
b = MyBlockingFeedStorage()
invalid_path = tmp_path / "invalid_path"
spider = self.get_test_spider({"FEED_TEMPDIR": str(invalid_path)})
spider = get_test_spider({"FEED_TEMPDIR": str(invalid_path)})
with pytest.raises(OSError, match="Not a Directory:"):
b.open(spider=spider)
@ -364,7 +364,7 @@ class TestS3FeedStorage:
assert storage.access_key == "access_key"
assert storage.secret_key == "secret_key"
assert storage.region_name == region_name
assert storage.s3_client._client_config.region_name == region_name
assert storage.s3_client._client_config.region_name == region_name # type: ignore[attr-defined]
def test_from_crawler_without_acl(self):
settings = {
@ -406,7 +406,7 @@ class TestS3FeedStorage:
)
assert storage.access_key == "access_key"
assert storage.secret_key == "secret_key"
assert storage.s3_client._client_config.region_name == "us-east-1"
assert storage.s3_client._client_config.region_name == "us-east-1" # type: ignore[attr-defined]
def test_from_crawler_with_acl(self):
settings = {
@ -447,7 +447,7 @@ class TestS3FeedStorage:
assert storage.access_key == "access_key"
assert storage.secret_key == "secret_key"
assert storage.region_name == region_name
assert storage.s3_client._client_config.region_name == region_name
assert storage.s3_client._client_config.region_name == region_name # type: ignore[attr-defined]
def test_init_without_max_pool_connections(self) -> None:
storage = S3FeedStorage("s3://mybucket/export.csv", "access_key", "secret_key")
@ -607,7 +607,7 @@ class TestGCSFeedStorage:
def test_parse_empty_acl(self):
pytest.importorskip("google.cloud.storage")
settings = {"GCS_PROJECT_ID": "123", "FEED_STORAGE_GCS_ACL": ""}
settings: dict[str, Any] = {"GCS_PROJECT_ID": "123", "FEED_STORAGE_GCS_ACL": ""}
crawler = get_crawler(settings_dict=settings)
storage = GCSFeedStorage.from_crawler(crawler, "gs://mybucket/export.csv")
assert storage.acl is None

View File

@ -2,6 +2,7 @@ from __future__ import annotations
import warnings
from abc import ABC, abstractmethod
from typing import TYPE_CHECKING, Any
import pytest
@ -10,16 +11,27 @@ from scrapy.exceptions import ScrapyDeprecationWarning
from scrapy.extensions.feedexport import FeedExporter
from scrapy.utils.test import get_crawler
if TYPE_CHECKING:
from collections.abc import Callable
from scrapy.crawler import Crawler
class TestURIParams(ABC):
spider_name = "uri_params_spider"
deprecated_options = False
@abstractmethod
def build_settings(self, uri="file:///tmp/foobar", uri_params=None):
def build_settings(
self,
uri: str = "file:///tmp/foobar",
uri_params: Callable[..., dict[str, Any] | None] | None = None,
) -> dict[str, Any]:
raise NotImplementedError
def _crawler_feed_exporter(self, settings):
def _crawler_feed_exporter(
self, settings: dict[str, Any]
) -> tuple[Crawler, FeedExporter]:
if self.deprecated_options:
with pytest.warns(
ScrapyDeprecationWarning,
@ -29,6 +41,7 @@ class TestURIParams(ABC):
else:
crawler = get_crawler(settings_dict=settings)
feed_exporter = crawler.get_extension(FeedExporter)
assert feed_exporter is not None
return crawler, feed_exporter
def test_default(self):
@ -116,8 +129,12 @@ class TestURIParams(ABC):
class TestURIParamsSetting(TestURIParams):
deprecated_options = True
def build_settings(self, uri="file:///tmp/foobar", uri_params=None):
extra_settings = {}
def build_settings(
self,
uri: str = "file:///tmp/foobar",
uri_params: Callable[..., dict[str, Any] | None] | None = None,
) -> dict[str, Any]:
extra_settings: dict[str, Any] = {}
if uri_params:
extra_settings["FEED_URI_PARAMS"] = uri_params
return {
@ -129,8 +146,12 @@ class TestURIParamsSetting(TestURIParams):
class TestURIParamsFeedOption(TestURIParams):
deprecated_options = False
def build_settings(self, uri="file:///tmp/foobar", uri_params=None):
options = {
def build_settings(
self,
uri: str = "file:///tmp/foobar",
uri_params: Callable[..., dict[str, Any] | None] | None = None,
) -> dict[str, Any]:
options: dict[str, Any] = {
"format": "jl",
}
if uri_params:

View File

@ -1,4 +1,5 @@
from abc import ABCMeta
from typing import Any
from unittest import mock
import pytest
@ -7,9 +8,6 @@ from scrapy.item import Field, Item, ItemMeta
class TestItem:
def assertSortedEqual(self, first, second, msg=None):
assert sorted(first) == sorted(second), msg
def test_simple(self):
class TestItem(Item):
name = Field()
@ -98,16 +96,16 @@ class TestItem:
i = TestItem()
with pytest.raises(AttributeError):
i.name = "john"
i.name = "john" # type: ignore[assignment]
def test_custom_methods(self):
class TestItem(Item):
name = Field()
def get_name(self):
def get_name(self) -> Any:
return self["name"]
def change_name(self, name):
def change_name(self, name: str) -> None:
self["name"] = name
i = TestItem()
@ -121,40 +119,40 @@ class TestItem:
def test_metaclass(self):
class TestItem(Item):
name = Field()
keys = Field()
values = Field()
keys = Field() # type: ignore[assignment]
values = Field() # type: ignore[assignment]
i = TestItem()
i["name"] = "John"
assert list(i.keys()) == ["name"]
assert list(i.values()) == ["John"]
assert list(i.keys()) == ["name"] # type: ignore[operator]
assert list(i.values()) == ["John"] # type: ignore[operator]
i["keys"] = "Keys"
i["values"] = "Values"
self.assertSortedEqual(list(i.keys()), ["keys", "values", "name"])
self.assertSortedEqual(list(i.values()), ["Keys", "Values", "John"])
assert sorted(i.keys()) == ["keys", "name", "values"] # type: ignore[operator]
assert sorted(i.values()) == ["John", "Keys", "Values"] # type: ignore[operator]
def test_metaclass_with_fields_attribute(self):
class TestItem(Item):
fields = {"new": Field(default="X")}
item = TestItem(new="New")
self.assertSortedEqual(list(item.keys()), ["new"])
self.assertSortedEqual(list(item.values()), ["New"])
assert list(item.keys()) == ["new"]
assert list(item.values()) == ["New"]
def test_fields_order(self):
class TestItem(Item):
name = Field()
keys = Field()
values = Field()
keys = Field() # type: ignore[assignment]
values = Field() # type: ignore[assignment]
assert list(TestItem.fields) == ["name", "keys", "values"]
def test_fields_order_inheritance(self):
class ParentItem(Item):
name = Field()
keys = Field()
values = Field()
keys = Field() # type: ignore[assignment]
values = Field() # type: ignore[assignment]
class TestItem(ParentItem):
extra = Field()
@ -169,16 +167,16 @@ class TestItem:
def test_metaclass_inheritance(self):
class ParentItem(Item):
name = Field()
keys = Field()
values = Field()
keys = Field() # type: ignore[assignment]
values = Field() # type: ignore[assignment]
class TestItem(ParentItem):
keys = Field()
i = TestItem()
i["keys"] = 3
assert list(i.keys()) == ["keys"]
assert list(i.values()) == [3]
assert list(i.keys()) == ["keys"] # type: ignore[operator]
assert list(i.values()) == [3] # type: ignore[operator]
def test_metaclass_multiple_inheritance_simple(self):
class A(Item):
@ -314,7 +312,7 @@ class TestItemMeta:
def f(self):
# For rationale of this see:
# https://github.com/python/cpython/blob/ee1a81b77444c6715cbe610e951c655b6adab88b/Lib/test/test_super.py#L222
return __class__
return __class__ # type: ignore[name-defined]
MyItem()

View File

@ -9,6 +9,7 @@ from w3lib import __version__ as w3lib_version
from scrapy.http import HtmlResponse, XmlResponse
from scrapy.link import Link
from scrapy.linkextractors import lxmlhtml
from scrapy.linkextractors.lxmlhtml import LxmlLinkExtractor, LxmlParserLinkExtractor
from tests import get_testdata
@ -798,6 +799,25 @@ class Base:
class TestLxmlLinkExtractor(Base.TestLinkExtractorBase):
extractor_cls = LxmlLinkExtractor
def test_canonicalize_once_per_link(self, monkeypatch):
canonicalize_url = lxmlhtml.canonicalize_url
calls = []
def counting_canonicalize_url(url, *args, **kwargs):
calls.append(url)
return canonicalize_url(url, *args, **kwargs)
monkeypatch.setattr(lxmlhtml, "canonicalize_url", counting_canonicalize_url)
response = HtmlResponse(
"https://example.com",
body=b"".join(b'<a href="/p?b=2&a=1#f%d">x</a>' % i for i in range(10)),
)
lx = self.extractor_cls(canonicalize=True)
assert lx.extract_links(response) == [
Link(url="https://example.com/p?a=1&b=2", text="x")
]
assert len(calls) == 10
def test_link_restrict_text(self):
html = b"""
<a href="http://example.org/item1.html">Pic of a cat</a>

View File

@ -79,7 +79,7 @@ class TestBasicItemLoader:
class InitializationTestMixin:
item_class: type | None = None
item_class: type
def test_keep_single_value(self):
"""Loaded item should contain values from the initial item"""
@ -311,7 +311,7 @@ class TestSelectortemLoader:
def test_init_method_with_base_response(self):
"""Selector should be None after initialization"""
response = Response("https://scrapy.org")
l = ProcessorItemLoader(response=response)
l = ProcessorItemLoader(response=response) # type: ignore[arg-type]
assert l.selector is None
def test_init_method_with_response(self):
@ -461,6 +461,7 @@ class TestSubselectorLoader:
l = NestedItemLoader(response=self.response)
nl = l.nested_xpath("//header")
assert nl.selector is not None
nl.add_xpath("name", "div/text()")
nl.add_css("name_div", "#id")
nl.add_value("name_value", nl.selector.xpath('div[@id = "id"]/text()').getall())
@ -476,6 +477,7 @@ class TestSubselectorLoader:
def test_nested_css(self):
l = NestedItemLoader(response=self.response)
nl = l.nested_css("header")
assert nl.selector is not None
nl.add_xpath("name", "div/text()")
nl.add_css("name_div", "#id")
nl.add_value("name_value", nl.selector.xpath('div[@id = "id"]/text()').getall())

View File

@ -23,7 +23,10 @@ if TYPE_CHECKING:
class MediaDownloadSpider(SimpleSpider):
name = "mediadownload"
def _process_url(self, url):
media_key: str
media_urls_key: str
def _process_url(self, url: str) -> str:
return url
def parse(self, response):
@ -44,14 +47,15 @@ class MediaDownloadSpider(SimpleSpider):
class BrokenLinksMediaDownloadSpider(MediaDownloadSpider):
name = "brokenmedia"
def _process_url(self, url):
def _process_url(self, url: str) -> str:
return url + ".foo"
class RedirectedMediaDownloadSpider(MediaDownloadSpider):
name = "redirectedmedia"
def _process_url(self, url):
def _process_url(self, url: str) -> str:
assert self.mockserver
return add_or_replace_parameter(
self.mockserver.url("/redirect-to"), "goto", url
)

View File

@ -1,6 +1,7 @@
import base64
import dataclasses
import logging
import mimetypes
import random
import re
import time
@ -22,6 +23,7 @@ from itemadapter import ItemAdapter
from twisted.internet.defer import Deferred
from twisted.python.failure import Failure
from scrapy.crawler import Crawler
from scrapy.exceptions import IgnoreRequest, NotConfigured
from scrapy.http import Request, Response
from scrapy.item import Field, Item
@ -75,7 +77,7 @@ class DeferredFSFilesStore(FSFilesStore):
"""A simple store with persist_file() returning a deferred."""
def persist_file(self, path, buf, info, meta=None, headers=None):
deferred = Deferred()
deferred: Deferred[None] = Deferred()
# short-hand super() doesn't work in nested functions
parent_persist_file = super().persist_file
@ -95,8 +97,14 @@ class TestFilesPipeline:
def teardown_method(self):
rmtree(self.tempdir)
def _create_pipeline(self, pipeline_cls: type[FilesPipeline]) -> FilesPipeline:
crawler = get_crawler(DefaultSpider, {"FILES_STORE": self.tempdir})
def _create_pipeline(
self,
pipeline_cls: type[FilesPipeline],
settings: dict[str, Any] | None = None,
) -> FilesPipeline:
crawler = get_crawler(
DefaultSpider, {"FILES_STORE": self.tempdir, **(settings or {})}
)
crawler.spider = crawler._create_spider()
crawler.engine = MagicMock(download_async=mocked_download_func)
pipeline = pipeline_cls.from_crawler(crawler)
@ -152,7 +160,7 @@ class TestFilesPipeline:
file_path(
Request("http://www.dorma.co.uk/images/product_details/2532"),
response=Response("http://www.dorma.co.uk/images/product_details/2532"),
info=object(),
info=object(), # type: ignore[arg-type]
)
== "full/244e0dd7d96a3b7b01f54eded250c9e272577aa1"
)
@ -383,7 +391,7 @@ class TestFilesPipeline:
"""
class CustomFilesPipeline(FilesPipeline):
def file_path(self, request, response=None, info=None, item=None):
def file_path(self, request, response=None, info=None, item=None) -> str:
return f"full/{item.get('path')}"
file_path = CustomFilesPipeline.from_crawler(
@ -393,6 +401,47 @@ class TestFilesPipeline:
request = Request("http://example.com")
assert file_path(request, item=item) == "full/path-to-store-file"
@coroutine_test
async def test_file_path_from_response(self) -> None:
"""file_path() may build the path out of the response, e.g. to get the
file extension from a response header, as long as FILES_EXPIRES is 0 to
disable the up-to-date check, which runs before the download and hence
cannot reach the same path."""
class ContentTypeFilesPipeline(FilesPipeline):
def file_path(self, request, response=None, info=None, *, item=None):
path = super().file_path(request, response, info, item=item)
if response is None:
return path
content_type = response.headers["Content-Type"].decode()
return path + (mimetypes.guess_extension(content_type) or "")
item_url = "http://example.com/download?id=1"
item = _create_item_with_files(item_url)
pipeline = self._create_pipeline(ContentTypeFilesPipeline, {"FILES_EXPIRES": 0})
request = _prepare_request_object(
item_url, headers={"Content-Type": "application/pdf"}
)
with (
mock.patch.object(FilesPipeline, "inc_stats", return_value=True),
# A fresh file at the response-less path is ignored thanks to
# FILES_EXPIRES being 0.
mock.patch.object(
FSFilesStore,
"stat_file",
return_value={"checksum": "abc", "last_modified": time.time()},
),
mock.patch.object(
FilesPipeline, "get_media_requests", return_value=[request]
),
):
result = await pipeline.process_item(item)
file_info = result["files"][0]
assert file_info["status"] == "downloaded"
assert file_info["path"].endswith(".pdf")
assert (Path(self.tempdir) / file_info["path"]).read_bytes() == b"data"
def test_media_failed_filtered_request(
self, caplog: pytest.LogCaptureFixture
) -> None:
@ -476,7 +525,7 @@ class TestFilesPipeline:
item["file_urls"] = bad_type
with pytest.raises(TypeError, match="file_urls must be a list of URLs"):
list(pipeline.get_media_requests(item, None))
list(pipeline.get_media_requests(item, None)) # type: ignore[arg-type]
class TestFilesPipelineFieldsMixin(ABC):
@ -491,10 +540,10 @@ class TestFilesPipelineFieldsMixin(ABC):
pipeline = FilesPipeline.from_crawler(
get_crawler(None, {"FILES_STORE": tmp_path})
)
requests = list(pipeline.get_media_requests(item, None))
requests = list(pipeline.get_media_requests(item, None)) # type: ignore[arg-type]
assert requests[0].url == url
results = [(True, {"url": url})]
item = pipeline.item_completed(results, item, None)
item = pipeline.item_completed(results, item, None) # type: ignore[arg-type]
files = ItemAdapter(item).get("files")
assert files == [results[0][1]]
assert isinstance(item, self.item_class)
@ -512,10 +561,10 @@ class TestFilesPipelineFieldsMixin(ABC):
},
)
)
requests = list(pipeline.get_media_requests(item, None))
requests = list(pipeline.get_media_requests(item, None)) # type: ignore[arg-type]
assert requests[0].url == url
results = [(True, {"url": url})]
item = pipeline.item_completed(results, item, None)
item = pipeline.item_completed(results, item, None) # type: ignore[arg-type]
custom_files = ItemAdapter(item).get("custom_files")
assert custom_files == [results[0][1]]
assert isinstance(item, self.item_class)
@ -581,8 +630,10 @@ class TestFilesPipelineCustomSettings:
("FILES_RESULT_FIELD", "FILES_RESULT_FIELD", "files_result_field"),
}
def _generate_fake_settings(self, tmp_path, prefix=None):
def random_string():
def _generate_fake_settings(
self, tmp_path: Path, prefix: str | None = None
) -> dict[str, Any]:
def random_string() -> str:
return "".join([chr(random.randint(97, 123)) for _ in range(10)])
settings = {
@ -599,7 +650,7 @@ class TestFilesPipelineCustomSettings:
for k, v in settings.items()
}
def _generate_fake_pipeline(self):
def _generate_fake_pipeline(self) -> type[FilesPipeline]:
class UserDefinedFilePipeline(FilesPipeline):
EXPIRES = 1001
FILES_URLS_FIELD = "alfa"
@ -739,14 +790,14 @@ class TestFilesPipelineCustomSettings:
def test_file_pipeline_using_pathlike_objects(self, tmp_path):
class CustomFilesPipelineWithPathLikeDir(FilesPipeline):
def file_path(self, request, response=None, info=None, *, item=None):
return Path("subdir") / Path(request.url).name
def file_path(self, request, response=None, info=None, *, item=None) -> str:
return str(Path("subdir") / Path(request.url).name)
pipeline = CustomFilesPipelineWithPathLikeDir.from_crawler(
get_crawler(None, {"FILES_STORE": tmp_path})
)
request = Request("http://example.com/image01.jpg")
assert pipeline.file_path(request) == Path("subdir/image01.jpg")
assert pipeline.file_path(request) == str(Path("subdir/image01.jpg"))
class TestFSFilesStore:
@ -1092,7 +1143,7 @@ class TestFTPFileStore:
store.port,
store.username,
store.password,
store.USE_ACTIVE_MODE,
bool(store.USE_ACTIVE_MODE),
)
assert data == content
@ -1130,10 +1181,18 @@ def _create_item_with_files(*files: str) -> ItemWithFiles:
return item
def _prepare_request_object(item_url: str, flags: list[str] | None = None) -> Request:
def _prepare_request_object(
item_url: str,
flags: list[str] | None = None,
headers: dict[str, str] | None = None,
) -> Request:
return Request(
item_url,
meta={"response": Response(item_url, status=200, body=b"data", flags=flags)},
meta={
"response": Response(
item_url, status=200, body=b"data", flags=flags, headers=headers
)
},
)
@ -1160,7 +1219,7 @@ class TestBuildFromCrawler:
_from_crawler_called = False
@classmethod
def from_crawler(cls, crawler):
def from_crawler(cls, crawler: Crawler) -> "Pipeline":
settings = crawler.settings
store_uri = settings["FILES_STORE"]
o = cls(store_uri, crawler=crawler)
@ -1179,7 +1238,7 @@ def test_files_pipeline_raises_notconfigured_when_files_store_invalid(store):
settings = Settings()
settings.clear()
settings.set("FILES_STORE", store, priority="cmdline")
crawler = get_crawler(settings_dict=settings)
crawler = get_crawler(settings_dict=dict(settings))
with pytest.raises(NotConfigured):
FilesPipeline.from_crawler(crawler)

View File

@ -91,7 +91,7 @@ class TestImagesPipeline:
file_path(
Request("http://www.dorma.co.uk/images/product_details/2532"),
response=Response("http://www.dorma.co.uk/images/product_details/2532"),
info=object(),
info=DUMMY_SPIDER_INFO,
)
== "full/244e0dd7d96a3b7b01f54eded250c9e272577aa1.jpg"
)
@ -120,7 +120,7 @@ class TestImagesPipeline:
Request("file:///tmp/some.name/foo"),
name,
response=Response("file:///tmp/some.name/foo"),
info=object(),
info=DUMMY_SPIDER_INFO,
)
== "thumbs/50/850233df65a5b83361798f532f1fc549cd13cbe9.jpg"
)
@ -133,7 +133,7 @@ class TestImagesPipeline:
class CustomImagesPipeline(ImagesPipeline):
def thumb_path(
self, request, thumb_id, response=None, info=None, item=None
):
) -> str:
return f"thumb/{thumb_id}/{item.get('path')}"
thumb_path = CustomImagesPipeline.from_crawler(
@ -159,11 +159,23 @@ class TestImagesPipeline:
req = Request(url="https://dev.mydeco.com/mydeco.gif")
with pytest.raises(ImageException):
next(self.pipeline.get_images(response=resp1, request=req, info=object()))
next(
self.pipeline.get_images(
response=resp1, request=req, info=DUMMY_SPIDER_INFO
)
)
with pytest.raises(ImageException):
next(self.pipeline.get_images(response=resp2, request=req, info=object()))
next(
self.pipeline.get_images(
response=resp2, request=req, info=DUMMY_SPIDER_INFO
)
)
with pytest.raises(ImageException):
next(self.pipeline.get_images(response=resp3, request=req, info=object()))
next(
self.pipeline.get_images(
response=resp3, request=req, info=DUMMY_SPIDER_INFO
)
)
def test_get_images(self):
self.pipeline.min_width = 0
@ -176,7 +188,7 @@ class TestImagesPipeline:
req = Request(url="https://dev.mydeco.com/mydeco.gif")
get_images_gen = self.pipeline.get_images(
response=resp, request=req, info=object()
response=resp, request=req, info=DUMMY_SPIDER_INFO
)
path, new_im, new_buf = next(get_images_gen)
@ -201,7 +213,7 @@ class TestImagesPipeline:
req = Request(url="https://dev.mydeco.com/mydeco.gif")
get_images_gen = self.pipeline.get_images(
response=resp, request=req, info=object()
response=resp, request=req, info=DUMMY_SPIDER_INFO
)
path, new_im, _ = next(get_images_gen)
@ -230,7 +242,7 @@ class TestImagesPipeline:
def test_convert_image(self):
SIZE = (100, 100)
# straight forward case: RGB and JPEG
COLOUR = (0, 127, 255)
COLOUR: tuple[int, ...] = (0, 127, 255)
im, buf = _create_image("JPEG", "RGB", SIZE, COLOUR)
converted, converted_buf = self.pipeline.convert_image(im, response_body=buf)
assert converted.mode == "RGB"
@ -296,7 +308,7 @@ class TestImagesPipeline:
item["image_urls"] = bad_type
with pytest.raises(TypeError, match="image_urls must be a list of URLs"):
list(pipeline.get_media_requests(item, None))
list(pipeline.get_media_requests(item, DUMMY_SPIDER_INFO))
class TestImagesPipelineFieldsMixin(ABC):
@ -311,10 +323,10 @@ class TestImagesPipelineFieldsMixin(ABC):
pipeline = ImagesPipeline.from_crawler(
get_crawler(None, {"IMAGES_STORE": "s3://example/images/"})
)
requests = list(pipeline.get_media_requests(item, None))
requests = list(pipeline.get_media_requests(item, DUMMY_SPIDER_INFO))
assert requests[0].url == url
results = [(True, {"url": url})]
item = pipeline.item_completed(results, item, None)
results: Any = [(True, {"url": url})]
item = pipeline.item_completed(results, item, DUMMY_SPIDER_INFO)
images = ItemAdapter(item).get("images")
assert images == [results[0][1]]
assert isinstance(item, self.item_class)
@ -332,10 +344,10 @@ class TestImagesPipelineFieldsMixin(ABC):
},
)
)
requests = list(pipeline.get_media_requests(item, None))
requests = list(pipeline.get_media_requests(item, DUMMY_SPIDER_INFO))
assert requests[0].url == url
results = [(True, {"url": url})]
item = pipeline.item_completed(results, item, None)
results: Any = [(True, {"url": url})]
item = pipeline.item_completed(results, item, DUMMY_SPIDER_INFO)
custom_images = ItemAdapter(item).get("custom_images")
assert custom_images == [results[0][1]]
assert isinstance(item, self.item_class)
@ -410,13 +422,15 @@ class TestImagesPipelineCustomSettings:
"IMAGES_RESULT_FIELD": "images",
}
def _generate_fake_settings(self, tmp_path, prefix=None):
def _generate_fake_settings(
self, tmp_path: Path, prefix: str | None = None
) -> dict[str, Any]:
"""
:param prefix: string for setting keys
:return: dictionary of image pipeline settings
"""
def random_string():
def random_string() -> str:
return "".join([chr(random.randint(97, 123)) for _ in range(10)])
settings = {
@ -439,7 +453,7 @@ class TestImagesPipelineCustomSettings:
for k, v in settings.items()
}
def _generate_fake_pipeline_subclass(self):
def _generate_fake_pipeline_subclass(self) -> type[ImagesPipeline]:
"""
:return: ImagePipeline class will all uppercase attributes set.
"""

View File

@ -1,6 +1,7 @@
from __future__ import annotations
import logging
from typing import TYPE_CHECKING, Any, cast
from unittest.mock import MagicMock
import pytest
@ -10,7 +11,12 @@ from scrapy import signals
from scrapy.exceptions import ScrapyDeprecationWarning
from scrapy.http import Request, Response
from scrapy.pipelines.files import FileException
from scrapy.pipelines.media import MediaPipeline, _MediaRequestFiltered
from scrapy.pipelines.media import (
FileInfo,
FileInfoOrError,
MediaPipeline,
_MediaRequestFiltered,
)
from scrapy.utils.defer import _defer_sleep_async
from scrapy.utils.log import failure_to_exc_info
from scrapy.utils.signal import disconnect_all
@ -19,21 +25,48 @@ from scrapy.utils.test import get_crawler
from tests.utils.decorators import coroutine_test
from tests.utils.media_pipelines import mocked_download_func
if TYPE_CHECKING:
from collections.abc import Awaitable
from twisted.internet.defer import Deferred
from scrapy.crawler import Crawler
class UserDefinedPipeline(MediaPipeline):
def media_to_download(self, request, info, *, item=None):
pass
def media_to_download(
self, request: Request, info: MediaPipeline.SpiderInfo, *, item: Any = None
) -> Deferred[FileInfo | None] | None:
return None
def get_media_requests(self, item, info):
pass
def get_media_requests(
self, item: Any, info: MediaPipeline.SpiderInfo
) -> list[Request]:
return []
def media_downloaded(self, response, request, info, *, item=None):
return {}
def media_downloaded(
self,
response: Response,
request: Request,
info: MediaPipeline.SpiderInfo,
*,
item: Any = None,
) -> FileInfo | Awaitable[FileInfo]:
return cast("FileInfo", {})
def media_failed(self, failure, request, info):
def media_failed(
self, failure: Failure, request: Request, info: MediaPipeline.SpiderInfo
) -> Failure:
failure.raiseException()
def file_path(self, request, response=None, info=None, *, item=None):
def file_path(
self,
request: Request,
response: Response | None = None,
info: MediaPipeline.SpiderInfo | None = None,
*,
item: Any = None,
) -> str:
return ""
@ -48,8 +81,14 @@ class TestBaseMediaPipeline:
self.pipe = self.pipeline_class.from_crawler(crawler)
self.pipe.open_spider()
self.info = self.pipe.spiderinfo
assert crawler.request_fingerprinter is not None
self.fingerprint = crawler.request_fingerprinter.fingerprint
@property
def mocked_pipe(self) -> MockedMediaPipeline:
assert isinstance(self.pipe, MockedMediaPipeline)
return self.pipe
def teardown_method(self):
for name, signal in vars(signals).items():
if not name.startswith("_"):
@ -121,11 +160,13 @@ class TestBaseMediaPipeline:
# When calling the method that caches the Request's result ...
self.pipe._cache_result_and_execute_waiters(failure, fp, info)
# ... it should store the Twisted Failure ...
assert info.downloaded[fp] == failure
downloaded = info.downloaded[fp]
assert downloaded == failure
# ... encapsulating the original FileException ...
assert info.downloaded[fp].value == file_exc
assert isinstance(downloaded, Failure)
assert downloaded.value == file_exc
# ... but it should not store the StopIteration exception on its context
context = getattr(info.downloaded[fp].value, "__context__", None)
context = getattr(downloaded.value, "__context__", None)
assert context is None
def test_default_item_completed(self, caplog: pytest.LogCaptureFixture) -> None:
@ -134,7 +175,7 @@ class TestBaseMediaPipeline:
# Check that failures are logged by default
fail = Failure(Exception())
results = [(True, 1), (False, fail)]
results: Any = [(True, 1), (False, fail)]
caplog.clear()
new_item = self.pipe.item_completed(results, item, self.info)
@ -158,7 +199,7 @@ class TestBaseMediaPipeline:
by item_completed(), as they are not download errors."""
item = {"name": "name"}
fail = Failure(_MediaRequestFiltered("Filtered offsite request"))
results = [(True, 1), (False, fail)]
results: Any = [(True, 1), (False, fail)]
with caplog.at_level(logging.DEBUG):
new_item = self.pipe.item_completed(results, item, self.info)
@ -174,29 +215,44 @@ class TestBaseMediaPipeline:
class MockedMediaPipeline(UserDefinedPipeline):
def __init__(self, *args, crawler=None, **kwargs):
def __init__(self, *args: Any, crawler: Crawler, **kwargs: Any):
super().__init__(*args, crawler=crawler, **kwargs)
self._mockcalled = []
self._mockcalled: list[str] = []
def media_to_download(self, request, info, *, item=None):
def media_to_download(
self, request: Request, info: MediaPipeline.SpiderInfo, *, item: Any = None
) -> Deferred[FileInfo | None] | None:
self._mockcalled.append("media_to_download")
if "result" in request.meta:
return request.meta.get("result")
return super().media_to_download(request, info)
def get_media_requests(self, item, info):
def get_media_requests(
self, item: Any, info: MediaPipeline.SpiderInfo
) -> list[Request]:
self._mockcalled.append("get_media_requests")
return item.get("requests")
return item.get("requests") # type: ignore[no-any-return]
def media_downloaded(self, response, request, info, *, item=None):
def media_downloaded(
self,
response: Response,
request: Request,
info: MediaPipeline.SpiderInfo,
*,
item: Any = None,
) -> FileInfo | Awaitable[FileInfo]:
self._mockcalled.append("media_downloaded")
return super().media_downloaded(response, request, info)
def media_failed(self, failure, request, info):
def media_failed(
self, failure: Failure, request: Request, info: MediaPipeline.SpiderInfo
) -> Failure:
self._mockcalled.append("media_failed")
return super().media_failed(failure, request, info)
def item_completed(self, results, item, info):
def item_completed(
self, results: list[FileInfoOrError], item: Any, info: MediaPipeline.SpiderInfo
) -> Any:
self._mockcalled.append("item_completed")
item = super().item_completed(results, item, info)
item["results"] = results
@ -204,7 +260,14 @@ class MockedMediaPipeline(UserDefinedPipeline):
class AsyncMediaDownloadedPipeline(MockedMediaPipeline):
async def media_downloaded(self, response, request, info, *, item=None):
async def media_downloaded( # type: ignore[override]
self,
response: Response,
request: Request,
info: MediaPipeline.SpiderInfo,
*,
item: Any = None,
) -> FileInfo | Awaitable[FileInfo]:
return super().media_downloaded(response, request, info)
@ -212,7 +275,7 @@ class TestMediaPipeline(TestBaseMediaPipeline):
pipeline_class = MockedMediaPipeline
def _errback(self, result):
self.pipe._mockcalled.append("request_errback")
self.mocked_pipe._mockcalled.append("request_errback")
return result
@coroutine_test
@ -226,7 +289,7 @@ class TestMediaPipeline(TestBaseMediaPipeline):
item = {"requests": req}
new_item = await self.pipe.process_item(item)
assert new_item["results"] == [(True, {})]
assert self.pipe._mockcalled == [
assert self.mocked_pipe._mockcalled == [
"get_media_requests",
"media_to_download",
"media_downloaded",
@ -248,7 +311,7 @@ class TestMediaPipeline(TestBaseMediaPipeline):
assert new_item["results"][0][0] is False
assert isinstance(new_item["results"][0][1], Failure)
assert new_item["results"][0][1].value == exc
assert self.pipe._mockcalled == [
assert self.mocked_pipe._mockcalled == [
"get_media_requests",
"media_to_download",
"media_failed",
@ -270,7 +333,7 @@ class TestMediaPipeline(TestBaseMediaPipeline):
assert new_item["results"][1][0] is False
assert isinstance(new_item["results"][1][1], Failure)
assert new_item["results"][1][1].value == exc
m = self.pipe._mockcalled
m = self.mocked_pipe._mockcalled
# only once
assert m[0] == "get_media_requests" # first hook called
assert m.count("get_media_requests") == 1
@ -294,7 +357,7 @@ class TestMediaPipeline(TestBaseMediaPipeline):
# returns iterable of Requests
req1 = Request("http://url1")
req2 = Request("http://url2")
item = {"requests": iter([req1, req2])}
item = {"requests": iter([req1, req2])} # type: ignore[dict-item]
new_item = await self.pipe.process_item(item)
assert new_item is item
assert self.fingerprint(req1) in self.info.downloaded
@ -304,7 +367,7 @@ class TestMediaPipeline(TestBaseMediaPipeline):
async def test_results_are_cached_across_multiple_items(self):
rsp1 = Response("http://url1")
req1 = Request("http://url1", meta={"response": rsp1})
item = {"requests": req1}
item: dict[str, Any] = {"requests": req1}
new_item = await self.pipe.process_item(item)
assert new_item is item
assert new_item["results"] == [(True, {})]
@ -335,7 +398,7 @@ class TestMediaPipeline(TestBaseMediaPipeline):
new_item = await self.pipe.process_item({"requests": req2})
assert new_item["results"][0][0] is False
assert new_item["results"][0][1].value is exc
assert self.pipe._mockcalled.count("media_to_download") == 1
assert self.mocked_pipe._mockcalled.count("media_to_download") == 1
@coroutine_test
async def test_cached_failure_calls_errback(self):
@ -347,13 +410,13 @@ class TestMediaPipeline(TestBaseMediaPipeline):
)
def errback(failure):
self.pipe._mockcalled.append("request_errback")
self.mocked_pipe._mockcalled.append("request_errback")
return {"recovered": failure.value}
req = Request("http://url1", errback=errback)
new_item = await self.pipe.process_item({"requests": req})
assert new_item["results"] == [(True, {"recovered": exc})]
assert self.pipe._mockcalled.count("request_errback") == 1
assert self.mocked_pipe._mockcalled.count("request_errback") == 1
@coroutine_test
async def test_results_are_cached_for_requests_of_single_item(self):
@ -362,14 +425,14 @@ class TestMediaPipeline(TestBaseMediaPipeline):
req2 = Request(
req1.url, meta={"response": Response("http://donot.download.me")}
)
item = {"requests": [req1, req2]}
item: dict[str, Any] = {"requests": [req1, req2]}
new_item = await self.pipe.process_item(item)
assert new_item is item
assert new_item["results"] == [(True, {}), (True, {})]
@coroutine_test
async def test_wait_if_request_is_downloading(self):
def _check_downloading(response):
def _check_downloading(response: Response) -> Response:
fp = self.fingerprint(req1)
assert fp in self.info.downloading
assert fp in self.info.waiting
@ -398,7 +461,7 @@ class TestMediaPipeline(TestBaseMediaPipeline):
item = {"requests": req}
new_item = await self.pipe.process_item(item)
assert new_item["results"] == [(True, "ITSME")]
assert self.pipe._mockcalled == [
assert self.mocked_pipe._mockcalled == [
"get_media_requests",
"media_to_download",
"item_completed",
@ -422,7 +485,9 @@ class TestAsyncMediaDownloaded(TestMediaPipeline):
class TestMediaPipelineAllowRedirectSettings:
def _assert_request_no3xx(self, pipeline_class, settings):
def _assert_request_no3xx(
self, pipeline_class: type[MediaPipeline], settings: dict[str, Any]
) -> None:
pipe = pipeline_class(crawler=get_crawler(None, settings))
request = Request("http://url")
pipe._modify_media_request(request)
@ -477,7 +542,7 @@ class TestBuildFromCrawler:
self._init_called = True
@classmethod
def from_crawler(cls, crawler):
def from_crawler(cls, crawler: Crawler) -> Pipeline:
settings = crawler.settings
store_uri = settings["FILES_STORE"]
o = cls(store_uri, settings=settings, crawler=crawler)
@ -493,9 +558,10 @@ class TestBuildFromCrawler:
def test_has_from_crawler(self):
class Pipeline(UserDefinedPipeline):
_from_crawler_called = False
store_uri: str
@classmethod
def from_crawler(cls, crawler):
def from_crawler(cls, crawler: Crawler) -> Pipeline:
settings = crawler.settings
o = super().from_crawler(crawler)
o._from_crawler_called = True
@ -509,7 +575,9 @@ class TestBuildFromCrawler:
class MediaFailedNonePipeline(MockedMediaPipeline):
def media_failed(self, failure, request, info):
def media_failed( # type: ignore[override]
self, failure: Failure, request: Request, info: MediaPipeline.SpiderInfo
) -> None:
self._mockcalled.append("media_failed")
@ -524,7 +592,7 @@ class TestMediaFailedNone(TestBaseMediaPipeline):
req = Request("http://url1", meta={"response": Exception("foo")})
new_item = await self.pipe.process_item({"requests": req})
assert new_item["results"] == [(True, None)]
assert self.pipe._mockcalled == [
assert self.mocked_pipe._mockcalled == [
"get_media_requests",
"media_to_download",
"media_failed",
@ -533,7 +601,9 @@ class TestMediaFailedNone(TestBaseMediaPipeline):
class MediaFailedFailurePipeline(MockedMediaPipeline):
def media_failed(self, failure, request, info):
def media_failed(
self, failure: Failure, request: Request, info: MediaPipeline.SpiderInfo
) -> Failure:
self._mockcalled.append("media_failed")
return failure # deprecated
@ -544,7 +614,7 @@ class TestMediaFailedFailure(TestBaseMediaPipeline):
pipeline_class = MediaFailedFailurePipeline
def _errback(self, result):
self.pipe._mockcalled.append("request_errback")
self.mocked_pipe._mockcalled.append("request_errback")
return result
@coroutine_test
@ -565,7 +635,7 @@ class TestMediaFailedFailure(TestBaseMediaPipeline):
assert new_item["results"][0][0] is False
assert isinstance(new_item["results"][0][1], Failure)
assert new_item["results"][0][1].value == exc
assert self.pipe._mockcalled == [
assert self.mocked_pipe._mockcalled == [
"get_media_requests",
"media_to_download",
"media_failed",

View File

@ -47,7 +47,7 @@ class DeferredPipeline:
return succeed(None)
def process_item(self, item):
d = Deferred()
d: Deferred[Any] = Deferred()
d.addCallback(self.cb)
d.callback(item)
return d
@ -55,7 +55,7 @@ class DeferredPipeline:
class AsyncDefPipeline:
async def process_item(self, item):
d = Deferred()
d: Deferred[Any] = Deferred()
call_later(0, d.callback, None)
await maybe_deferred_to_future(d)
item["pipeline_passed"] = True
@ -64,7 +64,7 @@ class AsyncDefPipeline:
class AsyncDefAsyncioPipeline:
async def process_item(self, item):
d = Deferred()
d: Deferred[Any] = Deferred()
loop = asyncio.get_event_loop()
loop.call_later(0, d.callback, None)
await deferred_to_future(d)
@ -75,12 +75,12 @@ class AsyncDefAsyncioPipeline:
class AsyncDefNotAsyncioPipeline:
async def process_item(self, item):
d1 = Deferred()
d1: Deferred[Any] = Deferred()
from twisted.internet import reactor
reactor.callLater(0, d1.callback, None)
await d1
d2 = Deferred()
d2: Deferred[Any] = Deferred()
reactor.callLater(0, d2.callback, None)
await maybe_deferred_to_future(d2)
item["pipeline_passed"] = True
@ -120,6 +120,8 @@ class OpenSpiderExceptionAsyncPipeline:
class ItemSpider(Spider):
name = "itemspider"
mockserver: MockServer
async def start(self):
yield Request(self.mockserver.url("/status?n=200"))

View File

@ -60,6 +60,19 @@ def test_process_spider_output(mw: DepthMiddleware, stats: StatsCollector) -> No
assert rdm == 1
def test_depth_reset(mw: DepthMiddleware, stats: StatsCollector) -> None:
resp = Response("https://example.com")
resp.request = Request("https://example.com", meta={"depth": 5})
result = [Request("https://example.com", meta={"depth_reset": True})]
out = list(mw.process_spider_output(resp, result))
assert out == result
assert out[0].meta["depth"] == 0
assert "depth_reset" not in out[0].meta
assert stats.get_value("request_depth_count/0") == 1
def test_process_spider_output_no_response(
mw: DepthMiddleware, stats: StatsCollector
) -> None:

View File

@ -5,7 +5,7 @@ import logging
import re
import sys
from io import StringIO
from typing import TYPE_CHECKING, Any
from typing import TYPE_CHECKING, Any, cast
import pytest
from twisted.python.failure import Failure
@ -16,6 +16,7 @@ from scrapy.utils.log import (
StreamLogger,
TopLevelFormatter,
failure_to_exc_info,
logformatter_adapter,
)
from scrapy.utils.test import get_crawler
from tests.spiders import LogSpider
@ -24,6 +25,7 @@ if TYPE_CHECKING:
from collections.abc import Generator, Mapping, MutableMapping
from scrapy.crawler import Crawler
from scrapy.logformatter import LogFormatterResult
class TestFailureToExcInfo:
@ -311,3 +313,36 @@ class TestLoggingWithExtra:
assert log_contents["message"] == log_message
assert self.regex_pattern.match(log_contents["spider"])
assert log_contents["important_info"] == extra["important_info"]
class TestLogformatterAdapter:
@staticmethod
def _log(caplog: pytest.LogCaptureFixture, logkws: LogFormatterResult) -> str:
with caplog.at_level(logging.INFO):
logging.getLogger(__name__).log(*logformatter_adapter(logkws))
return caplog.records[-1].getMessage()
@pytest.mark.parametrize("args", [None, {}, ()])
def test_empty_args(
self,
caplog: pytest.LogCaptureFixture,
args: dict[str, Any] | tuple[Any, ...] | None,
) -> None:
logkws = cast(
"LogFormatterResult",
{"level": logging.INFO, "msg": "90% done", "args": args},
)
assert self._log(caplog, logkws) == "90% done"
@pytest.mark.parametrize(
("msg", "args"),
[("%(pct)d%% done", {"pct": 90}), ("%d%% done", (90,))],
)
def test_args(
self,
caplog: pytest.LogCaptureFixture,
msg: str,
args: dict[str, Any] | tuple[Any, ...],
) -> None:
logkws: LogFormatterResult = {"level": logging.INFO, "msg": msg, "args": args}
assert self._log(caplog, logkws) == "90% done"

View File

@ -628,6 +628,7 @@ class TestHttpBase(ABC):
"Expected to receive 5 bytes which is larger than download warn size (4)"
in caplog.text
)
assert caplog.text.count("download warn size (4)") == 1
@coroutine_test
async def test_download_with_warnsize_no_content_length(

13
tox.ini
View File

@ -44,6 +44,7 @@ deps =
pygments
pytest
pytest-cov >= 7.0.0
pytest-timeout
pytest-xdist
sybil >= 1.3.0 # https://github.com/cjw296/sybil/issues/20#issuecomment-605433422
pytest-twisted >= 1.14.3
@ -136,6 +137,8 @@ deps =
pytest==8.4.0
Protego==0.1.15
Twisted==21.7.0
brotli==1.2.0; implementation_name != "pypy"
brotlicffi==1.2.0.0; implementation_name == "pypy"
cryptography==37.0.0
cssselect==0.9.1
httpx2==2.0.0
@ -170,8 +173,6 @@ deps =
Twisted[http2]
boto3
bpython # optional for shell wrapper tests
brotli >= 1.2.0; implementation_name != "pypy" # optional for HTTP compress downloader middleware tests
brotlicffi >= 1.2.0.0; implementation_name == "pypy" # optional for HTTP compress downloader middleware tests
google-cloud-storage
httpx2[http2,socks]
ipython
@ -188,8 +189,6 @@ deps =
Twisted[http2]==21.7.0
boto3==1.20.0
bpython==0.7.1
brotli==1.2.0; implementation_name != "pypy"
brotlicffi==1.2.0.0; implementation_name == "pypy"
google-cloud-storage==1.29.0
httpx2[http2,socks]==2.0.0
ipython==8.15.0
@ -214,7 +213,8 @@ deps =
#
# Pillow and uvloop build from source, and need the libjpeg headers and
# autotools respectively. robotexclusionrulesparser has no public repository,
# so it stays at its latest release.
# and PyDispatcher has seen no commit since 2023-10-23, so they stay at their
# latest release.
commands_pre =
uv pip install --python {envpython} --no-deps --reinstall \
git+https://github.com/twisted/twisted \
@ -235,7 +235,6 @@ commands_pre =
git+https://github.com/john-kurkowski/tldextract \
git+https://github.com/scrapy/w3lib \
git+https://github.com/zopefoundation/zope.interface \
git+https://github.com/mcfletch/pydispatcher \
git+https://github.com/boto/boto3 \
git+https://github.com/bpython/bpython \
git+https://github.com/google/brotli \
@ -290,7 +289,6 @@ commands =
basepython = pypy3
deps =
{[testenv:extra-deps]deps}
commands = {[testenv:pypy3]commands}
[testenv:min-pypy3]
basepython = pypy3.11
@ -300,6 +298,7 @@ deps =
pytest==8.4.0
Protego==0.1.15
Twisted==21.7.0
brotlicffi==1.2.0.0
cryptography==44.0.2
cssselect==0.9.1
itemadapter==0.1.0