The HTML specification states that link type keywords like 'nofollow'
are ASCII case-insensitive. Sites using rel="NoFollow" or rel="NOFOLLOW"
were incorrectly treated as follow links, causing Scrapy to crawl pages
it should skip.
Fix: add .lower() before the token split so all casing variants of
'nofollow' are correctly recognized.
Co-authored-by: JSap0914 <JSap0914@users.noreply.github.com>
When JOBDIR is enabled, requests are serialized to disk with pickle, so
the objects stored in a request's cb_kwargs and meta are deep-copied on
the round trip. Callbacks then receive copies rather than the original
objects, which is easy to miss and can silently break code that relies on
sharing mutable state. Add a note to the request serialization section of
the jobs docs and a cross-referenced caution to the Request.cb_kwargs
attribute docs.
Closes#6120
- Add xtractmime to the mypy ignore_missing_imports overrides (no py.typed
marker in that library)
- Narrow Content-Disposition header via a local variable so mypy can see
it is non-None inside the if-block
- Remove a redundant local re-import of HtmlResponse/TextResponse inside
open_in_browser (they are already imported at module level, fixing
pylint W0404 reimported)
- Add Any annotations to _unmark() in test_responsetypes.py so mypy does
not raise no-untyped-call when it is called from a typed context
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
`identity` in `Content-Encoding` means no transformation (RFC 7231),
so it should not be treated as a compression encoding when sniffing
the MIME type. Previously it was mapped to `application/identity`,
causing `get_response_class` to return `Response` for HTML bodies,
which then broke `response.replace(cls=Response)` on `TextResponse`
objects (encoding is not a valid `Response.__init__` kwarg).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* Add support for HTTP/2 and SOCKS proxies to HttpxDownloadHandler.
* Update the docs.
* Trim the tables.
* Restore lost wording.
* Handlers docs improvements and fixes.
* Allow configuring the log level of the retry give-up message
The "Gave up retrying ..." message was always logged at ERROR, which
inflates the log_count/ERROR stat even when giving up on a request is
expected (e.g. broad crawls hitting dead hosts).
Add a RETRY_GIVE_UP_LOG_LEVEL setting, a give_up_log_level argument to
get_retry_request(), and a give_up_log_level request meta key to override
it per request. The value accepts a level name ("WARNING") or number
(logging.WARNING). The default ("ERROR") preserves the previous behaviour.
Fixes#5297, fixes#4622
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Address Adrian's review feedback: simplify and reorganize docs
- Move give_up_log_level reqmeta section before max_retry_times (alphabetical order)
- Simplify give_up_log_level section in request-response.rst (brief, links to setting)
- Simplify RETRY_GIVE_UP_LOG_LEVEL setting docs in downloader-middleware.rst
- Change 'When initialized' to 'When set' in max_retry_times section
- Docs now follow pattern of linking to complementary setting/meta key rather than duplicating information
Per Adrian's feedback: keep docs concise and cross-link setting ↔ meta key
* Address Adrian's code review feedback on RETRY_GIVE_UP_LOG_LEVEL feature
- Fix alphabetical ordering of give_up_log_level in request-response.rst
- Remove circular references: change 'see X for details' to 'see also X' pattern
- Simplify docstring for give_up_log_level parameter (4 lines → 2 lines)
- Update test domains from www.scrapytest.org to example.com
* Apply suggestion from @AdrianAtZyte
* Minor changes
* Fix example formatting.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Adrian <adrian@zyte.com>
Co-authored-by: Andrey Rakhmatullin <wrar@wrar.name>
* docs: switch `scrapy.Item` examples to dataclasses
* make serializer doc generic
* use modern type hints
* docs/spiders: switch TestItem consumer snippets to attribute access
Since the TestItem migration to @dataclass, the existing
item["id"] = ... assignments would raise TypeError on copy-paste.
Switch to item.id = ... to match the new dataclass declaration.
The snippets sit under .. skip: next so docs-tests still pass either
way, but the change keeps the examples runnable for readers.