Mohammad Hassan
24cafa80bc
Merge f7382561cf into ad43bf0c56
2026-08-15 11:31:49 -05:00
Adrian
a22523bfd3
Add a HTTP2_MAX_FRAME_SIZE setting ( #7988 )
2026-08-14 14:19:06 +05:00
Adrian
8bb06bf00b
Stop marking the Twisted-based HTTP/2 download handler as experimental, and privatize its API ( #7986 )
2026-08-12 23:33:38 +05:00
Adrian
65b37286cc
Document the ITEM_PROCESSOR setting ( #7983 )
2026-08-12 20:34:25 +05:00
Adrian
8ce041e528
:issue: → :gh: ( #7974 )
2026-08-10 19:30:48 +02:00
Adrian
f1694269d8
Add FTPS support to the FTP feed export storage ( #7953 )
2026-08-10 11:33:58 +02:00
Adrian
fe96c1f54b
Serialize dates and times as ISO 8601 in ScrapyJSONEncoder ( #7918 )
2026-08-10 14:07:42 +05:00
Adrian
4cc2356637
Implement browser-like bad header handling for the default download handler ( #7806 )
...
* Implement browser-like bad header handling for the default download handler
* Remove dead code
2026-08-10 13:59:44 +05:00
Adrian
c285f4cb18
Support CloseSpider during spider startup ( #7905 )
2026-08-10 10:06:54 +02:00
Adrian
508bd7faec
Type late Crawler attributes as always set instead of None ( #7882 )
2026-08-10 09:07:34 +02:00
Adrian
09c918115b
Document response parsing memory use in the security page ( #7930 )
2026-08-10 08:43:57 +02:00
Mohammad Hassan
f7382561cf
Do not fail to disable a component with an unimportable path
...
`get_component_priority_dict_with_base()` normalizes every key of a
component priority dictionary by calling `load_object()` on it, including
keys whose value is `None`, i.e. components the user is disabling.
`normalize_key()` only guarded against `(NameError, TypeError, ValueError)`,
the three exceptions `load_object()` raises itself, but not against the
`ImportError` raised by the `import_module()` call it makes. A stale entry
disabling a component whose module no longer exists, e.g.
`{"scrapy.webservice.WebService": None}`, therefore aborted settings
resolution with `ModuleNotFoundError` instead of being ignored.
Add `ImportError` to the guard so that an unimportable key falls back to its
raw string form. Since its value is `None`, it is then dropped by the
existing `if v is not None` filter, and disabling by import path keeps
working. Enabled components with an unimportable path are unaffected:
`MiddlewareManager.from_crawler()` calls `load_object()` on them again and
still fails there.
Fixes #7820
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 20:08:58 -04:00
Adrian
ae68786210
Document how to customize media pipeline file names from responses ( #7909 )
2026-08-09 23:08:59 +05:00
Adrian
5270f3cf99
Add a depth_reset request metadata key ( #7913 )
2026-08-09 23:07:10 +05:00
Adrian
d41bfaec05
Document how to disallow subdomains of allowed domains ( #7903 )
2026-08-09 23:05:43 +05:00
Adrian
ba29ce8d3b
Add a docs section on using download_async() from a downloader middleware ( #7872 )
2026-08-09 23:02:44 +05:00
Adrian
fc082aa914
Clarify the docs about reactor settings ( #7880 )
2026-08-09 23:01:31 +05:00
Adrian
bee31890a2
Report exceptions from Spider.start() ( #7884 )
2026-08-09 22:56:03 +05:00
Adrian
482a02d30c
Add a cookies documentation page ( #7947 )
2026-08-09 22:41:32 +05:00
Adrian
786ab10494
Document how to write a custom item exporter ( #7931 )
2026-08-09 22:31:53 +05:00
Adrian
609f64c55d
Make brotli a hard dependency ( #7929 )
2026-08-09 19:24:14 +02:00
Adrian
5427080f48
Create a page on optimization ( #7938 )
2026-08-09 18:37:24 +02:00
Adrian
a6c017c2ca
Advise setting an identifying user agent ( #7890 )
2026-08-09 16:59:22 +05:00
Adrian
18ed0c0f7c
Fall back to the response encoding in TextResponse.json() ( #7897 )
2026-08-09 16:58:09 +05:00
Adrian
050a8cf159
Improve the docs about delaying start request iteration ( #7883 )
2026-08-09 16:57:08 +05:00
Adrian
59ce27afdd
Send the bytes_received and headers_received signals over HTTP/2 ( #7896 )
2026-08-09 16:55:02 +05:00
Adrian
7c797968a3
Docs: clarify the handling of exceptions raised in errbacks ( #7898 )
2026-08-09 16:53:33 +05:00
Adrian
a18d58d7b5
Document that process_spider_output receives a lazy result ( #7939 )
2026-08-09 12:26:15 +02:00
Adrian
9e84112221
Cover update_vars() in the shell docs ( #7889 )
2026-08-09 14:25:14 +05:00
Adrian
0b2d220197
Improve docs for multi-spider runs ( #7907 )
2026-08-09 14:19:50 +05:00
Adrian
0e324f3d4a
Let spiders change allowed_domains at run time ( #7912 )
2026-08-09 14:17:44 +05:00
Adrian
4a69e48f0f
Log the first depth-limited link only ( #7916 )
2026-08-09 14:08:49 +05:00
Adrian
81d12c6eb8
Document the Referer caveat of DEFAULT_REQUEST_HEADERS ( #7917 )
2026-08-09 14:08:06 +05:00
Adrian
5b4888a0b1
Document that signal handler order is undefined ( #7941 )
2026-08-09 14:07:23 +05:00
Adrian
91b70e4db4
Improve the docs about scrapy parse --pipelines ( #7876 )
2026-08-04 17:30:51 +02:00
Adrian
639fac78b3
Cover the signature change of scrape_func ( #7875 )
2026-08-04 20:02:52 +05:00
Adrian
6e2081ca41
Ask custom download handlers not to use engine.download_async() ( #7871 )
2026-08-04 14:41:25 +05:00
Adrian
54da6c88aa
Deprecate the download_delay spider attribute, and fix the suggested replacement for max_concurrent_requests ( #7833 )
...
* Deprecate the download_delay and max_concurrent_requests spider attributes
* Fix the deprecation entry of max_concurrent_requests
2026-08-03 23:46:41 +05:00
Adrian
a9f1770306
Docs: sort and compact the component-settings list ( #7862 )
2026-08-03 13:03:28 +02:00
Adrian
e83c709574
Docs: a job directory belongs to one Scrapy version ( #7861 )
2026-08-03 12:44:48 +02:00
Adrian
24de06bcd3
CI: Install dependencies with uv ( #7838 )
...
* CI: Install dependencies with uv
* Setup GitHub Actions hardening
* CI: Use uv for mitmproxy, benchmarks and cache keys
* Fix mitmproxy install on PyPy
2026-07-31 23:02:52 +05:00
Adrian
14478e3f24
Support CONCURRENT_REQUESTS = 0 for unlimited concurrency ( #7840 )
2026-07-31 20:19:07 +05:00
Adrian
1f03fbc17e
Add scrapy.utils.asyncio.sleep() ( #7843 )
2026-07-31 18:59:38 +05:00
Adrian
6cefaa5434
Add a stats reference ( #7814 )
2026-07-31 15:43:40 +05:00
Adrian
746bc7548d
Clarify crawl vs runspider in help and docs ( #7832 )
2026-07-31 12:17:34 +05:00
Adrian
37661508db
Add RobotParser.crawl_delay() and a robots_parsed signal ( #7830 )
...
* Add RobotParser.crawl_delay() and a robots_parsed signal
* Improve test coverage
2026-07-31 11:59:52 +05:00
Adrian
434fd1154a
Improve pre-crawler setting docs ( #7835 )
2026-07-31 11:38:33 +05:00
Adrian
f02a99fe71
Add doc sections for callbacks and errbacks ( #7821 )
2026-07-30 20:15:15 +05:00
Adrian
433603e6ca
Add AWS_MAX_POOL_CONNECTIONS ( #7794 )
2026-07-30 15:51:45 +02:00
Adrian
98696efa80
Export item fields in declaration order ( #7824 )
2026-07-30 15:46:45 +05:00