Scrapy, a fast high-level web crawling & scraping framework for Python.
Go to file
Dev Kumar 901ec3f08c Make redirect dupefilter bypass chain-aware
Extend the self-redirect dupefilter fix to cover the whole redirect
chain, not just the immediate predecessor request. A redirect chain
that loops back to an earlier request in the same chain (A -> B -> C
-> A) previously still got silently dropped, because only the
immediate hop's fingerprint was compared. Track the set of
fingerprints seen so far in the chain via request.meta and bypass the
dupe filter if the new target matches any of them.

Refs #1225
2026-07-21 07:49:15 -05:00
.github Use external mitmproxy. (#7720) 2026-07-09 17:02:05 +02:00
docs Modernize code examples, drop docs for nonexistent `MEMDEBUG_NOTIFY` (#7737) 2026-07-21 12:22:04 +02:00
extras chore: simplify some code, get rid of nested fns where it's makes sence + pylint (#7401) 2026-05-04 23:32:46 +05:00
scrapy Make redirect dupefilter bypass chain-aware 2026-07-21 07:49:15 -05:00
sep Add sphinx-lint. (#6920) 2025-06-28 01:37:20 +02:00
tests Test restructuring (#7736) 2026-07-21 12:44:56 +02:00
tests_typing Enable mypy warn_return_any. (#7492) 2026-05-04 22:24:47 +02:00
.git-blame-ignore-revs Update tool versions (#7127) 2025-10-27 14:11:31 +01:00
.gitattributes Maybe the problem is not in the code after all 2020-08-13 06:35:09 +02:00
.gitignore Add llms.txt and llms-full.txt generation (#7380) 2026-04-06 10:24:21 +02:00
.pre-commit-config.yaml Update tools (#7724) 2026-07-07 17:27:26 +05:00
.readthedocs.yml Add Python 3.14 to CI. (#6604) 2026-05-12 23:49:37 +05:00
AUTHORS Scrapinghub → Zyte 2021-02-02 15:03:20 +01:00
CITATION.cff Add CITATION.cff (#7519) 2026-05-14 13:07:13 +02:00
CODE_OF_CONDUCT.md Update Code of Conduct to Contributor Covenant v2.1 2022-10-28 02:13:37 +02:00
CONTRIBUTING.md Be consistent with domain used for links to documentation website 2019-01-31 01:28:53 -03:00
INSTALL.md Update and rename INSTALL to INSTALL.md 2022-10-06 19:58:48 +02:00
LICENSE added oxford commas to LICENSE 2018-06-01 21:48:43 -03:00
NEWS added NEWS file pointing to docs/news.rst 2012-04-28 23:32:51 -03:00
README.rst Remove Python 3.9 support (#7121) 2025-10-27 12:37:49 +01:00
SECURITY.md Bump version: 2.16.0 → 2.17.0 2026-07-07 12:26:52 +02:00
codecov.yml codecov config: disable project check, tweak PR comments 2017-05-19 00:01:27 +05:00
conftest.py Use external mitmproxy. (#7720) 2026-07-09 17:02:05 +02:00
pyproject.toml Use external mitmproxy. (#7720) 2026-07-09 17:02:05 +02:00
tox.ini Rewrite GCSFilesStore tests to use mocking. (#7727) 2026-07-21 11:43:36 +02:00

README.rst

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>

Scrapy is a web scraping framework to extract structured data from websites. It is cross-platform, and requires Python 3.10+. It is maintained by Zyte (formerly Scrapinghub) and many other contributors.

Install with:

System Message: WARNING/2 (<stdin>, line 52)

Cannot analyze code. Pygments package not found.

.. code:: bash

    pip install scrapy

And follow the documentation to learn how to use it.

If you wish to contribute, see Contributing.

</html>