Scrapy, a fast high-level web crawling & scraping framework for Python.
Go to file
Andrey Rakhmatullin 7d5b189c11
Fix getting annotations for _parse_sitemap() at the runtime. (#6671)
* Fix getting annotations for _parse_sitemap() at the runtime.

* Split off the callback annotations test.
2025-02-14 20:40:06 +05:00
.github Drop PyPy 3.9, add a pypy3-extra-deps CI job. (#6613) 2025-01-23 09:22:18 +01:00
artwork artwork/README.rst: add missing articles (#5827) 2023-02-14 09:42:43 +01:00
docs Support dark mode in the documentation (#6653) 2025-02-06 18:07:07 +01:00
extras Run and fix linkcheck. (#6524) 2024-11-04 11:40:07 +01:00
scrapy Fix getting annotations for _parse_sitemap() at the runtime. (#6671) 2025-02-14 20:40:06 +05:00
sep Fix some comments (#6285) 2024-03-11 10:03:06 +01:00
tests Fix getting annotations for _parse_sitemap() at the runtime. (#6671) 2025-02-14 20:40:06 +05:00
tests_typing Drop Python 3.8 Support (#6472) 2024-10-16 10:03:16 +02:00
.git-blame-ignore-revs chore: fix some typos in comments (#6317) 2024-04-17 10:56:26 +02:00
.gitattributes Maybe the problem is not in the code after all 2020-08-13 06:35:09 +02:00
.gitignore refact: add Osx DS_Store file to gitignore 2022-10-03 00:48:12 -03:00
.pre-commit-config.yaml Bump ruff, switch from black to ruff-format (#6631) 2025-01-27 11:07:09 +01:00
.readthedocs.yml chore(docs): refactor config (#6623) 2025-01-20 12:18:30 +01:00
AUTHORS Scrapinghub → Zyte 2021-02-02 15:03:20 +01:00
CODE_OF_CONDUCT.md Update Code of Conduct to Contributor Covenant v2.1 2022-10-28 02:13:37 +02:00
CONTRIBUTING.md Be consistent with domain used for links to documentation website 2019-01-31 01:28:53 -03:00
INSTALL.md Update and rename INSTALL to INSTALL.md 2022-10-06 19:58:48 +02:00
LICENSE added oxford commas to LICENSE 2018-06-01 21:48:43 -03:00
MANIFEST.in Integrating configs into pyproject.toml (#6547) 2024-11-19 19:21:15 +05:00
NEWS added NEWS file pointing to docs/news.rst 2012-04-28 23:32:51 -03:00
README.rst Run and fix linkcheck. (#6524) 2024-11-04 11:40:07 +01:00
SECURITY.md Bump version: 2.11.2 → 2.12.0 2024-11-18 13:08:05 +05:00
codecov.yml codecov config: disable project check, tweak PR comments 2017-05-19 00:01:27 +05:00
conftest.py Made path absolute to enable running pytest from a different directory. (#6567) 2024-12-09 11:01:00 +01:00
pyproject.toml Refactor downloader tests (#6647) 2025-02-03 20:11:47 +05:00
tox.ini Refactor downloader tests (#6647) 2025-02-03 20:11:47 +05:00

README.rst

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>
https://scrapy.org/img/scrapylogo.png

Scrapy

Overview

Scrapy is a BSD-licensed fast high-level web crawling and web scraping framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

Scrapy is maintained by Zyte (formerly Scrapinghub) and many other contributors.

Check the Scrapy homepage at https://scrapy.org for more information, including a list of features.

Requirements

  • Python 3.9+
  • Works on Linux, Windows, macOS, BSD

Install

The quick way:

System Message: WARNING/2 (<stdin>, line 70)

Cannot analyze code. Pygments package not found.

.. code:: bash

    pip install scrapy

See the install section in the documentation at https://docs.scrapy.org/en/latest/intro/install.html for more details.

Documentation

Documentation is available online at https://docs.scrapy.org/ and in the docs directory.

Releases

You can check https://docs.scrapy.org/en/latest/news.html for the release notes.

Community (blog, twitter, mail list, IRC)

See https://scrapy.org/community/ for details.

Contributing

See https://docs.scrapy.org/en/master/contributing.html for details.

Code of Conduct

Please note that this project is released with a Contributor Code of Conduct.

By participating in this project you agree to abide by its terms. Please report unacceptable behavior to opensource@zyte.com.

Companies using Scrapy

See https://scrapy.org/companies/ for a list.

Commercial Support

See https://scrapy.org/support/ for details.

</html>