Scrapy, a fast high-level web crawling & scraping framework for Python.
Go to file
Andrey Rakhmatullin d85c39f5bc
Deprecation removals. (#6500)
* Deprecation removals.

* Clean up the default pytest filterwarnings.

* Remove test_get_images_old().

* Redo boto-requiring test filtering.

* Remove an unused function.

* Improve the Crawler.crawl() error message.

* Fix the test.
2024-10-31 18:06:22 +05:00
.github Merge remote-tracking branch 'origin/master' into py313 2024-10-16 14:50:12 +05:00
artwork artwork/README.rst: add missing articles (#5827) 2023-02-14 09:42:43 +01:00
docs Deprecation removals. (#6500) 2024-10-31 18:06:22 +05:00
extras Deprecation removals. (#6500) 2024-10-31 18:06:22 +05:00
scrapy Deprecation removals. (#6500) 2024-10-31 18:06:22 +05:00
sep Fix some comments (#6285) 2024-03-11 10:03:06 +01:00
tests Deprecation removals. (#6500) 2024-10-31 18:06:22 +05:00
tests_typing Drop Python 3.8 Support (#6472) 2024-10-16 10:03:16 +02:00
.bandit.yml Bandit: allow-list lxml usages (#6265) 2024-03-01 16:02:03 +01:00
.bumpversion.cfg Merge 2.11.2 changes (#6363) 2024-05-14 18:54:11 +02:00
.coveragerc Skip coverage checks for TYPE_CHECKING blocks. 2024-05-05 22:55:21 +05:00
.flake8 Add flake8-type-checking. (#6413) 2024-06-25 10:20:59 +02:00
.git-blame-ignore-revs chore: fix some typos in comments (#6317) 2024-04-17 10:56:26 +02:00
.gitattributes Maybe the problem is not in the code after all 2020-08-13 06:35:09 +02:00
.gitignore refact: add Osx DS_Store file to gitignore 2022-10-03 00:48:12 -03:00
.isort.cfg fix .isort.cfg 2023-01-27 14:59:08 -06:00
.pre-commit-config.yaml Remove --keep-runtime-typing from pyupgrade. 2024-10-17 21:26:02 +05:00
.readthedocs.yml Bump the Python version for RTD. 2024-07-11 12:25:13 +05:00
AUTHORS Scrapinghub → Zyte 2021-02-02 15:03:20 +01:00
CODE_OF_CONDUCT.md Update Code of Conduct to Contributor Covenant v2.1 2022-10-28 02:13:37 +02:00
CONTRIBUTING.md Be consistent with domain used for links to documentation website 2019-01-31 01:28:53 -03:00
INSTALL.md Update and rename INSTALL to INSTALL.md 2022-10-06 19:58:48 +02:00
LICENSE added oxford commas to LICENSE 2018-06-01 21:48:43 -03:00
MANIFEST.in Update MANIFEST.in. 2024-05-08 00:39:05 +05:00
NEWS added NEWS file pointing to docs/news.rst 2012-04-28 23:32:51 -03:00
README.rst Drop Python 3.8 Support (#6472) 2024-10-16 10:03:16 +02:00
SECURITY.md Add a SECURITY.md file (#6051) 2024-02-20 10:50:16 +01:00
codecov.yml codecov config: disable project check, tweak PR comments 2017-05-19 00:01:27 +05:00
conftest.py Deprecation removals. (#6500) 2024-10-31 18:06:22 +05:00
pylintrc Fix and re-enable unnecessary-comprehension and use-dict-literal pylint tags 2024-02-28 16:14:08 -03:00
pytest.ini Deprecation removals. (#6500) 2024-10-31 18:06:22 +05:00
setup.cfg Fix and remove most of the entries from the mypy ignore list (#6137) 2023-11-07 09:34:35 +01:00
setup.py Merge remote-tracking branch 'origin/master' into py313 2024-10-16 14:50:12 +05:00
tox.ini Deprecation removals. (#6500) 2024-10-31 18:06:22 +05:00

README.rst

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>
https://scrapy.org/img/scrapylogo.png

Scrapy

Overview

Scrapy is a BSD-licensed fast high-level web crawling and web scraping framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

Scrapy is maintained by Zyte (formerly Scrapinghub) and many other contributors.

Check the Scrapy homepage at https://scrapy.org for more information, including a list of features.

Requirements

  • Python 3.9+
  • Works on Linux, Windows, macOS, BSD

Install

The quick way:

System Message: WARNING/2 (<stdin>, line 70)

Cannot analyze code. Pygments package not found.

.. code:: bash

    pip install scrapy

See the install section in the documentation at https://docs.scrapy.org/en/latest/intro/install.html for more details.

Documentation

Documentation is available online at https://docs.scrapy.org/ and in the docs directory.

Releases

You can check https://docs.scrapy.org/en/latest/news.html for the release notes.

Community (blog, twitter, mail list, IRC)

See https://scrapy.org/community/ for details.

Contributing

See https://docs.scrapy.org/en/master/contributing.html for details.

Code of Conduct

Please note that this project is released with a Contributor Code of Conduct.

By participating in this project you agree to abide by its terms. Please report unacceptable behavior to opensource@zyte.com.

Companies using Scrapy

See https://scrapy.org/companies/ for a list.

Commercial Support

See https://scrapy.org/support/ for details.

</html>