Scrapy, a fast high-level web crawling & scraping framework for Python.
Go to file
Adrián Chaves e72de11f55 Add super 2024-03-12 09:29:10 +01:00
.github Re-enable uvloop tests on 3.12 (#6098) 2023-11-16 19:16:31 +04:00
artwork artwork/README.rst: add missing articles (#5827) 2023-02-14 09:42:43 +01:00
docs Full typing for scrapy/extensions, part 2. (#6279) 2024-03-11 10:09:09 +01:00
extras Upgrade CI tools 2023-02-02 06:37:40 +01:00
scrapy Full typing for scrapy/extensions, part 2. (#6279) 2024-03-11 10:09:09 +01:00
sep Fix some comments (#6285) 2024-03-11 10:03:06 +01:00
tests Add super 2024-03-12 09:29:10 +01:00
.bandit.yml Bandit: allow-list lxml usages (#6265) 2024-03-01 16:02:03 +01:00
.bumpversion.cfg Add a SECURITY.md file (#6051) 2024-02-20 10:50:16 +01:00
.coveragerc added disable_warnings instruction to .coveragerc 2023-01-31 14:28:08 -08:00
.flake8 Bump black. 2024-02-28 14:30:38 +05:00
.git-blame-ignore-revs ignoring last changes made by black for blame 2022-12-29 12:51:05 -03:00
.gitattributes Maybe the problem is not in the code after all 2020-08-13 06:35:09 +02:00
.gitignore refact: add Osx DS_Store file to gitignore 2022-10-03 00:48:12 -03:00
.isort.cfg fix .isort.cfg 2023-01-27 14:59:08 -06:00
.pre-commit-config.yaml Bump bandit, flake8 and isort. 2024-02-28 14:30:38 +05:00
.readthedocs.yml Use Python 3.11 as the default in CI (#5696) 2022-10-27 14:00:36 +02:00
AUTHORS Scrapinghub → Zyte 2021-02-02 15:03:20 +01:00
CODE_OF_CONDUCT.md Update Code of Conduct to Contributor Covenant v2.1 2022-10-28 02:13:37 +02:00
CONTRIBUTING.md Be consistent with domain used for links to documentation website 2019-01-31 01:28:53 -03:00
INSTALL.md Update and rename INSTALL to INSTALL.md 2022-10-06 19:58:48 +02:00
LICENSE added oxford commas to LICENSE 2018-06-01 21:48:43 -03:00
MANIFEST.in Add py.typed 2023-09-21 13:06:12 +02:00
NEWS added NEWS file pointing to docs/news.rst 2012-04-28 23:32:51 -03:00
README.rst Updated README.rst (#6144) 2023-11-16 19:17:52 +04:00
SECURITY.md Add a SECURITY.md file (#6051) 2024-02-20 10:50:16 +01:00
codecov.yml codecov config: disable project check, tweak PR comments 2017-05-19 00:01:27 +05:00
conftest.py Skip more non-test files during discovery. 2023-07-22 17:54:55 +04:00
pylintrc Fix and re-enable unnecessary-comprehension and use-dict-literal pylint tags 2024-02-28 16:14:08 -03:00
pytest.ini Simplify skipping uvloop tests. 2023-07-22 17:44:37 +04:00
setup.cfg Fix and remove most of the entries from the mypy ignore list (#6137) 2023-11-07 09:34:35 +01:00
setup.py Use defusedxml.xmlrpc 2024-02-27 17:08:13 -03:00
tox.ini Use brotlicffi for PyPy 2024-03-05 21:30:20 -03:00

README.rst

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>
https://scrapy.org/img/scrapylogo.png

Scrapy

Overview

Scrapy is a BSD-licensed fast high-level web crawling and web scraping framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

Scrapy is maintained by Zyte (formerly Scrapinghub) and many other contributors.

Check the Scrapy homepage at https://scrapy.org for more information, including a list of features.

Requirements

  • Python 3.8+
  • Works on Linux, Windows, macOS, BSD

Install

The quick way:

System Message: WARNING/2 (<stdin>, line 70)

Cannot analyze code. Pygments package not found.

.. code:: bash

    pip install scrapy

See the install section in the documentation at https://docs.scrapy.org/en/latest/intro/install.html for more details.

Documentation

Documentation is available online at https://docs.scrapy.org/ and in the docs directory.

Releases

You can check https://docs.scrapy.org/en/latest/news.html for the release notes.

Community (blog, twitter, mail list, IRC)

See https://scrapy.org/community/ for details.

Contributing

See https://docs.scrapy.org/en/master/contributing.html for details.

Code of Conduct

Please note that this project is released with a Contributor Code of Conduct.

By participating in this project you agree to abide by its terms. Please report unacceptable behavior to opensource@zyte.com.

Companies using Scrapy

See https://scrapy.org/companies/ for a list.

Commercial Support

See https://scrapy.org/support/ for details.

</html>