Scrapy, a fast high-level web crawling & scraping framework for Python.
Go to file
Andrey Rakhmatullin fc1a83e7c4 Full typing for scrapy/item.py. 2024-04-29 19:17:29 +05:00
.github Undo an unintended change 2024-02-02 14:04:28 +01:00
artwork
…
docs Remove the auto-generated copyright years from the docs footer. (#6322) 2024-04-29 09:39:22 +02:00
extras
…
scrapy Full typing for scrapy/item.py. 2024-04-29 19:17:29 +05:00
sep Fix some comments (#6285) 2024-03-11 10:03:06 +01:00
tests Merge pull request #6221 from jxlil/fix/LxmlLinkExtractor 2024-04-19 19:57:11 +05:00
.bandit.yml Bandit: allow-list lxml usages (#6265) 2024-03-01 16:02:03 +01:00
.bumpversion.cfg Add a SECURITY.md file (#6051) 2024-02-20 10:50:16 +01:00
.coveragerc
…
.flake8 Bump black. 2024-02-28 14:30:38 +05:00
.git-blame-ignore-revs chore: fix some typos in comments (#6317) 2024-04-17 10:56:26 +02:00
.gitattributes
…
.gitignore
…
.isort.cfg
…
.pre-commit-config.yaml Bump bandit, flake8 and isort. 2024-02-28 14:30:38 +05:00
.readthedocs.yml
…
AUTHORS
…
CODE_OF_CONDUCT.md
…
CONTRIBUTING.md
…
INSTALL.md
…
LICENSE
…
MANIFEST.in
…
NEWS
…
README.rst Updated README.rst (#6144) 2023-11-13 20:13:10 +01:00
SECURITY.md Add a SECURITY.md file (#6051) 2024-02-20 10:50:16 +01:00
codecov.yml
…
conftest.py Remove tests/requirements.txt and refactor extra deps (#6272) 2024-03-13 07:22:48 +01:00
pylintrc Fix and re-enable unnecessary-comprehension and use-dict-literal pylint tags 2024-02-28 16:14:08 -03:00
pytest.ini Remove slow leftovers 2024-02-02 14:06:45 +01:00
setup.cfg Fix and remove most of the entries from the mypy ignore list (#6137) 2023-11-07 09:34:35 +01:00
setup.py Use defusedxml.xmlrpc 2024-02-27 17:08:13 -03:00
tox.ini Full typing for scrapy/extensions, part 3. (#6325) 2024-04-29 09:43:45 +02:00

README.rst

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>
https://scrapy.org/img/scrapylogo.png

Scrapy

Overview

Scrapy is a BSD-licensed fast high-level web crawling and web scraping framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

Scrapy is maintained by Zyte (formerly Scrapinghub) and many other contributors.

Check the Scrapy homepage at https://scrapy.org for more information, including a list of features.

Requirements

  • Python 3.8+
  • Works on Linux, Windows, macOS, BSD

Install

The quick way:

System Message: WARNING/2 (<stdin>, line 70)

Cannot analyze code. Pygments package not found.

.. code:: bash

    pip install scrapy

See the install section in the documentation at https://docs.scrapy.org/en/latest/intro/install.html for more details.

Documentation

Documentation is available online at https://docs.scrapy.org/ and in the docs directory.

Releases

You can check https://docs.scrapy.org/en/latest/news.html for the release notes.

Community (blog, twitter, mail list, IRC)

See https://scrapy.org/community/ for details.

Contributing

See https://docs.scrapy.org/en/master/contributing.html for details.

Code of Conduct

Please note that this project is released with a Contributor Code of Conduct.

By participating in this project you agree to abide by its terms. Please report unacceptable behavior to opensource@zyte.com.

Companies using Scrapy

See https://scrapy.org/companies/ for a list.

Commercial Support

See https://scrapy.org/support/ for details.

</html>