Scrapy, a fast high-level web crawling & scraping framework for Python.
Go to file
Mikhail Korobov dc95ecbe25 DOC use autodocs for selectors; document more methods and attributes; suggest get/getall 2018-09-12 18:36:25 +05:00
artwork Use https for external links wherever possible in docs 2017-10-26 23:33:45 +05:30
debian Use https links wherever possible 2017-10-28 16:24:40 +05:30
docs DOC use autodocs for selectors; document more methods and attributes; suggest get/getall 2018-09-12 18:36:25 +05:00
extras Use https for external links wherever possible in docs 2017-10-26 23:33:45 +05:30
scrapy DOC use autodocs for selectors; document more methods and attributes; suggest get/getall 2018-09-12 18:36:25 +05:00
sep Fix link for 'XPath and XSLT with lxml' 2017-10-28 16:34:49 +05:30
tests TST update tests to use get/getall/attrib instead of extract 2018-09-12 17:57:27 +05:00
.bumpversion.cfg Bump version: 1.4.0 → 1.5.0 2017-12-30 02:09:52 +05:00
.coveragerc remove ancient modules kept only for error messages 2018-07-06 03:23:37 +05:00
.gitignore add pytest temp files to gitignore 2018-09-12 17:57:27 +05:00
.travis.yml Try to get python3.7 by using xenial base and sudo 2018-07-09 13:12:47 +03:00
AUTHORS added Nicolas Ramirez to AUTHORS 2013-03-14 12:44:39 -03:00
CODE_OF_CONDUCT.md minor grammatical fixes in CODE_OF_CONDUCT.md 2018-06-01 21:48:43 -03:00
CONTRIBUTING.md Use https links wherever possible 2017-10-28 16:24:40 +05:30
INSTALL Use https links wherever possible 2017-10-28 16:24:40 +05:30
LICENSE added oxford commas to LICENSE 2018-06-01 21:48:43 -03:00
MANIFEST.in Ignore explicitly compiled python files. 2016-11-08 20:52:32 -03:00
Makefile.buildbot Generated version as pep440 and dpkg compatible 2015-06-16 00:16:09 +00:00
NEWS added NEWS file pointing to docs/news.rst 2012-04-28 23:32:51 -03:00
README.rst Drop support for EOL Python 3.3 2017-12-19 17:59:05 +02:00
appveyor.yml Cache pip cache and do not rebuild tags on appveyor and travis 2018-08-15 01:35:01 -03:00
codecov.yml codecov config: disable project check, tweak PR comments 2017-05-19 00:01:27 +05:00
conftest.py remove ancient modules kept only for error messages 2018-07-06 03:23:37 +05:00
pytest.ini Don't collect tests by their class name 2015-05-04 18:10:04 -03:00
requirements-py2.txt require parsel 1.5+ 2018-09-12 17:57:27 +05:00
requirements-py3.txt require parsel 1.5+ 2018-09-12 17:57:27 +05:00
setup.cfg Build universal wheels 2016-03-01 11:00:20 +01:00
setup.py require parsel 1.5+ 2018-09-12 17:57:27 +05:00
tox.ini Merge branch 'master' of https://github.com/patiences/scrapy into patiences-master 2018-07-09 11:59:22 +03:00

README.rst

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>

Scrapy

Overview

Scrapy is a fast high-level web crawling and web scraping framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

For more information including a list of features check the Scrapy homepage at: https://scrapy.org

Requirements

  • Python 2.7 or Python 3.4+
  • Works on Linux, Windows, Mac OSX, BSD

Install

The quick way:

pip install scrapy

For more details see the install section in the documentation: https://doc.scrapy.org/en/latest/intro/install.html

Documentation

Documentation is available online at https://doc.scrapy.org/ and in the docs directory.

Releases

You can find release notes at https://doc.scrapy.org/en/latest/news.html

Community (blog, twitter, mail list, IRC)

See https://scrapy.org/community/

Contributing

See https://doc.scrapy.org/en/master/contributing.html

Code of Conduct

Please note that this project is released with a Contributor Code of Conduct (see https://github.com/scrapy/scrapy/blob/master/CODE_OF_CONDUCT.md).

By participating in this project you agree to abide by its terms. Please report unacceptable behavior to opensource@scrapinghub.com.

Companies using Scrapy

See https://scrapy.org/companies/

Commercial Support

See https://scrapy.org/support/

</html>