Scrapy, a fast high-level web crawling & scraping framework for Python.
Go to file
Elias Dorneles 3ac3ac4d92 docs: update data flow description and image (fixes: #2278)
This fixes the explanation to use Requests instead of URLs,
which is what actually happens, and is also consistent with the
new tutorial, which already explains how URLs become Request objects.

I've also changed the "loop", jumping from 9 to step 2.
2016-09-28 16:38:45 -03:00
artwork added artwork files properly now 2012-03-20 10:46:45 -03:00
debian Merge pull request #934 from Dineshs91/zsh-support 2015-07-31 03:18:19 +05:00
docs docs: update data flow description and image (fixes: #2278) 2016-09-28 16:38:45 -03:00
extras Merge pull request #934 from Dineshs91/zsh-support 2015-07-31 03:18:19 +05:00
scrapy Merge pull request #2229 from ahlinc/fix_shell_completion 2016-09-22 11:15:30 -03:00
sep Spelling fixes 2015-12-13 19:39:48 -08:00
tests Merge pull request #2243 from pawelmhm/image-pipeline-2198 2016-09-19 18:43:52 +02:00
.bumpversion.cfg Allow more pre-releases with bumpversion 2016-04-21 16:51:17 +02:00
.coveragerc Add coverage report trough codecov.io 2015-08-13 13:56:24 -03:00
.gitignore add coverage files to gitignore 2015-08-26 01:58:33 +05:00
.travis.yml Remove "precise" test env from Travis-CI config 2016-09-01 11:17:53 +02:00
AUTHORS added Nicolas Ramirez to AUTHORS 2013-03-14 12:44:39 -03:00
CODE_OF_CONDUCT.md Add Code of Conduct Version 1.3.0 from http://contributor-covenant.org/ 2016-01-15 18:01:04 +01:00
CONTRIBUTING.md Put a blurb about support channels in CONTRIBUTING 2015-07-24 01:48:43 +00:00
INSTALL fix link to online installation instructions 2012-10-02 12:26:14 +01:00
LICENSE mv scrapy/trunk to root as part of svn2hg migration 2009-05-06 15:55:17 -03:00
MANIFEST.in ENH: include tests/ to source distribution in MANIFEST.in 2015-06-25 23:00:00 -04:00
Makefile.buildbot Generated version as pep440 and dpkg compatible 2015-06-16 00:16:09 +00:00
NEWS added NEWS file pointing to docs/news.rst 2012-04-28 23:32:51 -03:00
README.rst Remove download stats badge 2016-08-01 12:16:50 -03:00
conftest.py Simplify if statement 2016-01-18 07:45:36 +01:00
pytest.ini Don't collect tests by their class name 2015-05-04 18:10:04 -03:00
requirements-py3.txt Bump w3lib version dependency in setup.py 2016-04-29 10:29:37 +02:00
requirements.txt Use w3lib.url.canonicalize_url() from w3lib 1.15.0 2016-08-16 17:42:16 +05:30
setup.cfg Build universal wheels 2016-03-01 11:00:20 +01:00
setup.py Use w3lib.url.canonicalize_url() from w3lib 1.15.0 2016-08-16 17:42:16 +05:30
tox.ini Add Debian Jessie test env 2016-09-01 10:19:49 +02:00

README.rst

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>

Scrapy

Overview

Scrapy is a fast high-level web crawling and web scraping framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

For more information including a list of features check the Scrapy homepage at: http://scrapy.org

Requirements

  • Python 2.7 or Python 3.3+
  • Works on Linux, Windows, Mac OSX, BSD

Install

The quick way:

pip install scrapy

For more details see the install section in the documentation: http://doc.scrapy.org/en/latest/intro/install.html

Releases

You can download the latest stable and development releases from: http://scrapy.org/download/

Documentation

Documentation is available online at http://doc.scrapy.org/ and in the docs directory.

Community (blog, twitter, mail list, IRC)

See http://scrapy.org/community/

Contributing

Please note that this project is released with a Contributor Code of Conduct (see https://github.com/scrapy/scrapy/blob/master/CODE_OF_CONDUCT.md).

By participating in this project you agree to abide by its terms. Please report unacceptable behavior to opensource@scrapinghub.com.

See http://doc.scrapy.org/en/master/contributing.html

Companies using Scrapy

See http://scrapy.org/companies/

Commercial Support

See http://scrapy.org/support/

</html>