Scrapy, a fast high-level web crawling & scraping framework for Python.
Go to file
Pablo Hoffman dd13dfe82b Raise error when settings module is missing.
Previously, it failed silently if an ImportError was caught when trying
to import the scrapy settings module. This not only happened when the
scrapy settings module itself was missing, but also when it tried to
import a missing module, which made the whole thing a bad idea.

A side effect of this change (not required, but for simplification) is
that we no longer support the default "scrapy_settings" name for the
scrapy settings module, but this was never used afaik.
2012-10-03 12:31:19 -03:00
.travis add precise to travis-ci build enviroments 2012-05-17 09:03:28 -03:00
artwork added artwork files properly now 2012-03-20 10:46:45 -03:00
bin return trial status as bin/runtests.sh exit value. #118 2012-04-20 16:40:55 -03:00
debian require w3lib 1.2 or greater 2012-05-15 17:11:51 -03:00
docs added spider contracts to release notes and warn that its API is still subject to change 2012-09-29 03:06:30 -03:00
extras simplified installation guide to only mention pip/easy_install mechanism, and provide hints for Windows users 2012-08-29 15:37:05 -03:00
scrapy Raise error when settings module is missing. 2012-10-03 12:31:19 -03:00
scrapyd simplified code for finished jobs times, updated docs for scrapyd 2012-09-12 11:10:49 +03:00
sep added sep directory with Scrapy Enhancement Proposal imported from old Trac site 2012-03-20 10:15:00 -03:00
.coveragerc Added rules to Makefile.buildbot for generating coverage reports 2010-12-15 11:13:45 -02:00
.gitignore SEP-017 contracts: various minor changes 2012-08-29 18:31:10 +02:00
.hgtags Added tag 0.10.3 for changeset 803efdb19e0b 2010-09-29 13:34:56 -03:00
.travis.yml add precise to travis-ci build enviroments 2012-05-17 09:03:28 -03:00
AUTHORS updated maintainer to scrapinghub 2012-05-02 03:25:35 -03:00
CONTRIBUTING.md renamed CONTRIBUTING to CONTRIBUTING.md so that links are rendered as links in github 2012-09-19 13:58:58 -03:00
INSTALL fix link to online installation instructions 2012-10-02 12:26:14 +01:00
LICENSE mv scrapy/trunk to root as part of svn2hg migration 2009-05-06 15:55:17 -03:00
MANIFEST.in removed obsolete entries from MANIFEST.in 2012-05-17 01:09:48 -03:00
Makefile.buildbot moved coverage report script to extras/ 2010-12-15 12:08:59 -02:00
NEWS added NEWS file pointing to docs/news.rst 2012-04-28 23:32:51 -03:00
README.rst add travis-ci build status to README 2012-05-17 09:07:36 -03:00
setup.cfg simplified installation guide to only mention pip/easy_install mechanism, and provide hints for Windows users 2012-08-29 15:37:05 -03:00
setup.py require w3lib 1.2 or greater 2012-05-15 17:11:51 -03:00
tox.ini Add tox.ini for tox (http://tox.testrun.org/) 2012-06-28 12:43:12 -07:00

README.rst

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>

Scrapy

https://secure.travis-ci.org/scrapy/scrapy.png?branch=master

Overview

Scrapy is a fast high-level screen scraping and web crawling framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

For more information including a list of features check the Scrapy homepage at: http://scrapy.org

Requirements

  • Python 2.6 or up
  • Works on Linux, Windows, Mac OSX, BSD

Install

The quick way:

pip install scrapy

For more details see the install section in the documentation: http://doc.scrapy.org/en/latest/intro/install.html

Releases

You can download the latest stable and development releases from: http://scrapy.org/download/

Documentation

Documentation is available online at http://doc.scrapy.org/ and in the docs directory.

Community (blog, twitter, mail list, IRC)

See http://scrapy.org/community/

Companies using Scrapy

See http://scrapy.org/companies/

Commercial Support

See http://scrapy.org/support/

</html>