Scrapy, a fast high-level web crawling & scraping framework for Python.
Go to file
Capi Etheriel bc17e9d412 Adds HtmlCSSSelector and XmlCSSSelector classes, cssselect as optional dependency.
Ported .get() from _Element and .text_content() from HTMLMixin

Add CSS selectors to scrapy shell

Documenting CSS Selectors: Constructing selectors

Documenting CSS Selectors: Using Selectors

Make CSS Selectors a default feature.

Adds XPath powers to CSS Selectors and some syntactic sugar.

Removes methods copied over from lxml.html.HtmlMixin.

Updating docs to use new CSS Selector super powers.

Documenting CSS Selectors: Regular Expressions

Moving section after Nesting section, since it mentions it.

Documenting CSS Selectors: Nesting Selectors

Fix XPath specificity in lxml.selector.CSSSelectorMixin.text

Cleaning up unused stuff from cssel.py

Changing the behavior of lxml.selector.CSSSelectorMixin.text.

Concatenating all of the descendant text nodes is more useful
than returning it in pieces (there's xpath() if you need that).

Documenting CSS Selectors: CSS Selector objects

Documenting CSS Selectors: CSSSelectorList objects

Documenting CSS Selectors: HtmlCSSSelector objects

Documenting CSS Selectors: XmlCSSSelector objects

Fixing some documentations typos and errors

Enforcing the 80-char width lines

Tidying up CSS selectors and CSSSelectorMixin objects

Adding some missing references in documentation.

Fixing lxml.selector.CSSSelectorList.text
2013-10-10 18:23:15 -02:00
.travis mock is required for running tests since 45ff6ec28a 2013-09-18 00:02:23 +06:00
artwork added artwork files properly now 2012-03-20 10:46:45 -03:00
bin remove scrapyd, it was migrated to its own repository 2013-02-06 05:24:07 +00:00
debian update copyright notes 2013-05-16 15:05:52 -03:00
docs Adds HtmlCSSSelector and XmlCSSSelector classes, cssselect as optional dependency. 2013-10-10 18:23:15 -02:00
extras get scrapy version from package data 2013-02-06 11:44:26 -02:00
scrapy Adds HtmlCSSSelector and XmlCSSSelector classes, cssselect as optional dependency. 2013-10-10 18:23:15 -02:00
sep sep-019: fixed typo 2013-03-12 18:44:09 -03:00
.coveragerc Added rules to Makefile.buildbot for generating coverage reports 2010-12-15 11:13:45 -02:00
.gitignore Instead of extending from HttpCachePolicy, following the same approach used for storage selection 2012-12-28 16:11:47 +01:00
.travis.yml no need to test latest versions of dependencies on python 2.6 2013-08-27 17:38:28 -03:00
AUTHORS added Nicolas Ramirez to AUTHORS 2013-03-14 12:44:39 -03:00
CONTRIBUTING.md renamed CONTRIBUTING to CONTRIBUTING.md so that links are rendered as links in github 2012-09-19 13:58:58 -03:00
INSTALL fix link to online installation instructions 2012-10-02 12:26:14 +01:00
LICENSE mv scrapy/trunk to root as part of svn2hg migration 2009-05-06 15:55:17 -03:00
MANIFEST.in get scrapy version from package data 2013-02-06 11:44:26 -02:00
Makefile.buildbot Fix permission and set umask before generating sdist tarball 2013-09-03 23:22:17 -03:00
NEWS added NEWS file pointing to docs/news.rst 2012-04-28 23:32:51 -03:00
README.rst added pypi version badge to README 2013-10-03 12:47:22 -03:00
requirements.txt updated required twisted version to 10.0 2013-10-01 14:07:38 -03:00
setup.cfg remove no longer existent examples from doc_files used in bdist_rpm. closes GH-417 2013-10-08 15:18:45 -02:00
setup.py Adds HtmlCSSSelector and XmlCSSSelector classes, cssselect as optional dependency. 2013-10-10 18:23:15 -02:00
tox.ini tox.ini: disable sitepackages on windows, as a compiler is often not available 2013-08-12 18:59:41 -03:00

README.rst

<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en"> <head> </head>

Scrapy

https://badge.fury.io/py/Scrapy.png https://secure.travis-ci.org/scrapy/scrapy.png?branch=master

Overview

Scrapy is a fast high-level screen scraping and web crawling framework, used to crawl websites and extract structured data from their pages. It can be used for a wide range of purposes, from data mining to monitoring and automated testing.

For more information including a list of features check the Scrapy homepage at: http://scrapy.org

Requirements

  • Python 2.6 or up
  • Works on Linux, Windows, Mac OSX, BSD

Install

The quick way:

pip install scrapy

For more details see the install section in the documentation: http://doc.scrapy.org/en/latest/intro/install.html

Releases

You can download the latest stable and development releases from: http://scrapy.org/download/

Documentation

Documentation is available online at http://doc.scrapy.org/ and in the docs directory.

Community (blog, twitter, mail list, IRC)

See http://scrapy.org/community/

Companies using Scrapy

See http://scrapy.org/companies/

Commercial Support

See http://scrapy.org/support/

</html>