Scrapy shell refactoring

This commit is contained in:
Pablo Hoffman 2010-08-20 11:26:14 -03:00
parent 7858244dca
commit 1d3b9e2ca8
8 changed files with 204 additions and 156 deletions

View File

@ -16,14 +16,12 @@ Scrapy provides a built-in mechanism for extracting data (called
all, they're just parsing libraries which can be imported and used from any
Python code.
In other words, comparing `BeautifulSoup`_ or `lxml`_ to Scrapy is like
comparing `urllib`_ or `urlparse`_ to `Django`_ (a popular Python web
application framework).
In other words, comparing `BeautifulSoup`_ (or `lxml`_) to Scrapy is like
comparing `jinja2`_ to `Django`_.
.. _BeautifulSoup: http://www.crummy.com/software/BeautifulSoup/
.. _lxml: http://codespeak.net/lxml/
.. _urllib: http://docs.python.org/library/urllib.html
.. _urlparse: http://docs.python.org/library/urlparse.html
.. _jinja2: http://jinja.pocoo.org/2/
.. _Django: http://www.djangoproject.com
Does Scrapy work with Python 3.0?
@ -81,6 +79,14 @@ My Scrapy crawler has memory leaks. What can I do?
See :ref:`topics-leaks`.
Also, Python has a builtin memory leak issue which is described in
:ref:`topics-leaks-without-leaks`.
How can I make Scrapy consume less memory?
------------------------------------------
See previous question.
Can I use Basic HTTP Authentication in my spiders?
--------------------------------------------------
@ -101,7 +107,8 @@ Scrapy comes with a built-in, fully functional project to scrape the `Google
Directory`_. You can find it in the ``examples/googledir`` directory of the
Scrapy distribution.
Also, there is a public repository of spiders called `Community Spiders`_.
Also, there's a site for sharing code snippets (spiders, middlewares,
extensions) called `Scrapy snippets`_.
Finally, you can find some example code for performing not-so-trivial tasks in
the `Scrapy Recipes`_ page.
@ -109,12 +116,13 @@ the `Scrapy Recipes`_ page.
.. _Google Directory: http://www.google.com/dirhp
.. _Community Spiders: http://dev.scrapy.org/wiki/CommunitySpiders
.. _Scrapy Recipes: http://dev.scrapy.org/wiki/ScrapyRecipes
.. _Scrapy snippets: http://snippets.scrapy.org/
Can I run a spider without creating a project?
----------------------------------------------
Yes. You can use the ``runspider`` command. For example, if you have a spider
written in a ``my_spider.py`` file you can run it with::
Yes. You can use the :command:`runspider` command. For example, if you have a
spider written in a ``my_spider.py`` file you can run it with::
scrapy runspider my_spider.py
@ -131,15 +139,6 @@ domains outside the ones covered by the spider.
For more info see:
:class:`~scrapy.contrib.spidermiddleware.offsite.OffsiteMiddleware`.
How can I make Scrapy consume less memory?
------------------------------------------
There's a whole documentation section about this subject, please see:
:ref:`topics-leaks`.
Also, Python has a builtin memory leak issue which is described in
:ref:`topics-leaks-without-leaks`.
What is the recommended way to deploy a Scrapy crawler in production?
---------------------------------------------------------------------
@ -191,3 +190,12 @@ higher) in your spider::
Or by setting a global download delay in your project with the
:setting:`DOWNLOAD_DELAY` setting.
Can I call ``pdb.set_trace()`` from my spiders to debug them?
-------------------------------------------------------------
Yes, but you can also use the Scrapy shell which allows you too quickly analyze
(and even modify) the response being processed by your spider, which is, quite
often, more useful than plain old ``pdb.set_trace()``.
For more info see :ref:`topics-shell-inspect-response`.

View File

@ -149,17 +149,16 @@ To put our spider to work, go to the project's top level directory and run::
The ``crawl dmoz.org`` command runs the spider for the ``dmoz.org`` domain. You
will get an output similar to this::
[-] Log opened.
[dmoz] INFO: Enabled extensions: ...
[dmoz] INFO: Enabled scheduler middlewares: ...
[dmoz] INFO: Enabled downloader middlewares: ...
[dmoz] INFO: Enabled spider middlewares: ...
[dmoz] INFO: Enabled item pipelines: ...
[dmoz.org] INFO: Spider opened
[dmoz.org] DEBUG: Crawled <http://www.dmoz.org/Computers/Programming/Languages/Python/Resources/> from <None>
[dmoz.org] DEBUG: Crawled <http://www.dmoz.org/Computers/Programming/Languages/Python/Books/> from <None>
[dmoz.org] INFO: Spider closed (finished)
[-] Main loop terminated.
2008-08-20 03:51:13-0300 [scrapy] INFO: Started project: dmoz
2008-08-20 03:51:13-0300 [dmoz] INFO: Enabled extensions: ...
2008-08-20 03:51:13-0300 [dmoz] INFO: Enabled scheduler middlewares: ...
2008-08-20 03:51:13-0300 [dmoz] INFO: Enabled downloader middlewares: ...
2008-08-20 03:51:13-0300 [dmoz] INFO: Enabled spider middlewares: ...
2008-08-20 03:51:13-0300 [dmoz] INFO: Enabled item pipelines: ...
2008-08-20 03:51:14-0300 [dmoz.org] INFO: Spider opened
2008-08-20 03:51:14-0300 [dmoz.org] DEBUG: Crawled <http://www.dmoz.org/Computers/Programming/Languages/Python/Resources/> from <None>
2008-08-20 03:51:14-0300 [dmoz.org] DEBUG: Crawled <http://www.dmoz.org/Computers/Programming/Languages/Python/Books/> from <None>
2008-08-20 03:51:14-0300 [dmoz.org] INFO: Spider closed (finished)
Pay attention to the lines containing ``[dmoz.org]``, which corresponds to
our spider (identified by the domain ``"dmoz.org"``). You can see a log line
@ -250,20 +249,18 @@ This is what the shell looks like::
[ ... Scrapy log here ... ]
Available objects:
2010-08-19 21:45:59-0300 [default] INFO: Spider closed (finished)
xxs <XmlXPathSelector (http://www.dmoz.org/Computers/Programming/Languages/Python/Books/) xpath=None>
url http://www.dmoz.org/Computers/Programming/Languages/Python/Books/
request <GET http://www.dmoz.org/Computers/Programming/Languages/Python/Books/>
spider <BaseSpider 'default' at 0x1b6c2d0>
response <200 http://www.dmoz.org/Computers/Programming/Languages/Python/Books/>
hxs <HtmlXPathSelector (http://www.dmoz.org/Computers/Programming/Languages/Python/Books/) xpath=None>
item Item()
Convenient shortcuts:
shelp() Print this help
fetch(req_or_url) Fetch a new request or URL and update shell objects
view(response) View response in a browser
[s] Available Scrapy objects:
[s] 2010-08-19 21:45:59-0300 [default] INFO: Spider closed (finished)
[s] hxs <HtmlXPathSelector (http://www.dmoz.org/Computers/Programming/Languages/Python/Books/) xpath=None>
[s] item Item()
[s] request <GET http://www.dmoz.org/Computers/Programming/Languages/Python/Books/>
[s] response <200 http://www.dmoz.org/Computers/Programming/Languages/Python/Books/>
[s] spider <BaseSpider 'default' at 0x1b6c2d0>
[s] xxs <XmlXPathSelector (http://www.dmoz.org/Computers/Programming/Languages/Python/Books/) xpath=None>
[s] Useful shortcuts:
[s] shelp() Print this help
[s] fetch(req_or_url) Fetch a new request or URL and update shell objects
[s] view(response) View response in a browser
In [1]:

View File

@ -106,7 +106,8 @@ scrapy.log module
.. data:: INFO
Log level for informational messages (recommended level for production)
Log level for informational messages (recommended level for production
deployments)
.. data:: DEBUG

View File

@ -31,7 +31,8 @@ for more info.
Launch the shell
================
To launch the shell type::
To launch the Scrapy shell you can use the :command:`shell` command like
this::
scrapy shell <url>
@ -40,12 +41,12 @@ Where the ``<url>`` is the URL you want to scrape.
Using the shell
===============
The Scrapy shell is just a regular Python console (or `IPython` shell if you
have it available) which provides some additional functions available by
default (as shortcuts):
The Scrapy shell is just a regular Python console (or `IPython` console if you
have it available) which provides some additional shortcut functions for
convenience.
Built-in Shortcuts
------------------
Available Shortcuts
-------------------
* ``shelp()`` - print a help with the list of available objects and shortcuts
@ -60,8 +61,8 @@ Built-in Shortcuts
.. _<base> tag: http://www.w3schools.com/TAGS/tag_base.asp
Built-in Objects
----------------
Available Scrapy objects
-------------------------
The Scrapy shell automatically creates some convenient objects from the
downloaded page, like the :class:`~scrapy.http.Response` object and the
@ -70,8 +71,6 @@ content).
Those objects are:
* ``url`` - the URL being analyzed
* ``spider`` - the Spider which is known to handle the URL, or a
:class:`~scrapy.spider.BaseSpider` object if there is no spider is found for
the current URL
@ -108,23 +107,21 @@ First, we launch the shell::
scrapy shell http://scrapy.org --nolog
Then, the shell fetches the url (using the Scrapy downloader) and prints the
list of available objects and some help::
Then, the shell fetches the URL (using the Scrapy downloader) and prints the
list of available objects and useful shortcuts (you'll notice that these lines
all start with the ``[s]`` prefix)::
Fetching <http://scrapy.org>...
Available objects
xxs <XmlXPathSelector (http://scrapy.org) xpath=None>
url http://scrapy.org
request <http://scrapy.org>
spider <scrapy.spider.models.BaseSpider object at 0x2bed9d0>
hxs <HtmlXPathSelector (http://scrapy.org) xpath=None>
item Item()
response <http://scrapy.org>
Available shortcuts
shelp() Prints this help.
fetch(req_or_url) Fetch a new request or URL and update objects
view(response) View response in a browser
[s] Available objects
[s] hxs <HtmlXPathSelector (http://scrapy.org) xpath=None>
[s] item Item()
[s] request <http://scrapy.org>
[s] response <http://scrapy.org>
[s] spider <scrapy.spider.models.BaseSpider object at 0x2bed9d0>
[s] xxs <XmlXPathSelector (http://scrapy.org) xpath=None>
[s] Useful shortcuts:
[s] shelp() Prints this help.
[s] fetch(req_or_url) Fetch a new request or URL and update objects
[s] view(response) View response in a browser
>>>
@ -144,6 +141,8 @@ After that, we can stary playing with the objects::
>>>
.. _topics-shell-inspect-response:
Invoking the shell from spiders to inspect responses
====================================================

View File

@ -4,6 +4,8 @@ Scrapy Shell
See documentation in docs/topics/shell.rst
"""
from scrapy.core.manager import scrapymanager
from scrapy.core.queue import KeepAliveExecutionQueue
from scrapy.command import ScrapyCommand
from scrapy.shell import Shell
@ -28,5 +30,7 @@ class Command(ScrapyCommand):
def run(self, args, opts):
url = args[0] if args else None
shell = Shell(self.update_vars)
shell.start(url)
shell = Shell(scrapymanager, update_vars=self.update_vars, inthread=True)
shell.start(url=url).addBoth(lambda _: scrapymanager.stop())
scrapymanager.queue = KeepAliveExecutionQueue()
scrapymanager.start()

View File

@ -4,73 +4,69 @@ Scrapy Shell
See documentation in docs/topics/shell.rst
"""
import os
import urllib
import urlparse
import signal
from twisted.internet import reactor, threads
from twisted.python.failure import Failure
from scrapy import log
from scrapy.item import BaseItem
from scrapy.spider import BaseSpider, spiders
from scrapy.selector import XmlXPathSelector, HtmlXPathSelector
from scrapy.selector import XPathSelector, XmlXPathSelector, HtmlXPathSelector
from scrapy.utils.misc import load_object
from scrapy.utils.response import open_in_browser
from scrapy.utils.url import any_to_uri
from scrapy.utils.console import start_python_console
from scrapy.conf import settings
from scrapy.core.manager import scrapymanager
from scrapy.core.queue import KeepAliveExecutionQueue
from scrapy.http import Request, TextResponse
def relevant_var(varname):
return varname not in ['shelp', 'fetch', 'view', '__builtins__', 'In', \
'Out', 'help', 'namespace'] and not varname.startswith('_')
def parse_url(url):
"""Parse url which can be a direct path to a direct file"""
url = url.strip()
if url:
u = urlparse.urlparse(url)
if not u.scheme:
path = os.path.abspath(url).replace(os.sep, '/')
url = 'file://' + urllib.pathname2url(path)
u = urlparse.urlparse(url)
return url
from scrapy.http import Request, Response, TextResponse
class Shell(object):
def __init__(self, update_vars=None, nofetch=False):
self.vars = {}
self.update_vars = update_vars
self.item_class = load_object(settings['DEFAULT_ITEM_CLASS'])
self.nofetch = nofetch
relevant_classes = (BaseSpider, Request, Response, BaseItem, XPathSelector)
def fetch(self, request_or_url, print_help=False):
def __init__(self, crawler, update_vars=None, inthread=False):
self.crawler = crawler
self.vars = {}
self.update_vars = update_vars or (lambda x: None)
self.item_class = load_object(settings['DEFAULT_ITEM_CLASS'])
self.inthread = inthread
def start(self, *a, **kw):
# disable accidental Ctrl-C key press from shutting down the engine
signal.signal(signal.SIGINT, signal.SIG_IGN)
if self.inthread:
return threads.deferToThread(self._start, *a, **kw)
else:
self._start(*a, **kw)
def _start(self, url=None, request=None, response=None, spider=None):
if url:
self.fetch(url, spider)
elif request:
self.fetch(request, spider)
elif response:
request = response.request
self.populate_vars(request.url, response, request, spider)
start_python_console(self.vars)
def fetch(self, request_or_url, spider=None):
if isinstance(request_or_url, Request):
request = request_or_url
url = request.url
else:
url = parse_url(request_or_url)
url = any_to_uri(request_or_url)
request = Request(url, dont_filter=True)
spider = spiders.create_for_request(request, BaseSpider('default'), \
log_multiple=True)
print "Fetching %s..." % request
scrapymanager.engine.open_spider(spider)
if spider is None:
spider = spiders.create_for_request(request, BaseSpider('default'), \
log_multiple=True)
self.crawler.engine.open_spider(spider)
response = None
try:
response = threads.blockingCallFromThread(reactor, \
scrapymanager.engine.schedule, request, spider)
self.crawler.engine.schedule, request, spider)
except:
log.err(Failure(), "Error fetching response", spider=spider)
self.populate_vars(url, response, request, spider)
if print_help:
self.print_help()
else:
print "Done - use shelp() to see available objects"
def populate_vars(self, url=None, response=None, request=None, spider=None):
item = self.item_class()
@ -79,59 +75,35 @@ class Shell(object):
if isinstance(response, TextResponse):
self.vars['xxs'] = XmlXPathSelector(response)
self.vars['hxs'] = HtmlXPathSelector(response)
self.vars['url'] = url
self.vars['response'] = response
self.vars['request'] = request
self.vars['spider'] = spider
if not self.nofetch:
if self.inthread:
self.vars['fetch'] = self.fetch
self.vars['view'] = open_in_browser
self.vars['shelp'] = self.print_help
if self.update_vars:
self.update_vars(self.vars)
self.update_vars(self.vars)
self.print_help()
def print_help(self):
print
print "Available objects:"
for k, v in self.vars.iteritems():
if relevant_var(k):
print " %-10s %s" % (k, v)
print
print "Convenient shortcuts:"
print " shelp() Print this help"
if not self.nofetch:
print " fetch(req_or_url) Fetch a new request or URL and update shell objects"
print " view(response) View response in a browser"
print
self.p("Available Scrapy objects:")
for k, v in sorted(self.vars.iteritems()):
if self._is_relevant(v):
self.p(" %-10s %s" % (k, v))
self.p("Useful shortcuts:")
self.p(" shelp() Shell help (print this help)")
if self.inthread:
self.p(" fetch(req_or_url) Fetch request (or URL) and update local objects")
self.p(" view(response) View response in a browser")
def start(self, url):
# disable accidental Ctrl-C key press from shutting down the engine
signal.signal(signal.SIGINT, signal.SIG_IGN)
def p(self, line=''):
print "[s] %s" % line
reactor.callInThread(self._console_thread, url)
scrapymanager.queue = KeepAliveExecutionQueue()
scrapymanager.start()
def _is_relevant(self, value):
return isinstance(value, self.relevant_classes)
def inspect_response(self, response):
print
print "Scrapy Shell - inspecting response: %s" % response
print "Use shelp() to see available objects"
print
request = response.request
url = request.url
self.populate_vars(url, response, request)
start_python_console(self.vars)
def _console_thread(self, url=None):
self.populate_vars()
if url:
result = self.fetch(url, print_help=True)
else:
self.print_help()
start_python_console(self.vars)
reactor.callFromThread(scrapymanager.stop)
def inspect_response(response):
def inspect_response(response, spider=None):
"""Open a shell to inspect the given response"""
Shell(nofetch=True).inspect_response(response)
from scrapy.core.manager import scrapymanager
Shell(scrapymanager).start(response=response, spider=spider)

View File

@ -1,8 +1,9 @@
import os
import unittest
from scrapy.spider import BaseSpider
from scrapy.utils.url import url_is_from_any_domain, safe_url_string, safe_download_url, \
url_query_parameter, add_or_replace_parameter, url_query_cleaner, canonicalize_url, \
urljoin_rfc, url_is_from_spider
urljoin_rfc, url_is_from_spider, file_uri_to_path, path_to_file_uri, any_to_uri
class UrlUtilsTest(unittest.TestCase):
@ -273,6 +274,51 @@ class UrlUtilsTest(unittest.TestCase):
self.assertEqual(canonicalize_url(u'http://www.example.com/caf%E9-con-leche.htm'),
'http://www.example.com/caf%E9-con-leche.htm')
def test_path_to_file_uri(self):
if os.name == 'nt':
self.assertEqual(path_to_file_uri("C:\\windows\clock.avi"),
"file:///C|/windows/clock.avi")
else:
self.assertEqual(path_to_file_uri("/some/path.txt"),
"file:///some/path.txt")
fn = "test.txt"
x = path_to_file_uri(fn)
self.assert_(x.startswith('file:///'))
self.assertEqual(file_uri_to_path(x), os.path.abspath(fn))
def test_file_uri_to_path(self):
if os.name == 'nt':
self.assertEqual(file_uri_to_path("file:///C|/windows/clock.avi"),
"C:\\windows\clock.avi")
uri = "file:///C|/windows/clock.avi"
uri2 = path_to_file_uri(file_uri_to_path(uri))
self.assertEqual(uri, uri2)
else:
self.assertEqual(file_uri_to_path("file:///path/to/test.txt"),
"/path/to/test.txt")
self.assertEqual(file_uri_to_path("/path/to/test.txt"),
"/path/to/test.txt")
uri = "file:///path/to/test.txt"
uri2 = path_to_file_uri(file_uri_to_path(uri))
self.assertEqual(uri, uri2)
self.assertEqual(file_uri_to_path("test.txt"),
"test.txt")
def test_any_to_uri(self):
if os.name == 'nt':
self.assertEqual(any_to_uri("C:\\windows\clock.avi"),
"file:///C|/windows/clock.avi")
else:
self.assertEqual(any_to_uri("/some/path.txt"),
"file:///some/path.txt")
self.assertEqual(any_to_uri("file:///some/path.txt"),
"file:///some/path.txt")
self.assertEqual(any_to_uri("http://www.example.com/some/path.txt"),
"http://www.example.com/some/path.txt")
if __name__ == "__main__":
unittest.main()

View File

@ -3,6 +3,7 @@ This module contains general purpose URL functions not found in the standard
library.
"""
import os
import re
import urlparse
import urllib
@ -154,3 +155,23 @@ def canonicalize_url(url, keep_blank_values=True, keep_fragments=False, \
path = urllib.quote(urllib.unquote(path))
fragment = '' if not keep_fragments else fragment
return urlparse.urlunparse((scheme, netloc, path, params, query, fragment))
def path_to_file_uri(path):
"""Convert local filesystem path to legal File URIs as described in:
http://en.wikipedia.org/wiki/File_URI_scheme
"""
x = urllib.pathname2url(os.path.abspath(path))
return 'file:///%s' % x.lstrip('/')
def file_uri_to_path(uri):
"""Convert File URI to local filesystem path according to:
http://en.wikipedia.org/wiki/File_URI_scheme
"""
return urllib.url2pathname(urlparse.urlparse(uri).path)
def any_to_uri(uri_or_path):
"""If given a path name, return its File URI, otherwise return it
unmodified
"""
u = urlparse.urlparse(uri_or_path)
return uri_or_path if u.scheme else path_to_file_uri(uri_or_path)