..
__init__.py
mv scrapy/trunk to root as part of svn2hg migration
2009-05-06 15:55:17 -03:00
conf.py
Reversed scrapy.cfg lookup order so that the one in the current project has more precedence. Also added alternative system-wide location for windows.
2010-10-18 03:18:54 -02:00
console.py
scrapy.utils.python: fixed bug introduced when adding support for new IPython 0.11. refs #335
2011-08-09 01:38:31 -03:00
datatypes.py
Initial support for a persistent scheduler, to support pausing and resuming
2011-08-02 11:57:55 -03:00
decorator.py
replaced DeprecationWarning by a new ScrapyDeprecationWarning category, since the default DeprecationWarning is silenced on Python 2.7+
2011-08-08 10:39:53 -03:00
defer.py
Improve spider errors logging which were previously logged as confusing "Unhandled errors" - closes #196
2010-08-13 01:45:47 -03:00
deprecate.py
replaced DeprecationWarning by a new ScrapyDeprecationWarning category, since the default DeprecationWarning is silenced on Python 2.7+
2011-08-08 10:39:53 -03:00
display.py
"parse" command refactoring. This fixes #173 and renders #106 invalid.
2010-08-22 05:04:17 -03:00
encoding.py
Some improvements to Response encoding support:
2010-03-25 15:47:10 -03:00
engine.py
Another step towards singleton removal: deprecated crawler singleton import (from scrapy.project import crawler) by a new class method that extensions can implement to receive the crawler
2011-08-08 11:42:44 -03:00
ftp.py
Added new Feed exports extension with documentation and storage tests. Closes #197 .
2010-08-17 14:27:48 -03:00
gz.py
fixed bug handling truncated gzipped responses. closes #319
2011-06-06 18:25:14 -03:00
http.py
added new downloader middleware: ChunkedTransferMiddleware
2011-08-09 03:03:25 -03:00
httpobj.py
remove Url class and use str instead for Request and Response urls. Also added urlparse_cached function for achieving the same caching functionality provided by old Url class
2009-08-17 21:16:55 -03:00
iterators.py
fixed csviter bug when called with a Response, not TextResponse
2010-12-14 18:21:07 -02:00
job.py
Initial support for a persistent scheduler, to support pausing and resuming
2011-08-02 11:57:55 -03:00
jsonrpc.py
Check that arguments and keyword arguments are not passed simultaneously in jsonrpc_client_call()
2010-08-27 13:45:52 -03:00
markup.py
Moved several functions from scrapy.utils.{http,markup,multipart,response,url} (and their tests) to a new library called 'w3lib'. Scrapy will now depend on w3lib.
2011-04-18 22:37:19 -03:00
memory.py
fixed compatibility with python 2.5 and removed unused code
2011-01-05 12:03:54 -02:00
misc.py
ported internal scrapy.utils imports to w3lib
2011-04-19 01:33:52 -03:00
multipart.py
Moved several functions from scrapy.utils.{http,markup,multipart,response,url} (and their tests) to a new library called 'w3lib'. Scrapy will now depend on w3lib.
2011-04-18 22:37:19 -03:00
ossignal.py
Some enhancements to Scrapy core:
2009-09-03 08:27:48 -03:00
pqueue.py
fixed subtle bug in disk-based priority queues caused by serialization errors, and added tests
2011-09-02 09:40:52 -03:00
project.py
moved scrapy.utils.sqlite to scrapyd.sqlite
2011-08-27 01:20:57 -03:00
py26.py
Removed scrapy.xlib.simplejson module. Scrapy now requires simplejson if running on Python 2.5. Closes #289
2010-11-16 03:11:12 -02:00
py27.py
Added cached DNS resolver based on old caching resolver extension from scrapy.contrib.resolver. This new one is *not* an extension, it comes builtin and always enabled.
2011-07-27 03:45:15 -03:00
python.py
Downloader: Added support for dynamically adjusting download delay and maximum concurrent requests
2011-05-14 21:35:46 -03:00
queue.py
Some changes to persistent scheduler after some initial usage feedback:
2011-09-23 13:03:07 -03:00
reactor.py
more core refactoring including moving engine next request call logic to a separate class
2011-07-25 10:46:00 -03:00
reqser.py
Initial support for a persistent scheduler, to support pausing and resuming
2011-08-02 11:57:55 -03:00
request.py
Scheduler refactoring which introduces the following changes:
2011-06-06 03:16:56 -03:00
response.py
forked UnicodeDammit from BeautifulSoup to explicitly disable usage of chardet library
2011-07-20 17:41:53 -03:00
serialize.py
Another step towards singleton removal: deprecated crawler singleton import (from scrapy.project import crawler) by a new class method that extensions can implement to receive the crawler
2011-08-08 11:42:44 -03:00
signal.py
Added support for returning deferreds from (some) signal handlers. Closes #193
2010-08-14 21:10:37 -03:00
sitemap.py
scrapy.utils.sitemap: added one more case of parsing invalid sitemaps
2011-08-07 03:24:32 -03:00
spider.py
Added ISpiderManager interface and a test to verify the default SpiderManager comforms to it
2010-09-03 14:29:27 -03:00
template.py
Updated CAMELCASE_EXCLUDE_CHARS to also exclude digits (patch by Adam Wentz)
2011-07-06 20:11:11 -03:00
test.py
* Added lxml backend for XPath selectors. Closes #147
2010-10-25 14:47:10 -02:00
testproc.py
utils.testproc: make spawned process use the original CWD, instead of the temporary one created by twisted trial
2010-11-29 09:56:14 -02:00
testsite.py
* Added tests for shell/fetch/version commands ( closes #255 )
2010-11-28 18:14:45 -02:00
trackref.py
added trackref stats to memory debugger report. closes #272
2010-10-27 21:18:58 -02:00
txweb.py
Added new Scrapy service with support for:
2010-09-03 15:54:42 -03:00
url.py
Some Link extractor improvements:
2011-05-18 12:32:34 -03:00