diff --git a/docs/faq.rst b/docs/faq.rst index fde98c112..f0b5c758e 100644 --- a/docs/faq.rst +++ b/docs/faq.rst @@ -121,7 +121,7 @@ written in a ``my_spider.py`` file you can run it with:: I get "Filtered offsite request" messages. How can I fix them? -------------------------------------------------------------- -Those messages (logged with ``DEBUG`` level) don't necesarilly mean there is a +Those messages (logged with ``DEBUG`` level) don't necessarily mean there is a problem, so you may not need to fix them. Those message are thrown by the Offsite Spider Middleware, which is a spider @@ -130,3 +130,12 @@ domains outside the ones covered by the spider. For more info see: :class:`~scrapy.contrib.spidermiddleware.offsite.OffsiteMiddleware`. + +How can I make Scrapy consume less memory? +------------------------------------------ + +There's a whole documentation section about this subject, please see: +:ref:`topics-leaks`. + +Also, Python has a builtin memory leak issue which is described in +:ref:`topics-leaks-without-leaks`. diff --git a/docs/topics/leaks.rst b/docs/topics/leaks.rst index b9214a23d..bdb121baf 100644 --- a/docs/topics/leaks.rst +++ b/docs/topics/leaks.rst @@ -56,9 +56,9 @@ Debugging memory leaks with ``trackref`` memory leaks. It basically tracks the references to all live Requests, Responses, Item and Selector objects. -To active the ``trackref`` module, enable the :setting:`TRACK_REFS` setting. It -only imposes a minor performance impact so it should be OK for use it in -production environments. +To activate the ``trackref`` module, enable the :setting:`TRACK_REFS` setting. +It only imposes a minor performance impact so it should be OK for use it, even +in production environments. Once you have ``trackref`` enabled you can enter the telnet console and inspect how many objects (of the classes mentioned above) are currently alive using the @@ -70,6 +70,7 @@ how many objects (of the classes mentioned above) are currently alive using the >>> prefs() Live References + ExampleSpider 1 oldest: 15s ago HtmlResponse 10 oldest: 1s ago XPathSelector 2 oldest: 0s ago FormRequest 878 oldest: 7s ago @@ -82,6 +83,19 @@ looking at the oldest request or response. You can get the oldest object of each class using the :func:`~scrapy.utils.trackref.get_oldest` function like this (from the telnet console). +Which objects are tracked? +-------------------------- + +The objects tracked by ``trackrefs`` are all from these classes (and all its +subclasses): + +* ``scrapy.http.Request`` +* ``scrapy.http.Response`` +* ``scrapy.item.Item`` +* ``scrapy.selector.XPathSelector`` +* ``scrapy.spider.BaseSpider`` +* ``scrapy.selector.document.Libxml2Document`` + A real example -------------- @@ -106,6 +120,7 @@ references:: >>> prefs() Live References + SomenastySpider 1 oldest: 15s ago HtmlResponse 3890 oldest: 265s ago XPathSelector 2 oldest: 0s ago Request 3878 oldest: 250s ago @@ -124,6 +139,28 @@ to the ``somenastyspider.com`` spider. We can now go and check the code of that spider to discover the nasty line that is generating the leaks (passing response references inside requests). +If you want to iterate over all objects, instead of getting the oldest one, you +can use the :func:`iter_all` function:: + + >>> from scrapy.utils.trackref import iter_all + >>> [r.url for r in iter_all('HtmlResponse')] + ['http://www.somenastyspider.com/product.php?pid=123', + 'http://www.somenastyspider.com/product.php?pid=584', + ... + +Too many spiders? +----------------- + +If your project has too many spiders, the output of ``prefs()`` can be +difficult to read. For this reason, that function has a ``ignore`` argument +which can be used to ignore a particular class (and all its subclases). For +example, using:: + + >>> from scrapy.spider import BaseSpider + >>> prefs(ignore=BaseSpider) + +Won't show any live references to spiders. + .. module:: scrapy.utils.trackref :synopsis: Track references of live objects @@ -137,15 +174,25 @@ Here are the functions available in the :mod:`~scrapy.utils.trackref` module. Inherit from this class (instead of object) if you want to track live instances with the ``trackref`` module. -.. function:: print_live_refs(class_name) +.. function:: print_live_refs(class_name, ignore=NoneType) Print a report of live references, grouped by class name. + :param ignore: if given, all objects from the specified class (or tuple of + classes) will be ignored. + :type ignore: class or classes tuple + .. function:: get_oldest(class_name) Return the oldest object alive with the given class name, or ``None`` if none is found. Use :func:`print_live_refs` first to get a list of all - tracked live objects. + tracked live objects per class name. + +.. function:: iter_all(class_name) + + Return an iterator over all objects alive with the given class name, or + ``None`` if none is found. Use :func:`print_live_refs` first to get a list + of all tracked live objects per class name. .. _topics-leaks-guppy: @@ -161,7 +208,7 @@ you still have another resource: the `Guppy library`_. .. _Guppy library: http://pypi.python.org/pypi/guppy -If you use setuptools, you can install Guppy with the following command:: +If you use ``setuptools``, you can install Guppy with the following command:: easy_install guppy @@ -211,3 +258,34 @@ knowledge about Python internals. For more info about Guppy, refer to the .. _Guppy documentation: http://guppy-pe.sourceforge.net/ +.. _topics-leaks-without-leaks: + +Leaks without leaks +=================== + +Sometimes, you may notice that the memory usage of your Scrapy process will +only increase, but never decrease. Unfortunately, this could happen even +though neither Scrapy nor your project are leaking memory. This is due to a +(not so well) known problem of Python, which may not return released memory to +the operating system in some cases. For more information on this issue see: + +* `Python Memory Management `_ +* `Python Memory Management Part 2 `_ +* `Python Memory Management Part 3 `_ + +The improvements proposed by Evan Jones, which are detailed in `this paper`_, +got merged in Python 2.5, but the only reduce the problem, it doesn't fixes it +completely. To quote the paper: + + *Unfortunately, this patch can only free an arena if there are no more + objects allocated in it anymore. This means that fragmentation is a large + issue. An application could have many megabytes of free memory, scattered + throughout all the arenas, but it will be unable to free any of it. This is + a problem experienced by all memory allocators. The only way to solve it is + to move to a compacting garbage collector, which is able to move objects in + memory. This would require significant changes to the Python interpreter.* + +This problem will be fixed in future Scrapy releases, where we plan to adopt a +new process model and run spiders in a pool of recyclable sub-processes. + +.. _this paper: http://evanjones.ca/memoryallocator/ diff --git a/scrapy/utils/trackref.py b/scrapy/utils/trackref.py index a9df06946..6006c1bbe 100644 --- a/scrapy/utils/trackref.py +++ b/scrapy/utils/trackref.py @@ -14,6 +14,7 @@ import weakref from collections import defaultdict from time import time from operator import itemgetter +from types import NoneType from scrapy.conf import settings @@ -33,7 +34,7 @@ class object_ref(object): if not settings.getbool('TRACK_REFS'): object_ref = object -def print_live_refs(): +def print_live_refs(ignore=NoneType): if object_ref is object: print "The trackref module is disabled. Use TRACK_REFS setting to enable it." return @@ -43,6 +44,8 @@ def print_live_refs(): for cls, wdict in live_refs.iteritems(): if not wdict: continue + if issubclass(cls, ignore): + continue oldest = min(wdict.itervalues()) print "%-30s %6d oldest: %ds ago" % (cls.__name__, len(wdict), \ now-oldest) @@ -52,3 +55,8 @@ def get_oldest(class_name): if cls.__name__ == class_name: if wdict: return min(wdict.iteritems(), key=itemgetter(1))[0] + +def iter_all(class_name): + for cls, wdict in live_refs.iteritems(): + if cls.__name__ == class_name: + return wdict.iterkeys()