6.3 KiB
Telnet Console
System Message: ERROR/3 (<stdin>, line 7)
Unknown directive type "module".
.. module:: scrapy.management.telnet :synopsis: The Telnet Console
Scrapy comes with a built-in telnet console for inspecting and controlling a Scrapy running process. The telnet console is just a regular python shell running inside the Scrapy process, so you can do literally anything from it.
The telnet console is a :ref:`built-in Scrapy extension <ref-extensions>` which comes enabled by default, but you can also disable it if you want. For more information about the extension itself see :ref:`ref-extensions-telnetconsole`.
System Message: ERROR/3 (<stdin>, line 14); backlink
Unknown interpreted text role "ref".System Message: ERROR/3 (<stdin>, line 14); backlink
Unknown interpreted text role "ref".System Message: ERROR/3 (<stdin>, line 18)
Unknown directive type "highlight".
.. highlight:: none
How to access the telnet console
The telnet console listens in the TCP port defined in the :setting:`TELNETCONSOLE_PORT` setting, which defaults to 6023. To access the console you need to type:
System Message: ERROR/3 (<stdin>, line 23); backlink
Unknown interpreted text role "setting".telnet localhost 6023 >>>
You need the telnet program which comes installed by default in Windows, and most Linux distros.
Available aliases in the telnet console
The telnet console is like a regular Python shell running inside the Scrapy process, so you can do anything from it including imports, etc.
However, the telnet console comes with some default aliases defined for convenience:
engine: the Scrapy engine object (scrapy.core.engine.scrapyengine)
manager: the Scrapy manager object (scrapy.core.manager.scrapymanager)
extensions: the extensions object (scrapy.extension.extensions)
stats: the Scrapy stats object (scrapy.stats.stats)
settings: the Scrapy settings object (scrapy.conf.settings)
p: the pprint function (pprint.pprint)
hpy: for memory debugging (see :ref:`topics-telnetconsole-leaks`)
System Message: ERROR/3 (<stdin>, line 48); backlink
Unknown interpreted text role "ref".
Some example of using the telnet console
Here are some example tasks you can do with the telnet console:
View engine status
You can use the st() method of the Scrapy engine to quickly show its state using the telnet console:
telnet localhost 6023 >>> engine.st() Execution engine status datetime.now()-self.start_time : 0:00:09.051588 self.is_idle() : False self.scheduler.is_idle() : False len(self.scheduler.pending_requests) : 1 self.downloader.is_idle() : False len(self.downloader.sites) : 1 self.downloader.has_capacity() : True self.pipeline.is_idle() : False len(self.pipeline.domaininfo) : 1 len(self._scraping) : 1 example.com self.domain_is_idle(domain) : False self.closing.get(domain) : None self.scheduler.domain_has_pending_requests(domain) : True len(self.scheduler.pending_requests[domain]) : 97 len(self.downloader.sites[domain].queue) : 17 len(self.downloader.sites[domain].active) : 25 len(self.downloader.sites[domain].transferring) : 8 self.downloader.sites[domain].closing : False self.downloader.sites[domain].lastseen : 2009-06-23 15:20:16.563675 self.pipeline.domain_is_idle(domain) : True len(self.pipeline.domaininfo[domain]) : 0 len(self._scraping[domain]) : 0
Pause, resume and stop Scrapy engine
To pause:
telnet localhost 6023 >>> engine.pause() >>>
To resume:
telnet localhost 6023 >>> engine.unpause() >>>
To stop:
telnet localhost 6023 >>> engine.stop() Connection closed by foreign host.
How to debug memory leaks using the telnet console
The Telnet Console can be used to debug memory leaks, for example, if your Scrapy process is getting too big. You need the guppy module available. If you use setuptools, you can install it by typing:
easy_install guppy
Here's an example to view all Python objects available in the heap:
>>> x = hpy.heap()
>>> x.bytype
Partition of a set of 297033 objects. Total size = 52587824 bytes.
Index Count % Size % Cumulative % Type
0 22307 8 16423880 31 16423880 31 dict
1 122285 41 12441544 24 28865424 55 str
2 68346 23 5966696 11 34832120 66 tuple
3 227 0 5836528 11 40668648 77 unicode
4 2461 1 2222272 4 42890920 82 type
5 16870 6 2024400 4 44915320 85 function
6 13949 5 1673880 3 46589200 89 types.CodeType
7 13422 5 1653104 3 48242304 92 list
8 3735 1 1173680 2 49415984 94 _sre.SRE_Pattern
9 1209 0 456936 1 49872920 95 scrapy.http.headers.Headers
<1676 more rows. Type e.g. '_.more' to view.>
You can see that most space is used by dicts. Then, if you want to see from which attribute those dicts are referenced you can do:
>>> x.bytype[0].byvia
Partition of a set of 22307 objects. Total size = 16423880 bytes.
Index Count % Size % Cumulative % Referred Via:
0 10982 49 9416336 57 9416336 57 '.__dict__'
1 1820 8 2681504 16 12097840 74 '.__dict__', '.func_globals'
2 3097 14 1122904 7 13220744 80
3 990 4 277200 2 13497944 82 "['cookies']"
4 987 4 276360 2 13774304 84 "['cache']"
5 985 4 275800 2 14050104 86 "['meta']"
6 897 4 251160 2 14301264 87 '[2]'
7 1 0 196888 1 14498152 88 "['moduleDict']", "['modules']"
8 672 3 188160 1 14686312 89 "['cb_kwargs']"
9 27 0 155016 1 14841328 90 '[1]'
<333 more rows. Type e.g. '_.more' to view.>
As you can see, the guppy module is very powerful, but also requires some knowledge about Python internals. For more info about guppy read the guppy documentation.