6.8 KiB
Installation guide
This document describes how to install Scrapy in Linux, Windows and Mac OS X systems and it consists on the following 3 steps:
-
System Message: ERROR/3 (<stdin>, line 10); backlink
Unknown interpreted text role "ref".
-
System Message: ERROR/3 (<stdin>, line 11); backlink
Unknown interpreted text role "ref".
-
System Message: ERROR/3 (<stdin>, line 12); backlink
Unknown interpreted text role "ref".
Requirements
- Python 2.5 or 2.6 (3.x is not yet supported)
- Twisted 2.5.0, 8.0 or above (Windows users: you may need to install pywin32 because of this Twisted bug)
- libxml2 (2.6.28 or above is recommended)
Optional:
- pyopenssl (for HTTPS support, highly recommended)
- spidermonkey (for parsing Javascript)
Step 1. Install Python
Scrapy works with Python 2.5 or 2.6, you can get it at http://www.python.org/download/
System Message: ERROR/3 (<stdin>, line 44)
Unknown directive type "highlight".
.. highlight:: sh
Step 2. Install required libraries
The procedure for installing the required third party libraries depends on the platform and operating system you use.
Ubuntu/Debian
If you're running Ubuntu/Debian Linux run the following command as root:
apt-get install python-twisted python-libxml2
To install optional libraries:
apt-get install python-pyopenssl spidermonkey-bin
Arch Linux
If you are running Arch Linux run the following command as root:
pacman -S twisted libxml2
To install optional libraries:
pacman -S pyopenssl spidermonkey
Mac OS X
First, download Twisted for Mac.
Mac OS X ships an libxml2 version too old to be used by Scrapy. Also, by looking on the web it seems that installing libxml2 on MacOSX is a bit of a challenge. Here is a way to achieve this, though not acceptable on the long run:
Fetch the following libxml2 and libxslt packages:
Extract, build and install them both with:
./configure --with-python=/Library/Frameworks/Python.framework/Versions/2.5/ make sudo make install
Replacing /Library/Frameworks/Python.framework/Version/2.5/ with your current python framework location.
Install libxml2 Python bidings with:
cd libxml2-2.7.3/python sudo make install
The libraries and modules should be installed in something like /usr/local/lib/python2.5/site-packages. Add it to your PYTHONPATH and you are done.
Check the libxml2 library was installed propertly with:
python -c 'import libxml2'
Windows
Download and install:
- Twisted for Windows - you may need to install pywin32 because of this Twisted bug
- libxml2 for Windows
- PyOpenSSL for Windows
Step 3. Install Scrapy
We're working hard to get the first release of Scrapy out. In the meantime, please download the latest development version from the Subversion repository.
Just follow these steps:
3.1. Install Subversion
Make sure that you have Subversion installed, and that you can run its commands from a shell. (Enter svn help at a shell prompt to test this.)
3.2. Check out the Scrapy source code
By running the following command:
svn checkout http://svn.scrapy.org/scrapy/trunk/ scrapy-trunk
3.3. Install the Scrapy module
Install the Scrapy module by running the following commands:
cd scrapy-trunk python setup.py install
If you're on Unix-like systems (Linux, Mac, etc) you may need to run the second command with root privileges, for example by running:
sudo python setup.py install
Warning
In Windows, you may need to add the C:\Python25\Scripts folder to the system path by adding that directory to the PATH environment variable from the Control Panel.
Warning
Keep in mind that Scrapy is still being changed, as we haven't yet released the first stable version. So it's important that you keep updating the Subversion code periodically and reinstalling the Scrapy module. A more convenient way is to use Scrapy module without installing it (see below).
Use Scrapy without installing it
Another alternative is to use the Scrapy module without installing it which makes it easier to keep using the last Subversion code without having to reinstall it everytime you do a svn update.
You can do this by following the next steps:
Add Scrapy to your Python path
If you're on Linux, Mac or any Unix-like system, you can make a symbolic link to your system site-packages directory like this:
ln -s /path/to/scrapy-trunk/scrapy SITE-PACKAGES/scrapy
Where SITE-PACKAGES is the location of your system site-packages directory. To find this out execute the following:
python -c "from distutils.sysconfig import get_python_lib; print get_python_lib()"
Alternatively, you can define your PYTHONPATH environment variable so that it includes the scrapy-trunk directory. This solution also works on Windows systems, which don't support symbolic links. (Environment variables can be defined on Windows systems from the Control Panel).
Unix-like example:
PYTHONPATH=/path/to/scrapy-trunk
Windows example (from command line, but you should probably use the Control Panel):
set PYTHONPATH=C:\path\to\scrapy-trunk
Make the scrapy-admin.py script available
On Unix-like systems, create a symbolic link to the file scrapy-trunk/scrapy/bin/scrapy-admin.py in a directory on your system path, such as /usr/local/bin. For example:
ln -s `pwd`/scrapy-trunk/scrapy/bin/scrapy-admin.py /usr/local/bin
This simply lets you type scrapy-admin.py from within any directory, rather than having to qualify the command with the full path to the file.
On Windows systems, the same result can be achieved by copying the file scrapy-trunk/scrapy/bin/scrapy-admin.py to somewhere on your system path, for example C:\Python25\Scripts, which is customary for Python scripts.