diff --git a/scrapy/trunk/docs/intro/tutorial/tutorial1.rst b/scrapy/trunk/docs/intro/tutorial/tutorial1.rst index 0e4498caf..92753e90f 100644 --- a/scrapy/trunk/docs/intro/tutorial/tutorial1.rst +++ b/scrapy/trunk/docs/intro/tutorial/tutorial1.rst @@ -17,16 +17,11 @@ For starting a new project, enter the directory where you'd like your project to As long as Scrapy is well installed and the path is set, this should create a directory called "google" containing the following files: * *scrapy-ctl.py* - the project's control script. It's used for running the different tasks (like "genspider", "crawl" and "parse"). We'll talk more about this later. -* *scrapy_settings.py* - the project's settings file. -* *items.py* - were you define the different kinds of items you're going to scrape. -* *spiders* - directory where you'll later place your spiders. -* *templates* - directory containing some templates for newly created spiders, and where you can put your own. - -| Ok, now that you have your project's structure defined, the last thing to do is to set your PYTHONPATH to your project's directory. -| You can do this by adding this to your .bashrc file: - -:: - - $ export PYTHONPATH=/path/to/your/project +* *google/* - the projects' actual python module, you'll import your code from here. +* *google/items.py* - were you define the different kinds of items you're going to scrape. +* *google/pipelines.py* - were you define your item pipelines. +* *google/settings.py* - the project's settings file. +* *google/spiders/* - directory where you'll later place your spiders. +* *google/templates/* - directory containing some templates for newly created spiders, and where you can put your own. Now you can continue with the next part of the tutorial: :ref:`intro-tutorial2`. diff --git a/scrapy/trunk/docs/intro/tutorial/tutorial2.rst b/scrapy/trunk/docs/intro/tutorial/tutorial2.rst index 4465daa09..dcea72641 100644 --- a/scrapy/trunk/docs/intro/tutorial/tutorial2.rst +++ b/scrapy/trunk/docs/intro/tutorial/tutorial2.rst @@ -8,13 +8,13 @@ Ok, the time to write our first spider has come. Make sure that you're standing ./scrapy-ctl.py genspider google_directory google.com -This should create a file called google_directory.py under the *spiders* directory looking like this:: +This should create a file called google_directory.py under the *google/spiders* directory looking like this:: # -*- coding: utf8 -*- import re from scrapy.xpath import HtmlXPathSelector - from scrapy.item import ScrapedItem + from google.items import GoogleScrapedItem from scrapy.link.extractors import RegexLinkExtractor from scrapy.contrib.spiders import CrawlSpider, Rule @@ -28,7 +28,7 @@ This should create a file called google_directory.py under the *spiders* directo def parse_item(self, response): xs = HtmlXPathSelector(response) - i = ScrapedItem() + i = GoogleScrapedItem() #i.attribute('site_id', xs.x('//input[@id="sid"]/@value')) #i.attribute('name', xs.x('//div[@id="name"]')) #i.attribute('description', xs.x('//div[@id="description"]')) @@ -95,7 +95,7 @@ Until now, our spider would look something like:: # -*- coding: utf8 -*- from scrapy.xpath import HtmlXPathSelector - from scrapy.item import ScrapedItem + from google.items import GoogleScrapedItem from scrapy.link.extractors import RegexLinkExtractor from scrapy.contrib.spiders import CrawlSpider, Rule diff --git a/scrapy/trunk/docs/intro/tutorial/tutorial3.rst b/scrapy/trunk/docs/intro/tutorial/tutorial3.rst index f7d80f2df..6729d3136 100644 --- a/scrapy/trunk/docs/intro/tutorial/tutorial3.rst +++ b/scrapy/trunk/docs/intro/tutorial/tutorial3.rst @@ -35,7 +35,7 @@ Anyway, having said that, a possible *parse_category* could be:: links = hxs.x('//td[descendant::a[contains(@href, "#pagerank")]]/following-sibling::td/font') for link in links: - item = ScrapedItem() + item = GoogleScrapedItem() adaptor_pipe = [adaptors.extract, adaptors.Delist(''), adaptors.strip] item.set_adaptors({ 'name': adaptor_pipe, diff --git a/scrapy/trunk/docs/intro/tutorial/tutorial4.rst b/scrapy/trunk/docs/intro/tutorial/tutorial4.rst index 73753fdf8..a10308e0d 100644 --- a/scrapy/trunk/docs/intro/tutorial/tutorial4.rst +++ b/scrapy/trunk/docs/intro/tutorial/tutorial4.rst @@ -16,7 +16,7 @@ Let's see how would our spider end up looking like after applying this change:: # -*- coding: utf8 -*- from scrapy.xpath import HtmlXPathSelector - from scrapy.item import ScrapedItem + from google.items import GoogleScrapedItem from scrapy.contrib import adaptors from scrapy.contrib.spiders import CrawlSpider, Rule from scrapy.link.extractors import RegexLinkExtractor @@ -40,7 +40,7 @@ Let's see how would our spider end up looking like after applying this change:: links = hxs.x('//td[descendant::a[contains(@href, "#pagerank")]]/following-sibling::td/font') for link in links: - item = ScrapedItem() + item = GoogleScrapedItem() adaptor_pipe = [adaptors.extract, adaptors.Delist(''), adaptors.strip] item.set_adaptors({ 'name': adaptor_pipe, @@ -60,11 +60,8 @@ Let's see how would our spider end up looking like after applying this change:: With this code, our spider will crawl over Google's directory, and save each -link's name, description, and url to a file called 'scraped_items.csv'. +link's name, description, and url to a file called 'scraped_items.csv':: - Cool, huh? - -:: ./scrapy-ctl.py crawl google.com This is the end of the tutorial. If you'd like to know more about Scrapy and its use, please read the rest of the documentation.