diff --git a/.github/ISSUE_TEMPLATE/bug_report.md b/.github/ISSUE_TEMPLATE/bug_report.md index 8ca10109b..49089191f 100644 --- a/.github/ISSUE_TEMPLATE/bug_report.md +++ b/.github/ISSUE_TEMPLATE/bug_report.md @@ -12,7 +12,7 @@ The GitHub issue tracker's purpose is to deal with bug reports and feature reque Keep in mind that by filing an issue, you are expected to comply with Scrapy's Code of Conduct, including treating everyone with respect: https://github.com/scrapy/scrapy/blob/master/CODE_OF_CONDUCT.md -The following is a suggested template to structure your issue, you can find more guidelines at https://doc.scrapy.org/en/latest/contributing.html#reporting-bugs +The following is a suggested template to structure your issue, you can find more guidelines at https://docs.scrapy.org/en/latest/contributing.html#reporting-bugs --> diff --git a/.github/ISSUE_TEMPLATE/feature_request.md b/.github/ISSUE_TEMPLATE/feature_request.md index e05273fe2..7d9995de6 100644 --- a/.github/ISSUE_TEMPLATE/feature_request.md +++ b/.github/ISSUE_TEMPLATE/feature_request.md @@ -12,7 +12,7 @@ The GitHub issue tracker's purpose is to deal with bug reports and feature reque Keep in mind that by filing an issue, you are expected to comply with Scrapy's Code of Conduct, including treating everyone with respect: https://github.com/scrapy/scrapy/blob/master/CODE_OF_CONDUCT.md -The following is a suggested template to structure your pull request, you can find more guidelines at https://doc.scrapy.org/en/latest/contributing.html#writing-patches and https://doc.scrapy.org/en/latest/contributing.html#submitting-patches +The following is a suggested template to structure your pull request, you can find more guidelines at https://docs.scrapy.org/en/latest/contributing.html#writing-patches and https://docs.scrapy.org/en/latest/contributing.html#submitting-patches --> diff --git a/docs/README.rst b/docs/README.rst index 36dd5aea4..a900e3e72 100644 --- a/docs/README.rst +++ b/docs/README.rst @@ -7,13 +7,11 @@ Scrapy documentation quick start guide This file provides a quick guide on how to compile the Scrapy documentation. -Setup the environment ---------------------- +Set up the environment +---------------------- -To compile the documentation you need Sphinx Python library. To install it -and all its dependencies run the following command from this dir - -:: +To compile the documentation you need the Sphinx Python library. To install it +and all its dependencies, run the following command from this directory:: pip install -r requirements.txt @@ -21,8 +19,8 @@ and all its dependencies run the following command from this dir Compile the documentation ------------------------- -To compile the documentation (to classic HTML output) run the following command -from this dir:: +To compile the documentation (to classic HTML output), run the following +command from this directory:: make html @@ -32,7 +30,7 @@ Documentation will be generated (in HTML format) inside the ``build/html`` dir. View the documentation ---------------------- -To view the documentation run the following command:: +To view the documentation, run the following command:: make htmlview @@ -43,7 +41,7 @@ This command will fire up your default browser and open the main page of your Start over ---------- -To clean up all generated documentation files and start from scratch run:: +To clean up all generated documentation files and start from scratch, run:: make clean @@ -53,15 +51,15 @@ Keep in mind that this command won't touch any documentation source files. Recreating documentation on the fly ----------------------------------- -There is a way to recreate the doc automatically when you make changes, you -need to install watchdog (``pip install watchdog``) and then use:: +You can recreate the documentation automatically when you make changes. Install +watchdog (``pip install watchdog``) and then run:: make watch Alternative method using tox ---------------------------- -To compile the documentation to HTML run the following command:: +To compile the documentation to HTML, run the following command:: tox -e docs diff --git a/docs/contributing.rst b/docs/contributing.rst index 3976d34c2..e56e121ef 100644 --- a/docs/contributing.rst +++ b/docs/contributing.rst @@ -6,7 +6,7 @@ Contributing to Scrapy .. important:: - Double check that you are reading the most recent version of this document + Double-check that you are reading the most recent version of this document at https://docs.scrapy.org/en/master/contributing.html By participating in this project you agree to abide by the terms of our @@ -216,14 +216,13 @@ has been validated and proven useful. Alternatively, you can start a conversation in the `Scrapy subreddit`_ to discuss your idea first. Sometimes there is an existing pull request for the problem you'd like to -solve, which is stalled for some reason. Often the pull request is in a -right direction, but changes are requested by Scrapy maintainers, and the -original pull request author hasn't had time to address them. -In this case consider picking up this pull request: open -a new pull request with all commits from the original pull request, as well as -additional changes to address the raised issues. Doing so helps a lot; it is -not considered rude as long as the original author is acknowledged by keeping -his/her commits. +solve, which is stalled for some reason. Often the pull request is in the right +direction, but changes are requested by Scrapy maintainers, and the original +pull request author hasn't had time to address them. In this case consider +picking up this pull request: open a new pull request with all commits from the +original pull request, as well as additional changes to address the raised +issues. Doing so helps a lot; it is not considered rude as long as the original +author is acknowledged by keeping their commits. You can pull an existing pull request to a local branch by running ``git fetch upstream pull/$PR_NUMBER/head:$BRANCH_NAME_TO_CREATE`` @@ -233,9 +232,9 @@ with a name of the branch you want to create locally). See also: https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/reviewing-changes-in-pull-requests/checking-out-pull-requests-locally#modifying-an-inactive-pull-request-locally. When writing GitHub pull requests, try to keep titles short but descriptive. -E.g. For bug #411: "Scrapy hangs if an exception raises in start_requests" -prefer "Fix hanging when exception occurs in start_requests (#411)" -instead of "Fix for #411". Complete titles make it easy to skim through +For example, for bug #411: "Scrapy hangs if an exception raises in +start_requests" prefer "Fix hanging when exception occurs in start_requests +(#411)" instead of "Fix for #411". Complete titles make it easy to skim through the issue tracker. Finally, try to keep aesthetic changes (:pep:`8` compliance, unused imports @@ -271,7 +270,7 @@ commit. .. _pre-commit: https://pre-commit.com/ -After your create a local clone of your fork of the Scrapy repository: +After you create a local clone of your fork of the Scrapy repository: #. `Install pre-commit `_. diff --git a/docs/faq.rst b/docs/faq.rst index fdf0582de..4fab00152 100644 --- a/docs/faq.rst +++ b/docs/faq.rst @@ -110,7 +110,7 @@ My Scrapy crawler has memory leaks. What can I do? See :ref:`topics-leaks`. -Also, Python has a builtin memory leak issue which is described in +Also, Python has a built-in memory leak issue which is described in :ref:`topics-leaks-without-leaks`. How can I make Scrapy consume less memory? @@ -128,8 +128,8 @@ middleware with a :ref:`custom downloader middleware ` that requires less memory. For example: - If your domain names are similar enough, use your own regular expression - instead joining the strings in :attr:`~scrapy.Spider.allowed_domains` into - a complex regular expression. + instead of joining the strings in :attr:`~scrapy.Spider.allowed_domains` + into a complex regular expression. - If you can meet the installation requirements, use pyre2_ instead of Python’s re_ to compile your URL-filtering regular expression. See @@ -274,8 +274,8 @@ Parsing big feeds with XPath selectors can be problematic since they need to build the DOM of the entire feed in memory, and this can be quite slow and consume a lot of memory. -In order to avoid parsing all the entire feed at once in memory, you can use -the :func:`~scrapy.utils.iterators.xmliter_lxml` and +In order to avoid parsing the entire feed at once in memory, you can use the +:func:`~scrapy.utils.iterators.xmliter_lxml` and :func:`~scrapy.utils.iterators.csviter` functions. In fact, this is what :class:`~scrapy.spiders.XMLFeedSpider` uses. @@ -311,8 +311,8 @@ Should I use spider arguments or settings to configure my spider? ----------------------------------------------------------------- Both :ref:`spider arguments ` and :ref:`settings ` -can be used to configure your spider. There is no strict rule that mandates to -use one or the other, but settings are more suited for parameters that, once +can be used to configure your spider. There is no strict rule that mandates +using one or the other, but settings are more suited for parameters that, once set, don't change much, while spider arguments are meant to change more often, even on each spider run and sometimes are required for the spider to run at all (for example, to set the start url of a spider). diff --git a/docs/index.rst b/docs/index.rst index 71047f9ef..480068139 100644 --- a/docs/index.rst +++ b/docs/index.rst @@ -21,16 +21,16 @@ Having trouble? We'd like to help! * Try the :doc:`FAQ ` -- it's got answers to some common questions. * Looking for specific information? Try the :ref:`genindex` or :ref:`modindex`. -* Ask or search questions in `StackOverflow using the scrapy tag`_. +* Ask or search questions on `Stack Overflow using the scrapy tag`_. * Ask or search questions in the `Scrapy subreddit`_. * Search for questions on the archives of the `scrapy-users mailing list`_. -* Ask a question in the `#scrapy IRC channel`_, +* Ask a question in the `#scrapy IRC channel`_. * Report bugs with Scrapy in our `issue tracker`_. * Join the Discord community `Scrapy Discord`_. .. _scrapy-users mailing list: https://groups.google.com/forum/#!forum/scrapy-users .. _Scrapy subreddit: https://www.reddit.com/r/scrapy/ -.. _StackOverflow using the scrapy tag: https://stackoverflow.com/tags/scrapy +.. _Stack Overflow using the scrapy tag: https://stackoverflow.com/tags/scrapy .. _#scrapy IRC channel: irc://irc.freenode.net/scrapy .. _issue tracker: https://github.com/scrapy/scrapy/issues .. _Scrapy Discord: https://discord.com/invite/mv3yErfpvq @@ -132,7 +132,7 @@ Built-in services topics/telnetconsole :doc:`topics/logging` - Learn how to use Python's built-in logging on Scrapy. + Learn how to use Python's built-in logging in Scrapy. :doc:`topics/stats` Collect statistics about your scraping crawler. @@ -180,7 +180,7 @@ Solving specific problems Get familiar with some Scrapy common practices. :doc:`topics/broad-crawls` - Tune Scrapy for crawling a lot domains in parallel. + Tune Scrapy for crawling a lot of domains in parallel. :doc:`topics/developer-tools` Learn how to scrape with your browser's developer tools. @@ -195,7 +195,7 @@ Solving specific problems Download files and/or images associated with your scraped items. :doc:`topics/deploy` - Deploying your Scrapy spiders and run them in a remote server. + Deploy your Scrapy spiders and run them on a remote server. :doc:`topics/autothrottle` Adjust crawl rate dynamically based on load. @@ -246,7 +246,7 @@ Extending Scrapy Customize the input and output of your spiders. :doc:`topics/extensions` - Extend Scrapy with your custom functionality + Extend Scrapy with your custom functionality. :doc:`topics/signals` See all available signals and how to work with them. @@ -255,14 +255,14 @@ Extending Scrapy Understand the scheduler component. :doc:`topics/exporters` - Quickly export your scraped items to a file (XML, CSV, etc). + Quickly export your scraped items to a file (XML, CSV, etc.). :doc:`topics/components` Learn the common API and some good practices when building custom Scrapy components. :doc:`topics/api` - Use it on extensions and middlewares to extend Scrapy functionality. + Use it in extensions and middlewares to extend Scrapy functionality. All the rest diff --git a/docs/intro/examples.rst b/docs/intro/examples.rst index edff894c6..f72373577 100644 --- a/docs/intro/examples.rst +++ b/docs/intro/examples.rst @@ -5,16 +5,16 @@ Examples ======== The best way to learn is with examples, and Scrapy is no exception. For this -reason, there is an example Scrapy project named quotesbot_, that you can use to +reason, there is an example Scrapy project named quotesbot_ that you can use to play and learn more about Scrapy. It contains two spiders for -https://quotes.toscrape.com, one using CSS selectors and another one using XPath -expressions. +https://quotes.toscrape.com, one using CSS selectors and another one using +XPath expressions. The quotesbot_ project is available at: https://github.com/scrapy/quotesbot. You can find more information about it in the project's README. -If you're familiar with git, you can checkout the code. Otherwise you can -download the project as a zip file by clicking -`here `_. +If you're familiar with git, you can check out the code. Otherwise, you can +download the project as a zip file by clicking `here +`_. .. _quotesbot: https://github.com/scrapy/quotesbot diff --git a/docs/intro/install.rst b/docs/intro/install.rst index 488a66f36..265fb2753 100644 --- a/docs/intro/install.rst +++ b/docs/intro/install.rst @@ -25,15 +25,15 @@ To install Scrapy using ``conda``, run:: conda install -c conda-forge scrapy -Alternatively, if you’re already familiar with installation of Python packages, -you can install Scrapy and its dependencies from PyPI with:: +Alternatively, if you’re already familiar with installing Python packages, you +can install Scrapy and its dependencies from PyPI with:: pip install Scrapy We strongly recommend that you install Scrapy in :ref:`a dedicated virtualenv `, to avoid conflicting with your system packages. -Note that sometimes this may require solving compilation issues for some Scrapy +Note that this may sometimes require solving compilation issues for some Scrapy dependencies depending on your operating system, so be sure to check the :ref:`intro-install-platform-notes`. @@ -47,7 +47,7 @@ Things that are good to know Scrapy is written in pure Python and depends on a few key Python packages (among others): * `lxml`_, an efficient XML and HTML parser -* `parsel`_, an HTML/XML data extraction library written on top of lxml, +* `parsel`_, an HTML/XML data extraction library written on top of lxml * `w3lib`_, a multi-purpose helper for dealing with URLs and web page encodings * `twisted`_, an asynchronous networking framework * `cryptography`_ and `pyOpenSSL`_, to deal with various network-level security needs @@ -73,14 +73,14 @@ Using a virtual environment (recommended) TL;DR: We recommend installing Scrapy inside a virtual environment on all platforms. -Python packages can be installed either globally (a.k.a system wide), -or in user-space. We do not recommend installing Scrapy system wide. +Python packages can be installed either globally (a.k.a. system-wide), or in +user-space. We do not recommend installing Scrapy system-wide. -Instead, we recommend that you install Scrapy within a so-called -"virtual environment" (:mod:`venv`). -Virtual environments allow you to not conflict with already-installed Python -system packages (which could break some of your system tools and scripts), -and still install packages normally with ``pip`` (without ``sudo`` and the likes). +Instead, we recommend that you install Scrapy within a so-called "virtual +environment" (:mod:`venv`). Virtual environments allow you to avoid conflicts +with already-installed Python system packages (which could break some of your +system tools and scripts), and still install packages normally with ``pip`` +(without ``sudo`` or similar tools). See :ref:`tut-venv` on how to create your virtual environment. @@ -92,7 +92,7 @@ below for non-Python dependencies that you may need to install beforehand). .. _intro-install-platform-notes: -Platform specific installation notes +Platform-specific installation notes ==================================== .. _intro-install-windows: @@ -120,11 +120,12 @@ To install Scrapy on Windows using ``pip``: #. Under the Workloads section, select **C++ build tools**. -#. Check the installation details and make sure following packages are selected as optional components: +#. Check the installation details and make sure the following packages are + selected as optional components: - * **MSVC** (e.g MSVC v142 - VS 2019 C++ x64/x86 build tools (v14.23) ) + * **MSVC** (e.g. MSVC v142 - VS 2019 C++ x64/x86 build tools (v14.23) ) - * **Windows SDK** (e.g Windows 10 SDK (10.0.18362.0)) + * **Windows SDK** (e.g. Windows 10 SDK (10.0.18362.0)) #. Install the Visual Studio Build Tools. @@ -140,8 +141,8 @@ twisted and pyOpenSSL, and is compatible with recent Ubuntu distributions. But it should support older versions of Ubuntu too, like Ubuntu 14.04, albeit with potential issues with TLS connections. -**Don't** use the ``python-scrapy`` package provided by Ubuntu, they are -typically too old and slow to catch up with the latest Scrapy release. +**Don't** use the ``python-scrapy`` package provided by Ubuntu; it is typically +too old and slow to catch up with the latest Scrapy release. To install Scrapy on Ubuntu (or Ubuntu-based) systems, you need to install @@ -187,8 +188,8 @@ solutions: * Install `homebrew`_ following the instructions in https://brew.sh/ * Update your ``PATH`` variable to state that homebrew packages should be - used before system packages (Change ``.bashrc`` to ``.zshrc`` accordingly - if you're using `zsh`_ as default shell):: + used before system packages (change ``.bashrc`` to ``.zshrc`` accordingly + if you're using `zsh`_ as the default shell):: echo "export PATH=/usr/local/bin:/usr/local/sbin:$PATH" >> ~/.bashrc @@ -196,7 +197,7 @@ solutions: source ~/.bashrc - * Install python:: + * Install Python:: brew install python @@ -206,7 +207,7 @@ solutions: This method is a workaround for the above macOS issue, but it's an overall good practice for managing dependencies and can complement the first method. -After any of these workarounds you should be able to install Scrapy:: +After any of these workarounds, you should be able to install Scrapy:: pip install Scrapy @@ -218,14 +219,14 @@ We recommend using the latest PyPy version. For PyPy3, only Linux installation was tested. Most Scrapy dependencies now have binary wheels for CPython, but not for PyPy. -This means that these dependencies will be built during installation. -On macOS, you are likely to face an issue with building the Cryptography -dependency. The solution to this problem is described -`here `_, +This means that these dependencies will be built during installation. On macOS, +you are likely to face an issue with building the Cryptography dependency. The +solution to this problem is described `here +`_, that is to ``brew install openssl`` and then export the flags that this command -recommends (only needed when installing Scrapy). Installing on Linux has no special -issues besides installing build dependencies. -Installing Scrapy with PyPy on Windows is not tested. +recommends (only needed when installing Scrapy). Installing Scrapy on Linux has +no special issues beyond installing the build dependencies. Installing Scrapy +with PyPy on Windows has not been tested. You can check that Scrapy is installed correctly by running ``scrapy bench``. If this command gives errors such as diff --git a/docs/intro/overview.rst b/docs/intro/overview.rst index d05e46551..94e274448 100644 --- a/docs/intro/overview.rst +++ b/docs/intro/overview.rst @@ -4,13 +4,13 @@ Scrapy at a glance ================== -Scrapy (/ˈskreɪpaɪ/) is an application framework for crawling web sites and extracting -structured data which can be used for a wide range of useful applications, like -data mining, information processing or historical archival. +Scrapy (/ˈskreɪpaɪ/) is an application framework for crawling websites and +extracting structured data which can be used for a wide range of useful +applications, like data mining, information processing or historical archival. Even though Scrapy was originally designed for `web scraping`_, it can also be used to extract data using APIs (such as `Amazon Associates Web Services`_) or -as a general purpose web crawler. +as a general-purpose web crawler. Walk-through of an example spider @@ -19,7 +19,7 @@ Walk-through of an example spider In order to show you what Scrapy brings to the table, we'll walk you through an example of a Scrapy Spider using the simplest way to run a spider. -Here's the code for a spider that scrapes famous quotes from website +Here's the code for a spider that scrapes famous quotes from the website https://quotes.toscrape.com, following the pagination: .. code-block:: python @@ -49,8 +49,9 @@ and run the spider using the :command:`runspider` command:: scrapy runspider quotes_spider.py -o quotes.jsonl -When this finishes you will have in the ``quotes.jsonl`` file a list of the -quotes in JSON Lines format, containing the text and author, which will look like this:: +When this finishes, you will have a list of the quotes in JSON Lines format in +the ``quotes.jsonl`` file, containing the text and author, which will look like +this:: {"author": "Jane Austen", "text": "\u201cThe person, be it gentleman or lady, who has not pleasure in a good novel, must be intolerably stupid.\u201d"} {"author": "Steve Martin", "text": "\u201cA day without sunshine is like, you know, night.\u201d"} @@ -73,19 +74,19 @@ look for a link to the next page and schedule another request using the same ``parse`` method as callback. Here you will notice one of the main advantages of Scrapy: requests are -:ref:`scheduled and processed asynchronously `. This -means that Scrapy doesn't need to wait for a request to be finished and -processed, it can send another request or do other things in the meantime. This -also means that other requests can keep going even if a request fails or an -error happens while handling it. +:ref:`scheduled and processed asynchronously `. This means +that Scrapy doesn't need to wait for a request to be finished and processed; it +can send another request or do other things in the meantime. This also means +that other requests can keep going even if a request fails or an error happens +while handling it. While this enables you to do very fast crawls (sending multiple concurrent requests at the same time, in a fault-tolerant way) Scrapy also gives you control over the politeness of the crawl through :ref:`a few settings `. You can do things like setting a download delay between -each request, limiting the amount of concurrent requests per domain or per IP, and -even :ref:`using an auto-throttling extension ` that tries -to figure these settings out automatically. +each request, limiting the number of concurrent requests per domain or per IP, +and even :ref:`using an auto-throttling extension ` that +tries to figure these settings out automatically. .. note:: @@ -116,7 +117,7 @@ scraping easy and efficient, such as: multiple formats (JSON, CSV, XML) and storing them in multiple backends (FTP, S3, local filesystem) -* Robust encoding support and auto-detection, for dealing with foreign, +* Robust encoding support and auto-detection for dealing with foreign, non-standard and broken encoding declarations. * :ref:`Strong extensibility support `, allowing you to plug @@ -137,10 +138,10 @@ scraping easy and efficient, such as: console running inside your Scrapy process, to introspect and debug your crawler -* Plus other goodies like reusable spiders to crawl sites from `Sitemaps`_ and - XML/CSV feeds, a media pipeline for :ref:`automatically downloading images - ` (or any other media) associated with the scraped - items, a caching DNS resolver, and much more! +* Plus other goodies such as reusable spiders to crawl sites from `Sitemaps`_ + and XML/CSV feeds, a media pipeline for :ref:`automatically downloading + images ` (or any other media) associated with the + scraped items, a caching DNS resolver, and much more! What's next? ============ diff --git a/docs/intro/tutorial.rst b/docs/intro/tutorial.rst index c4e04364b..406f848ae 100644 --- a/docs/intro/tutorial.rst +++ b/docs/intro/tutorial.rst @@ -15,7 +15,7 @@ This tutorial will walk you through these tasks: 1. Creating a new Scrapy project 2. Writing a :ref:`spider ` to crawl a site and extract data 3. Exporting the scraped data using the command line -4. Changing spider to recursively follow links +4. Changing the spider to recursively follow links 5. Using spider arguments Scrapy is written in Python_. The more you learn about Python, the more you @@ -76,10 +76,11 @@ This will create a ``tutorial`` directory with the following contents:: Our first Spider ================ -Spiders are classes that you define and that Scrapy uses to scrape information from a website -(or a group of websites). They must subclass :class:`~scrapy.Spider` and define the initial -requests to be made, and optionally, how to follow links in pages and parse the downloaded -page content to extract data. +Spiders are classes that you define and that Scrapy uses to scrape information +from a website (or a group of websites). They must subclass +:class:`~scrapy.Spider` and define the initial requests to be made and, +optionally, how to follow links in pages and parse the downloaded page content +to extract data. This is the code for our first Spider. Save it in a file named ``quotes_spider.py`` under the ``tutorial/spiders`` directory in your project: @@ -137,9 +138,9 @@ To put our spider to work, go to the project's top level directory and run:: scrapy crawl quotes -This command runs the spider named ``quotes`` that we've just added, that -will send some requests for the ``quotes.toscrape.com`` domain. You will get an output -similar to this:: +This command runs the spider named ``quotes`` that we've just added, which will +send some requests for the ``quotes.toscrape.com`` domain. You will get an +output similar to this:: ... (omitted for brevity) 2016-12-16 21:24:05 [scrapy.core.engine] INFO: Spider opened @@ -400,8 +401,8 @@ like this: -Let's open up scrapy shell and play a bit to find out how to extract the data -we want:: +Let's open up the Scrapy shell and play a bit to find out how to extract the +data we want:: scrapy shell 'https://quotes.toscrape.com' @@ -668,14 +669,14 @@ As a shortcut for creating Request objects you can use if next_page is not None: yield response.follow(next_page, callback=self.parse) -Unlike scrapy.Request, ``response.follow`` supports relative URLs directly - no -need to call urljoin. Note that ``response.follow`` just returns a Request -instance; you still have to yield this Request. +Unlike :class:`scrapy.Request`, ``response.follow`` supports relative URLs +directly - no need to call urljoin. Note that ``response.follow`` just returns +a Request instance; you still have to yield this Request. .. skip: start -You can also pass a selector to ``response.follow`` instead of a string; -this selector should extract necessary attributes: +You can also pass a selector to ``response.follow`` instead of a string; this +selector should extract the necessary attributes: .. code-block:: python @@ -741,7 +742,7 @@ this time for scraping author information: } This spider will start from the main page, it will follow all the links to the -authors pages calling the ``parse_author`` callback for each of them, and also +author pages calling the ``parse_author`` callback for each of them, and also the pagination links with the ``parse`` callback as we saw before. Here we're passing callbacks to @@ -749,8 +750,8 @@ Here we're passing callbacks to arguments to make the code shorter; it also works for :class:`~scrapy.Request`. -The ``parse_author`` callback defines a helper function to extract and cleanup the -data from a CSS query and yields the Python dict with the author data. +The ``parse_author`` callback defines a helper function to extract and clean up +the data from a CSS query and yields the Python dict with the author data. Another interesting thing this spider demonstrates is that, even if there are many quotes from the same author, we don't need to worry about visiting the diff --git a/docs/news.rst b/docs/news.rst index 2a769bc1c..cb4f05903 100644 --- a/docs/news.rst +++ b/docs/news.rst @@ -150,7 +150,7 @@ Scrapy 2.13.2 (2025-06-09) when it's closed before being fully initialized. (:issue:`6857`, :issue:`6867`) -- Improved the README, updated the Scrapy logo in it. +- Improved the README and updated the Scrapy logo in it. (:issue:`6831`, :issue:`6833`, :issue:`6839`) - Restricted the Twisted version used in tests to below 25.5.0, as some tests @@ -169,14 +169,14 @@ Scrapy 2.13.2 (2025-06-09) Scrapy 2.13.1 (2025-05-28) -------------------------- -- Give callback requests precedence over start requests when priority values +- Gave callback requests precedence over start requests when priority values are the same. - This makes changes from 2.13.0 to start request handling more intuitive and - backward compatible. For scenarios where all requests have the same + This makes the changes from 2.13.0 to start request handling more intuitive + and backward compatible. For scenarios where all requests have the same priorities, in 2.13.0 all start requests were sent before the first - callback request. In 2.13.1, same as in 2.12 and lower, start requests are - only sent when there are not enough pending callback requests to reach + callback request. In 2.13.1, the same as in 2.12 and lower, start requests + are only sent when there are not enough pending callback requests to reach concurrency limits. (:issue:`6828`) diff --git a/docs/topics/addons.rst b/docs/topics/addons.rst index dba14e749..0b07b2485 100644 --- a/docs/topics/addons.rst +++ b/docs/topics/addons.rst @@ -4,7 +4,7 @@ Add-ons ======= -Scrapy's add-on system is a framework which unifies managing and configuring +Scrapy's add-on system is a framework that unifies managing and configuring components that extend Scrapy's core functionality, such as middlewares, extensions, or pipelines. It provides users with a plug-and-play experience in Scrapy extension management, and grants extensive configuration control to @@ -71,7 +71,7 @@ usually best to leave its priority unchanged. For example, when editing a If the ``update_settings`` method raises :exc:`scrapy.exceptions.NotConfigured`, the add-on will be skipped. This makes -it easy to enable an add-on only when some conditions are met. +it easy to enable an add-on only when certain conditions are met. Fallbacks --------- diff --git a/docs/topics/api.rst b/docs/topics/api.rst index 19082d9d7..f49531f42 100644 --- a/docs/topics/api.rst +++ b/docs/topics/api.rst @@ -22,9 +22,9 @@ functionality into Scrapy. :synopsis: The Scrapy crawler The Extension Manager is responsible for loading and keeping track of installed -extensions and it's configured through the :setting:`EXTENSIONS` setting which -contains a dictionary of all available extensions and their order similar to -how you :ref:`configure the downloader middlewares +extensions, and it's configured through the :setting:`EXTENSIONS` setting, +which contains a dictionary of all available extensions and their order similar +to how you :ref:`configure the downloader middlewares `. .. autoclass:: Crawler @@ -89,7 +89,7 @@ how you :ref:`configure the downloader middlewares The execution engine, which coordinates the core crawling logic between the scheduler, downloader and spiders. - Some extension may want to access the Scrapy engine, to inspect or + Some extensions may want to access the Scrapy engine to inspect or modify the downloader and scheduler behaviour, although this is an advanced use and this API is not yet stable. diff --git a/docs/topics/architecture.rst b/docs/topics/architecture.rst index e8c510ea5..3f0adedfd 100644 --- a/docs/topics/architecture.rst +++ b/docs/topics/architecture.rst @@ -62,9 +62,9 @@ this: :ref:`Spider Middleware ` (see :meth:`~scrapy.spidermiddlewares.SpiderMiddleware.process_spider_output`). -8. The :ref:`Engine ` sends processed items to - :ref:`Item Pipelines `, then send processed Requests to - the :ref:`Scheduler ` and asks for possible next Requests +8. The :ref:`Engine ` sends processed items to :ref:`Item + Pipelines `, then sends processed Requests to the + :ref:`Scheduler ` and asks for possible next Requests to crawl. 9. The process repeats (from step 3) until there are no more requests from the diff --git a/docs/topics/benchmarking.rst b/docs/topics/benchmarking.rst index e8ddec00c..9dcb45e89 100644 --- a/docs/topics/benchmarking.rst +++ b/docs/topics/benchmarking.rst @@ -78,8 +78,8 @@ You should see an output like this:: That tells you that Scrapy is able to crawl about 3000 pages per minute in the hardware where you run it. Note that this is a very simple spider intended to follow links, any custom spider you write will probably do more stuff which -results in slower crawl rates. How slower depends on how much your spider does -and how well it's written. +results in slower crawl rates. How much slower depends on how much your spider +does and how well it's written. Use scrapy-bench_ for more complex benchmarking. diff --git a/docs/topics/broad-crawls.rst b/docs/topics/broad-crawls.rst index b4d058754..d23c528ab 100644 --- a/docs/topics/broad-crawls.rst +++ b/docs/topics/broad-crawls.rst @@ -13,7 +13,7 @@ In addition to this "focused crawl", there is another common type of crawling which covers a large (potentially unlimited) number of domains, and is only limited by time or other arbitrary constraint, rather than stopping when the domain was crawled to completion or when there are no more requests to perform. -These are called "broad crawls" and is the typical crawlers employed by search +These are called "broad crawls" and are the typical crawlers employed by search engines. These are some common properties often found in broad crawls: @@ -44,9 +44,9 @@ efficient broad crawl. Use the right :setting:`SCHEDULER_PRIORITY_QUEUE` ================================================= -Scrapy’s default scheduler priority queue is ``'scrapy.pqueues.ScrapyPriorityQueue'``. -It works best during single-domain crawl. It does not work well with crawling -many different domains in parallel +Scrapy’s default scheduler priority queue is +``'scrapy.pqueues.ScrapyPriorityQueue'``. It works best during a single-domain +crawl. It does not work well when crawling many different domains in parallel. To apply the recommended priority queue use: @@ -87,9 +87,9 @@ Increase Twisted IO thread pool maximum size Currently Scrapy does DNS resolution in a blocking way with usage of thread pool. With higher concurrency levels the crawling could be slow or even fail -hitting DNS resolver timeouts. Possible solution to increase the number of -threads handling DNS queries. The DNS queue will be processed faster speeding -up establishing of connection and crawling overall. +hitting DNS resolver timeouts. A possible solution is to increase the number of +threads handling DNS queries. The DNS queue will be processed faster, speeding +up connection establishment and the crawl overall. To increase maximum thread pool size use: @@ -100,10 +100,11 @@ To increase maximum thread pool size use: Setup your own DNS ================== -If you have multiple crawling processes and single central DNS, it can act -like DoS attack on the DNS server resulting to slow down of entire network or -even blocking your machines. To avoid this setup your own DNS server with -local cache and upstream to some large DNS like OpenDNS or Verizon. +If you have multiple crawling processes and a single central DNS, it can act +like a DoS attack on the DNS server, resulting in a slowdown of the entire +network or even blocking your machines. To avoid this, set up your own DNS +server with a local cache and upstream to some large DNS like OpenDNS or +Verizon. Reduce log level ================ @@ -124,10 +125,10 @@ To set the log level use: Disable cookies =============== -Disable cookies unless you *really* need. Cookies are often not needed when -doing broad crawls (search engine crawlers ignore them), and they improve -performance by saving some CPU cycles and reducing the memory footprint of your -Scrapy crawler. +Disable cookies unless you *really* need them. Cookies are often not needed +when doing broad crawls (search engine crawlers ignore them), and disabling +them improves performance by saving some CPU cycles and reducing the memory +footprint of your Scrapy crawler. To disable cookies use: @@ -138,10 +139,10 @@ To disable cookies use: Disable retries =============== -Retrying failed HTTP requests can slow down the crawls substantially, especially -when sites causes are very slow (or fail) to respond, thus causing a timeout -error which gets retried many times, unnecessarily, preventing crawler capacity -to be reused for other domains. +Retrying failed HTTP requests can slow down the crawls substantially, +especially when sites are very slow (or fail) to respond, thus causing a +timeout error that gets retried many times, unnecessarily preventing crawler +capacity from being reused for other domains. To disable retries use: @@ -167,8 +168,8 @@ Disable redirects Consider disabling redirects, unless you are interested in following them. When doing broad crawls it's common to save redirects and resolve them when -revisiting the site at a later crawl. This also help to keep the number of -request constant per crawl batch, otherwise redirect loops may cause the +revisiting the site at a later crawl. This also helps to keep the number of +requests constant per crawl batch, otherwise redirect loops may cause the crawler to dedicate too many resources on any specific domain. To disable redirects use: diff --git a/docs/topics/commands.rst b/docs/topics/commands.rst index 4994fe1d6..996fd3ede 100644 --- a/docs/topics/commands.rst +++ b/docs/topics/commands.rst @@ -10,7 +10,7 @@ Scrapy is controlled through the ``scrapy`` command-line tool, to be referred to here as the "Scrapy tool" to differentiate it from the sub-commands, which we just call "commands" or "Scrapy commands". -The Scrapy tool provides several commands, for multiple purposes, and each one +The Scrapy tool provides several commands for multiple purposes, and each one accepts a different set of arguments and options. (The ``scrapy deploy`` command has been removed in 1.0 in favor of the @@ -65,7 +65,7 @@ structure by default, similar to this:: ... The directory where the ``scrapy.cfg`` file resides is known as the *project -root directory*. That file contains the name of the python module that defines +root directory*. That file contains the name of the Python module that defines the project settings. Here is an example: .. code-block:: ini @@ -105,7 +105,7 @@ for ``scrapy`` to use:: Using the ``scrapy`` tool ========================= -You can start by running the Scrapy tool with no arguments and it will print +You can start by running the Scrapy tool with no arguments, and it will print some usage help and the available commands:: Scrapy X.Y - no active project @@ -119,8 +119,8 @@ some usage help and the available commands:: [...] The first line will print the currently active project if you're inside a -Scrapy project. In this example it was run from outside a project. If run from inside -a project it would have printed something like this:: +Scrapy project. In this example, it was run from outside a project. If run from +inside a project, it would have printed something like this:: Scrapy X.Y - project: myproject @@ -137,8 +137,8 @@ project:: scrapy startproject myproject [project_dir] -That will create a Scrapy project under the ``project_dir`` directory. -If ``project_dir`` wasn't specified, ``project_dir`` will be the same as ``myproject``. +That will create a Scrapy project under the ``project_dir`` directory. If +``project_dir`` isn't specified, it defaults to ``myproject``. Next, you go inside the new project directory:: @@ -159,7 +159,8 @@ For example, to create a new spider:: Some Scrapy commands (like :command:`crawl`) must be run from inside a Scrapy project. See the :ref:`commands reference ` below for more -information on which commands must be run from inside projects, and which not. +information on which commands must be run from inside projects and which ones +do not. Also keep in mind that some commands may have slightly different behaviours when running them from inside projects. For example, the fetch command will use @@ -186,7 +187,7 @@ And you can see all available commands with:: There are two kinds of commands, those that only work from inside a Scrapy project (Project-specific commands) and those that also work without an active Scrapy project (Global commands), though they may behave slightly differently -when run from inside a project (as they would use the project overridden +when run from inside a project (as they would use the project-overridden settings). Global commands: @@ -218,8 +219,7 @@ startproject * Requires project: *no* Creates a new Scrapy project named ``project_name``, under the ``project_dir`` -directory. -If ``project_dir`` wasn't specified, ``project_dir`` will be the same as ``project_name``. +directory. If ``project_dir`` isn't specified, it defaults to ``project_name``. Usage example:: @@ -236,7 +236,10 @@ genspider .. versionadded:: 2.6.0 The ability to pass a URL instead of a domain. -Creates a new spider in the current folder or in the current project's ``spiders`` folder, if called from inside a project. The ```` parameter is set as the spider's ``name``, while ```` is used to generate the ``allowed_domains`` and ``start_urls`` spider's attributes. +Creates a new spider in the current folder or in the current project's +``spiders`` folder, if called from inside a project. The ```` parameter +becomes the spider's ``name``, while ```` is used to generate +the spider's ``allowed_domains`` and ``start_urls`` attributes. Usage example:: @@ -253,9 +256,9 @@ Usage example:: $ scrapy genspider -t crawl scrapyorg scrapy.org Created spider 'scrapyorg' using template 'crawl' -This is just a convenient shortcut command for creating spiders based on -pre-defined templates, but certainly not the only way to create spiders. You -can just create the spider source code files yourself, instead of using this +This command is just a convenient shortcut for creating spiders based on +pre-defined templates, but it's certainly not the only way to create spiders; +you can create the spider source code files yourself instead of using this command. .. command:: crawl @@ -274,9 +277,13 @@ Supported options: * ``-a NAME=VALUE``: set a spider argument (may be repeated) -* ``--output FILE`` or ``-o FILE``: append scraped items to the end of FILE (use - for stdout). To define the output format, set a colon at the end of the output URI (i.e. ``-o FILE:FORMAT``) +* ``--output FILE`` or ``-o FILE``: append scraped items to the end of FILE + (use ``-`` for stdout). To define the output format, add a colon at the end + of the output URI (for example, ``-o FILE:FORMAT``) -* ``--overwrite-output FILE`` or ``-O FILE``: dump scraped items into FILE, overwriting any existing file. To define the output format, set a colon at the end of the output URI (i.e. ``-O FILE:FORMAT``) +* ``--overwrite-output FILE`` or ``-O FILE``: dump scraped items into FILE, + overwriting any existing file. To define the output format, add a colon at + the end of the output URI (for example, ``-O FILE:FORMAT``) Usage examples:: @@ -284,10 +291,10 @@ Usage examples:: [ ... myspider starts crawling ... ] $ scrapy crawl -o myfile:csv myspider - [ ... myspider starts crawling and appends the result to the file myfile in csv format ... ] + [ ... myspider starts crawling and appends the result to the file myfile in CSV format ... ] $ scrapy crawl -O myfile:json myspider - [ ... myspider starts crawling and saves the result in myfile in json format overwriting the original content... ] + [ ... myspider starts crawling and saves the result in myfile in JSON format, overwriting the original content ... ] .. command:: check @@ -349,7 +356,7 @@ Edit the given spider using the editor defined in the ``EDITOR`` environment variable or (if unset) the :setting:`EDITOR` setting. This command is provided only as a convenient shortcut for the most common -case, the developer is of course free to choose any tool or IDE to write and +case; the developer is of course free to choose any tool or IDE to write and debug spiders. Usage example:: @@ -373,7 +380,7 @@ attribute which overrides the User Agent, it will use that one. So this command can be used to "see" how your spider would fetch a certain page. -If used outside a project, no particular per-spider behaviour would be applied +If used outside a project, no particular per-spider behaviour will be applied, and it will just use the default Scrapy downloader settings. Supported options: @@ -431,10 +438,10 @@ shell * Syntax: ``scrapy shell [url]`` * Requires project: *no* -Starts the Scrapy shell for the given URL (if given) or empty if no URL is -given. Also supports UNIX-style local file paths, either relative with -``./`` or ``../`` prefixes or absolute file paths. -See :ref:`topics-shell` for more info. +Starts the Scrapy shell for the given URL (if provided) or leaves it empty if +no URL is given. It also supports UNIX-style local file paths, either relative +with ``./`` or ``../`` prefixes or absolute file paths. See :ref:`topics-shell` +for more info. Supported options: @@ -479,16 +486,18 @@ Supported options: * ``--spider=SPIDER``: bypass spider autodetection and force use of specific spider -* ``--a NAME=VALUE``: set spider argument (may be repeated) +* ``-a NAME=VALUE``: set spider argument (may be repeated) * ``--callback`` or ``-c``: spider method to use as callback for parsing the response -* ``--meta`` or ``-m``: additional request meta that will be passed to the callback - request. This must be a valid json string. Example: --meta='{"foo" : "bar"}' +* ``--meta`` or ``-m``: additional request meta that will be passed to the + callback request. This must be a valid JSON string. Example: + ``--meta='{"foo": "bar"}'`` -* ``--cbkwargs``: additional keyword arguments that will be passed to the callback. - This must be a valid json string. Example: --cbkwargs='{"foo" : "bar"}' +* ``--cbkwargs``: additional keyword arguments that will be passed to the + callback. This must be a valid JSON string. Example: ``--cbkwargs='{"foo": + "bar"}'`` * ``--pipelines``: process items through pipelines @@ -540,8 +549,8 @@ settings Get the value of a Scrapy setting. -If used inside a project it'll show the project setting value, otherwise it'll -show the default Scrapy value for that setting. +If used inside a project, it'll show the project setting value; otherwise, +it'll show the default Scrapy value for that setting. Example usage:: @@ -605,7 +614,7 @@ spider or a special internal one: * :command:`view` They use an internal instance of :class:`scrapy.crawler.AsyncCrawlerProcess` or -:class:`scrapy.crawler.CrawlerProcess` for this. In most cases this detail +:class:`scrapy.crawler.CrawlerProcess` for this. In most cases, this detail shouldn't matter to the user running the command, but when the user :ref:`needs a non-default Twisted reactor `, it may be important. @@ -621,9 +630,9 @@ project-level setting is set to :ref:`the asyncio reactor ` `) and :ref:`the setting of the spider being run ` is set to :ref:`a different one `, because :class:`~scrapy.crawler.AsyncCrawlerProcess` only supports the asyncio reactor. -In this case you should set the :setting:`FORCE_CRAWLER_PROCESS` setting to +In this case, you should set the :setting:`FORCE_CRAWLER_PROCESS` setting to ``True`` (at the project level or via the command line) so that Scrapy uses -:class:`~scrapy.crawler.CrawlerProcess` which supports all reactors. +:class:`~scrapy.crawler.CrawlerProcess`, which supports all reactors. Custom project commands ======================= @@ -658,7 +667,7 @@ You can also add Scrapy commands from an external library by adding a ``scrapy.commands`` section in the entry points of the library ``setup.py`` file. -The following example adds ``my_command`` command: +The following example adds the ``my_command`` command: .. skip: next diff --git a/docs/topics/components.rst b/docs/topics/components.rst index 56f8c6498..6585f679b 100644 --- a/docs/topics/components.rst +++ b/docs/topics/components.rst @@ -97,8 +97,8 @@ For example: print("log is enabled!") Components do not need to declare their custom settings programmatically. -However, they should document them, so that users know they exist and how to -use them. +However, they should document them so that users know they exist and how to use +them. It is a good practice to prefix custom settings with the name of the component, to avoid collisions with custom settings of other existing (or future) @@ -114,8 +114,8 @@ initialization, and if ``False``, raise :exc:`~scrapy.exceptions.NotConfigured`. When choosing a name for a custom setting, it is also a good idea to have a -look at the names of :ref:`built-in settings `, to try to -maintain consistency with them. +look at the names of :ref:`built-in settings ` to maintain +consistency with them. .. _enforce-component-requirements: @@ -135,7 +135,7 @@ In the case of :ref:`downloader middlewares `, `, and :ref:`spider middlewares `, you should raise :exc:`~scrapy.exceptions.NotConfigured`, passing a description of the issue as -a parameter to the exception so that it is printed in the logs, for the user to +a parameter to the exception so that it is printed in the logs for the user to see. For other components, feel free to raise whatever other exception feels right to you; for example, :exc:`RuntimeError` would make sense for a Scrapy version mismatch, while :exc:`ValueError` may be better if the issue is the @@ -156,7 +156,7 @@ If your requirement is a minimum Scrapy version, you may use if parse_version(scrapy.__version__) < parse_version("2.7"): raise RuntimeError( f"{MyComponent.__qualname__} requires Scrapy 2.7 or " - f"later, which allow defining the process_spider_output " + f"later, which allows defining the process_spider_output " f"method of spider middlewares as an asynchronous " f"generator." ) @@ -168,7 +168,7 @@ The following function can be used to create an instance of a component class: .. autofunction:: scrapy.utils.misc.build_from_crawler -The following function can also be useful when implementing a component, to +The following function can also be useful when implementing a component to report the import path of the component class, e.g. when reporting problems: .. autofunction:: scrapy.utils.python.global_object_name diff --git a/docs/topics/contracts.rst b/docs/topics/contracts.rst index 61aef4bbb..9e9caea2c 100644 --- a/docs/topics/contracts.rst +++ b/docs/topics/contracts.rst @@ -1,16 +1,16 @@ .. _topics-contracts: -================= -Spiders Contracts -================= +================ +Spider Contracts +================ -Testing spiders can get particularly annoying and while nothing prevents you -from writing unit tests the task gets cumbersome quickly. Scrapy offers an +Testing spiders can get particularly annoying, and while nothing prevents you +from writing unit tests, the task gets cumbersome quickly. Scrapy offers an integrated way of testing your spiders by the means of contracts. -This allows you to test each callback of your spider by hardcoding a sample url -and check various constraints for how the callback processes the response. Each -contract is prefixed with an ``@`` and included in the docstring. See the +This allows you to test each callback of your spider by hardcoding a sample URL +and checking various constraints for how the callback processes the response. +Each contract is prefixed with an ``@`` and included in the docstring. See the following example: .. code-block:: python @@ -73,7 +73,7 @@ Use the :command:`check` command to run the contract checks. Custom Contracts ================ -If you find you need more power than the built-in Scrapy contracts you can +If you find you need more power than the built-in Scrapy contracts, you can create and load your own contracts in the project by using the :setting:`SPIDER_CONTRACTS` setting: @@ -100,17 +100,17 @@ override three methods: .. method:: Contract.adjust_request_args(args) - This receives a ``dict`` as an argument containing default arguments - for request object. :class:`~scrapy.Request` is used by default, - but this can be changed with the ``request_cls`` attribute. - If multiple contracts in chain have this attribute defined, the last one is used. + This receives a ``dict`` containing default arguments for the request + object. :class:`~scrapy.Request` is used by default, but you can change + it with the ``request_cls`` attribute. If multiple contracts in the + chain define this attribute, the last one is used. - Must return the same or a modified version of it. + It must return the same dictionary or a modified version. .. method:: Contract.pre_process(response) This allows hooking in various checks on the response received from the - sample request, before it's being passed to the callback. + sample request, before it is passed to the callback. .. method:: Contract.post_process(output) @@ -134,10 +134,8 @@ response received: class HasHeaderContract(Contract): - """ - Demo contract which checks the presence of a custom header - @has_header X-CustomHeader - """ + """Demo contract that checks the presence of a custom header: + @has_header X-CustomHeader""" name = "has_header" @@ -152,8 +150,8 @@ Detecting check runs ==================== When ``scrapy check`` is running, the ``SCRAPY_CHECK`` environment variable is -set to the ``true`` string. You can use :data:`os.environ` to perform any change to -your spiders or your settings when ``scrapy check`` is used: +set to the ``true`` string. You can use :data:`os.environ` to make any changes +to your spiders or your settings when ``scrapy check`` is used: .. code-block:: python diff --git a/docs/topics/coroutines.rst b/docs/topics/coroutines.rst index fdd3dcd87..070b57426 100644 --- a/docs/topics/coroutines.rst +++ b/docs/topics/coroutines.rst @@ -69,16 +69,16 @@ hence use coroutine syntax (e.g. ``await``, ``async for``, ``async with``): Using Deferred-based APIs ========================= -In addition to native coroutine APIs Scrapy has some APIs that return a +In addition to native coroutine APIs, Scrapy has some APIs that return a :class:`~twisted.internet.defer.Deferred` object or take a user-supplied function that returns a :class:`~twisted.internet.defer.Deferred` object. These APIs are also asynchronous but don't yet support native ``async def`` syntax. -In the future we plan to add support for the ``async def`` syntax to these APIs -or replace them with other APIs where changing the existing ones is +In the future, we plan to add support for the ``async def`` syntax to these +APIs or replace them with other APIs where changing the existing ones is possible. The following Scrapy methods return :class:`~twisted.internet.defer.Deferred` -objects (this list is not complete as it only includes methods that we think +objects (this list is not complete, as it only includes methods that we think may be useful for user code): - :class:`scrapy.crawler.Crawler`: @@ -151,7 +151,7 @@ return coroutines are listed in :ref:`coroutine-support`): - ``stat_file()`` -In most cases you can use these APIs in code that otherwise uses coroutines, by +In most cases, you can use these APIs in code that otherwise uses coroutines by wrapping a :class:`~twisted.internet.defer.Deferred` object into a :class:`~asyncio.Future` object or vice versa. See :ref:`asyncio-await-dfd` for more information about this. @@ -212,8 +212,8 @@ becomes: Coroutines may be used to call asynchronous code. This includes other coroutines, functions that return Deferreds and functions that return -:term:`awaitable objects ` such as :class:`~asyncio.Future`. -This means you can use many useful Python libraries providing such code: +:term:`awaitable objects ` such as :class:`~asyncio.Future`. This +means you can use many useful Python libraries that provide such code: .. skip: next .. code-block:: python @@ -235,8 +235,8 @@ This means you can use many useful Python libraries providing such code: # ... use response and additional_data to yield items and requests .. note:: Many libraries that use coroutines, such as `aio-libs`_, require the - :mod:`asyncio` loop and to use them you need to - :doc:`enable asyncio support in Scrapy`. + :mod:`asyncio` loop, and to use them you need to :doc:`enable asyncio + support in Scrapy`. .. note:: If you want to ``await`` on Deferreds while using the asyncio reactor, you need to :ref:`wrap them`. @@ -341,7 +341,7 @@ synchronous ``process_spider_output`` method gets a synchronous iterable as its - The whole output of the previous ``Request`` callback or ``process_spider_output`` method is awaited at this point. -- If an exception raises while awaiting the output of the previous +- If an exception is raised while awaiting the output of the previous ``Request`` callback or ``process_spider_output`` method, none of that output will be processed. @@ -442,8 +442,8 @@ For example: to define their ``process_spider_output`` method as an asynchronous generator. -Since 2.13.0, Scrapy provides a base class, +Since version 2.13.0, Scrapy provides a base class, :class:`~scrapy.spidermiddlewares.base.BaseSpiderMiddleware`, which implements the ``process_spider_output()`` and ``process_spider_output_async()`` methods, -so instead of duplicating the processing code you can override the +so instead of duplicating the processing code, you can override the ``get_processed_request()`` and/or the ``get_processed_item()`` method. diff --git a/docs/topics/debug.rst b/docs/topics/debug.rst index 988e37bbd..d45e33f0e 100644 --- a/docs/topics/debug.rst +++ b/docs/topics/debug.rst @@ -5,7 +5,7 @@ Debugging Spiders ================= This document explains the most common techniques for debugging spiders. -Consider the following Scrapy spider below: +Consider the following Scrapy spider: .. skip: next .. code-block:: python @@ -40,19 +40,19 @@ Consider the following Scrapy spider below: # populate more `item` fields return item -Basically this is a simple spider which parses two pages of items (the -start_urls). Items also have a details page with additional information, so we -use the ``cb_kwargs`` functionality of :class:`~scrapy.Request` to pass a -partially populated item. +Basically, this is a simple spider that parses two pages of items (the start +URLs). Items also have a details page with additional information, so we use +the ``cb_kwargs`` functionality of :class:`~scrapy.Request` to pass a partially +populated item. Parse Command ============= The most basic way of checking the output of your spider is to use the -:command:`parse` command. It allows to check the behaviour of different parts -of the spider at the method level. It has the advantage of being flexible and -simple to use, but does not allow debugging code inside a method. +:command:`parse` command. It allows you to check the behaviour of different +parts of the spider at the method level. It has the advantage of being flexible +and simple to use, but it does not allow debugging code inside a method. .. highlight:: none @@ -90,7 +90,7 @@ Using the ``--verbose`` or ``-v`` option we can see the status at each depth lev # Requests ----------------------------------------------------------------- [] -Checking items scraped from a single start_url, can also be easily achieved +Checking items scraped from a single start URL can also be easily achieved using:: $ scrapy parse --spider=myspider -d 3 'http://example.com/page1' @@ -101,10 +101,10 @@ using:: Scrapy Shell ============ -While the :command:`parse` command is very useful for checking behaviour of a -spider, it is of little help to check what happens inside a callback, besides -showing the response received and the output. How to debug the situation when -``parse_details`` sometimes receives no item? +While the :command:`parse` command is very useful for checking the behaviour of +a spider, it is of little help when checking what happens inside a callback +besides showing the response received and the output. How do you debug the +situation when ``parse_details`` sometimes receives no item? .. highlight:: python @@ -129,7 +129,7 @@ See also: :ref:`topics-shell-inspect-response`. Open in browser =============== -Sometimes you just want to see how a certain response looks in a browser, you +Sometimes you just want to see how a certain response looks in a browser; you can use the :func:`~scrapy.utils.response.open_in_browser` function for that: .. autofunction:: scrapy.utils.response.open_in_browser @@ -140,7 +140,7 @@ Logging Logging is another useful option for getting information about your spider run. Although not as convenient, it comes with the advantage that the logs will be -available in all future runs should they be necessary again: +available in all future runs should you need them again: .. code-block:: python diff --git a/docs/topics/deploy.rst b/docs/topics/deploy.rst index f3515b4be..7217fc6a3 100644 --- a/docs/topics/deploy.rst +++ b/docs/topics/deploy.rst @@ -24,8 +24,8 @@ Deploying to a Scrapyd Server `Scrapyd`_ is an open source application to run Scrapy spiders. It provides a server with HTTP API, capable of running and monitoring Scrapy spiders. -To deploy spiders to Scrapyd, you can use the scrapyd-deploy tool provided by -the `scrapyd-client`_ package. Please refer to the `scrapyd-deploy +To deploy spiders to Scrapyd, you can use the ``scrapyd-deploy`` tool provided +by the `scrapyd-client`_ package. Please refer to the `scrapyd-deploy documentation`_ for more information. Scrapyd is maintained by some of the Scrapy developers. @@ -38,12 +38,12 @@ Deploying to Zyte Scrapy Cloud `Zyte Scrapy Cloud`_ is a hosted, cloud-based service by Zyte_, the company behind Scrapy. -Zyte Scrapy Cloud removes the need to setup and monitor servers and provides a +Zyte Scrapy Cloud removes the need to set up and monitor servers and provides a nice UI to manage spiders and review scraped items, logs and stats. -To deploy spiders to Zyte Scrapy Cloud you can use the `shub`_ command line -tool. -Please refer to the `Zyte Scrapy Cloud documentation`_ for more information. +To deploy spiders to Zyte Scrapy Cloud, you can use the ``shub`` command-line +tool. Please refer to the `Zyte Scrapy Cloud documentation`_ for more +information. Zyte Scrapy Cloud is compatible with Scrapyd and one can switch between them as needed - the configuration is read from the ``scrapy.cfg`` file diff --git a/docs/topics/developer-tools.rst b/docs/topics/developer-tools.rst index 3d3f04793..5ba662946 100644 --- a/docs/topics/developer-tools.rst +++ b/docs/topics/developer-tools.rst @@ -4,10 +4,10 @@ Using your browser's Developer Tools for scraping ================================================= -Here is a general guide on how to use your browser's Developer Tools -to ease the scraping process. Today almost all browsers come with -built in `Developer Tools`_ and although we will use Firefox in this -guide, the concepts are applicable to any other browser. +Here is a general guide on how to use your browser's Developer Tools to ease +the scraping process. Today almost all browsers come with built-in `Developer +Tools`_, and although we will use Firefox in this guide, the concepts are +applicable to any other browser. In this guide we'll introduce the basic tools to use from a browser's Developer Tools by scraping `quotes.toscrape.com`_. @@ -19,9 +19,9 @@ Caveats with inspecting the live browser DOM Since Developer Tools operate on a live browser DOM, what you'll actually see when inspecting the page source is not the original HTML, but a modified one -after applying some browser clean up and executing JavaScript code. Firefox, -in particular, is known for adding ```` elements to tables. Scrapy, on -the other hand, does not modify the original page HTML, so you won't be able to +after applying some browser clean up and executing JavaScript code. Firefox, in +particular, is known for adding ```` elements to tables. Scrapy, on the +other hand, does not modify the original page HTML, so you won't be able to extract any data if you use ```` in your XPath expressions. Therefore, you should keep in mind the following things: @@ -42,17 +42,16 @@ Inspecting a website ==================== By far the most handy feature of the Developer Tools is the `Inspector` -feature, which allows you to inspect the underlying HTML code of -any webpage. To demonstrate the Inspector, let's look at the -`quotes.toscrape.com`_-site. +feature, which allows you to inspect the underlying HTML code of any webpage. +To demonstrate the Inspector, let's look at the `quotes.toscrape.com`_ site. On the site we have a total of ten quotes from various authors with specific tags, as well as the Top Ten Tags. Let's say we want to extract all the quotes on this page, without any meta-information about authors, tags, etc. -Instead of viewing the whole source code for the page, we can simply right click -on a quote and select ``Inspect Element (Q)``, which opens up the `Inspector`. -In it you should see something like this: +Instead of viewing the whole source code for the page, we can simply +right-click on a quote and select ``Inspect Element (Q)``, which opens up the +`Inspector`. In it you should see something like this: .. image:: _images/inspector_01.png :width: 777 @@ -76,10 +75,10 @@ anywhere. The advantage of the `Inspector` is that it automatically expands and collapses sections and tags of a webpage, which greatly improves readability. You can -expand and collapse a tag by clicking on the arrow in front of it or by double -clicking directly on the tag. If we expand the ``span`` tag with the ``class= -"text"`` we will see the quote-text we clicked on. The `Inspector` lets you -copy XPaths to selected elements. Let's try it out. +expand and collapse a tag by clicking on the arrow in front of it or by +double-clicking directly on the tag. If we expand the ``span`` tag with the +``class="text"`` attribute, we see the quote text we clicked on. The +`Inspector` lets you copy XPaths to selected elements. Let's try it out. First open the Scrapy shell at https://quotes.toscrape.com/ in a terminal: @@ -99,17 +98,17 @@ Then, back to your web browser, right-click on the ``span`` tag, select >>> response.xpath("/html/body/div/div[2]/div[1]/div[1]/span[1]/text()").getall() ['“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”'] -Adding ``text()`` at the end we are able to extract the first quote with this -basic selector. But this XPath is not really that clever. All it does is -go down a desired path in the source code starting from ``html``. So let's -see if we can refine our XPath a bit: +Adding ``text()`` at the end, we are able to extract the first quote with this +basic selector. But this XPath is not really that clever. All it does is go +down a desired path in the source code starting from ``html``. So let's see if +we can refine our XPath a bit: -If we check the `Inspector` again we'll see that directly beneath our -expanded ``div`` tag we have nine identical ``div`` tags, each with the -same attributes as our first. If we expand any of them, we'll see the same -structure as with our first quote: Two ``span`` tags and one ``div`` tag. We can -expand each ``span`` tag with the ``class="text"`` inside our ``div`` tags and -see each quote: +If we check the `Inspector` again we'll see that directly beneath our expanded +``div`` tag we have nine identical ``div`` tags, each with the same attributes +as our first. If we expand any of them, we'll see the same structure as with +our first quote: two ``span`` tags and one ``div`` tag. We can expand each +``span`` tag with the ``class="text"`` inside our ``div`` tags and see each +quote: .. code-block:: html @@ -145,14 +144,14 @@ source code or directly scrolling to an element you selected. Let's demonstrate a use case: Say you want to find the ``Next`` button on the page. Type ``Next`` into the -search bar on the top right of the `Inspector`. You should get two results. -The first is a ``li`` tag with the ``class="next"``, the second the text -of an ``a`` tag. Right click on the ``a`` tag and select ``Scroll into View``. -If you hover over the tag, you'll see the button highlighted. From here -we could easily create a :ref:`Link Extractor ` to -follow the pagination. On a simple site such as this, there may not be -the need to find an element visually but the ``Scroll into View`` function -can be quite useful on complex sites. +search bar on the top-right of the `Inspector`. You should get two results. The +first is a ``li`` tag with the ``class="next"``, the second the text of an +``a`` tag. Right-click on the ``a`` tag and select ``Scroll into View``. If you +hover over the tag, you'll see the button highlighted. From here we could +easily create a :ref:`Link Extractor ` to follow the +pagination. On a simple site such as this, there may not be the need to find an +element visually but the ``Scroll into View`` function can be quite useful on +complex sites. Note that the search bar can also be used to search for and test CSS selectors. For example, you could search for ``span.text`` to find @@ -161,19 +160,19 @@ exactly the ``span`` tag with the ``class="text"`` in the page. .. _topics-network-tool: -The Network-tool +The Network tool ================ -While scraping you may come across dynamic webpages where some parts -of the page are loaded dynamically through multiple requests. While -this can be quite tricky, the `Network`-tool in the Developer Tools -greatly facilitates this task. To demonstrate the Network-tool, let's -take a look at the page `quotes.toscrape.com/scroll`_. +While scraping you may come across dynamic webpages where some parts of the +page are loaded dynamically through multiple requests. While this can be quite +tricky, the `Network` tool in the Developer Tools greatly facilitates this +task. To demonstrate the Network tool, let's take a look at the page +`quotes.toscrape.com/scroll`_. -The page is quite similar to the basic `quotes.toscrape.com`_-page, -but instead of the above-mentioned ``Next`` button, the page -automatically loads new quotes when you scroll to the bottom. We -could go ahead and try out different XPaths directly, but instead -we'll check another quite useful command from the Scrapy shell: +The page is quite similar to the basic `quotes.toscrape.com`_ page, but instead +of the above-mentioned ``Next`` button, the page automatically loads new quotes +when you scroll to the bottom. We could go ahead and try out different XPaths +directly, but instead we'll check another quite useful command from the Scrapy +shell: .. skip: next @@ -192,12 +191,11 @@ bar with the word ``Loading...``. :height: 296 :alt: Response from quotes.toscrape.com/scroll -The ``view(response)`` command let's us view the response our -shell or later our spider receives from the server. Here we see -that some basic template is loaded which includes the title, -the login-button and the footer, but the quotes are missing. This -tells us that the quotes are being loaded from a different request -than ``quotes.toscrape/scroll``. +The ``view(response)`` command lets us view the response our shell or later our +spider receives from the server. Here we see that some basic template is loaded +which includes the title, the login button and the footer, but the quotes are +missing. This tells us that the quotes are being loaded from a different +request than ``quotes.toscrape.com/scroll``. If you click on the ``Network`` tab, you will probably only see two entries. The first thing we do is enable persistent logs by @@ -218,30 +216,29 @@ Here we see every request that has been made when reloading the page and can inspect each request and its response. So let's find out where our quotes are coming from: -First click on the request with the name ``scroll``. On the right -you can now inspect the request. In ``Headers`` you'll find details -about the request headers, such as the URL, the method, the IP-address, -and so on. We'll ignore the other tabs and click directly on ``Response``. +First click on the request with the name ``scroll``. On the right you can now +inspect the request. In ``Headers`` you'll find details about the request +headers, such as the URL, the method, the IP address, and so on. We'll ignore +the other tabs and click directly on ``Response``. -What you should see in the ``Preview`` pane is the rendered HTML-code, -that is exactly what we saw when we called ``view(response)`` in the -shell. Accordingly the ``type`` of the request in the log is ``html``. -The other requests have types like ``css`` or ``js``, but what -interests us is the one request called ``quotes?page=1`` with the -type ``json``. +What you should see in the ``Preview`` pane is the rendered HTML code, exactly +what we saw when we called ``view(response)`` in the shell. Accordingly, the +``type`` of the request in the log is ``html``. The other requests have types +like ``css`` or ``js``, but what interests us is the one request called +``quotes?page=1`` with the type ``json``. If we click on this request, we see that the request URL is -``https://quotes.toscrape.com/api/quotes?page=1`` and the response -is a JSON-object that contains our quotes. We can also right-click -on the request and open ``Open in new tab`` to get a better overview. +``https://quotes.toscrape.com/api/quotes?page=1`` and the response is a JSON +object that contains our quotes. We can also right-click on the request and +open ``Open in new tab`` to get a better overview. .. image:: _images/network_03.png :width: 777 :height: 375 - :alt: JSON-object returned from the quotes.toscrape API + :alt: JSON object returned from the quotes.toscrape API -With this response we can now easily parse the JSON-object and -also request each page to get every quote on the site: +With this response we can now easily parse the JSON object and also request +each page to get every quote on the site: .. code-block:: python @@ -264,15 +261,14 @@ also request each page to get every quote on the site: url = f"https://quotes.toscrape.com/api/quotes?page={self.page}" yield scrapy.Request(url=url, callback=self.parse) -This spider starts at the first page of the quotes-API. With each -response, we parse the ``response.text`` and assign it to ``data``. -This lets us operate on the JSON-object like on a Python dictionary. -We iterate through the ``quotes`` and print out the ``quote["text"]``. -If the handy ``has_next`` element is ``true`` (try loading -`quotes.toscrape.com/api/quotes?page=10`_ in your browser or a -page-number greater than 10), we increment the ``page`` attribute -and ``yield`` a new request, inserting the incremented page-number -into our ``url``. +This spider starts at the first page of the quotes-API. With each response, we +parse the ``response.text`` and assign it to ``data``. This lets us operate on +the JSON object like on a Python dictionary. We iterate through the ``quotes`` +and print out the ``quote["text"]``. If the handy ``has_next`` element is +``true`` (try loading `quotes.toscrape.com/api/quotes?page=10`_ in your browser +or a page number greater than 10), we increment the ``page`` attribute and +``yield`` a new request, inserting the incremented page number into our +``url``. .. _requests-from-curl: @@ -306,11 +302,11 @@ function to get a dictionary with the equivalent arguments: Note that to translate a cURL command into a Scrapy request, you may use `curl2scrapy `_. -As you can see, with a few inspections in the `Network`-tool we -were able to easily replicate the dynamic requests of the scrolling -functionality of the page. Crawling dynamic pages can be quite -daunting and pages can be very complex, but it (mostly) boils down -to identifying the correct request and replicating it in your spider. +As you can see, with a few inspections in the `Network` tool we were able to +easily replicate the dynamic requests of the scrolling functionality of the +page. Crawling dynamic pages can be quite daunting and pages can be very +complex, but it (mostly) boils down to identifying the correct request and +replicating it in your spider. .. _Developer Tools: https://en.wikipedia.org/wiki/Web_development_tools .. _quotes.toscrape.com: https://quotes.toscrape.com diff --git a/docs/topics/downloader-middleware.rst b/docs/topics/downloader-middleware.rst index 6f36acc68..c2d729670 100644 --- a/docs/topics/downloader-middleware.rst +++ b/docs/topics/downloader-middleware.rst @@ -5,7 +5,7 @@ Downloader Middleware ===================== The downloader middleware is a framework of hooks into Scrapy's -request/response processing. It's a light, low-level system for globally +request/response processing. It's a light, low-level system for globally altering Scrapy's requests and responses. .. _topics-downloader-middleware-setting: @@ -42,9 +42,10 @@ middleware performs a different action and your middleware could depend on some previous (or subsequent) middleware being applied. If you want to disable a built-in middleware (the ones defined in -:setting:`DOWNLOADER_MIDDLEWARES_BASE` and enabled by default) you must define it -in your project's :setting:`DOWNLOADER_MIDDLEWARES` setting and assign ``None`` -as its value. For example, if you want to disable the user-agent middleware: +:setting:`DOWNLOADER_MIDDLEWARES_BASE` and enabled by default) you must define +it in your project's :setting:`DOWNLOADER_MIDDLEWARES` setting and assign +``None`` as its value. For example, if you want to disable the user-agent +middleware: .. code-block:: python @@ -54,7 +55,7 @@ as its value. For example, if you want to disable the user-agent middleware: } Finally, keep in mind that some middlewares may need to be enabled through a -particular setting. See each middleware documentation for more info. +particular setting. See each middleware's documentation for more info. .. _topics-downloader-middleware-custom: @@ -79,14 +80,16 @@ defines one or more of these methods: :class:`~scrapy.http.Response` object, return a :class:`~scrapy.Request` object, or raise :exc:`~scrapy.exceptions.IgnoreRequest`. - If it returns ``None``, Scrapy will continue processing this request, executing all - other middlewares until, finally, the appropriate downloader handler is called - the request performed (and its response downloaded). + If it returns ``None``, Scrapy will continue processing this request, + executing all other middlewares until, finally, the appropriate + downloader handler is called and the request is performed (and its + response downloaded). - If it returns a :class:`~scrapy.http.Response` object, Scrapy won't bother - calling *any* other :meth:`process_request` or :meth:`process_exception` methods, - or the appropriate download function; it'll return that response. The :meth:`process_response` - methods of installed middleware is always called on every response. + If it returns a :class:`~scrapy.http.Response` object, Scrapy won't + bother calling *any* other :meth:`process_request` or + :meth:`process_exception` methods, or the appropriate download function; + it'll return that response. The :meth:`process_response` methods of + installed middleware are always called on every response. If it returns a :class:`~scrapy.Request` object, Scrapy will stop calling :meth:`process_request` methods and reschedule the returned request. Once the newly returned @@ -182,10 +185,10 @@ CookiesMiddleware sends them back on subsequent requests (from that spider), just like web browsers do. - .. caution:: When non-UTF8 encoded byte sequences are passed to a - :class:`~scrapy.Request`, the ``CookiesMiddleware`` will log - a warning. Refer to :ref:`topics-logging-advanced-customization` - to customize the logging behaviour. + .. caution:: When non-UTF-8 encoded byte sequences are passed to a + :class:`~scrapy.Request`, the ``CookiesMiddleware`` will log a warning. + Refer to :ref:`topics-logging-advanced-customization` + to customize the logging behavior. .. caution:: Cookies set via the ``Cookie`` header are not considered by the :ref:`cookies-mw`. If you need to set cookies for a request, use the @@ -307,23 +310,24 @@ HttpAuthMiddleware .. class:: HttpAuthMiddleware This middleware authenticates all requests generated from certain spiders - using `Basic access authentication`_ (aka. HTTP auth). + using `Basic access authentication`_ (aka HTTP auth). To enable HTTP authentication for a spider, set the ``http_user`` and ``http_pass`` spider attributes to the authentication data and the ``http_auth_domain`` spider attribute to the domain which requires this - authentication (its subdomains will be also handled in the same way). - You can set ``http_auth_domain`` to ``None`` to enable the - authentication for all requests but you risk leaking your authentication - credentials to unrelated domains. + authentication (its subdomains will also be handled in the same way). You + can set ``http_auth_domain`` to ``None`` to enable the authentication for + all requests but you risk leaking your authentication credentials to + unrelated domains. .. warning:: - In previous Scrapy versions HttpAuthMiddleware sent the authentication + In previous Scrapy versions, HttpAuthMiddleware sent the authentication data with all requests, which is a security problem if the spider - makes requests to several different domains. Currently if the + makes requests to several different domains. Currently, if the ``http_auth_domain`` attribute is not set, the middleware will use the domain of the first request, which will work for some spiders but not - for others. In the future the middleware will produce an error instead. + for others. In the future, the middleware will produce an error + instead. Example: @@ -351,8 +355,9 @@ HttpCacheMiddleware .. class:: HttpCacheMiddleware - This middleware provides low-level cache to all HTTP requests and responses. - It has to be combined with a cache storage backend as well as a cache policy. + This middleware provides a low-level cache to all HTTP requests and + responses. It has to be combined with a cache storage backend as well as a + cache policy. Scrapy ships with the following HTTP cache storage backends: @@ -372,7 +377,8 @@ HttpCacheMiddleware .. reqmeta:: dont_cache - You can also avoid caching a response on every policy using :reqmeta:`dont_cache` meta key equals ``True``. + You can also avoid caching a response, regardless of the policy, by setting + the :reqmeta:`dont_cache` meta key to ``True``. .. module:: scrapy.extensions.httpcache :noindex: @@ -384,10 +390,10 @@ Dummy policy (default) .. class:: DummyPolicy - This policy has no awareness of any HTTP Cache-Control directives. - Every request and its corresponding response are cached. When the same - request is seen again, the response is returned without transferring - anything from the Internet. + This policy has no awareness of any HTTP Cache-Control directives. Every + request and its corresponding response are cached. When the same request is + seen again, the response is returned without transferring anything from the + Internet. The Dummy policy is useful for testing spiders faster (without having to wait for downloads every time) and for trying your spider offline, @@ -576,8 +582,8 @@ HTTPCACHE_DIR Default: ``'httpcache'`` The directory to use for storing the (low-level) HTTP cache. If empty, the HTTP -cache will be disabled. If a relative path is given, is taken relative to the -project data dir. For more info see: :ref:`topics-project-structure`. +cache will be disabled. If a relative path is given, it is taken relative to +the project data dir. For more info see: :ref:`topics-project-structure`. .. setting:: HTTPCACHE_IGNORE_HTTP_CODES @@ -586,7 +592,7 @@ HTTPCACHE_IGNORE_HTTP_CODES Default: ``[]`` -Don't cache response with these HTTP codes. +Don't cache responses with these HTTP codes. .. setting:: HTTPCACHE_IGNORE_MISSING @@ -671,10 +677,10 @@ Default: ``[]`` List of Cache-Control directives in responses to be ignored. -Sites often set "no-store", "no-cache", "must-revalidate", etc., but get -upset at the traffic a spider can generate if it actually respects those -directives. This allows to selectively ignore Cache-Control directives -that are known to be unimportant for the sites being crawled. +Sites often set "no-store", "no-cache", "must-revalidate", etc., but get upset +at the traffic a spider can generate if it actually respects those directives. +This allows you to selectively ignore Cache-Control directives that are known +to be unimportant for the sites being crawled. We assume that the spider will not issue Cache-Control directives in requests unless it actually needs them, so directives in requests are @@ -735,9 +741,10 @@ HttpProxyMiddleware * ``no_proxy`` You can also set the meta key ``proxy`` per-request, to a value like - ``http://some_proxy_server:port`` or ``http://username:password@some_proxy_server:port``. - Keep in mind this value will take precedence over ``http_proxy``/``https_proxy`` - environment variables, and it will also ignore ``no_proxy`` environment variable. + ``http://some_proxy_server:port`` or + ``http://username:password@some_proxy_server:port``. Keep in mind that this + value will take precedence over ``http_proxy``/``https_proxy`` environment + variables, and it will also ignore ``no_proxy`` environment variable. HttpProxyMiddleware settings ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ @@ -882,7 +889,8 @@ MetaRefreshMiddleware .. class:: MetaRefreshMiddleware - This middleware handles redirection of requests based on meta-refresh html tag. + This middleware handles redirection of requests based on the HTML + meta-refresh tag. The :class:`MetaRefreshMiddleware` can be configured through the following settings (see the settings documentation for more info): @@ -891,9 +899,10 @@ settings (see the settings documentation for more info): * :setting:`METAREFRESH_IGNORE_TAGS` * :setting:`METAREFRESH_MAXDELAY` -This middleware obey :setting:`REDIRECT_MAX_TIMES` setting, :reqmeta:`dont_redirect`, -:reqmeta:`redirect_urls` and :reqmeta:`redirect_reasons` request meta keys as described -for :class:`RedirectMiddleware` +This middleware obeys the :setting:`REDIRECT_MAX_TIMES` setting, +:reqmeta:`dont_redirect`, :reqmeta:`redirect_urls` and +:reqmeta:`redirect_reasons` request meta keys as described for +:class:`RedirectMiddleware` MetaRefreshMiddleware settings @@ -1002,7 +1011,7 @@ RETRY_HTTP_CODES Default: ``[500, 502, 503, 504, 522, 524, 408, 429]`` Which HTTP response codes to retry. Other errors (DNS lookup issues, -connections lost, etc) are always retried. +connections lost, etc.) are always retried. In some cases you may want to add 400 to :setting:`RETRY_HTTP_CODES` because it is a common code used to indicate server overload. It is not included by diff --git a/docs/topics/dynamic-content.rst b/docs/topics/dynamic-content.rst index 18b3ce244..cac48a00e 100644 --- a/docs/topics/dynamic-content.rst +++ b/docs/topics/dynamic-content.rst @@ -85,16 +85,16 @@ It might be enough to yield a :class:`~scrapy.Request` with the same HTTP method and URL. However, you may also need to reproduce the body, headers and form parameters (see :class:`~scrapy.FormRequest`) of that request. -As all major browsers allow to export the requests in curl_ format, Scrapy -incorporates the method :meth:`~scrapy.Request.from_curl` to generate an equivalent -:class:`~scrapy.Request` from a cURL command. To get more information -visit :ref:`request from curl ` inside the network -tool section. +As all major browsers allow you to export requests in curl_ format, Scrapy +incorporates the method :meth:`~scrapy.Request.from_curl` to generate an +equivalent :class:`~scrapy.Request` from a cURL command. To get more +information visit :ref:`request from curl ` inside the +network tool section. Once you get the expected response, you can :ref:`extract the desired data from it `. -You can reproduce any request with Scrapy. However, some times reproducing all +You can reproduce any request with Scrapy. However, sometimes reproducing all necessary requests may not seem efficient in developer time. If that is your case, and crawling speed is not a major concern for you, you can alternatively consider :ref:`using a headless browser `. @@ -271,8 +271,8 @@ The following is a simple snippet to illustrate its usage within a Scrapy spider However, using `playwright-python`_ directly as in the above example -circumvents most of the Scrapy components (middlewares, dupefilter, etc). -We recommend using `scrapy-playwright`_ for a better integration. +circumvents most of the Scrapy components (middlewares, dupefilter, etc.). We +recommend using `scrapy-playwright`_ for a better integration. .. _AJAX: https://en.wikipedia.org/wiki/Ajax_%28programming%29 .. _CSS: https://en.wikipedia.org/wiki/Cascading_Style_Sheets diff --git a/docs/topics/email.rst b/docs/topics/email.rst index 1d7bad787..61920b4fb 100644 --- a/docs/topics/email.rst +++ b/docs/topics/email.rst @@ -8,12 +8,11 @@ Sending e-mail :synopsis: Email sending facility Although Python makes sending e-mails relatively easy via the :mod:`smtplib` -library, Scrapy provides its own facility for sending e-mails which is very -easy to use and it's implemented using :doc:`Twisted non-blocking IO +library, Scrapy provides its own facility for sending e-mails that is very easy +to use and is implemented using :doc:`Twisted non-blocking IO `, to avoid interfering with the non-blocking -IO of the crawler. It also provides a simple API for sending attachments and -it's very easy to configure, with a few :ref:`settings -`. +IO of the crawler. It also provides a simple API for sending attachments and is +very easy to configure, with a few :ref:`settings `. Quick example ============= @@ -27,7 +26,7 @@ the standard ``__init__`` method: mailer = MailSender() -Or you can instantiate it passing a :class:`scrapy.Crawler` instance, which +Or you can instantiate it by passing a :class:`scrapy.Crawler` instance, which will respect the :ref:`settings `: .. skip: start @@ -50,8 +49,8 @@ And here is how to use it to send an e-mail (without attachments): MailSender class reference ========================== -The MailSender :ref:`components ` is the preferred class to -use for sending emails from Scrapy, as it uses :doc:`Twisted non-blocking IO +The MailSender :ref:`component ` is the preferred class to +use for sending e-mails from Scrapy, as it uses :doc:`Twisted non-blocking IO `, like the rest of the framework. .. class:: MailSender(smtphost=None, mailfrom=None, smtpuser=None, smtppass=None, smtpport=None) @@ -67,9 +66,9 @@ use for sending emails from Scrapy, as it uses :doc:`Twisted non-blocking IO :param smtpuser: the SMTP user. If omitted, the :setting:`MAIL_USER` setting will be used. If not given, no SMTP authentication will be performed. - :type smtphost: str or bytes + :type smtpuser: str or bytes - :param smtppass: the SMTP pass for authentication. + :param smtppass: the SMTP password for authentication. :type smtppass: str or bytes :param smtpport: the SMTP port to connect to @@ -83,7 +82,7 @@ use for sending emails from Scrapy, as it uses :doc:`Twisted non-blocking IO .. method:: send(to, subject, body, cc=None, attachs=(), mimetype='text/plain', charset=None) - Send email to the given recipients. + Send an e-mail to the given recipients. :param to: the e-mail recipients as a string or as a list of strings :type to: str or list @@ -98,10 +97,10 @@ use for sending emails from Scrapy, as it uses :doc:`Twisted non-blocking IO :type body: str :param attachs: an iterable of tuples ``(attach_name, mimetype, - file_object)`` where ``attach_name`` is a string with the name that will - appear on the e-mail's attachment, ``mimetype`` is the mimetype of the - attachment and ``file_object`` is a readable file object with the - contents of the attachment + file_object)`` where ``attach_name`` is a string with the name that + will appear on the e-mail attachment, ``mimetype`` is the mimetype of + the attachment, and ``file_object`` is a readable file object with + the contents of the attachment :type attachs: collections.abc.Iterable :param mimetype: the MIME type of the e-mail @@ -116,9 +115,9 @@ use for sending emails from Scrapy, as it uses :doc:`Twisted non-blocking IO Mail settings ============= -These settings define the default ``__init__`` method values of the :class:`MailSender` -class, and can be used to configure e-mail notifications in your project without -writing any code (for those extensions and code that uses :class:`MailSender`). +These settings define the default arguments passed to :class:`MailSender` and +can be used to configure e-mail notifications in your project without writing +any code (for those extensions and code that use :class:`MailSender`). .. setting:: MAIL_FROM @@ -154,8 +153,8 @@ MAIL_USER Default: ``None`` -User to use for SMTP authentication. If disabled no SMTP authentication will be -performed. +User to use for SMTP authentication. If disabled, no SMTP authentication will +be performed. .. setting:: MAIL_PASS @@ -173,7 +172,8 @@ MAIL_TLS Default: ``False`` -Enforce using STARTTLS. STARTTLS is a way to take an existing insecure connection, and upgrade it to a secure connection using SSL/TLS. +Enforce using STARTTLS. STARTTLS is a way to take an existing insecure +connection and upgrade it to a secure connection using SSL/TLS. .. setting:: MAIL_SSL @@ -182,4 +182,4 @@ MAIL_SSL Default: ``False`` -Enforce connecting using an SSL encrypted connection +Enforce connecting using an SSL-encrypted connection. diff --git a/docs/topics/exceptions.rst b/docs/topics/exceptions.rst index 0b572ff95..58ba0ad44 100644 --- a/docs/topics/exceptions.rst +++ b/docs/topics/exceptions.rst @@ -93,12 +93,12 @@ signal handler to indicate that no further bytes should be downloaded for a resp The ``fail`` boolean parameter controls which method will handle the resulting response: -* If ``fail=True`` (default), the request errback is called. The response object is - available as the ``response`` attribute of the ``StopDownload`` exception, - which is in turn stored as the ``value`` attribute of the received - :class:`~twisted.python.failure.Failure` object. This means that in an errback - defined as ``def errback(self, failure)``, the response can be accessed though - ``failure.value.response``. +* If ``fail=True`` (default), the request errback is called. The response + object is available as the ``response`` attribute of the ``StopDownload`` + exception, which is in turn stored as the ``value`` attribute of the received + :class:`~twisted.python.failure.Failure` object. This means that in an + errback defined as ``def errback(self, failure)``, the response can be + accessed through ``failure.value.response``. * If ``fail=False``, the request callback is called instead. diff --git a/docs/topics/exporters.rst b/docs/topics/exporters.rst index d3186ef5c..96c9658f6 100644 --- a/docs/topics/exporters.rst +++ b/docs/topics/exporters.rst @@ -11,34 +11,34 @@ Once you have scraped your items, you often want to persist or export those items, to use the data in some other application. That is, after all, the whole purpose of the scraping process. -For this purpose Scrapy provides a collection of Item Exporters for different +For this purpose, Scrapy provides a collection of Item Exporters for different output formats, such as XML, CSV or JSON. Using Item Exporters ==================== If you are in a hurry, and just want to use an Item Exporter to output scraped -data see the :ref:`topics-feed-exports`. Otherwise, if you want to know how +data, see the :ref:`topics-feed-exports`. Otherwise, if you want to know how Item Exporters work or need more custom functionality (not covered by the default exports), continue reading below. -In order to use an Item Exporter, you must instantiate it with its required -args. Each Item Exporter requires different arguments, so check each exporter -documentation to be sure, in :ref:`topics-exporters-reference`. After you have +In order to use an Item Exporter, you must instantiate it with its required +args. Each Item Exporter requires different arguments, so check each exporter's +documentation in :ref:`topics-exporters-reference` to be sure. After you have instantiated your exporter, you have to: -1. call the method :meth:`~BaseItemExporter.start_exporting` in order to -signal the beginning of the exporting process +1. Call the method :meth:`~BaseItemExporter.start_exporting` in order to + signal the beginning of the exporting process -2. call the :meth:`~BaseItemExporter.export_item` method for each item you want -to export +2. Call the :meth:`~BaseItemExporter.export_item` method for each item you want + to export -3. and finally call the :meth:`~BaseItemExporter.finish_exporting` to signal -the end of the exporting process +3. Finally, call the :meth:`~BaseItemExporter.finish_exporting` method to + signal the end of the exporting process -Here you can see an :doc:`Item Pipeline ` which uses multiple -Item Exporters to group scraped items to different files according to the -value of one of their fields: +Here you can see an :doc:`Item Pipeline ` that uses multiple +Item Exporters to distribute scraped items into different files according to +the value of one of their fields: .. code-block:: python @@ -159,9 +159,10 @@ BaseItemExporter defining what fields to export, whether to export empty fields, or which encoding to use. - These features can be configured through the ``__init__`` method arguments which - populate their respective instance attributes: :attr:`fields_to_export`, - :attr:`export_empty_fields`, :attr:`encoding`, :attr:`indent`. + These features can be configured through the ``__init__`` method arguments, + which populate their respective instance attributes: + :attr:`fields_to_export`, :attr:`export_empty_fields`, :attr:`encoding`, + :attr:`indent`. .. versionadded:: 2.0 The *dont_fail* parameter. @@ -263,8 +264,9 @@ XmlItemExporter Exports items in XML format to the specified file object. - :param file: the file-like object to use for exporting the data. Its ``write`` method should - accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) + :param file: the file-like object to use for exporting the data. Its + ``write`` method should accept ``bytes`` (a disk file opened in binary + mode, an ``io.BytesIO`` object, etc) :param root_element: The name of root element in the exported XML. :type root_element: str @@ -297,7 +299,7 @@ XmlItemExporter Item(name=['John', 'Doe'], age='23') - Would be serialized as:: + It would be serialized as:: @@ -320,8 +322,9 @@ CsvItemExporter CSV columns, their order and their column names. The :attr:`export_empty_fields` attribute has no effect on this exporter. - :param file: the file-like object to use for exporting the data. Its ``write`` method should - accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) + :param file: the file-like object to use for exporting the data. Its + ``write`` method should accept ``bytes`` (a disk file opened in binary + mode, an ``io.BytesIO`` object, etc) :param include_headers_line: If enabled, makes the exporter output a header line with the field names taken from @@ -355,8 +358,9 @@ PickleItemExporter Exports items in pickle format to the given file-like object. - :param file: the file-like object to use for exporting the data. Its ``write`` method should - accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) + :param file: the file-like object to use for exporting the data. Its + ``write`` method should accept ``bytes`` (a disk file opened in binary + mode, an ``io.BytesIO`` object, etc) :param protocol: The pickle protocol to use. :type protocol: int @@ -366,7 +370,7 @@ PickleItemExporter The additional keyword arguments of this ``__init__`` method are passed to the :class:`BaseItemExporter` ``__init__`` method. - Pickle isn't a human readable format, so no output examples are provided. + Pickle isn't a human-readable format, so no output examples are provided. PprintItemExporter ------------------ @@ -375,8 +379,9 @@ PprintItemExporter Exports items in pretty print format to the specified file object. - :param file: the file-like object to use for exporting the data. Its ``write`` method should - accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) + :param file: the file-like object to use for exporting the data. Its + ``write`` method should accept ``bytes`` (a disk file opened in binary + mode, an ``io.BytesIO`` object, etc) The additional keyword arguments of this ``__init__`` method are passed to the :class:`BaseItemExporter` ``__init__`` method. @@ -399,8 +404,9 @@ JsonItemExporter arguments to the :class:`~json.JSONEncoder` ``__init__`` method, so you can use any :class:`~json.JSONEncoder` ``__init__`` method argument to customize this exporter. - :param file: the file-like object to use for exporting the data. Its ``write`` method should - accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) + :param file: the file-like object to use for exporting the data. Its + ``write`` method should accept ``bytes`` (a disk file opened in binary + mode, an ``io.BytesIO`` object, etc) A typical output of this exporter would be:: @@ -409,10 +415,10 @@ JsonItemExporter .. _json-with-large-data: - .. warning:: JSON is very simple and flexible serialization format, but it - doesn't scale well for large amounts of data since incremental (aka. + .. warning:: JSON is a very simple and flexible serialization format, but it + doesn't scale well for large amounts of data because incremental (aka stream-mode) parsing is not well supported (if at all) among JSON parsers - (on any language), and most of them just parse the entire object in + in any language, and most of them just parse the entire object in memory. If you want the power and simplicity of JSON with a more stream-friendly format, consider using :class:`JsonLinesItemExporter` instead, or splitting the output in multiple chunks. @@ -428,8 +434,9 @@ JsonLinesItemExporter the :class:`~json.JSONEncoder` ``__init__`` method, so you can use any :class:`~json.JSONEncoder` ``__init__`` method argument to customize this exporter. - :param file: the file-like object to use for exporting the data. Its ``write`` method should - accept ``bytes`` (a disk file opened in binary mode, a ``io.BytesIO`` object, etc) + :param file: the file-like object to use for exporting the data. Its + ``write`` method should accept ``bytes`` (a disk file opened in binary + mode, an ``io.BytesIO`` object, etc) A typical output of this exporter would be:: diff --git a/docs/topics/extensions.rst b/docs/topics/extensions.rst index 0fbd22a27..c3c356932 100644 --- a/docs/topics/extensions.rst +++ b/docs/topics/extensions.rst @@ -48,8 +48,8 @@ tasks triggered by them. Sample extension ---------------- -Here we will implement a simple extension to illustrate the concepts described -in the previous section. This extension will log a message every time: +Here we implement a simple extension to illustrate the concepts described in +the previous section. This extension will log a message every time: * a spider is opened * a spider is closed @@ -58,7 +58,7 @@ in the previous section. This extension will log a message every time: The extension will be enabled through the ``MYEXT_ENABLED`` setting and the number of items will be specified through the ``MYEXT_ITEMCOUNT`` setting. -Here is the code of such extension: +Here is the code for that extension: .. code-block:: python @@ -158,7 +158,7 @@ Provides a telnet console for getting into a Python interpreter inside the currently running Scrapy process, which can be very useful for debugging. The telnet console must be enabled by the :setting:`TELNETCONSOLE_ENABLED` -setting, and the server will listen in the port specified in +setting, and the server will listen on the port specified in :setting:`TELNETCONSOLE_PORT`. .. _topics-extensions-ref-memusage: @@ -180,7 +180,7 @@ Monitors the memory used by the Scrapy process that runs the spider and: The notification e-mails can be triggered when a certain warning value is reached (:setting:`MEMUSAGE_WARNING_MB`) and when the maximum value is reached -(:setting:`MEMUSAGE_LIMIT_MB`) which will also cause the spider to be closed +(:setting:`MEMUSAGE_LIMIT_MB`), which will also cause the spider to be closed and the Scrapy process to be terminated. This extension is enabled by the :setting:`MEMUSAGE_ENABLED` setting and @@ -202,7 +202,8 @@ Memory debugger extension An extension for debugging memory usage. It collects information about: * objects uncollected by the Python garbage collector -* objects left alive that shouldn't. For more info, see :ref:`topics-leaks-trackrefs` +* objects left alive that shouldn't be. For more info, see + :ref:`topics-leaks-trackrefs` To enable this extension, turn on the :setting:`MEMDEBUG_ENABLED` setting. The info will be stored in the stats. @@ -255,8 +256,8 @@ settings: .. note:: - When a certain closing condition is met, requests which are - currently in the downloader queue (up to :setting:`CONCURRENT_REQUESTS` + When a certain closing condition is met, requests that are currently in the + downloader queue (up to :setting:`CONCURRENT_REQUESTS` requests) are still processed. .. setting:: CLOSESPIDER_TIMEOUT @@ -268,8 +269,8 @@ Default: ``0`` An integer which specifies a number of seconds. If the spider remains open for more than that number of seconds, it will be automatically closed with the -reason ``closespider_timeout``. If zero (or non set), spiders won't be closed by -timeout. +reason ``closespider_timeout``. If zero (or not set), spiders won't be closed +by timeout. .. setting:: CLOSESPIDER_TIMEOUT_NO_ITEM @@ -280,8 +281,8 @@ Default: ``0`` An integer which specifies a number of seconds. If the spider has not produced any items in the last number of seconds, it will be closed with the reason -``closespider_timeout_no_item``. If zero (or non set), spiders won't be closed -regardless if it hasn't produced any items. +``closespider_timeout_no_item``. If zero (or not set), spiders won't be closed +even if they have not produced any items. .. setting:: CLOSESPIDER_ITEMCOUNT @@ -291,9 +292,9 @@ CLOSESPIDER_ITEMCOUNT Default: ``0`` An integer which specifies a number of items. If the spider scrapes more than -that amount and those items are passed by the item pipeline, the -spider will be closed with the reason ``closespider_itemcount``. -If zero (or non set), spiders won't be closed by number of passed items. +that amount and those items are passed by the item pipeline, the spider will be +closed with the reason ``closespider_itemcount``. If zero (or not set), spiders +won't be closed by the number of passed items. .. setting:: CLOSESPIDER_PAGECOUNT @@ -304,7 +305,7 @@ Default: ``0`` An integer which specifies the maximum number of responses to crawl. If the spider crawls more than that, the spider will be closed with the reason -``closespider_pagecount``. If zero (or non set), spiders won't be closed by +``closespider_pagecount``. If zero (or not set), spiders won't be closed by the number of crawled responses. .. setting:: CLOSESPIDER_PAGECOUNT_NO_ITEM @@ -318,7 +319,7 @@ An integer which specifies the maximum number of consecutive responses to crawl without items scraped. If the spider crawls more consecutive responses than that and no items are scraped in the meantime, the spider will be closed with the reason ``closespider_pagecount_no_item``. If zero (or not set), spiders won't be -closed by number of crawled responses with no items. +closed by the number of crawled responses with no items. .. setting:: CLOSESPIDER_ERRORCOUNT @@ -329,8 +330,8 @@ Default: ``0`` An integer which specifies the maximum number of errors to receive before closing the spider. If the spider generates more than that number of errors, -it will be closed with the reason ``closespider_errorcount``. If zero (or non -set), spiders won't be closed by number of errors. +it will be closed with the reason ``closespider_errorcount``. If zero (or not +set), spiders won't be closed by the number of errors. StatsMailer extension ~~~~~~~~~~~~~~~~~~~~~ @@ -341,7 +342,7 @@ StatsMailer extension .. class:: StatsMailer This simple extension can be used to send a notification e-mail every time a -domain has finished scraping, including the Scrapy stats collected. The email +domain has finished scraping, including the Scrapy stats collected. The e-mail will be sent to all recipients specified in the :setting:`STATSMAILER_RCPTS` setting. diff --git a/docs/topics/feed-exports.rst b/docs/topics/feed-exports.rst index 2184f2d0e..f2db4f0ca 100644 --- a/docs/topics/feed-exports.rst +++ b/docs/topics/feed-exports.rst @@ -5,9 +5,9 @@ Feed exports ============ One of the most frequently required features when implementing scrapers is -being able to store the scraped data properly and, quite often, that means -generating an "export file" with the scraped data (commonly called "export -feed") to be consumed by other systems. +properly storing the scraped data and, quite often, that means generating an +"export file" with the scraped data (commonly called an "export feed") to be +consumed by other systems. Scrapy provides this functionality out of the box with the Feed Exports, which allows you to generate feeds with the scraped items, using multiple @@ -31,7 +31,7 @@ For serializing the scraped data, the feed exports use the :ref:`Item exporters - :ref:`topics-feed-format-csv` - :ref:`topics-feed-format-xml` -But you can also extend the supported format through the +But you can also extend the supported formats through the :setting:`FEED_EXPORTERS` setting. .. _topics-feed-format-json: @@ -98,9 +98,9 @@ Marshal Storages ======== -When using the feed exports you define where to store the feed using one or multiple URIs_ -(through the :setting:`FEEDS` setting). The feed exports supports multiple -storage backend types which are defined by the URI scheme. +When using the feed exports you define where to store the feed using one or +multiple URIs_ (through the :setting:`FEEDS` setting). The feed exports support +multiple storage backend types that are defined by the URI scheme. The storages backends supported out of the box are: @@ -126,8 +126,8 @@ being created. These parameters are: - ``%(time)s`` - gets replaced by a timestamp when the feed is being created - ``%(name)s`` - gets replaced by the spider name -Any other named parameter gets replaced by the spider attribute of the same -name. For example, ``%(site_id)s`` would get replaced by the ``spider.site_id`` +Any other named parameter is replaced by the spider attribute of the same name. +For example, ``%(site_id)s`` would be replaced by the ``spider.site_id`` attribute the moment the feed is being created. Here are some examples to illustrate: @@ -160,16 +160,16 @@ The feeds are stored in the local filesystem. - Example URI: ``file:///tmp/export.csv`` - Required external libraries: none -Note that for the local filesystem storage (only) you can omit the scheme if -you specify an absolute path like ``/tmp/export.csv`` (Unix systems only). -Alternatively you can also use a :class:`pathlib.Path` object. +Note that for the local filesystem storage you can omit the scheme if you +specify an absolute path like ``/tmp/export.csv`` (Unix systems only). +Alternatively, you can use a :class:`pathlib.Path` object. .. _topics-feed-storage-ftp: FTP --- -The feeds are stored in a FTP server. +The feeds are stored on an FTP server. - URI scheme: ``ftp`` - Example URI: ``ftp://user:pass@ftp.example.com/path/to/export.csv`` diff --git a/docs/topics/item-pipeline.rst b/docs/topics/item-pipeline.rst index e67cf06c8..7756597f1 100644 --- a/docs/topics/item-pipeline.rst +++ b/docs/topics/item-pipeline.rst @@ -7,10 +7,10 @@ Item Pipeline After an item has been scraped by a spider, it is sent to the Item Pipeline which processes it through several components that are executed sequentially. -Each item pipeline component (sometimes referred as just "Item Pipeline") is a -Python class that implements a simple method. They receive an item and perform -an action over it, also deciding if the item should continue through the -pipeline or be dropped and no longer processed. +Each item pipeline component (sometimes referred to as just "Item Pipeline") is +a Python class that implements a simple method. Each component receives an item +and performs an action on it while deciding whether the item should continue +through the pipeline or be dropped and no longer processed. Typical uses of item pipelines are: @@ -28,7 +28,7 @@ implement the following method: .. method:: process_item(self, item) - This method is called for every item pipeline component. + Scrapy calls this method for every item processed by the pipeline component. `item` is an :ref:`item object `, see :ref:`supporting-item-types`. @@ -42,7 +42,7 @@ implement the following method: :param item: the scraped item :type item: :ref:`item object ` -Additionally, they may also implement the following methods: +Additionally, a component may implement the following methods: .. method:: open_spider(self) @@ -61,8 +61,8 @@ Price validation and dropping items with no prices Let's take a look at the following hypothetical pipeline that adjusts the ``price`` attribute for those items that do not include VAT -(``price_excludes_vat`` attribute), and drops those items which don't -contain a price: +(``price_excludes_vat`` attribute), and drops those items that don't contain a +price: .. code-block:: python @@ -109,16 +109,16 @@ format: self.file.write(line) return item -.. note:: The purpose of JsonWriterPipeline is just to introduce how to write - item pipelines. If you really want to store all scraped items into a JSON - file you should use the :ref:`Feed exports `. +.. note:: The JsonWriterPipeline example simply introduces how to write item + pipelines. If you really want to store all scraped items in a JSON file, + you should use the :ref:`Feed exports `. Write items to MongoDB ---------------------- -In this example we'll write items to MongoDB_ using pymongo_. -MongoDB address and database name are specified in Scrapy settings; -MongoDB collection is named after item class. +In this example, we'll write items to MongoDB_ using pymongo_. The MongoDB +address and database name are specified in the Scrapy settings; the MongoDB +collection is named after the item class. The main point of this example is to show how to :ref:`get the crawler ` and how to clean up the resources properly. @@ -161,13 +161,13 @@ The main point of this example is to show how to :ref:`get the crawler .. _ScreenshotPipeline: -Take screenshot of item ------------------------ +Take a screenshot of an item +---------------------------- This example demonstrates how to use :doc:`coroutine syntax ` in the :meth:`process_item` method. -This item pipeline makes a request to a locally-running instance of Splash_ to +This item pipeline makes a request to a locally running instance of Splash_ to render a screenshot of the item URL. After the request response is downloaded, the item pipeline saves the screenshot to a file and adds the filename to the item. @@ -184,12 +184,12 @@ item. class ScreenshotPipeline: - """Pipeline that uses Splash to render screenshot of - every Scrapy item.""" + """Pipeline that uses Splash to render a screenshot of every Scrapy + item.""" SPLASH_URL = "http://localhost:8050/render.png?url={}" - def __init__(crawler): + def __init__(self, crawler): self.crawler = crawler @classmethod @@ -204,16 +204,17 @@ item. response = await self.crawler.engine.download_async(request) if response.status != 200: - # Error happened, return item. + # An error occurred, so return the item. return item - # Save screenshot to file, filename will be hash of url. + # Save the screenshot to a file; the filename is the hash of the + # URL. url = adapter["url"] url_hash = hashlib.md5(url.encode("utf8")).hexdigest() filename = f"{url_hash}.png" Path(filename).write_bytes(response.body) - # Store filename in item. + # Store the filename in the item. adapter["screenshot_filename"] = filename return item @@ -222,9 +223,9 @@ item. Duplicates filter ----------------- -A filter that looks for duplicate items, and drops those items that were -already processed. Let's say that our items have a unique id, but our spider -returns multiples items with the same id: +This filter looks for duplicate items and drops those that were already +processed. Let's say that our items have a unique id, but our spider returns +multiple items with the same id: .. code-block:: python diff --git a/docs/topics/items.rst b/docs/topics/items.rst index 3588d033e..0feb00deb 100644 --- a/docs/topics/items.rst +++ b/docs/topics/items.rst @@ -8,7 +8,7 @@ Items :synopsis: Item and Field classes The main goal in scraping is to extract structured data from unstructured -sources, typically, web pages. :ref:`Spiders ` may return the +sources, typically web pages. :ref:`Spiders ` may return the extracted data as `items`, Python objects that define key-value pairs. Scrapy supports :ref:`multiple types of items `. When you create an @@ -21,9 +21,10 @@ receives an item, your code should :ref:`work for any item type Item Types ========== -Scrapy supports the following types of items, via the `itemadapter`_ library: +Scrapy supports the following types of items via the `itemadapter`_ library: :ref:`dictionaries `, :ref:`Item objects `, -:ref:`dataclass objects `, and :ref:`attrs objects `. +:ref:`dataclass objects `, and :ref:`attrs objects +`. .. _itemadapter: https://github.com/scrapy/itemadapter @@ -49,16 +50,16 @@ make it the most feature-complete item type: :class:`Item` objects replicate the standard :class:`dict` API, including its ``__init__`` method. -:class:`Item` allows the defining of field names, so that: +:class:`Item` allows you to define field names, so that: -- :class:`KeyError` is raised when using undefined field names (i.e. - prevents typos going unnoticed) +- :class:`KeyError` is raised when using undefined field names (i.e. prevents + typos from going unnoticed) - :ref:`Item exporters ` can export all fields by default even if the first scraped object does not have values for all of them -:class:`Item` also allows the defining of field metadata, which can be used to +:class:`Item` also lets you define field metadata, which can be used to :ref:`customize serialization `. :mod:`trackref` tracks :class:`Item` objects to help find memory leaks @@ -82,11 +83,12 @@ Dataclass objects .. versionadded:: 2.2 -:func:`~dataclasses.dataclass` allows the defining of item classes with field names, -so that :ref:`item exporters ` can export all fields by -default even if the first scraped object does not have values for all of them. +:func:`~dataclasses.dataclass` allows you to define item classes with field +names, so that :ref:`item exporters ` can export all fields +by default even if the first scraped object does not have values for all of +them. -Additionally, ``dataclass`` items also allow you to: +Additionally, ``dataclass`` items allow you to: * define the type and default value of each defined field. @@ -114,18 +116,18 @@ attr.s objects .. versionadded:: 2.2 -:func:`attr.s` allows the defining of item classes with field names, -so that :ref:`item exporters ` can export all fields by -default even if the first scraped object does not have values for all of them. +:func:`attr.s` allows you to define item classes with field names, so that +:ref:`item exporters ` can export all fields by default even +if the first scraped object does not have values for all of them. -Additionally, ``attr.s`` items also allow to: +Additionally, ``attr.s`` items allow you to: * define the type and default value of each defined field. * define custom field :ref:`metadata `, which can be used to :ref:`customize serialization `. -In order to use this type, the :doc:`attrs package ` needs to be installed. +To use this type, the :doc:`attrs package ` needs to be installed. Example: @@ -163,9 +165,9 @@ Item subclasses are declared using a simple class definition syntax and tags = scrapy.Field() last_updated = scrapy.Field(serializer=str) -.. note:: Those familiar with `Django`_ will notice that Scrapy Items are - declared similar to `Django Models`_, except that Scrapy Items are much - simpler as there is no concept of different field types. +.. note:: Those familiar with `Django`_ will notice that Scrapy items are + declared similarly to `Django Models`_, except that Scrapy items are much + simpler as there is no concept of different field types. .. _Django: https://www.djangoproject.com/ .. _Django Models: https://docs.djangoproject.com/en/dev/topics/db/models/ @@ -177,19 +179,19 @@ Declaring fields ---------------- :class:`Field` objects are used to specify metadata for each field. For -example, the serializer function for the ``last_updated`` field illustrated in -the example above. +example, they can store the serializer function for the ``last_updated`` field +illustrated above. You can specify any kind of metadata for each field. There is no restriction on -the values accepted by :class:`Field` objects. For this same -reason, there is no reference list of all available metadata keys. Each key -defined in :class:`Field` objects could be used by a different component, and -only those components know about it. You can also define and use any other -:class:`Field` key in your project too, for your own needs. The main goal of -:class:`Field` objects is to provide a way to define all field metadata in one -place. Typically, those components whose behaviour depends on each field use -certain field keys to configure that behaviour. You must refer to their -documentation to see which metadata keys are used by each component. +the values accepted by :class:`Field` objects. For this same reason, there is +no reference list of all available metadata keys. Each key defined in +:class:`Field` objects could be used by a different component, and only those +components know about it. You can also define and use any other :class:`Field` +key in your project too, for your own needs. The main goal of :class:`Field` +objects is to provide a way to define all field metadata in one place. +Typically, those components whose behavior depends on each field use certain +field keys to configure that behavior. You must refer to their documentation to +see which metadata keys are used by each component. It's important to note that the :class:`Field` objects used to declare the item do not stay assigned as class attributes. Instead, they can be accessed through @@ -217,7 +219,7 @@ Working with Item objects .. skip: start Here are some examples of common tasks performed with items, using the -``Product`` item :ref:`declared above `. You will +``Product`` item :ref:`declared above `. You will notice the API is very similar to the :class:`dict` API. Creating items @@ -341,7 +343,9 @@ Creating dicts from items: >>> dict(product) # create a dict from all populated values {'price': 1000, 'name': 'Desktop PC'} - Creating items from dicts: +Creating items from dicts: + +.. code-block:: pycon >>> Product({"name": "Laptop PC", "price": 1500}) Product(price=1500, name='Laptop PC') @@ -355,8 +359,8 @@ Creating dicts from items: Extending Item subclasses ------------------------- -You can extend Items (to add more fields or to change some metadata for some -fields) by declaring a subclass of your original Item. +You can extend items (to add more fields or to change some metadata for some +fields) by declaring a subclass of your original item. For example: @@ -387,7 +391,7 @@ Supporting All Item Types In code that receives an item, such as methods of :ref:`item pipelines ` or :ref:`spider middlewares -`, it is a good practice to use the +`, it is good practice to use the :class:`~itemadapter.ItemAdapter` class to write code that works for any supported item type. diff --git a/docs/topics/jobs.rst b/docs/topics/jobs.rst index 50bcaa6d6..6aa20f211 100644 --- a/docs/topics/jobs.rst +++ b/docs/topics/jobs.rst @@ -7,7 +7,7 @@ Jobs: pausing and resuming crawls Sometimes, for big sites, it's desirable to pause crawls and be able to resume them later. -Scrapy supports this functionality out of the box by providing the following +Scrapy supports this functionality out of the box through the following facilities: * a scheduler that persists scheduled requests on disk @@ -20,12 +20,11 @@ facilities: Job directory ============= -To enable persistence support you just need to define a *job directory* through -the ``JOBDIR`` setting. This directory will be for storing all required data to -keep the state of a single job (i.e. a spider run). It's important to note that -this directory must not be shared by different spiders, or even different -jobs/runs of the same spider, as it's meant to be used for storing the state of -a *single* job. +To enable persistence support, define a *job directory* through the ``JOBDIR`` +setting. This directory stores all required data to keep the state of a single +job (i.e. a spider run). It's important to note that this directory must not be +shared with different spiders or with other runs of the same spider, as it's +meant to store the state of a *single* job. How to use it ============= @@ -34,7 +33,7 @@ To start a spider with persistence support enabled, run it like this:: scrapy crawl somespider -s JOBDIR=crawls/somespider-1 -Then, you can stop the spider safely at any time (by pressing Ctrl-C or sending +Then you can stop the spider safely at any time (by pressing Ctrl-C or sending a signal), and resume it later by issuing the same command:: scrapy crawl somespider -s JOBDIR=crawls/somespider-1 @@ -44,11 +43,11 @@ a signal), and resume it later by issuing the same command:: Keeping persistent state between batches ======================================== -Sometimes you'll want to keep some persistent spider state between pause/resume -batches. You can use the ``spider.state`` attribute for that, which should be a -dict. There's :ref:`a built-in extension ` -that takes care of serializing, storing and loading that attribute from the job -directory, when the spider starts and stops. +Sometimes you'll want to keep persistent spider state between pause/resume +batches. Use the ``spider.state`` attribute for that. It should be a dict. +There's :ref:`a built-in extension ` that +takes care of serializing, storing and loading that attribute from the job +directory when the spider starts and stops. Here's an example of a callback that uses the spider state (other spider code is omitted for brevity): @@ -62,14 +61,14 @@ is omitted for brevity): Persistence gotchas =================== -There are a few things to keep in mind if you want to be able to use the Scrapy -persistence support: +There are a few things to keep in mind if you want to use Scrapy's persistence +support: -Cookies expiration ------------------- +Cookie expiration +----------------- -Cookies may expire. So, if you don't resume your spider quickly the requests -scheduled may no longer work. This won't be an issue if your spider doesn't rely +Cookies may expire, so if you don't resume your spider quickly, the scheduled +requests may no longer work. This won't be an issue if your spider doesn't rely on cookies. diff --git a/docs/topics/leaks.rst b/docs/topics/leaks.rst index e61f33aed..8f9022358 100644 --- a/docs/topics/leaks.rst +++ b/docs/topics/leaks.rst @@ -5,55 +5,53 @@ Debugging memory leaks ====================== In Scrapy, objects such as requests, responses and items have a finite -lifetime: they are created, used for a while, and finally destroyed. +lifetime: they are created, used for a while and finally destroyed. -From all those objects, the Request is probably the one with the longest -lifetime, as it stays waiting in the Scheduler queue until it's time to process -it. For more info see :ref:`topics-architecture`. +Of all those objects, the Request usually has the longest lifetime because it +waits in the Scheduler queue until it's time to process it. For more +information, see :ref:`topics-architecture`. -As these Scrapy objects have a (rather long) lifetime, there is always the risk -of accumulating them in memory without releasing them properly and thus causing -what is known as a "memory leak". +Because these Scrapy objects have a comparatively long lifetime, there is +always the risk of accumulating them in memory without releasing them properly +and thus causing what is known as a "memory leak". -To help debugging memory leaks, Scrapy provides a built-in mechanism for -tracking objects references called :ref:`trackref `, -and you can also use a third-party library called :ref:`muppy -` for more advanced memory debugging (see below for more -info). Both mechanisms must be used from the :ref:`Telnet Console -`. +To help debug memory leaks, Scrapy provides a built-in mechanism for tracking +object references called :ref:`trackref `, and you can +also use a third-party library called :ref:`muppy ` for +more advanced memory debugging (see below for more information). Both +mechanisms must be used from the :ref:`Telnet Console `. Common causes of memory leaks ============================= It happens quite often (sometimes by accident, sometimes on purpose) that the -Scrapy developer passes objects referenced in Requests (for example, using the -:attr:`~scrapy.Request.cb_kwargs` or :attr:`~scrapy.Request.meta` -attributes or the request callback function) and that effectively bounds the -lifetime of those referenced objects to the lifetime of the Request. This is, -by far, the most common cause of memory leaks in Scrapy projects, and a quite -difficult one to debug for newcomers. +Scrapy developer passes objects referenced in requests—for example, through the +:attr:`~scrapy.Request.cb_kwargs` or :attr:`~scrapy.Request.meta` attributes or +the request callback function—and that effectively ties the lifetime of those +referenced objects to the lifetime of the request. This is, by far, the most +common cause of memory leaks in Scrapy projects, and it can be quite difficult +for newcomers to debug. -In big projects, the spiders are typically written by different people and some -of those spiders could be "leaking" and thus affecting the rest of the other -(well-written) spiders when they get to run concurrently, which, in turn, -affects the whole crawling process. +In big projects, spiders are typically written by different people, and some of +those spiders could be "leaking" and thus affecting the rest of the +(well-written) spiders when they run concurrently, which, in turn, affects the +whole crawling process. -The leak could also come from a custom middleware, pipeline or extension that -you have written, if you are not releasing the (previously allocated) resources -properly. For example, allocating resources on :signal:`spider_opened` -but not releasing them on :signal:`spider_closed` may cause problems if -you're running :ref:`multiple spiders per process `. +The leak could also come from a custom middleware, pipeline, or extension that +you wrote if you are not releasing the previously allocated resources properly. +For example, allocating resources on :signal:`spider_opened` but not releasing +them on :signal:`spider_closed` may cause problems if you're running +:ref:`multiple spiders per process `. Too Many Requests? ------------------ By default Scrapy keeps the request queue in memory; it includes -:class:`~scrapy.Request` objects and all objects -referenced in Request attributes (e.g. in :attr:`~scrapy.Request.cb_kwargs` -and :attr:`~scrapy.Request.meta`). -While not necessarily a leak, this can take a lot of memory. Enabling -:ref:`persistent job queue ` could help keeping memory usage -in control. +:class:`~scrapy.Request` objects and all objects referenced in request +attributes (for example, :attr:`~scrapy.Request.cb_kwargs` and +:attr:`~scrapy.Request.meta`). While this is not necessarily a leak, it can +consume a lot of memory. Enabling the :ref:`persistent job queue ` +can help keep memory usage under control. .. _topics-leaks-trackrefs: @@ -62,9 +60,9 @@ Debugging memory leaks with ``trackref`` .. skip: start -:mod:`trackref` is a module provided by Scrapy to debug the most common cases of -memory leaks. It basically tracks the references to all live Request, -Response, Item, Spider and Selector objects. +:mod:`trackref` is a module provided by Scrapy to debug the most common cases +of memory leaks. It tracks the references to all live Request, Response, Item, +Spider and Selector objects. You can enter the telnet console and inspect how many objects (of the classes mentioned above) are currently alive using the ``prefs()`` function which is an @@ -83,16 +81,15 @@ alias to the :func:`~scrapy.utils.trackref.print_live_refs` function:: FormRequest 878 oldest: 7s ago As you can see, that report also shows the "age" of the oldest object in each -class. If you're running multiple spiders per process chances are you can -figure out which spider is leaking by looking at the oldest request or response. -You can get the oldest object of each class using the +class. If you're running multiple spiders per process, chances are you can +figure out which spider is leaking by looking at the oldest request or +response. You can get the oldest object of each class using the :func:`~scrapy.utils.trackref.get_oldest` function (from the telnet console). Which objects are tracked? -------------------------- -The objects tracked by ``trackrefs`` are all from these classes (and all its -subclasses): +``trackref`` tracks objects from these classes (and all their subclasses): * :class:`scrapy.Request` * :class:`scrapy.http.Response` @@ -109,9 +106,9 @@ Suppose we have some spider with a line similar to this one:: return Request(f"http://www.somenastyspider.com/product.php?pid={product_id}", callback=self.parse, cb_kwargs={'referer': response}) -That line is passing a response reference inside a request which effectively -ties the response lifetime to the requests' one, and that would definitely -cause memory leaks. +That line passes a response reference inside a request, which effectively ties +the response lifetime to the request's lifetime, and that will definitely cause +memory leaks. Let's see how we can discover the cause (without knowing it a priori, of course) by using the ``trackref`` tool. @@ -132,10 +129,10 @@ references: The fact that there are so many live responses (and that they're so old) is definitely suspicious, as responses should have a relatively short lifetime -compared to Requests. The number of responses is similar to the number -of requests, so it looks like they are tied in a some way. We can now go -and check the code of the spider to discover the nasty line that is -generating the leaks (passing response references inside requests). +compared to requests. The number of responses is similar to the number of +requests, so it looks like they are tied in some way. We can now check the code +of the spider to discover the line that is generating the leaks (passing +response references inside requests). Sometimes extra information about live objects can be helpful. Let's check the oldest response: @@ -147,8 +144,8 @@ Let's check the oldest response: >>> r.url 'http://www.somenastyspider.com/product.php?pid=123' -If you want to iterate over all objects, instead of getting the oldest one, you -can use the :func:`scrapy.utils.trackref.iter_all` function: +If you want to iterate over all objects instead of getting only the oldest one, +you can use the :func:`scrapy.utils.trackref.iter_all` function: .. code-block:: pycon @@ -161,11 +158,10 @@ can use the :func:`scrapy.utils.trackref.iter_all` function: Too many spiders? ----------------- -If your project has too many spiders executed in parallel, -the output of :func:`prefs` can be difficult to read. -For this reason, that function has a ``ignore`` argument which can be used to -ignore a particular class (and all its subclasses). For -example, this won't show any live references to spiders: +If your project has too many spiders executed in parallel, the output of +:func:`prefs` can be difficult to read. For this reason, that function has an +``ignore`` argument that you can use to omit a particular class (and all its +subclasses). For example, this won't show any live references to spiders: .. code-block:: pycon @@ -212,14 +208,13 @@ Here are the functions available in the :mod:`~scrapy.utils.trackref` module. Debugging memory leaks with muppy ================================= -``trackref`` provides a very convenient mechanism for tracking down memory -leaks, but it only keeps track of the objects that are more likely to cause -memory leaks. However, there are other cases where the memory leaks could come -from other (more or less obscure) objects. If this is your case, and you can't -find your leaks using ``trackref``, you still have another resource: the muppy -library. +``trackref`` provides a convenient mechanism for tracking down memory leaks, +but it only keeps track of the objects that are more likely to cause them. +However, sometimes leaks come from other (more or less obscure) objects. If +that happens and you can't find your leaks using ``trackref``, you still have +another resource: the muppy library. -You can use muppy from `Pympler`_. +muppy is available as part of `Pympler`_. .. _Pympler: https://pypi.org/project/Pympler/ @@ -227,8 +222,8 @@ If you use ``pip``, you can install muppy with the following command:: pip install Pympler -Here's an example to view all Python objects available in -the heap using muppy: +Here's an example that shows all Python objects available in the heap using +muppy: .. skip: start .. code-block:: pycon @@ -260,7 +255,7 @@ the heap using muppy: .. skip: end -For more info about muppy, refer to the `muppy documentation`_. +For more information about muppy, refer to the `muppy documentation`_. .. _muppy documentation: https://pythonhosted.org/Pympler/muppy.html @@ -269,10 +264,10 @@ For more info about muppy, refer to the `muppy documentation`_. Leaks without leaks =================== -Sometimes, you may notice that the memory usage of your Scrapy process will -only increase, but never decrease. Unfortunately, this could happen even -though neither Scrapy nor your project are leaking memory. This is due to a -(not so well) known problem of Python, which may not return released memory to +Sometimes you may notice that the memory usage of your Scrapy process only +increases and never decreases. Unfortunately, this could happen even though +neither Scrapy nor your project are leaking memory. This is due to a +not-so-well-known problem in Python, which may not return released memory to the operating system in some cases. For more information on this issue see: * `Python Memory Management `_ @@ -293,6 +288,6 @@ completely. To quote the paper: .. _this paper: https://www.evanjones.ca/memoryallocator/ -To keep memory consumption reasonable you can split the job into several -smaller jobs or enable :ref:`persistent job queue ` -and stop/start spider from time to time. +To keep memory consumption reasonable, you can split the job into several +smaller jobs or enable the :ref:`persistent job queue ` and stop +and start the spider from time to time. diff --git a/docs/topics/link-extractors.rst b/docs/topics/link-extractors.rst index f9744ed16..b591880ed 100644 --- a/docs/topics/link-extractors.rst +++ b/docs/topics/link-extractors.rst @@ -16,9 +16,10 @@ list of matching :class:`~scrapy.link.Link` objects from a Link extractors are used in :class:`~scrapy.spiders.CrawlSpider` spiders through a set of :class:`~scrapy.spiders.Rule` objects. -You can also use link extractors in regular spiders. For example, you can instantiate -:class:`LinkExtractor ` into a class -variable in your spider, and use it from your spider callbacks: +You can also use link extractors in regular spiders. For example, you can +instantiate :class:`LinkExtractor +` as a class variable in your +spider and use it from your spider callbacks: .. code-block:: python @@ -53,22 +54,22 @@ LxmlLinkExtractor options. It is implemented using lxml's robust HTMLParser. :param allow: a single regular expression (or list of regular expressions) - that the (absolute) urls must match in order to be extracted. If not + that the (absolute) URLs must match in order to be extracted. If not given (or empty), it will match all links. :type allow: str or list :param deny: a single regular expression (or list of regular expressions) - that the (absolute) urls must match in order to be excluded (i.e. not + that the (absolute) URLs must match in order to be excluded (i.e. not extracted). It has precedence over the ``allow`` parameter. If not - given (or empty) it won't exclude any links. + given (or empty), it won't exclude any links. :type deny: str or list - :param allow_domains: a single value or a list of string containing - domains which will be considered for extracting the links + :param allow_domains: a single value or a list of strings containing + domains that will be considered for extracting the links. :type allow_domains: str or list :param deny_domains: a single value or a list of strings containing - domains which won't be considered for extracting the links + domains that won't be considered for extracting the links. :type deny_domains: str or list :param deny_extensions: a single value or list of strings containing @@ -82,15 +83,14 @@ LxmlLinkExtractor ``iso``, ``tar``, ``tar.gz``, ``webm``, and ``xz``. :type deny_extensions: list - :param restrict_xpaths: is an XPath (or list of XPath's) which defines - regions inside the response where links should be extracted from. - If given, only the text selected by those XPath will be scanned for - links. + :param restrict_xpaths: an XPath (or list of XPaths) that defines regions + inside the response where links should be extracted. If given, only the + text selected by those XPaths will be scanned for links. :type restrict_xpaths: str or list - :param restrict_css: a CSS selector (or list of selectors) which defines - regions inside the response where links should be extracted from. - Has the same behaviour as ``restrict_xpaths``. + :param restrict_css: a CSS selector (or list of selectors) that defines + regions inside the response where links should be extracted. It has the + same behaviour as ``restrict_xpaths``. :type restrict_css: str or list :param restrict_text: a single regular expression (or list of regular expressions) @@ -103,25 +103,25 @@ LxmlLinkExtractor Defaults to ``('a', 'area')``. :type tags: str or list - :param attrs: an attribute or list of attributes which should be considered when looking + :param attrs: an attribute or list of attributes that should be considered when looking for links to extract (only for those tags specified in the ``tags`` - parameter). Defaults to ``('href',)`` + parameter). Defaults to ``('href',)``. :type attrs: list - :param canonicalize: canonicalize each extracted url (using - w3lib.url.canonicalize_url). Defaults to ``False``. - Note that canonicalize_url is meant for duplicate checking; - it can change the URL visible at server side, so the response can be - different for requests with canonicalized and raw URLs. If you're - using LinkExtractor to follow links it is more robust to - keep the default ``canonicalize=False``. + :param canonicalize: canonicalize each extracted URL (using + w3lib.url.canonicalize_url). Defaults to ``False``. Note that + canonicalize_url is meant for duplicate checking; it can change the URL + visible at the server side, so the response can be different for + requests with canonicalized and raw URLs. If you're using LinkExtractor + to follow links, it is more robust to keep the default + ``canonicalize=False``. :type canonicalize: bool :param unique: whether duplicate filtering should be applied to extracted links. :type unique: bool - :param process_value: a function which receives each value extracted from + :param process_value: a function that receives each value extracted from the tag and attributes scanned and can modify the value and return a new one, or return ``None`` to ignore the link altogether. If not given, ``process_value`` defaults to ``lambda x: x``. @@ -146,12 +146,13 @@ LxmlLinkExtractor :type process_value: collections.abc.Callable :param strip: whether to strip whitespaces from extracted attributes. - According to HTML5 standard, leading and trailing whitespaces - must be stripped from ``href`` attributes of ````, ```` - and many other elements, ``src`` attribute of ````, ``