Remove trailing whitespace

This commit is contained in:
Adrián Chaves 2025-03-11 11:56:44 +01:00
parent 2accaa4af4
commit bee74fb753
26 changed files with 201 additions and 197 deletions

View File

@ -22,7 +22,7 @@ jobs:
- uses: actions/setup-python@v5 - uses: actions/setup-python@v5
with: with:
python-version: "3.13" python-version: "3.13"
- run: | - run: |
python -m pip install --upgrade build python -m pip install --upgrade build
python -m build python -m build
- name: Publish to PyPI - name: Publish to PyPI

View File

@ -11,3 +11,7 @@ repos:
- id: blacken-docs - id: blacken-docs
additional_dependencies: additional_dependencies:
- black==24.10.0 - black==24.10.0
- repo: https://github.com/pre-commit/pre-commit-hooks
rev: v5.0.0
hooks:
- id: trailing-whitespace

View File

@ -1,6 +1,6 @@
.. image:: https://scrapy.org/img/scrapylogo.png .. image:: https://scrapy.org/img/scrapylogo.png
:target: https://scrapy.org/ :target: https://scrapy.org/
====== ======
Scrapy Scrapy
====== ======

View File

@ -16,13 +16,13 @@
</div> </div>
<div class="col-md-4"> <div class="col-md-4">
<p> <p>
<a href="/login">Login</a> <a href="/login">Login</a>
</p> </p>
</div> </div>
</div> </div>
<div class="row"> <div class="row">
<div class="col-md-8"> <div class="col-md-8">
@ -34,16 +34,16 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="change,deep-thoughts,thinking,world" / > <meta class="keywords" itemprop="keywords" content="change,deep-thoughts,thinking,world" / >
<a class="tag" href="/tag/change/page/1/">change</a> <a class="tag" href="/tag/change/page/1/">change</a>
<a class="tag" href="/tag/deep-thoughts/page/1/">deep-thoughts</a> <a class="tag" href="/tag/deep-thoughts/page/1/">deep-thoughts</a>
<a class="tag" href="/tag/thinking/page/1/">thinking</a> <a class="tag" href="/tag/thinking/page/1/">thinking</a>
<a class="tag" href="/tag/world/page/1/">world</a> <a class="tag" href="/tag/world/page/1/">world</a>
</div> </div>
</div> </div>
@ -54,12 +54,12 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="abilities,choices" / > <meta class="keywords" itemprop="keywords" content="abilities,choices" / >
<a class="tag" href="/tag/abilities/page/1/">abilities</a> <a class="tag" href="/tag/abilities/page/1/">abilities</a>
<a class="tag" href="/tag/choices/page/1/">choices</a> <a class="tag" href="/tag/choices/page/1/">choices</a>
</div> </div>
</div> </div>
@ -70,18 +70,18 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="inspirational,life,live,miracle,miracles" / > <meta class="keywords" itemprop="keywords" content="inspirational,life,live,miracle,miracles" / >
<a class="tag" href="/tag/inspirational/page/1/">inspirational</a> <a class="tag" href="/tag/inspirational/page/1/">inspirational</a>
<a class="tag" href="/tag/life/page/1/">life</a> <a class="tag" href="/tag/life/page/1/">life</a>
<a class="tag" href="/tag/live/page/1/">live</a> <a class="tag" href="/tag/live/page/1/">live</a>
<a class="tag" href="/tag/miracle/page/1/">miracle</a> <a class="tag" href="/tag/miracle/page/1/">miracle</a>
<a class="tag" href="/tag/miracles/page/1/">miracles</a> <a class="tag" href="/tag/miracles/page/1/">miracles</a>
</div> </div>
</div> </div>
@ -92,16 +92,16 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="aliteracy,books,classic,humor" / > <meta class="keywords" itemprop="keywords" content="aliteracy,books,classic,humor" / >
<a class="tag" href="/tag/aliteracy/page/1/">aliteracy</a> <a class="tag" href="/tag/aliteracy/page/1/">aliteracy</a>
<a class="tag" href="/tag/books/page/1/">books</a> <a class="tag" href="/tag/books/page/1/">books</a>
<a class="tag" href="/tag/classic/page/1/">classic</a> <a class="tag" href="/tag/classic/page/1/">classic</a>
<a class="tag" href="/tag/humor/page/1/">humor</a> <a class="tag" href="/tag/humor/page/1/">humor</a>
</div> </div>
</div> </div>
@ -112,12 +112,12 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="be-yourself,inspirational" / > <meta class="keywords" itemprop="keywords" content="be-yourself,inspirational" / >
<a class="tag" href="/tag/be-yourself/page/1/">be-yourself</a> <a class="tag" href="/tag/be-yourself/page/1/">be-yourself</a>
<a class="tag" href="/tag/inspirational/page/1/">inspirational</a> <a class="tag" href="/tag/inspirational/page/1/">inspirational</a>
</div> </div>
</div> </div>
@ -128,14 +128,14 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="adulthood,success,value" / > <meta class="keywords" itemprop="keywords" content="adulthood,success,value" / >
<a class="tag" href="/tag/adulthood/page/1/">adulthood</a> <a class="tag" href="/tag/adulthood/page/1/">adulthood</a>
<a class="tag" href="/tag/success/page/1/">success</a> <a class="tag" href="/tag/success/page/1/">success</a>
<a class="tag" href="/tag/value/page/1/">value</a> <a class="tag" href="/tag/value/page/1/">value</a>
</div> </div>
</div> </div>
@ -146,12 +146,12 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="life,love" / > <meta class="keywords" itemprop="keywords" content="life,love" / >
<a class="tag" href="/tag/life/page/1/">life</a> <a class="tag" href="/tag/life/page/1/">life</a>
<a class="tag" href="/tag/love/page/1/">love</a> <a class="tag" href="/tag/love/page/1/">love</a>
</div> </div>
</div> </div>
@ -162,16 +162,16 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="edison,failure,inspirational,paraphrased" / > <meta class="keywords" itemprop="keywords" content="edison,failure,inspirational,paraphrased" / >
<a class="tag" href="/tag/edison/page/1/">edison</a> <a class="tag" href="/tag/edison/page/1/">edison</a>
<a class="tag" href="/tag/failure/page/1/">failure</a> <a class="tag" href="/tag/failure/page/1/">failure</a>
<a class="tag" href="/tag/inspirational/page/1/">inspirational</a> <a class="tag" href="/tag/inspirational/page/1/">inspirational</a>
<a class="tag" href="/tag/paraphrased/page/1/">paraphrased</a> <a class="tag" href="/tag/paraphrased/page/1/">paraphrased</a>
</div> </div>
</div> </div>
@ -182,10 +182,10 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="misattributed-eleanor-roosevelt" / > <meta class="keywords" itemprop="keywords" content="misattributed-eleanor-roosevelt" / >
<a class="tag" href="/tag/misattributed-eleanor-roosevelt/page/1/">misattributed-eleanor-roosevelt</a> <a class="tag" href="/tag/misattributed-eleanor-roosevelt/page/1/">misattributed-eleanor-roosevelt</a>
</div> </div>
</div> </div>
@ -196,73 +196,73 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="humor,obvious,simile" / > <meta class="keywords" itemprop="keywords" content="humor,obvious,simile" / >
<a class="tag" href="/tag/humor/page/1/">humor</a> <a class="tag" href="/tag/humor/page/1/">humor</a>
<a class="tag" href="/tag/obvious/page/1/">obvious</a> <a class="tag" href="/tag/obvious/page/1/">obvious</a>
<a class="tag" href="/tag/simile/page/1/">simile</a> <a class="tag" href="/tag/simile/page/1/">simile</a>
</div> </div>
</div> </div>
<nav> <nav>
<ul class="pager"> <ul class="pager">
<li class="next"> <li class="next">
<a href="/page/2/">Next <span aria-hidden="true">&rarr;</span></a> <a href="/page/2/">Next <span aria-hidden="true">&rarr;</span></a>
</li> </li>
</ul> </ul>
</nav> </nav>
</div> </div>
<div class="col-md-4 tags-box"> <div class="col-md-4 tags-box">
<h2>Top Ten tags</h2> <h2>Top Ten tags</h2>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 28px" href="/tag/love/">love</a> <a class="tag" style="font-size: 28px" href="/tag/love/">love</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 26px" href="/tag/inspirational/">inspirational</a> <a class="tag" style="font-size: 26px" href="/tag/inspirational/">inspirational</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 26px" href="/tag/life/">life</a> <a class="tag" style="font-size: 26px" href="/tag/life/">life</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 24px" href="/tag/humor/">humor</a> <a class="tag" style="font-size: 24px" href="/tag/humor/">humor</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 22px" href="/tag/books/">books</a> <a class="tag" style="font-size: 22px" href="/tag/books/">books</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 14px" href="/tag/reading/">reading</a> <a class="tag" style="font-size: 14px" href="/tag/reading/">reading</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 10px" href="/tag/friendship/">friendship</a> <a class="tag" style="font-size: 10px" href="/tag/friendship/">friendship</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 8px" href="/tag/friends/">friends</a> <a class="tag" style="font-size: 8px" href="/tag/friends/">friends</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 8px" href="/tag/truth/">truth</a> <a class="tag" style="font-size: 8px" href="/tag/truth/">truth</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 6px" href="/tag/simile/">simile</a> <a class="tag" style="font-size: 6px" href="/tag/simile/">simile</a>
</span> </span>
</div> </div>
</div> </div>

View File

@ -16,13 +16,13 @@
</div> </div>
<div class="col-md-4"> <div class="col-md-4">
<p> <p>
<a href="/login">Login</a> <a href="/login">Login</a>
</p> </p>
</div> </div>
</div> </div>
<div class="row"> <div class="row">
<div class="col-md-8"> <div class="col-md-8">
@ -34,16 +34,16 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="change,deep-thoughts,thinking,world" / > <meta class="keywords" itemprop="keywords" content="change,deep-thoughts,thinking,world" / >
<a class="tag" href="/tag/change/page/1/">change</a> <a class="tag" href="/tag/change/page/1/">change</a>
<a class="tag" href="/tag/deep-thoughts/page/1/">deep-thoughts</a> <a class="tag" href="/tag/deep-thoughts/page/1/">deep-thoughts</a>
<a class="tag" href="/tag/thinking/page/1/">thinking</a> <a class="tag" href="/tag/thinking/page/1/">thinking</a>
<a class="tag" href="/tag/world/page/1/">world</a> <a class="tag" href="/tag/world/page/1/">world</a>
</div> </div>
</div> </div>
@ -54,12 +54,12 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="abilities,choices" / > <meta class="keywords" itemprop="keywords" content="abilities,choices" / >
<a class="tag" href="/tag/abilities/page/1/">abilities</a> <a class="tag" href="/tag/abilities/page/1/">abilities</a>
<a class="tag" href="/tag/choices/page/1/">choices</a> <a class="tag" href="/tag/choices/page/1/">choices</a>
</div> </div>
</div> </div>
@ -70,18 +70,18 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="inspirational,life,live,miracle,miracles" / > <meta class="keywords" itemprop="keywords" content="inspirational,life,live,miracle,miracles" / >
<a class="tag" href="/tag/inspirational/page/1/">inspirational</a> <a class="tag" href="/tag/inspirational/page/1/">inspirational</a>
<a class="tag" href="/tag/life/page/1/">life</a> <a class="tag" href="/tag/life/page/1/">life</a>
<a class="tag" href="/tag/live/page/1/">live</a> <a class="tag" href="/tag/live/page/1/">live</a>
<a class="tag" href="/tag/miracle/page/1/">miracle</a> <a class="tag" href="/tag/miracle/page/1/">miracle</a>
<a class="tag" href="/tag/miracles/page/1/">miracles</a> <a class="tag" href="/tag/miracles/page/1/">miracles</a>
</div> </div>
</div> </div>
@ -92,16 +92,16 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="aliteracy,books,classic,humor" / > <meta class="keywords" itemprop="keywords" content="aliteracy,books,classic,humor" / >
<a class="tag" href="/tag/aliteracy/page/1/">aliteracy</a> <a class="tag" href="/tag/aliteracy/page/1/">aliteracy</a>
<a class="tag" href="/tag/books/page/1/">books</a> <a class="tag" href="/tag/books/page/1/">books</a>
<a class="tag" href="/tag/classic/page/1/">classic</a> <a class="tag" href="/tag/classic/page/1/">classic</a>
<a class="tag" href="/tag/humor/page/1/">humor</a> <a class="tag" href="/tag/humor/page/1/">humor</a>
</div> </div>
</div> </div>
@ -112,12 +112,12 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="be-yourself,inspirational" / > <meta class="keywords" itemprop="keywords" content="be-yourself,inspirational" / >
<a class="tag" href="/tag/be-yourself/page/1/">be-yourself</a> <a class="tag" href="/tag/be-yourself/page/1/">be-yourself</a>
<a class="tag" href="/tag/inspirational/page/1/">inspirational</a> <a class="tag" href="/tag/inspirational/page/1/">inspirational</a>
</div> </div>
</div> </div>
@ -128,14 +128,14 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="adulthood,success,value" / > <meta class="keywords" itemprop="keywords" content="adulthood,success,value" / >
<a class="tag" href="/tag/adulthood/page/1/">adulthood</a> <a class="tag" href="/tag/adulthood/page/1/">adulthood</a>
<a class="tag" href="/tag/success/page/1/">success</a> <a class="tag" href="/tag/success/page/1/">success</a>
<a class="tag" href="/tag/value/page/1/">value</a> <a class="tag" href="/tag/value/page/1/">value</a>
</div> </div>
</div> </div>
@ -146,12 +146,12 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="life,love" / > <meta class="keywords" itemprop="keywords" content="life,love" / >
<a class="tag" href="/tag/life/page/1/">life</a> <a class="tag" href="/tag/life/page/1/">life</a>
<a class="tag" href="/tag/love/page/1/">love</a> <a class="tag" href="/tag/love/page/1/">love</a>
</div> </div>
</div> </div>
@ -162,16 +162,16 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="edison,failure,inspirational,paraphrased" / > <meta class="keywords" itemprop="keywords" content="edison,failure,inspirational,paraphrased" / >
<a class="tag" href="/tag/edison/page/1/">edison</a> <a class="tag" href="/tag/edison/page/1/">edison</a>
<a class="tag" href="/tag/failure/page/1/">failure</a> <a class="tag" href="/tag/failure/page/1/">failure</a>
<a class="tag" href="/tag/inspirational/page/1/">inspirational</a> <a class="tag" href="/tag/inspirational/page/1/">inspirational</a>
<a class="tag" href="/tag/paraphrased/page/1/">paraphrased</a> <a class="tag" href="/tag/paraphrased/page/1/">paraphrased</a>
</div> </div>
</div> </div>
@ -182,10 +182,10 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="misattributed-eleanor-roosevelt" / > <meta class="keywords" itemprop="keywords" content="misattributed-eleanor-roosevelt" / >
<a class="tag" href="/tag/misattributed-eleanor-roosevelt/page/1/">misattributed-eleanor-roosevelt</a> <a class="tag" href="/tag/misattributed-eleanor-roosevelt/page/1/">misattributed-eleanor-roosevelt</a>
</div> </div>
</div> </div>
@ -196,73 +196,73 @@
</span> </span>
<div class="tags"> <div class="tags">
Tags: Tags:
<meta class="keywords" itemprop="keywords" content="humor,obvious,simile" / > <meta class="keywords" itemprop="keywords" content="humor,obvious,simile" / >
<a class="tag" href="/tag/humor/page/1/">humor</a> <a class="tag" href="/tag/humor/page/1/">humor</a>
<a class="tag" href="/tag/obvious/page/1/">obvious</a> <a class="tag" href="/tag/obvious/page/1/">obvious</a>
<a class="tag" href="/tag/simile/page/1/">simile</a> <a class="tag" href="/tag/simile/page/1/">simile</a>
</div> </div>
</div> </div>
<nav> <nav>
<ul class="pager"> <ul class="pager">
<li class="next"> <li class="next">
<a href="/page/2/">Next <span aria-hidden="true">&rarr;</span></a> <a href="/page/2/">Next <span aria-hidden="true">&rarr;</span></a>
</li> </li>
</ul> </ul>
</nav> </nav>
</div> </div>
<div class="col-md-4 tags-box"> <div class="col-md-4 tags-box">
<h2>Top Ten tags</h2> <h2>Top Ten tags</h2>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 28px" href="/tag/love/">love</a> <a class="tag" style="font-size: 28px" href="/tag/love/">love</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 26px" href="/tag/inspirational/">inspirational</a> <a class="tag" style="font-size: 26px" href="/tag/inspirational/">inspirational</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 26px" href="/tag/life/">life</a> <a class="tag" style="font-size: 26px" href="/tag/life/">life</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 24px" href="/tag/humor/">humor</a> <a class="tag" style="font-size: 24px" href="/tag/humor/">humor</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 22px" href="/tag/books/">books</a> <a class="tag" style="font-size: 22px" href="/tag/books/">books</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 14px" href="/tag/reading/">reading</a> <a class="tag" style="font-size: 14px" href="/tag/reading/">reading</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 10px" href="/tag/friendship/">friendship</a> <a class="tag" style="font-size: 10px" href="/tag/friendship/">friendship</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 8px" href="/tag/friends/">friends</a> <a class="tag" style="font-size: 8px" href="/tag/friends/">friends</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 8px" href="/tag/truth/">truth</a> <a class="tag" style="font-size: 8px" href="/tag/truth/">truth</a>
</span> </span>
<span class="tag-item"> <span class="tag-item">
<a class="tag" style="font-size: 6px" href="/tag/simile/">simile</a> <a class="tag" style="font-size: 6px" href="/tag/simile/">simile</a>
</span> </span>
</div> </div>
</div> </div>

View File

@ -410,7 +410,7 @@ How can I make a blank request?
------------------------------- -------------------------------
.. code-block:: python .. code-block:: python
from scrapy import Request from scrapy import Request

View File

@ -111,7 +111,7 @@ Once you've installed `Anaconda`_ or `Miniconda`_, install Scrapy with::
To install Scrapy on Windows using ``pip``: To install Scrapy on Windows using ``pip``:
.. warning:: .. warning::
This installation method requires “Microsoft Visual C++” for installing some This installation method requires “Microsoft Visual C++” for installing some
Scrapy dependencies, which demands significantly more disk space than Anaconda. Scrapy dependencies, which demands significantly more disk space than Anaconda.
#. Download and execute `Microsoft C++ Build Tools`_ to install the Visual Studio Installer. #. Download and execute `Microsoft C++ Build Tools`_ to install the Visual Studio Installer.
@ -123,7 +123,7 @@ To install Scrapy on Windows using ``pip``:
#. Check the installation details and make sure following packages are selected as optional components: #. Check the installation details and make sure following packages are selected as optional components:
* **MSVC** (e.g MSVC v142 - VS 2019 C++ x64/x86 build tools (v14.23) ) * **MSVC** (e.g MSVC v142 - VS 2019 C++ x64/x86 build tools (v14.23) )
* **Windows SDK** (e.g Windows 10 SDK (10.0.18362.0)) * **Windows SDK** (e.g Windows 10 SDK (10.0.18362.0))
#. Install the Visual Studio Build Tools. #. Install the Visual Studio Build Tools.

View File

@ -292,7 +292,7 @@ As an alternative, you could've written:
>>> response.css("title::text")[0].get() >>> response.css("title::text")[0].get()
'Quotes to Scrape' 'Quotes to Scrape'
Accessing an index on a :class:`~scrapy.selector.SelectorList` instance will Accessing an index on a :class:`~scrapy.selector.SelectorList` instance will
raise an :exc:`IndexError` exception if there are no results: raise an :exc:`IndexError` exception if there are no results:
.. code-block:: pycon .. code-block:: pycon
@ -302,8 +302,8 @@ raise an :exc:`IndexError` exception if there are no results:
... ...
IndexError: list index out of range IndexError: list index out of range
You might want to use ``.get()`` directly on the You might want to use ``.get()`` directly on the
:class:`~scrapy.selector.SelectorList` instance instead, which returns ``None`` :class:`~scrapy.selector.SelectorList` instance instead, which returns ``None``
if there are no results: if there are no results:
.. code-block:: pycon .. code-block:: pycon

View File

@ -934,10 +934,10 @@ Modified requirements
Backward-incompatible changes Backward-incompatible changes
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
- The value of the :setting:`FEED_STORE_EMPTY` setting is now ``True`` - The value of the :setting:`FEED_STORE_EMPTY` setting is now ``True``
instead of ``False``. In earlier Scrapy versions empty files were created instead of ``False``. In earlier Scrapy versions empty files were created
even when this setting was ``False`` (which was a bug that is now fixed), even when this setting was ``False`` (which was a bug that is now fixed),
so the new default should keep the old behavior. (:issue:`872`, so the new default should keep the old behavior. (:issue:`872`,
:issue:`5847`) :issue:`5847`)
Deprecation removals Deprecation removals

View File

@ -88,7 +88,7 @@ how you :ref:`configure the downloader middlewares
The execution engine, which coordinates the core crawling logic The execution engine, which coordinates the core crawling logic
between the scheduler, downloader and spiders. between the scheduler, downloader and spiders.
Some extension may want to access the Scrapy engine, to inspect or Some extension may want to access the Scrapy engine, to inspect or
modify the downloader and scheduler behaviour, although this is an modify the downloader and scheduler behaviour, although this is an
advanced use and this API is not yet stable. advanced use and this API is not yet stable.

View File

@ -87,8 +87,8 @@ of the system, and triggering events when certain actions occur. See the
Scheduler Scheduler
--------- ---------
The :ref:`scheduler <topics-scheduler>` receives requests from the engine and The :ref:`scheduler <topics-scheduler>` receives requests from the engine and
enqueues them for feeding them later (also to the engine) when the engine enqueues them for feeding them later (also to the engine) when the engine
requests them. requests them.
.. _component-downloader: .. _component-downloader:

View File

@ -224,7 +224,7 @@ BaseItemExporter
.. [1] Not all exporters respect the specified field order. .. [1] Not all exporters respect the specified field order.
.. [2] When using :ref:`item objects <item-types>` that do not expose .. [2] When using :ref:`item objects <item-types>` that do not expose
all their possible fields, exporters that do not support exporting all their possible fields, exporters that do not support exporting
a different subset of fields per item will only export the fields a different subset of fields per item will only export the fields
found in the first item exported. found in the first item exported.
.. attribute:: export_empty_fields .. attribute:: export_empty_fields

View File

@ -256,14 +256,14 @@ Spider state extension
Manages spider state data by loading it before a crawl and saving it after. Manages spider state data by loading it before a crawl and saving it after.
Give a value to the :setting:`JOBDIR` setting to enable this extension. Give a value to the :setting:`JOBDIR` setting to enable this extension.
When enabled, this extension manages the :attr:`~scrapy.Spider.state` When enabled, this extension manages the :attr:`~scrapy.Spider.state`
attribute of your :class:`~scrapy.Spider` instance: attribute of your :class:`~scrapy.Spider` instance:
- When your spider closes (:signal:`spider_closed`), the contents of its - When your spider closes (:signal:`spider_closed`), the contents of its
:attr:`~scrapy.Spider.state` attribute are serialized into a file named :attr:`~scrapy.Spider.state` attribute are serialized into a file named
``spider.state`` in the :setting:`JOBDIR` folder. ``spider.state`` in the :setting:`JOBDIR` folder.
- When your spider opens (:signal:`spider_opened`), if a previously-generated - When your spider opens (:signal:`spider_opened`), if a previously-generated
``spider.state`` file exists in the :setting:`JOBDIR` folder, it is loaded ``spider.state`` file exists in the :setting:`JOBDIR` folder, it is loaded
into the :attr:`~scrapy.Spider.state` attribute. into the :attr:`~scrapy.Spider.state` attribute.
@ -291,8 +291,8 @@ settings:
.. note:: .. note::
When a certain closing condition is met, requests which are When a certain closing condition is met, requests which are
currently in the downloader queue (up to :setting:`CONCURRENT_REQUESTS` currently in the downloader queue (up to :setting:`CONCURRENT_REQUESTS`
requests) are still processed. requests) are still processed.
.. setting:: CLOSESPIDER_TIMEOUT .. setting:: CLOSESPIDER_TIMEOUT

View File

@ -180,7 +180,7 @@ FTP supports two different connection modes: `active or passive
mode by default. To use the active connection mode instead, set the mode by default. To use the active connection mode instead, set the
:setting:`FEED_STORAGE_FTP_ACTIVE` setting to ``True``. :setting:`FEED_STORAGE_FTP_ACTIVE` setting to ``True``.
The default value for the ``overwrite`` key in the :setting:`FEEDS` for this The default value for the ``overwrite`` key in the :setting:`FEEDS` for this
storage backend is: ``True``. storage backend is: ``True``.
.. caution:: The value ``True`` in ``overwrite`` will cause you to lose the .. caution:: The value ``True`` in ``overwrite`` will cause you to lose the
@ -222,7 +222,7 @@ feeds using these settings:
- :setting:`AWS_ENDPOINT_URL` - :setting:`AWS_ENDPOINT_URL`
- :setting:`AWS_REGION_NAME` - :setting:`AWS_REGION_NAME`
The default value for the ``overwrite`` key in the :setting:`FEEDS` for this The default value for the ``overwrite`` key in the :setting:`FEEDS` for this
storage backend is: ``True``. storage backend is: ``True``.
.. caution:: The value ``True`` in ``overwrite`` will cause you to lose the .. caution:: The value ``True`` in ``overwrite`` will cause you to lose the
@ -255,7 +255,7 @@ You can set a *Project ID* and *Access Control List (ACL)* through the following
- :setting:`FEED_STORAGE_GCS_ACL` - :setting:`FEED_STORAGE_GCS_ACL`
- :setting:`GCS_PROJECT_ID` - :setting:`GCS_PROJECT_ID`
The default value for the ``overwrite`` key in the :setting:`FEEDS` for this The default value for the ``overwrite`` key in the :setting:`FEEDS` for this
storage backend is: ``True``. storage backend is: ``True``.
.. caution:: The value ``True`` in ``overwrite`` will cause you to lose the .. caution:: The value ``True`` in ``overwrite`` will cause you to lose the
@ -587,8 +587,8 @@ FEED_STORE_EMPTY
Default: ``True`` Default: ``True``
Whether to export empty feeds (i.e. feeds with no items). Whether to export empty feeds (i.e. feeds with no items).
If ``False``, and there are no items to export, no new files are created and If ``False``, and there are no items to export, no new files are created and
existing files are not modified, even if the :ref:`overwrite feed option existing files are not modified, even if the :ref:`overwrite feed option
<feed-options>` is enabled. <feed-options>` is enabled.
.. setting:: FEED_STORAGES .. setting:: FEED_STORAGES

View File

@ -266,9 +266,9 @@ e.g. in the spider's ``__init__`` method:
If you run this spider again then INFO messages from If you run this spider again then INFO messages from
``scrapy.spidermiddlewares.httperror`` logger will be gone. ``scrapy.spidermiddlewares.httperror`` logger will be gone.
You can also filter log records by :class:`~logging.LogRecord` data. For You can also filter log records by :class:`~logging.LogRecord` data. For
example, you can filter log records by message content using a substring or example, you can filter log records by message content using a substring or
a regular expression. Create a :class:`logging.Filter` subclass a regular expression. Create a :class:`logging.Filter` subclass
and equip it with a regular expression pattern to and equip it with a regular expression pattern to
filter out unwanted messages: filter out unwanted messages:
@ -284,8 +284,8 @@ filter out unwanted messages:
if match: if match:
return False return False
A project-level filter may be attached to the root A project-level filter may be attached to the root
handler created by Scrapy, this is a wieldy way to handler created by Scrapy, this is a wieldy way to
filter all loggers in different parts of the project filter all loggers in different parts of the project
(middlewares, spider, etc.): (middlewares, spider, etc.):
@ -301,7 +301,7 @@ filter all loggers in different parts of the project
for handler in logging.root.handlers: for handler in logging.root.handlers:
handler.addFilter(ContentFilter()) handler.addFilter(ContentFilter())
Alternatively, you may choose a specific logger Alternatively, you may choose a specific logger
and hide it without affecting other loggers: and hide it without affecting other loggers:
.. code-block:: python .. code-block:: python

View File

@ -414,7 +414,7 @@ class name. E.g. given pipeline class called MyPipeline you can set setting key:
and pipeline class MyPipeline will have expiration time set to 180. and pipeline class MyPipeline will have expiration time set to 180.
The last modified time from the file is used to determine the age of the file in days, The last modified time from the file is used to determine the age of the file in days,
which is then compared to the set expiration time to determine if the file is expired. which is then compared to the set expiration time to determine if the file is expired.
.. _topics-images-thumbnails: .. _topics-images-thumbnails:
@ -519,7 +519,7 @@ See here the methods that you can override in your custom Files Pipeline:
In addition to ``response``, this method receives the original In addition to ``response``, this method receives the original
:class:`request <scrapy.Request>`, :class:`request <scrapy.Request>`,
:class:`info <scrapy.pipelines.media.MediaPipeline.SpiderInfo>` and :class:`info <scrapy.pipelines.media.MediaPipeline.SpiderInfo>` and
:class:`item <scrapy.Item>` :class:`item <scrapy.Item>`
You can override this method to customize the download path of each file. You can override this method to customize the download path of each file.
@ -541,9 +541,9 @@ See here the methods that you can override in your custom Files Pipeline:
def file_path(self, request, response=None, info=None, *, item=None): def file_path(self, request, response=None, info=None, *, item=None):
return "files/" + PurePosixPath(urlparse_cached(request).path).name return "files/" + PurePosixPath(urlparse_cached(request).path).name
Similarly, you can use the ``item`` to determine the file path based on some item Similarly, you can use the ``item`` to determine the file path based on some item
property. property.
By default the :meth:`file_path` method returns By default the :meth:`file_path` method returns
``full/<request URL hash>.<extension>``. ``full/<request URL hash>.<extension>``.
@ -677,7 +677,7 @@ See here the methods that you can override in your custom Images Pipeline:
In addition to ``response``, this method receives the original In addition to ``response``, this method receives the original
:class:`request <scrapy.Request>`, :class:`request <scrapy.Request>`,
:class:`info <scrapy.pipelines.media.MediaPipeline.SpiderInfo>` and :class:`info <scrapy.pipelines.media.MediaPipeline.SpiderInfo>` and
:class:`item <scrapy.Item>` :class:`item <scrapy.Item>`
You can override this method to customize the download path of each file. You can override this method to customize the download path of each file.
@ -699,9 +699,9 @@ See here the methods that you can override in your custom Images Pipeline:
def file_path(self, request, response=None, info=None, *, item=None): def file_path(self, request, response=None, info=None, *, item=None):
return "files/" + PurePosixPath(urlparse_cached(request).path).name return "files/" + PurePosixPath(urlparse_cached(request).path).name
Similarly, you can use the ``item`` to determine the file path based on some item Similarly, you can use the ``item`` to determine the file path based on some item
property. property.
By default the :meth:`file_path` method returns By default the :meth:`file_path` method returns
``full/<request URL hash>.<extension>``. ``full/<request URL hash>.<extension>``.

View File

@ -309,7 +309,7 @@ Here are some tips to keep in mind when dealing with these kinds of sites:
services like `ProxyMesh`_. An open source alternative is `scrapoxy`_, a services like `ProxyMesh`_. An open source alternative is `scrapoxy`_, a
super proxy that you can attach your own proxies to. super proxy that you can attach your own proxies to.
* use a ban avoidance service, such as `Zyte API`_, which provides a `Scrapy * use a ban avoidance service, such as `Zyte API`_, which provides a `Scrapy
plugin <https://github.com/scrapy-plugins/scrapy-zyte-api>`__ and additional plugin <https://github.com/scrapy-plugins/scrapy-zyte-api>`__ and additional
features, like `AI web scraping <https://www.zyte.com/ai-web-scraping/>`__ features, like `AI web scraping <https://www.zyte.com/ai-web-scraping/>`__
If you are still unable to prevent your bot getting banned, consider contacting If you are still unable to prevent your bot getting banned, consider contacting

View File

@ -1309,7 +1309,7 @@ JsonResponse objects
.. class:: JsonResponse(url[, ...]) .. class:: JsonResponse(url[, ...])
The :class:`JsonResponse` class is a subclass of :class:`TextResponse` The :class:`JsonResponse` class is a subclass of :class:`TextResponse`
that is used when the response has a `JSON MIME type that is used when the response has a `JSON MIME type
<https://mimesniff.spec.whatwg.org/#json-mime-type>`_ in its `Content-Type` <https://mimesniff.spec.whatwg.org/#json-mime-type>`_ in its `Content-Type`
header. header.

View File

@ -559,7 +559,7 @@ For example, suppose you want to extract all ``<p>`` elements inside ``<div>``
elements. First, you would get all ``<div>`` elements: elements. First, you would get all ``<div>`` elements:
.. code-block:: pycon .. code-block:: pycon
>>> divs = response.xpath("//div") >>> divs = response.xpath("//div")
At first, you may be tempted to use the following approach, which is wrong, as At first, you may be tempted to use the following approach, which is wrong, as
@ -610,7 +610,7 @@ As it turns out, Scrapy selectors allow you to chain selectors, so most of the t
you can just select by class using CSS and then switch to XPath when needed: you can just select by class using CSS and then switch to XPath when needed:
.. code-block:: pycon .. code-block:: pycon
>>> from scrapy import Selector >>> from scrapy import Selector
>>> sel = Selector( >>> sel = Selector(
... text='<div class="hero shout"><time datetime="2014-07-23 19:00">Special date</time></div>' ... text='<div class="hero shout"><time datetime="2014-07-23 19:00">Special date</time></div>'
@ -1032,7 +1032,7 @@ whereas the CSS lookup is translated into XPath and thus runs more efficiently,
so performance-wise its uses are limited to situations that are not easily so performance-wise its uses are limited to situations that are not easily
described with CSS selectors. described with CSS selectors.
Parsel also simplifies adding your own XPath extensions with Parsel also simplifies adding your own XPath extensions with
:func:`~parsel.xpathfuncs.set_xpathfunc`. :func:`~parsel.xpathfuncs.set_xpathfunc`.
.. _topics-selectors-ref: .. _topics-selectors-ref:

View File

@ -379,8 +379,8 @@ The above example can also be written as follows:
def start_requests(self): def start_requests(self):
yield scrapy.Request(f"http://www.example.com/categories/{self.category}") yield scrapy.Request(f"http://www.example.com/categories/{self.category}")
If you are :ref:`running Scrapy from a script <run-from-script>`, you can If you are :ref:`running Scrapy from a script <run-from-script>`, you can
specify spider arguments when calling specify spider arguments when calling
:class:`CrawlerProcess.crawl <scrapy.crawler.CrawlerProcess.crawl>` or :class:`CrawlerProcess.crawl <scrapy.crawler.CrawlerProcess.crawl>` or
:class:`CrawlerRunner.crawl <scrapy.crawler.CrawlerRunner.crawl>`: :class:`CrawlerRunner.crawl <scrapy.crawler.CrawlerRunner.crawl>`:

View File

@ -86,7 +86,7 @@ Available Stats Collectors
Besides the basic :class:`StatsCollector` there are other Stats Collectors Besides the basic :class:`StatsCollector` there are other Stats Collectors
available in Scrapy which extend the basic Stats Collector. You can select available in Scrapy which extend the basic Stats Collector. You can select
which Stats Collector to use through the :setting:`STATS_CLASS` setting. The which Stats Collector to use through the :setting:`STATS_CLASS` setting. The
default Stats Collector used is the :class:`MemoryStatsCollector`. default Stats Collector used is the :class:`MemoryStatsCollector`.
.. currentmodule:: scrapy.statscollectors .. currentmodule:: scrapy.statscollectors

View File

@ -11,7 +11,7 @@ SEP-004: Library API
==================== ====================
.. note:: the library API has been implemented, but slightly different from .. note:: the library API has been implemented, but slightly different from
proposed in this SEP. You can run a Scrapy crawler inside a Twisted proposed in this SEP. You can run a Scrapy crawler inside a Twisted
reactor, but not outside it. reactor, but not outside it.
Introduction Introduction
============ ============

View File

@ -96,7 +96,7 @@ specified, else utf-8 is used) and returns a new unicode object. E.g:
``clean_spaces`` ``clean_spaces``
---------------- ----------------
Converts multispaces into single spaces for the given string. E.g: Converts multispaces into single spaces for the given string. E.g:
:: ::

View File

@ -73,8 +73,8 @@ Alternative Public API Proposal
- ``ItemLoader.get_stored_values()`` or ``ItemLoader.get_values()`` *(returns the ``ItemLoader values)* - ``ItemLoader.get_stored_values()`` or ``ItemLoader.get_values()`` *(returns the ``ItemLoader values)*
- ``ItemLoader.get_output_value()`` - ``ItemLoader.get_output_value()``
- ``ItemLoader.get_input_processor()`` or ``ItemLoader.get_in_processor()`` *(short version)* - ``ItemLoader.get_input_processor()`` or ``ItemLoader.get_in_processor()`` *(short version)*
- ``ItemLoader.get_output_processor()`` or ``ItemLoader.get_out_processor()`` *(short version)* - ``ItemLoader.get_output_processor()`` or ``ItemLoader.get_out_processor()`` *(short version)*
- ``ItemLoader.context`` - ``ItemLoader.context``

View File

@ -21,7 +21,7 @@ Current flaws and inconsistencies
2. Link extractors are inflexible and hard to maintain, link 2. Link extractors are inflexible and hard to maintain, link
processing/filtering is tightly coupled. (e.g. canonicalize) processing/filtering is tightly coupled. (e.g. canonicalize)
3. Isn't possible to crawl an url directly from command line because the Spider 3. Isn't possible to crawl an url directly from command line because the Spider
does not know which callback use. does not know which callback use.
These flaws will be corrected by the changes proposed in this SEP. These flaws will be corrected by the changes proposed in this SEP.
@ -55,7 +55,7 @@ Request Extractors
Request Extractors takes response object and determines which requests follow. Request Extractors takes response object and determines which requests follow.
This is an enhancement to ``LinkExtractors`` which returns urls (links), This is an enhancement to ``LinkExtractors`` which returns urls (links),
Request Extractors return Request objects. Request Extractors return Request objects.
Request Processors Request Processors
------------------ ------------------

View File

@ -200,7 +200,7 @@ the same spider:
# extract item from response # extract item from response
return item return item
The Spider Middleware that implements spider code The Spider Middleware that implements spider code
================================================= =================================================
There's gonna be one middleware that will take care of calling the proper There's gonna be one middleware that will take care of calling the proper
@ -625,7 +625,7 @@ Resolved:
not the original one (think of redirections), but it does carry the ``meta`` not the original one (think of redirections), but it does carry the ``meta``
of the original one. The original one may not be available anymore (in of the original one. The original one may not be available anymore (in
memory) if we're using a persistent scheduler., but in that case it would be memory) if we're using a persistent scheduler., but in that case it would be
the deserialized request from the persistent scheduler queue. the deserialized request from the persistent scheduler queue.
- No - this would make implementation more complex and we're not sure it's - No - this would make implementation more complex and we're not sure it's
really needed really needed