mirror of https://github.com/scrapy/scrapy.git
Clarify spider middleware handling of callback requests
Make explicit that process_spider_output() handles both items and requests from callbacks before they reach pipelines or scheduler.
This commit is contained in:
parent
4e175e7930
commit
71df815fcd
|
|
@ -229,14 +229,16 @@ Spider middlewares process data at two points in the lifecycle:
|
|||
|
||||
1. **Start processing**: Via ``process_start()`` before initial requests reach
|
||||
the scheduler (see :ref:`lifecycle-start-processing`)
|
||||
2. **Callback processing**: Via ``process_spider_input()`` and
|
||||
``process_spider_output()`` for responses and callback output
|
||||
2. **Callback processing**: Via ``process_spider_input()`` before the callback
|
||||
runs, and ``process_spider_output()`` after—processing both items and
|
||||
new requests before they reach pipelines or the scheduler
|
||||
|
||||
Common uses for callback processing include:
|
||||
|
||||
- Filtering responses (e.g., by HTTP status code or content type)
|
||||
- Handling spider exceptions
|
||||
- Modifying the items or requests yielded by callbacks
|
||||
- Modifying or filtering items before they reach pipelines
|
||||
- Modifying or filtering requests before they return to the scheduler
|
||||
|
||||
For details on writing spider middlewares, see :ref:`topics-spider-middleware`.
|
||||
|
||||
|
|
|
|||
Loading…
Reference in New Issue