Search code, repositories, users, issues, pull requests...

Values from the Referrer-Policy header of HTTP responses are no longer executed as Python callables. See the cwxj-rr6w-m6w7 security advisory for details.
In line with the standard, 301 redirects of POST requests are converted into GET requests.

Full Changelog

`v2.14.1`

Deprecate maybeDeferred_coro()
Pass the spider arg to custom stat collectors {open,close}_spider()
Replace deprecated Codecov CI action

Full Changelog

`v2.14.0`

More coroutine-based replacements for Deferred-based APIs
The default priority queue is now DownloaderAwarePriorityQueue
Dropped support for Python 3.9 and PyPy 3.10
Improved and documented the API for custom download handlers

Full changelog

`v2.13.4`

Fix for the CVE-2025-6176 security issue: improved protection against decompression bombs in HttpCompressionMiddleware for responses compressed using the br and deflate methods. Requires brotli >= 1.2.0.

Full changelog

`v2.13.3`

Changed the values for DOWNLOAD_DELAY (from 0 to 1) and CONCURRENT_REQUESTS_PER_DOMAIN (from 8 to 1) in the default project template.
Fixed several bugs in the engine initialization and exception handling logic.
Allowed running tests with Twisted 25.5.0+ again and fixed test failures with lxml 6.0.0.

`v2.13.2`

Fixed a bug introduced in Scrapy 2.13.0 that caused results of request errbacks to be ignored when the errback was called because of a downloader error.
Docs and error messages improvements related to the Scrapy 2.13.0 default reactor change.

`v2.13.1`

Give callback requests precedence over start requests when priority values are the same.

`v2.13.0`

The asyncio reactor is now enabled by default
Replaced start_requests() (sync) with start() (async) and changed how it is iterated.
Added the allow_offsite request meta key
Spider middlewares that don't support asynchronous spider output are deprecated
Added a base class for universal spider middlewares

`v2.12.0`

Dropped support for Python 3.8, added support for Python 3.13
start_requests can now yield items
Added scrapy.http.JsonResponse
Added the CLOSESPIDER_PAGECOUNT_NO_ITEM setting

`v2.11.2`

Mostly bug fixes, including security bug fixes.

`v2.11.1`

Security bug fixes.
Support for Twisted >= 23.8.0.
Documentation improvements.

`v2.11.0`

Spiders can now modify settings in their from_crawler methods, e.g. based on spider arguments.
Periodic logging of stats.
Bug fixes.

`v2.10.1`

Marked Twisted >= 23.8.0 as unsupported.

`v2.10.0`

Added Python 3.12 support, dropped Python 3.7 support.
The new add-ons framework simplifies configuring 3rd-party components that support it.
Exceptions to retry can now be configured.
Many fixes and improvements for feed exports.

`v2.9.0`

Per-domain download settings.
Compatibility with new cryptography and new parsel.
JMESPath selectors from the new parsel.
Bug fixes.

`v2.8.0`

This is a maintenance release, with minor features, bug fixes, and cleanups.

`v2.7.1`

Relaxed the restriction introduced in 2.6.2 so that the Proxy-Authentication header can again be set explicitly in certain cases, restoring compatibility with scrapy-zyte-smartproxy 2.1.0 and older
Bug fixes

`v2.7.0`

Added Python 3.11 support, dropped Python 3.6 support
Improved support for asynchronous callbacks
Asyncio support is enabled by default on new projects
Output names of item fields can now be arbitrary strings
Centralized request fingerprinting configuration is now possible

`v2.6.3`

Makes pip install Scrapy work again.

It required making changes to support pyOpenSSL 22.1.0. We had to drop support for SSLv3 as a result.

We also upgraded the minimum versions of some dependencies.

See the changelog.

`v2.6.2`

Fixes a security issue around HTTP proxy usage, and addresses a few regressions introduced in Scrapy 2.6.0.

See the changelog.

`v2.6.1`

Fixes a regression introduced in 2.6.0 that would unset the request method when following redirects.

`v2.6.0`

Security fixes for cookie handling (see details below)
Python 3.10 support
asyncio support is no longer considered experimental, and works out-of-the-box on Windows regardless of your Python version
Feed exports now support pathlib.Path output paths and per-feed item filtering and post-processing

Security bug fixes

When a Request object with cookies defined gets a redirect response causing a new Request object to be scheduled, the cookies defined in the original Request object are no longer copied into the new Request object.

If you manually set the Cookie header on a Request object and the domain name of the redirect URL is not an exact match for the domain of the URL of the original Request object, your Cookie header is now dropped from the new Request object.

The old behavior could be exploited by an attacker to gain access to your cookies. Please, see the cjvr-mfj7-j4j8 security advisory for more
information.

Note: It is still possible to enable the sharing of cookies between different domains with a shared domain suffix (e.g. example.com and any subdomain) by defining the shared domain suffix (e.g. example.com) as the cookie domain when defining your cookies. See the documentation of the Request class for more information.
When the domain of a cookie, either received in the Set-Cookie header of a response or defined in a Request object, is set to a public suffix <https://publicsuffix.org/>_, the cookie is now ignored unless the cookie domain is the same as the request domain.

The old behavior could be exploited by an attacker to inject cookies from a controlled domain into your cookiejar that could be sent to other domains not controlled by the attacker. Please, see the mfjm-vh54-3f96 security advisory for more information.

`v2.5.1`

Security bug fix:

If you use HttpAuthMiddleware (i.e. the http_user and http_pass spider attributes) for HTTP authentication, any request exposes your credentials to the request target.

To prevent unintended exposure of authentication credentials to unintended domains, you must now additionally set a new, additional spider attribute, http_auth_domain, and point it to the specific domain to which the authentication credentials must be sent.

If the http_auth_domain spider attribute is not set, the domain of the first request will be considered the HTTP authentication target, and authentication credentials will only be sent in requests targeting that domain.

If you need to send the same HTTP authentication credentials to multiple domains, you can use w3lib.http.basic_auth_header instead to set the value of the Authorization header of your requests.

If you really want your spider to send the same HTTP authentication credentials to any domain, set the http_auth_domain spider attribute to None.

Finally, if you are a user of scrapy-splash, know that this version of Scrapy breaks compatibility with scrapy-splash 0.7.2 and earlier. You will need to upgrade scrapy-splash to a greater version for it to continue to work.

`v2.5.0`

Official Python 3.9 support
Experimental HTTP/2 support
New get_retry_request() function to retry requests from spider callbacks
New headers_received signal that allows stopping downloads early
New Response.protocol attribute

`v2.4.1`