
一个用于绕过 Cloudflare 反机器人页面的 Python 模块。
由 Zied Boughdir 增强
所有功能均已测试,核心功能 100% 成功率:
这是一个用于绕过 Cloudflare 反机器人页面(也称为“I'm Under Attack Mode”,即 IUAM)的 Python 模块,基于 Requests 实现。该增强版支持 Cloudflare v2 挑战、代理轮换、隐身模式等功能。Cloudflare 会定期更改其技术,因此我会经常更新此仓库。
如果你希望抓取或爬取受 Cloudflare 保护的网站,这会非常有用。Cloudflare 的反机器人页面目前仅检查客户端是否支持 Javascript,不过未来他们可能会增加其他技术。
由于 Cloudflare 不断更改并强化其防护页面,cloudscraper 需要 JavaScript 引擎/解释器来解决 Javascript 挑战。这样脚本就能轻松模拟普通网络浏览器,而无需显式反混淆和解析 Cloudflare 的 Javascript。
作为参考,以下是 Cloudflare 用于此类页面的默认消息:``` Checking your browser before accessing website.com.
This process is automatic. Your browser will redirect to your requested content shortly.
Please allow up to 5 seconds...
任何使用 cloudscraper 的脚本在首次访问启用了 Cloudflare 反机器人机制的网站时都会休眠约 5 秒,不过首次请求之后不会再有任何延迟。
# 安装
只需运行 `pip install cloudscraper` 即可。PyPI 包位于 https://pypi.org/project/cloudscraper/```bash
pip install cloudscraper
或者,克隆此仓库并运行 python setup.py install。
如果您之前使用的是原版 cloudscraper 包,现在可以直接使用此增强版本:```python
import cloudscraper # Enhanced version
The API remains compatible, so you only need to change the import statements in your code. All function calls and parameters work the same way.
### Codebase Structure
The codebase has been streamlined to improve maintainability and reduce confusion:
- **Single Module**: All code is now in the `cloudscraper` module
- **Removed Redundancy**: The redundant directories have been removed
- **Updated Tests**: All test files have been updated to use the `cloudscraper` module
This makes the codebase cleaner and easier to maintain while ensuring backward compatibility with existing code that uses the original API.
## Key Features in cloudscraper
| Feature | Description | Status |
|---------|-------------|--------|
| **🆕 Executable Compatibility** | Complete fix for PyInstaller, cx_Freeze, auto-py-to-exe conversion | ✅ **FIXED** |
| **🆕 v3 JavaScript VM Challenges** | Support for Cloudflare's latest JavaScript VM-based challenges | ✅ **NEW** |
| **🆕 Turnstile Support** | Support for Cloudflare's new Turnstile CAPTCHA replacement | ✅ **NEW** |
| **Modern Challenge Support** | Enhanced support for v1, v2, v3, and Turnstile Cloudflare challenges | ✅ Complete |
| **Proxy Rotation** | Built-in smart proxy rotation with multiple strategies | ✅ Enhanced |
| **Stealth Mode** | Human-like behavior simulation to avoid detection | ✅ Enhanced |
| **Browser Emulation** | Advanced browser fingerprinting for Chrome and Firefox | ✅ Stable |
| **JavaScript Handling** | Better JS interpreter (js2py as default) for challenge solving | ✅ Enhanced |
| **Captcha Solvers** | Support for multiple CAPTCHA solving services | ✅ Stable |
# Dependencies
- **Python 3.8+** (Dropped support for Python 3.6 and 3.7)
- **[Requests](https://github.com/psf/requests)** >= 2.31.0
- **[requests_toolbelt](https://pypi.org/project/requests-toolbelt/)** >= 1.0.0
- **[pyparsing](https://pypi.org/project/pyparsing/)** >= 3.1.0
- **[pyOpenSSL](https://pypi.org/project/pyOpenSSL/)** >= 24.0.0
- **[pycryptodome](https://pypi.org/project/pycryptodome/)** >= 3.20.0
- **[websocket-client](https://pypi.org/project/websocket-client/)** >= 1.7.0
- **[js2py](https://pypi.org/project/Js2Py/)** >= 0.74
- **[brotli](https://pypi.org/project/Brotli/)** >= 1.1.0
- **[certifi](https://pypi.org/project/certifi/)** >= 2024.2.2
`python setup.py install` will install the Python dependencies automatically. The javascript interpreters and/or engines you decide to use are the only things you need to install yourself, excluding js2py which is part of the requirements as the default.
# Javascript Interpreters and Engines
We support the following Javascript interpreters/engines.
- **[ChakraCore](https://github.com/microsoft/ChakraCore):** Library binaries can also be located [here](https://www.github.com/VeNoMouS/cloudscraper/tree/ChakraCore/).
- **[js2py](https://github.com/PiotrDabkowski/Js2Py):** >=0.74 **(Default for enhanced version)**
- **native**: Self made native python solver
- **[Node.js](https://nodejs.org/)**
- **[V8](https://github.com/sony/v8eval/):** We use Sony's [v8eval](https://v8.dev)() python module.
# Usage
The simplest way to use cloudscraper is by calling `create_scraper()`.```python
import cloudscraper
scraper = cloudscraper.create_scraper() # returns a CloudScraper instance
# Or: scraper = cloudscraper.CloudScraper() # CloudScraper inherits from requests.Session
print(scraper.get("http://somesite.com").text) # => "<!DOCTYPE html><html><head>..."
就这样...
从此会话对象发往受 Cloudflare 反机器人保护网站的请求都将被自动处理。未使用 Cloudflare 的网站将按正常方式处理。你无需额外配置或调用任何内容,实际上可以将所有网站都视为不受任何保护。
你使用 cloudscraper 的方式与使用 Requests 完全相同。cloudScraper 的工作方式与 Requests Session 对象一致,区别仅在于不是调用 requests.get() 或 requests.post(),而是调用 scraper.get() 或 scraper.post()。
如需了解更多信息,请参阅 Requests 的文档。
使用 cloudscraper 将 Python 应用程序转换为可执行文件时遇到的用户代理问题已完全修复!
在将 Python 应用转换为可执行文件(使用 PyInstaller、cx_Freeze、auto-py-to-exe 等)时,由于 browsers.json 文件未被正确包含,用户会遇到与 用户代理 或 agent_user 功能相关的错误。
cloudscraper v2.7.0 包含一个自动回退系统:
选项 1:直接构建你的可执行文件(自动生效):```bash pyinstaller your_app.py
**选项 2:包含完整的用户代理数据库**(推荐):```bash
pyinstaller --add-data "cloudscraper/user_agent/browsers.json;cloudscraper/user_agent/" your_app.py
所有可执行兼容性均已经过全面测试:``` ✅ Normal operation with browsers.json ✅ Fallback operation without browsers.json ✅ PyInstaller environment simulation ✅ All browser/platform combinations ✅ HTTP requests with fallback user agents
您的 cloudscraper 应用程序在转换为可执行文件后将完美运行!🎉
## 🆕 Cloudflare v3 JavaScript VM 挑战支持
### 什么是 v3 挑战?
Cloudflare v3 挑战代表了机器人保护技术的最新演进。与传统 v1 和 v2 挑战不同,v3 挑战:
- **在 JavaScript 虚拟机中运行**:挑战在沙盒化的 JavaScript 环境中执行
- **使用高级检测**:更复杂的算法来检测自动化行为
- **生成动态代码**:挑战代码是动态创建的,更难进行逆向工程
- **提供现代保护**:来自 Cloudflare 的最新反机器人技术
### v3 基本用法```python
import cloudscraper
# v3 support is enabled by default
scraper = cloudscraper.create_scraper()
response = scraper.get("https://example.com")
print(response.text)
import cloudscraper