
CloudflareのアンチボットページをバイパスするPythonモジュール。
Zied Boughdir による拡張
すべての機能はコア機能で 100% の成功率 でテスト済み:
Requests を使用して実装された、Cloudflare のアンチボットページ (「I'm Under Attack Mode」または IUAM とも呼ばれる) をバイパスする Python モジュール。この拡張版には、Cloudflare v2 チャレンジ、プロキシローテーション、ステルスモードなどのサポートが含まれています。Cloudflare は技術を定期的に変更するため、このリポジトリは頻繁に更新されます。
これは、Cloudflare で保護されたウェブサイトをスクレイピングまたはクロールしたい場合に役立ちます。Cloudflare のアンチボットページは現在、クライアントが JavaScript をサポートしているかどうかのみをチェックしますが、将来的には追加の技術が導入される可能性があります。
Cloudflare は保護ページを継続的に変更・強化しているため、cloudscraper は JavaScript チャレンジを解決するために JavaScript エンジン/インタプリタを必要とします。これにより、Cloudflare の JavaScript を明示的に難読化解除して解析することなく、スクリプトが通常のウェブブラウザを簡単に偽装できるようになります。
参考までに、これは Cloudflare がこの種のページで使用するデフォルトのメッセージです:``` Checking your browser before accessing website.com.
This process is automatic. Your browser will redirect to your requested content shortly.
Please allow up to 5 seconds...
Any script using cloudscraper will sleep for ~5 seconds for the first visit to any site with Cloudflare anti-bots enabled, though no delay will occur after the first request.
cloudscraper を使用するスクリプトは、Cloudflare のアンチボットが有効なサイトへの最初の訪問時に約 5 秒間スリープしますが、最初のリクエスト以降は遅延は発生しません。
# インストール
`pip install cloudscraper` を実行するだけです。PyPI パッケージは https://pypi.org/project/cloudscraper/ にあります。```bash
pip install cloudscraper
または、このリポジトリをクローンして python setup.py install を実行してください。
以前に元の cloudscraper パッケージを使用していた場合は、この拡張版を直接使用できるようになりました:```python
import cloudscraper # Enhanced version
The API remains compatible, so you only need to change the import statements in your code. All function calls and parameters work the same way.
### Codebase Structure
The codebase has been streamlined to improve maintainability and reduce confusion:
- **Single Module**: All code is now in the `cloudscraper` module
- **Removed Redundancy**: The redundant directories have been removed
- **Updated Tests**: All test files have been updated to use the `cloudscraper` module
This makes the codebase cleaner and easier to maintain while ensuring backward compatibility with existing code that uses the original API.
## Key Features in cloudscraper
| Feature | Description | Status |
|---------|-------------|--------|
| **🆕 Executable Compatibility** | Complete fix for PyInstaller, cx_Freeze, auto-py-to-exe conversion | ✅ **FIXED** |
| **🆕 v3 JavaScript VM Challenges** | Support for Cloudflare's latest JavaScript VM-based challenges | ✅ **NEW** |
| **🆕 Turnstile Support** | Support for Cloudflare's new Turnstile CAPTCHA replacement | ✅ **NEW** |
| **Modern Challenge Support** | Enhanced support for v1, v2, v3, and Turnstile Cloudflare challenges | ✅ Complete |
| **Proxy Rotation** | Built-in smart proxy rotation with multiple strategies | ✅ Enhanced |
| **Stealth Mode** | Human-like behavior simulation to avoid detection | ✅ Enhanced |
| **Browser Emulation** | Advanced browser fingerprinting for Chrome and Firefox | ✅ Stable |
| **JavaScript Handling** | Better JS interpreter (js2py as default) for challenge solving | ✅ Enhanced |
| **Captcha Solvers** | Support for multiple CAPTCHA solving services | ✅ Stable |
# Dependencies
- **Python 3.8+** (Dropped support for Python 3.6 and 3.7)
- **[Requests](https://github.com/psf/requests)** >= 2.31.0
- **[requests_toolbelt](https://pypi.org/project/requests-toolbelt/)** >= 1.0.0
- **[pyparsing](https://pypi.org/project/pyparsing/)** >= 3.1.0
- **[pyOpenSSL](https://pypi.org/project/pyOpenSSL/)** >= 24.0.0
- **[pycryptodome](https://pypi.org/project/pycryptodome/)** >= 3.20.0
- **[websocket-client](https://pypi.org/project/websocket-client/)** >= 1.7.0
- **[js2py](https://pypi.org/project/Js2Py/)** >= 0.74
- **[brotli](https://pypi.org/project/Brotli/)** >= 1.1.0
- **[certifi](https://pypi.org/project/certifi/)** >= 2024.2.2
`python setup.py install` will install the Python dependencies automatically. The javascript interpreters and/or engines you decide to use are the only things you need to install yourself, excluding js2py which is part of the requirements as the default.
# Javascript Interpreters and Engines
We support the following Javascript interpreters/engines.
- **[ChakraCore](https://github.com/microsoft/ChakraCore):** Library binaries can also be located [here](https://www.github.com/VeNoMouS/cloudscraper/tree/ChakraCore/).
- **[js2py](https://github.com/PiotrDabkowski/Js2Py):** >=0.74 **(Default for enhanced version)**
- **native**: Self made native python solver
- **[Node.js](https://nodejs.org/)**
- **[V8](https://github.com/sony/v8eval/):** We use Sony's [v8eval](https://v8.dev)() python module.
# Usage
The simplest way to use cloudscraper is by calling `create_scraper()`.```python
import cloudscraper
scraper = cloudscraper.create_scraper() # returns a CloudScraper instance
# Or: scraper = cloudscraper.CloudScraper() # CloudScraper inherits from requests.Session
print(scraper.get("http://somesite.com").text) # => "<!DOCTYPE html><html><head>..."
それだけです...
Cloudflareのアンチボットによって保護されているWebサイトに対して、このセッションオブジェクトから行われるリクエストは自動的に処理されます。Cloudflareを使用していないWebサイトは通常どおりに扱われます。追加で設定や呼び出しを行う必要はなく、すべてのWebサイトをあたかも何の保護も受けていないかのように扱うことができます。
cloudscraperはRequestsとまったく同じように使用します。cloudScraperはRequestsのSessionオブジェクトと同様に動作します。requests.get()やrequests.post()を呼び出す代わりに、scraper.get()やscraper.post()を呼び出すだけです。
詳細については、Requestsのドキュメントを参照してください。
cloudscraperを使用するPythonアプリケーションを実行可能ファイルに変換する際のユーザーエージェントの問題は完全に修正されました!
Pythonアプリを実行可能ファイル(PyInstaller、cx_Freeze、auto-py-to-exeなど)に変換する際、browsers.jsonファイルが正しく含まれていなかったため、ユーザーエージェントまたはagent_user機能に関連するエラーが発生していました。
cloudscraper v2.7.0には自動フォールバックシステムが含まれています: