
Selenium Wire 已不再维护。感谢您的支持和所有贡献。
Selenium Wire 扩展了 Selenium 的 <https://www.selenium.dev/documentation/en/>_ Python 绑定,让您可以访问浏览器发出的底层请求。您可以像使用 Selenium 一样编写代码,但还能获得额外的 API,用于检查请求和响应,并即时对其进行修改。
.. image:: https://github.com/wkeeling/selenium-wire/workflows/build/badge.svg :target: https://github.com/wkeeling/selenium-wire/actions
.. image:: https://codecov.io/gh/wkeeling/selenium-wire/branch/master/graph/badge.svg :target: https://codecov.io/gh/wkeeling/selenium-wire
.. image:: https://img.shields.io/badge/python-3.7%2C%203.8%2C%203.9%2C%203.10-blue.svg :target: https://pypi.python.org/pypi/selenium-wire
.. image:: https://img.shields.io/pypi/v/selenium-wire.svg :target: https://pypi.python.org/pypi/selenium-wire
.. image:: https://img.shields.io/pypi/l/selenium-wire.svg :target: https://pypi.python.org/pypi/selenium-wire
.. image:: https://pepy.tech/badge/selenium-wire/month :target: https://pepy.tech/project/selenium-wire
简单示例~~~~~~~~~~~~~~
.. code:: python
from seleniumwire import webdriver # Import from seleniumwire
# Create a new instance of the Chrome driver
driver = webdriver.Chrome()
# Go to the Google home page
driver.get('https://www.google.com')
# Access requests via the `requests` attribute
for request in driver.requests:
if request.response:
print(
request.url,
request.response.status_code,
request.response.headers['Content-Type']
)
Prints:
.. code:: bash
https://www.google.com/ 200 text/html; charset=UTF-8
https://www.google.com/images/branding/googlelogo/2x/googlelogo_color_120x44dp.png 200 image/png
https://consent.google.com/status?continue=https://www.google.com&pc=s×tamp=1531511954&gl=GB 204 text/html; charset=utf-8
https://www.google.com/images/branding/googlelogo/2x/googlelogo_color_272x92dp.png 200 image/png
https://ssl.gstatic.com/gb/images/i2_2ec824b0.png 200 image/png
https://www.google.com/gen_204?s=webaft&t=aft&atyp=csi&ei=kgRJW7DBONKTlwTK77wQ&rt=wsrt.366,aft.58,prt.58 204 text/html; charset=UTF-8
...
Features
* Pure Python, user-friendly API
* HTTP and HTTPS requests captured
* Intercept requests and responses
* Modify headers, parameters, body content on the fly
* Capture websocket messages
* HAR format supported
* Proxy server support
Compatibilty
Table of Contents
- `安装`_
* `浏览器设置`_
* `OpenSSL`_
- `创建 Webdriver`_
- `访问请求`_
- `请求对象`_
- `响应对象`_
- `拦截请求和响应`_
* `示例:添加请求头`_
* `示例:替换现有请求头`_
* `示例:添加响应头`_
* `示例:添加请求参数`_
* `示例:更新 POST 请求体中的 JSON`_
* `示例:基本认证`_
* `示例:阻止请求`_
* `示例:模拟响应`_
* `取消设置拦截器`_
- `限制请求捕获`_
- `请求存储`_
* `内存存储`_
- `代理`_
* `SOCKS`_
* `动态切换`_
- `机器人检测`_
- `证书`_
* `使用您自己的证书`_
- `所有选项`_
- `许可证`_
安装~~~~~~~~~~~~
Install using pip:
.. code:: bash
pip install selenium-wire
If you get an error about not being able to build cryptography you may be running an old version of pip. Try upgrading pip with ``python -m pip install --upgrade pip`` and then re-run the above command.
Browser Setup
-------------
No specific configuration should be necessary except to ensure that you have downloaded the relevent webdriver executable for your browser and placed it somewhere on your system PATH.
- `Download <https://sites.google.com/chromium.org/driver/>`__ webdriver for Chrome
- `Download <https://github.com/mozilla/geckodriver/>`__ webdriver for Firefox
- `Download <https://developer.microsoft.com/en-us/microsoft-edge/tools/webdriver/>`__ webdriver for Edge
OpenSSL
-------
Selenium Wire requires OpenSSL for decrypting HTTPS requests. This is probably already installed on your system (you can check by running ``openssl version`` on the command line). If it's not installed you can install it with:
**Linux**
.. code:: bash
# For apt based Linux systems
sudo apt install openssl
# For RPM based Linux systems
sudo yum install openssl
# For Linux alpine
sudo apk add openssl
**MacOS**
.. code:: bash
brew install openssl
**Windows**
No installation is required.
Creating the Webdriver
确保从 seleniumwire 包导入 webdriver:
.. code:: python
from seleniumwire import webdriver
然后,就像直接使用 Selenium 一样实例化 webdriver。您可以传入任何所需的功能或浏览器特定选项——例如可执行文件路径、无头模式等。Selenium Wire 也有其 自己的选项_,可以通过 seleniumwire_options 属性传入。
.. code:: python
# Create the driver with no options (use defaults)
driver = webdriver.Chrome()
# Or create using browser specific options and/or seleniumwire_options options
driver = webdriver.Chrome(
options = webdriver.ChromeOptions(...),
seleniumwire_options={...}
)
.. _own options: #all-options
请注意,对于 webdriver 的子包,您仍应直接从 selenium 导入它们。例如,要导入 WebDriverWait:
.. code:: python
# Sub-packages of webdriver must still be imported from `selenium` itself
from selenium.webdriver.support.ui import WebDriverWait
远程 Webdriver
Selenium Wire 对远程 webdriver 客户端的支持有限。当您创建远程 webdriver 的实例时,需要指定运行 Selenium Wire 的机器(或容器)的主机名或 IP 地址。这使远程实例能够将其请求和响应传回 Selenium Wire。
.. code:: python
options = {
'addr': 'hostname_or_ip' # Address of the machine running Selenium Wire. Explicitly use 127.0.0.1 rather than localhost if remote session is running locally.
}
driver = webdriver.Remote(
command_executor='http://www.example.com',
seleniumwire_options=options
)
如果运行浏览器的机器需要使用不同的地址与运行 Selenium Wire 的机器通信,则需要手动配置浏览器。此问题 <https://github.com/wkeeling/selenium-wire/issues/220>_ 提供了更多详细信息。
访问请求~~~~~~~~~~~~~~~~~~
Selenium Wire captures all HTTP/HTTPS traffic made by the browser [1]_. The following attributes provide access to requests and responses.
driver.requests
The list of captured requests in chronological order.
driver.last_request
Convenience attribute for retrieving the most recently captured request. This is more efficient than using driver.requests[-1].
driver.wait_for_request(pat, timeout=10)
This method will wait until it sees a request matching a pattern. The pat attribute will be matched within the request URL. pat can be a simple substring or a regular expression. Note that driver.wait_for_request() doesn't make a request, it just waits for a previous request made by some other action and it will return the first request it finds. Also note that since pat can be a regular expression, you must escape special characters such as question marks with a slash. A TimeoutException is raised if no match is found within the timeout period.
For example, to wait for an AJAX request to return after a button is clicked:
.. code:: python
# Click a button that triggers a background request to https://server/api/products/12345/
button_element.click()
# Wait for the request/response to complete
request = driver.wait_for_request('/api/products/12345/')
driver.har
A JSON formatted HAR archive of HTTP transactions that have taken place. HAR capture is turned off by default and you must set the enable_har option_ to True before using driver.har.
driver.iter_requests()
Returns an iterator over captured requests. Useful when dealing with a large number of requests.
driver.request_interceptor
Used to set a request interceptor. See Intercepting Requests and Responses_.
driver.response_interceptor
Used to set a response interceptor.
Clearing Requests
To clear previously captured requests and HAR entries, use del:
.. code:: python
del driver.requests