Selenium Wire 已不再维护。感谢您的支持和所有贡献。
Selenium Wire 扩展了 Selenium 的 <https://www.selenium.dev/documentation/en/>_ Python 绑定,让您可以访问浏览器发出的底层请求。您可以像使用 Selenium 一样编写代码,但还能获得额外的 API,用于检查请求和响应,并即时对其进行修改。
.. image:: https://github.com/wkeeling/selenium-wire/workflows/build/badge.svg :target: https://github.com/wkeeling/selenium-wire/actions
.. image:: https://codecov.io/gh/wkeeling/selenium-wire/branch/master/graph/badge.svg :target: https://codecov.io/gh/wkeeling/selenium-wire
.. image:: https://img.shields.io/badge/python-3.7%2C%203.8%2C%203.9%2C%203.10-blue.svg :target: https://pypi.python.org/pypi/selenium-wire
.. image:: https://img.shields.io/pypi/v/selenium-wire.svg :target: https://pypi.python.org/pypi/selenium-wire
.. image:: https://img.shields.io/pypi/l/selenium-wire.svg :target: https://pypi.python.org/pypi/selenium-wire
.. image:: https://pepy.tech/badge/selenium-wire/month :target: https://pepy.tech/project/selenium-wire
简单示例~~~~~~~~~~~~~~
.. code:: python
from seleniumwire import webdriver # Import from seleniumwire
# Create a new instance of the Chrome driver
driver = webdriver.Chrome()
# Go to the Google home page
driver.get('https://www.google.com')
# Access requests via the `requests` attribute
for request in driver.requests:
if request.response:
print(
request.url,
request.response.status_code,
request.response.headers['Content-Type']
)
Prints:
.. code:: bash
https://www.google.com/ 200 text/html; charset=UTF-8
https://www.google.com/images/branding/googlelogo/2x/googlelogo_color_120x44dp.png 200 image/png
https://consent.google.com/status?continue=https://www.google.com&pc=s×tamp=1531511954&gl=GB 204 text/html; charset=utf-8
https://www.google.com/images/branding/googlelogo/2x/googlelogo_color_272x92dp.png 200 image/png
https://ssl.gstatic.com/gb/images/i2_2ec824b0.png 200 image/png
https://www.google.com/gen_204?s=webaft&t=aft&atyp=csi&ei=kgRJW7DBONKTlwTK77wQ&rt=wsrt.366,aft.58,prt.58 204 text/html; charset=UTF-8
...
Features
* Pure Python, user-friendly API
* HTTP and HTTPS requests captured
* Intercept requests and responses
* Modify headers, parameters, body content on the fly
* Capture websocket messages
* HAR format supported
* Proxy server support
Compatibilty
Table of Contents
- `安装`_
* `浏览器设置`_
* `OpenSSL`_
- `创建 Webdriver`_
- `访问请求`_
- `请求对象`_
- `响应对象`_
- `拦截请求和响应`_
* `示例:添加请求头`_
* `示例:替换现有请求头`_
* `示例:添加响应头`_
* `示例:添加请求参数`_
* `示例:更新 POST 请求体中的 JSON`_
* `示例:基本认证`_
* `示例:阻止请求`_
* `示例:模拟响应`_
* `取消设置拦截器`_
- `限制请求捕获`_
- `请求存储`_
* `内存存储`_
- `代理`_
* `SOCKS`_
* `动态切换`_
- `机器人检测`_
- `证书`_
* `使用您自己的证书`_
- `所有选项`_
- `许可证`_
安装~~~~~~~~~~~~
Install using pip:
.. code:: bash
pip install selenium-wire
If you get an error about not being able to build cryptography you may be running an old version of pip. Try upgrading pip with ``python -m pip install --upgrade pip`` and then re-run the above command.
Browser Setup
-------------
No specific configuration should be necessary except to ensure that you have downloaded the relevent webdriver executable for your browser and placed it somewhere on your system PATH.
- `Download <https://sites.google.com/chromium.org/driver/>`__ webdriver for Chrome
- `Download <https://github.com/mozilla/geckodriver/>`__ webdriver for Firefox
- `Download <https://developer.microsoft.com/en-us/microsoft-edge/tools/webdriver/>`__ webdriver for Edge
OpenSSL
-------
Selenium Wire requires OpenSSL for decrypting HTTPS requests. This is probably already installed on your system (you can check by running ``openssl version`` on the command line). If it's not installed you can install it with:
**Linux**
.. code:: bash
# For apt based Linux systems
sudo apt install openssl
# For RPM based Linux systems
sudo yum install openssl
# For Linux alpine
sudo apk add openssl
**MacOS**
.. code:: bash
brew install openssl
**Windows**
No installation is required.
Creating the Webdriver
确保从 seleniumwire 包导入 webdriver:
.. code:: python
from seleniumwire import webdriver
然后,就像直接使用 Selenium 一样实例化 webdriver。您可以传入任何所需的功能或浏览器特定选项——例如可执行文件路径、无头模式等。Selenium Wire 也有其 自己的选项_,可以通过 seleniumwire_options 属性传入。
.. code:: python
# Create the driver with no options (use defaults)
driver = webdriver.Chrome()
# Or create using browser specific options and/or seleniumwire_options options
driver = webdriver.Chrome(
options = webdriver.ChromeOptions(...),
seleniumwire_options={...}
)
.. _own options: #all-options
请注意,对于 webdriver 的子包,您仍应直接从 selenium 导入它们。例如,要导入 WebDriverWait:
.. code:: python
# Sub-packages of webdriver must still be imported from `selenium` itself
from selenium.webdriver.support.ui import WebDriverWait
远程 Webdriver
Selenium Wire 对远程 webdriver 客户端的支持有限。当您创建远程 webdriver 的实例时,需要指定运行 Selenium Wire 的机器(或容器)的主机名或 IP 地址。这使远程实例能够将其请求和响应传回 Selenium Wire。
.. code:: python
options = {
'addr': 'hostname_or_ip' # Address of the machine running Selenium Wire. Explicitly use 127.0.0.1 rather than localhost if remote session is running locally.
}
driver = webdriver.Remote(
command_executor='http://www.example.com',
seleniumwire_options=options
)
如果运行浏览器的机器需要使用不同的地址与运行 Selenium Wire 的机器通信,则需要手动配置浏览器。此问题 <https://github.com/wkeeling/selenium-wire/issues/220>_ 提供了更多详细信息。
访问请求~~~~~~~~~~~~~~~~~~
Selenium Wire captures all HTTP/HTTPS traffic made by the browser [1]_. The following attributes provide access to requests and responses.
driver.requests
The list of captured requests in chronological order.
driver.last_request
Convenience attribute for retrieving the most recently captured request. This is more efficient than using driver.requests[-1].
driver.wait_for_request(pat, timeout=10)
This method will wait until it sees a request matching a pattern. The pat attribute will be matched within the request URL. pat can be a simple substring or a regular expression. Note that driver.wait_for_request() doesn't make a request, it just waits for a previous request made by some other action and it will return the first request it finds. Also note that since pat can be a regular expression, you must escape special characters such as question marks with a slash. A TimeoutException is raised if no match is found within the timeout period.
For example, to wait for an AJAX request to return after a button is clicked:
.. code:: python
# Click a button that triggers a background request to https://server/api/products/12345/
button_element.click()
# Wait for the request/response to complete
request = driver.wait_for_request('/api/products/12345/')
driver.har
A JSON formatted HAR archive of HTTP transactions that have taken place. HAR capture is turned off by default and you must set the enable_har option_ to True before using driver.har.
driver.iter_requests()
Returns an iterator over captured requests. Useful when dealing with a large number of requests.
driver.request_interceptor
Used to set a request interceptor. See Intercepting Requests and Responses_.
driver.response_interceptor
Used to set a response interceptor.
Clearing Requests
To clear previously captured requests and HAR entries, use del:
.. code:: python
del driver.requests
.. [1] Selenium Wire ignores OPTIONS requests by default, as these are typically uninteresting and just add overhead. If you want to capture OPTIONS requests, you need to set the ignore_http_methods option_ to [].
.. _option: #all-options
Request Objects
Request objects have the following attributes.
``body``
The request body as ``bytes``. If the request has no body the value of ``body`` will be empty, i.e. ``b''``.
``cert``
Information about the server SSL certificate in dictionary format. Empty for non-HTTPS requests.
``date``
The datetime the request was made.
``headers``
A dictionary-like object of request headers. Headers are case-insensitive and duplicates are permitted. Asking for ``request.headers['user-agent']`` will return the value of the ``User-Agent`` header. If you wish to replace a header, make sure you delete the existing header first with ``del request.headers['header-name']``, otherwise you'll create a duplicate.
``host``
The request host, e.g. ``www.example.com``
``method``
The HTTP method, e.g. ``GET`` or ``POST`` etc.
``params``
A dictionary of request parameters. If a parameter with the same name appears more than once in the request, it's value in the dictionary will be a list.
``path``
The request path, e.g. ``/some/path/index.html``
``querystring``
The query string, e.g. ``foo=bar&spam=eggs``
``response``
The `response object`_ associated with the request. This will be ``None`` if the request has no response.
``url``
The request URL, e.g. ``https://www.example.com/some/path/index.html?foo=bar&spam=eggs``
``ws_messages``
Where the request is a websocket handshake request (normally with a URL starting ``wss://``) then ``ws_messages`` will contain a list of any websocket messages sent and received. See `WebSocketMessage Objects`_.
Request objects have the following methods.
``abort(error_code=403)``
Trigger immediate termination of the request with the supplied error code. For use within request interceptors. See `Example: Block a request`_.
``create_response(status_code, headers=(), body=b'')``
Create a response and return it without sending any data to the remote server. For use within request interceptors. See `Example: Mock a response`_.
.. _`response object`: #response-objects
WebSocketMessage Objects
------------------------
These objects represent websocket messages sent between the browser and server and vice versa. They are held in a list by ``request.ws_messages`` on websocket handshake requests. They have the following attributes.
``content``
The message content which may be either ``str`` or ``bytes``.
``date``
The datetime of the message.
``from_client``
``True`` when the message was sent by the client and ``False`` when sent by the server.
Response Objects
Response objects have the following attributes.
body
The response body as bytes. If the response has no body the value of body will be empty, i.e. b''. Sometimes the body may have been compressed by the server. You can prevent this with the disable_encoding option_. To manually decode an encoded response body you can do:
.. code:: python
from seleniumwire.utils import decode
body = decode(response.body, response.headers.get('Content-Encoding', 'identity'))
date
The datetime the response was received.
headers
A dictionary-like object of response headers. Headers are case-insensitive and duplicates are permitted. Asking for response.headers['content-length'] will return the value of the Content-Length header. If you wish to replace a header, make sure you delete the existing header first with del response.headers['header-name'], otherwise you'll create a duplicate.
reason
The reason phrase, e.g. OK or Not Found etc.
status_code
The status code of the response, e.g. 200 or 404 etc.
Intercepting Requests and Responses
除了捕获请求和响应之外,Selenium Wire 还允许你使用拦截器(interceptor)即时修改它们。拦截器是一个函数,当请求和响应通过 Selenium Wire 时,该函数会被调用。在拦截器内部,你可以按照你的需要修改请求和响应。
在你开始使用 driver 之前,请使用 ``driver.request_interceptor`` 和 ``driver.response_interceptor`` 属性来设置拦截器函数。请求拦截器应只接受一个参数,用于接收请求。响应拦截器应接受两个参数,分别用于原始请求和响应。
示例:添加请求头
-----------------------------
.. code:: python
def interceptor(request):
request.headers['New-Header'] = 'Some Value'
driver.request_interceptor = interceptor
driver.get(...)
# All requests will now contain New-Header
如何检查请求头是否正确设置?页面加载后,你可以使用 ``driver.requests`` 打印捕获到的请求头;或者将 webdriver 指向 https://httpbin.org/headers,它会将请求头回显到浏览器中,方便你查看。
示例:替换现有的请求头
-------------------------------------------
HTTP 请求允许重复的请求头名称,因此在设置替换请求头之前,你必须先用 ``del`` 删除现有的请求头,如下例所示;否则,将会存在两个同名请求头(``request.headers`` 是一个特殊的类字典对象,允许重复)。
.. code:: python
def interceptor(request):
del request.headers['Referer'] # Remember to delete the header first
request.headers['Referer'] = 'some_referer' # Spoof the referer
driver.request_interceptor = interceptor
driver.get(...)
# All requests will now use 'some_referer' for the referer
示例:添加响应头
------------------------------
.. code:: python
def interceptor(request, response): # A response interceptor takes two args
if request.url == 'https://server.com/some/path':
response.headers['New-Header'] = 'Some Value'
driver.response_interceptor = interceptor
driver.get(...)
# Responses from https://server.com/some/path will now contain New-Header
示例:添加请求参数
--------------------------------
请求参数的工作方式与请求头不同,它们是在设置到请求上时才被计算的。这意味着你需要先读取它们,然后更新,最后再写回——如下例所示。参数存储在一个普通的字典中,因此同名的参数会被覆盖。
.. code:: python
def interceptor(request):
params = request.params
params['foo'] = 'bar'
request.params = params
driver.request_interceptor = interceptor
driver.get(...)
# foo=bar will be added to all requests
示例:更新 POST 请求体中的 JSON
-----------------------------------------------
.. code:: python
import json
def interceptor(request):
if request.method == 'POST' and request.headers['Content-Type'] == 'application/json':
# The body is in bytes so convert to a string
body = request.body.decode('utf-8')
# Load the JSON
data = json.loads(body)
# Add a new property
data['foo'] = 'bar'
# Set the JSON back on the request
request.body = json.dumps(data).encode('utf-8')
# Update the content length
del request.headers['Content-Length']
request.headers['Content-Length'] = str(len(request.body))
driver.request_interceptor = interceptor
driver.get(...)
示例:基本认证
-----------------------------
如果某个站点需要用户名/密码,你可以使用请求拦截器为每个请求添加身份验证凭据。这样可以避免浏览器显示用户名/密码弹窗。
.. code:: python
import base64
auth = (
base64.encodebytes('my_username:my_password'.encode())
.decode()
.strip()
)
def interceptor(request):
if request.host == 'host_that_needs_auth':
request.headers['Authorization'] = f'Basic {auth}'
driver.request_interceptor = interceptor
driver.get(...)
# Credentials will be transmitted with every request to "host_that_needs_auth"
示例:阻止请求
------------------------
你可以使用 ``request.abort()`` 来阻止请求,并立即向浏览器发送响应。可以提供可选的错误代码。默认为 403(禁止)。
.. code:: python
def interceptor(request):
# Block PNG, JPEG and GIF images
if request.path.endswith(('.png', '.jpg', '.gif')):
request.abort()
driver.request_interceptor = interceptor
driver.get(...)
# Requests for PNG, JPEG and GIF images will result in a 403 Forbidden
示例:模拟响应
------------------------
你可以使用 ``request.create_response()`` 向浏览器发送自定义回复。不会向远程服务器发送任何数据。
.. code:: python
def interceptor(request):
if request.url == 'https://server.com/some/path':
request.create_response(
status_code=200,
headers={'Content-Type': 'text/html'}, # Optional headers dictionary
body='<html>Hello World!</html>' # Optional body
)
driver.request_interceptor = interceptor
driver.get(...)
# Requests to https://server.com/some/path will have their responses mocked
*还有其他你认为可能有用的示例吗?欢迎提交 PR。*
取消设置拦截器
--------------------
要取消设置拦截器,请使用 ``del``:
.. code:: python
del driver.request_interceptor
del driver.response_interceptor
限制请求捕获~~~~~~~~~~~~~~~~~~~~~~~~
Selenium Wire works by redirecting browser traffic through an internal proxy server it spins up in the background. As requests flow through the proxy they are intercepted and captured. Capturing requests can slow things down a little but there are a few things you can do to restrict what gets captured.
``driver.scopes``
This accepts a list of regular expressions that will match the URLs to be captured. It should be set on the driver before making any requests. When empty (the default) all URLs are captured.
.. code:: python
driver.scopes = [
'.*stackoverflow.*',
'.*github.*'
]
driver.get(...) # Start making requests
# Only request URLs containing "stackoverflow" or "github" will now be captured
Note that even if a request is out of scope and not captured, it will still travel through Selenium Wire.
``seleniumwire_options.disable_capture``
Use this option to switch off request capture. Requests will still pass through Selenium Wire and through any upstream proxy you have configured but they won't be intercepted or stored. Request interceptors will not execute.
.. code:: python
options = {
'disable_capture': True # Don't intercept/store any requests
}
driver = webdriver.Chrome(seleniumwire_options=options)
``seleniumwire_options.exclude_hosts``
Use this option to bypass Selenium Wire entirely. Any requests made to addresses listed here will go direct from the browser to the server without involving Selenium Wire. Note that if you've configured an upstream proxy then these requests will also bypass that proxy.
.. code:: python
options = {
'exclude_hosts': ['host1.com', 'host2.com'] # Bypass Selenium Wire for these hosts
}
driver = webdriver.Chrome(seleniumwire_options=options)
``request.abort()``
You can abort a request early by using ``request.abort()`` from within a `request interceptor`_. This will send an immediate response back to the client without the request travelling any further. You can use this mechanism to block certain types of requests (e.g. images) to improve page load performance.
.. code:: python
def interceptor(request):
# Block PNG, JPEG and GIF images
if request.path.endswith(('.png', '.jpg', '.gif')):
request.abort()
driver.request_interceptor = interceptor
driver.get(...) # Start making requests
.. _`request interceptor`: #intercepting-requests-and-responses
Request Storage
~~~~~~~~~~~~~~~
Captured requests and responses are stored in the system temp folder by default (that's ``/tmp`` on Linux and usually ``C:\Users\<username>\AppData\Local\Temp`` on Windows) in a sub-folder called ``.seleniumwire``. To change where the ``.seleniumwire`` folder gets created you can use the ``request_storage_base_dir`` option:
.. code:: python
options = {
'request_storage_base_dir': '/my/storage/folder' # .seleniumwire will get created here
}
driver = webdriver.Chrome(seleniumwire_options=options)
In-Memory Storage
-----------------
Selenium Wire also supports storing requests and responses in memory only, which may be useful in certain situations - e.g. if you're running short lived Docker containers and don't want the overhead of disk persistence. You can enable in-memory storage by setting the ``request_storage`` option to ``memory``:
.. code:: python
options = {
'request_storage': 'memory' # Store requests and responses in memory only
}
driver = webdriver.Chrome(seleniumwire_options=options)
If you're concerned about the amount of memory that may be consumed, you can restrict the number of requests that are stored with the ``request_storage_max_size`` option:
.. code:: python
options = {
'request_storage': 'memory',
'request_storage_max_size': 100 # Store no more than 100 requests in memory
}
driver = webdriver.Chrome(seleniumwire_options=options)
When the max size is reached, older requests are discarded as newer requests arrive. Keep in mind that if you restrict the number of requests being stored, requests may have disappeared from storage by the time you come to retrieve them with ``driver.requests`` or ``driver.wait_for_request()`` etc.
Proxies
~~~~~~~
If the site you are accessing sits behind a proxy server you can tell Selenium Wire about that proxy server in the options you pass to the webdriver.
The configuration takes the following format:
.. code:: python
options = {
'proxy': {
'http': 'http://192.168.10.100:8888',
'https': 'https://192.168.10.100:8888',
'no_proxy': 'localhost,127.0.0.1'
}
}
driver = webdriver.Chrome(seleniumwire_options=options)
To use HTTP Basic Auth with your proxy, specify the username and password in the URL:
.. code:: python
options = {
'proxy': {
'https': 'https://user:[email protected]:8888',
}
}
For authentication other than Basic, you can supply the full value for the ``Proxy-Authorization`` header using the ``custom_authorization`` option. For example, if your proxy used the Bearer scheme:
.. code:: python
options = {
'proxy': {
'https': 'https://192.168.10.100:8888', # No username or password used
'custom_authorization': 'Bearer mytoken123' # Custom Proxy-Authorization header value
}
}
More info on the ``Proxy-Authorization`` header can be found `here <https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Proxy-Authorization>`__.
The proxy configuration can also be loaded through environment variables called ``HTTP_PROXY``, ``HTTPS_PROXY`` and ``NO_PROXY``:
.. code:: bash
$ export HTTP_PROXY="http://192.168.10.100:8888"
$ export HTTPS_PROXY="https://192.168.10.100:8888"
$ export NO_PROXY="localhost,127.0.0.1"
SOCKS
-----
Using a SOCKS proxy is the same as using an HTTP based one but you set the scheme to ``socks5``:
.. code:: python
options = {
'proxy': {
'http': 'socks5://user:[email protected]:8888',
'https': 'socks5://user:[email protected]:8888',
'no_proxy': 'localhost,127.0.0.1'
}
}
driver = webdriver.Chrome(seleniumwire_options=options)
You can leave out the ``user`` and ``pass`` if your proxy doesn't require authentication.
As well as ``socks5``, the schemes ``socks4`` and ``socks5h`` are supported. Use ``socks5h`` when you want DNS resolution to happen on the proxy server rather than on the client.
**Using Selenium Wire with Tor**
See `this example <https://gist.github.com/woswos/38b921f0b82de009c12c6494db3f50c5>`_ if you want to run Selenium Wire with Tor.
Switching Dynamically
---------------------
If you want to change the proxy settings for an existing driver instance, use the ``driver.proxy`` attribute:
.. code:: python
driver.get(...) # Using some initial proxy
# Change the proxy
driver.proxy = {
'https': 'https://user:[email protected]:8888',
}
driver.get(...) # These requests will use the new proxy
To clear a proxy, set ``driver.proxy`` to an empty dict ``{}``.
This mechanism also supports the ``no_proxy`` and ``custom_authorization`` options.
Bot Detection
~~~~~~~~~~~~~
Selenium Wire will integrate with `undetected-chromedriver`_ if it finds it in your environment. This library will transparently modify ChromeDriver to prevent it from triggering anti-bot measures on websites.
.. _`undetected-chromedriver`: https://github.com/ultrafunkamsterdam/undetected-chromedriver
If you wish to take advantage of this make sure you have undetected_chromedriver installed:
.. code:: bash
pip install undetected-chromedriver
Then in your code, import the ``seleniumwire.undetected_chromedriver`` package:
.. code:: python
import seleniumwire.undetected_chromedriver as uc
chrome_options = uc.ChromeOptions()
driver = uc.Chrome(
options=chrome_options,
seleniumwire_options={}
)
Certificates
~~~~~~~~~~~~
Selenium Wire uses it's own root certificate to decrypt HTTPS traffic. It is not normally necessary for the browser to trust this certificate because Selenium Wire tells the browser to add it as an exception. This will allow the browser to function normally, but it will display a "Not Secure" message (and/or unlocked padlock) in the address bar. If you wish to get rid of this message you can install the root certificate manually.
You can download the root certificate `here <https://github.com/wkeeling/selenium-wire/raw/master/seleniumwire/ca.crt>`__. Once downloaded, navigate to "Certificates" in your browser settings and import the certificate in the "Authorities" section.
Using Your Own Certificate
--------------------------
If you would like to use your own root certificate you can supply the path to the certificate and the private key using the ``ca_cert`` and ``ca_key`` options.
If you do specify your own certificate, be sure to manually delete Selenium Wire's `temporary storage folder <#request-storage>`_. This will clear out any existing certificates that may have been cached from previous runs.
All Options
~~~~~~~~~~~
A summary of all options that can be passed to Selenium Wire via the ``seleniumwire_options`` webdriver attribute.
``addr``
The IP address or hostname of the machine running Selenium Wire. This defaults to 127.0.0.1. You may want to change this to the public IP of the machine (or container) if you're using the `remote webdriver`_.
.. code:: python
options = {
'addr': '192.168.0.10' # Use the public IP of the machine
}
driver = webdriver.Chrome(seleniumwire_options=options)
.. _`remote webdriver`: #creating-the-webdriver
``auto_config``
Whether Selenium Wire should auto-configure the browser for request capture. ``True`` by default.
``ca_cert``
The path to a root (CA) certificate if you prefer to use your own certificate rather than use the default.
.. code:: python
options = {
'ca_cert': '/path/to/ca.crt' # Use own root certificate
}
driver = webdriver.Chrome(seleniumwire_options=options)
``ca_key``
The path to the private key if you're using your own root certificate. The key must always be supplied when using your own certificate.
.. code:: python
options = {
'ca_key': '/path/to/ca.key' # Path to private key
}
driver = webdriver.Chrome(seleniumwire_options=options)
``disable_capture``
Disable request capture. When ``True`` nothing gets intercepted or stored. ``False`` by default.
.. code:: python
options = {
'disable_capture': True # Don't intercept/store any requests.
}
driver = webdriver.Chrome(seleniumwire_options=options)
``disable_encoding``
Ask the server to send back uncompressed data. ``False`` by default. When ``True`` this sets the ``Accept-Encoding`` header to ``identity`` for all outbound requests. Note that it won't always work - sometimes the server may ignore it.
.. code:: python
options = {
'disable_encoding': True # Ask the server not to compress the response
}
driver = webdriver.Chrome(seleniumwire_options=options)
``enable_har``
When ``True`` a HAR archive of HTTP transactions will be kept which can be retrieved with ``driver.har``. ``False`` by default.
.. code:: python
options = {
'enable_har': True # Capture HAR data, retrieve with driver.har
}
driver = webdriver.Chrome(seleniumwire_options=options)
``exclude_hosts``
A list of addresses for which Selenium Wire should be bypassed entirely. Note that if you have configured an upstream proxy then requests to excluded hosts will also bypass that proxy.
.. code:: python
options = {
'exclude_hosts': ['google-analytics.com'] # Bypass these hosts
}
driver = webdriver.Chrome(seleniumwire_options=options)
``ignore_http_methods``
A list of HTTP methods (specified as uppercase strings) that should be ignored by Selenium Wire and not captured. The default is ``['OPTIONS']`` which ignores all OPTIONS requests. To capture all request methods, set ``ignore_http_methods`` to an empty list:
.. code:: python
options = {
'ignore_http_methods': [] # Capture all requests, including OPTIONS requests
}
driver = webdriver.Chrome(seleniumwire_options=options)
``port``
The port number that Selenium Wire's backend listens on. You don't normally need to specify a port as a random port number is chosen automatically.
.. code:: python
options = {
'port': 9999 # Tell the backend to listen on port 9999 (not normally necessary to set this)
}
driver = webdriver.Chrome(seleniumwire_options=options)
``proxy``
The upstream `proxy server <https://github.com/wkeeling/selenium-wire#proxies>`__ configuration if you're using a proxy.
.. code:: python
options = {
'proxy': {
'http': 'http://user:[email protected]:8888',
'https': 'https://user:[email protected]:8889',
'no_proxy': 'localhost,127.0.0.1'
}
}
driver = webdriver.Chrome(seleniumwire_options=options)
``request_storage``
The type of storage to use. Selenium Wire defaults to disk based storage, but you can switch to in-memory storage by setting this option to ``memory``:
.. code:: python
options = {
'request_storage': 'memory' # Store requests and responses in memory only
}
driver = webdriver.Chrome(seleniumwire_options=options)
``request_storage_base_dir``
The base location where Selenium Wire stores captured requests and responses when using its default disk based storage. This defaults to the system temp folder (that's ``/tmp`` on Linux and usually ``C:\Users\<username>\AppData\Local\Temp`` on Windows). A sub-folder called ``.seleniumwire`` will get created here to store the captured data.
.. code:: python
options = {
'request_storage_base_dir': '/my/storage/folder' # .seleniumwire will get created here
}
driver = webdriver.Chrome(seleniumwire_options=options)
``request_storage_max_size``
The maximum number of requests to store when using in-memory storage. Unlimited by default. This option currently has no effect when using the default disk based storage.
.. code:: python
options = {
'request_storage': 'memory',
'request_storage_max_size': 100 # Store no more than 100 requests in memory
}
driver = webdriver.Chrome(seleniumwire_options=options)
``suppress_connection_errors``
Whether to suppress connection related tracebacks. ``True`` by default, meaning that harmless errors that sometimes occur at browser shutdown do not alarm users. When suppressed, the connection error message is logged at DEBUG level without a traceback. Set to ``False`` to allow exception propagation and see full tracebacks.
.. code:: python
options = {
'suppress_connection_errors': False # Show full tracebacks for any connection errors
}
driver = webdriver.Chrome(seleniumwire_options=options)
``verify_ssl``
Whether SSL certificates should be verified. ``False`` by default, which prevents errors with self-signed certificates.
.. code:: python
options = {
'verify_ssl': True # Verify SSL certificates but beware of errors with self-signed certificates
}
driver = webdriver.Chrome(seleniumwire_options=options)
License
~~~~~~~
MIT