
Sucessor do Undetected-Chromedriver. Oferecendo uma estrutura extremamente rápida para automação web, web scraping, bots e quaisquer outras ideias criativas que normalmente são dificultadas por sistemas irritantes de anti-bot como Captcha / CloudFlare / Imperva / hCaptcha.
A comunicação direta oferece resistência ainda melhor contra firewalls de aplicações web (WAF’s), enquanto o desempenho ganha um enorme impulso. Este módulo, ao contrário do undetected-chromedriver, é totalmente assíncrono.
O que torna este pacote diferente de outros pacotes conhecidos é a otimização para permanecer indetectado para a maioria das soluções anti-bot.
Outro ponto de foco é a usabilidade e a prototipagem rápida, então espere que muita coisa funcione -as is-,
com a maioria dos parâmetros de métodos tendo padrões de best practice (boas práticas).
Usando 1 ou 2 linhas, ele está pronto para funcionar, fornecendo configuração de boas práticas
por padrão. Ele limpa os arquivos criados (perfil) em seguida.
conhecido por funcionar com
Embora usabilidade e conveniência sejam importantes, também é fácil personalizar totalmente tudo usando toda a gama de domínios, métodos e eventos do CDP disponíveis.
Sem dependência do binário chromedriver ou do Selenium
Funcionando com 1 linha de código*
usa um perfil novo a cada execução e limpa ao sair
salva e carrega cookies em arquivo para não repetir etapas tediosas de login
tab.find("sometext")
tab.find_all("sometext")
tab.select("a[class*=something]")
tab.select_all("a[href] > div > img")
busca de elementos inteligente e eficiente, por seletor ou texto, incluindo conteúdo de iframes.
isso também pode ser usado como condição de espera para um elemento aparecer, pois tentará novamente
durante o período de <timeout> até encontrar. Portanto, um await tab.select('body') pode ser usado
como indicador de se uma página foi carregada.
o método find busca por texto, mas não retorna ingenuamente o primeiro
elemento correspondente; em vez disso, combina candidatos pela correspondência mais próxima em comprimento de texto (o mais curto vence),
isso faz com que buscas como tab.find('accept all') retornem o botão real de cookies em vez de
um script nos cabeçalhos
pode se conectar a uma sessão de depuração do chrome em execução
__repr__ descritivo para elementos, que representa o elemento como html
função utilitária para converter uma instância undetected_chromedriver.Chrome em execução em uma instância nodriver.Browser e continuar a partir daí
repleto de auxiliares e métodos utilitários para as operações mais usadas e importantes
Partes foram reescritas para usar conexões flat no protocolo.
Por quê?
- iframes são incluídos na maioria das operações.
- o objeto tab ganhou um novo método: await tab.get_frames()
que retornará Iframes inspecionáveis.
- find() incluirá iframes, para que você possa até mesmo procurar por "verify you are human" e
clicar na caixa de verificação em desafios de js.
Como isso exigiu bastante reescrita, teste minuciosamente, especialmente se você executa projetos grandes.
tab.xpath(selector, timeout=2.5)encontre nós usando o seletor xpath! veja tab xpath na documentação da API
tab.cf_verify()encontra a caixa de seleção e clica nela com sucesso isso só funciona quando NÃO está no modo expert. atualmente, apenas inglês embutido requer que o pacote opencv-python esteja instalado
tab.bypass_insecure_connection_warning()método de conveniência para o aviso de página insegura. por exemplo, quando um certificado é inválido.
tab.open_external_debugger()permite inspecionar a aba sem quebrar sua conexão
tab.get_local_storage()obtém o conteúdo do localstorage
tab.set_local_storage(dict)define o conteúdo do localstorage
tab.add_handler(someEvent, callback)o callback pode aceitar um único argumento (event) ou 2 argumentos (event, tab).
start(expert=True)faz alguns ajustes para usuários mais experientes. Desativa a segurança web e origin-trials, além de garantir que shadow-roots estejam sempre abertos. Porém, isso torna você mais detectável!
você precisa ter o chrome (ou algum navegador baseado em chromium) instalado, preferencialmente no local padrão, na máquina em que você usa este pacote.
ao executar em uma máquina headless, como AWS ou qualquer outro ambiente onde não há display, é melhor usar alguma ferramenta Xvfb para emular uma tela. alternativamente, este pacote pode ser usado no modo headless.
pip install nodriver
pip install -U nodriver
O objetivo deste projeto (assim como o undetected-chromedriver, em algum momento no passado) é manter tudo curto e simples, para que você possa abrir rapidamente um editor ou sessão interativa, digitar ou colar algumas linhas e começar.
import nodriver as uc
async def main():
browser = await uc.start()
page = await browser.get('https://www.nowsecure.nl')
... further code ...
if __name__ == '__main__':
# since asyncio.run never worked (for me)
uc.loop().run_until_complete(main())
Vou omitir o boilerplate assíncrono aqui
from nodriver import *
browser = await start(
headless=False,
user_data_dir="/path/to/existing/profile", # by specifying it, it won't be automatically cleaned up when finished
browser_executable_path="/path/to/some/other/browser",
browser_args=['--some-browser-arg=true', '--some-other-option'],
lang="en-US" # this could set iso-language-code in navigator, not recommended to change
)
tab = await browser.get('https://somewebsite.com')
Vou omitir o boilerplate assíncrono aqui
from nodriver import *
config = Config()
config.headless = False
config.user_data_dir="/path/to/existing/profile", # by specifying it, it won't be automatically cleaned up when finished
config.browser_executable_path="/path/to/some/other/browser",
config.browser_args=['--some-browser-arg=true', '--some-other-option'],
config.lang="en-US" # this could set iso-language-code in navigator, not recommended to change
)
import nodriver
async def main():
browser = await nodriver.start()
page = await browser.get('https://www.nowsecure.nl')
await page.save_screenshot()
await page.get_content()
await page.scroll_down(150)
elems = await page.select_all('*[src]')
for elem in elems:
await elem.flash()
page2 = await browser.get('https://twitter.com', new_tab=True)
page3 = await browser.get('https://github.com/ultrafunkamsterdam/nodriver', new_window=True)
for p in (page, page2, page3):
await p.bring_to_front()
await p.scroll_down(200)
await p # wait for events to be processed
await p.reload()
if p != page3:
await p.close()
if __name__ == '__main__':
# since asyncio.run never worked (for me)
uc.loop().run_until_complete(main())
automatizando a criação de conta no Twitter/X
A more concrete example, which can be found in the ./example/ folder,
shows a script to create a twitter account
```python
import random
import string
import logging
logging.basicConfig(level=30)
import nodriver as uc
months = [
"january",
"february",
"march",
"april",
"may",
"june",
"july",
"august",
"september",
"october",
"november",
"december",
]
async def main():
driver = await uc.start()
tab = await driver.get("https://twitter.com")
# wait for text to appear instead of a static number of seconds to wait
# this does not always work as expected, due to speed.
print('finding the "create account" button')
create_account = await tab.find("create account", best_match=True)
print('"create account" => click')
await create_account.click()
print("finding the email input field")
email = await tab.select("input[type=email]")
# sometimes, email field is not shown, because phone is being asked instead
# when this occurs, find the small text which says "use email instead"
if not email:
use_mail_instead = await tab.find("use email instead")
# and click it
await use_mail_instead.click()
# now find the email field again
email = await tab.select("input[type=email]")
randstr = lambda k: "".join(random.choices(string.ascii_letters, k=k))
# send keys to email field
print('filling in the "email" input field')
await email.send_keys("".join([randstr(8), "@", randstr(8), ".com"]))
# find the name input field
print("finding the name input field")
name = await tab.select("input[type=text]")
# again, send random text
print('filling in the "name" input field')
await name.send_keys(randstr(8))
# since there are 3 select fields on the tab, we can use unpacking
# to assign each field
print('finding the "month" , "day" and "year" fields in 1 go')
sel_month, sel_day, sel_year = await tab.select_all("select")
# await sel_month.focus()
print('filling in the "month" input field')
await sel_month.send_keys(months[random.randint(0, 11)].title())
# await sel_day.focus()
# i don't want to bother with month-lengths and leap years
print('filling in the "day" input field')
await sel_day.send_keys(str(random.randint(0, 28)))
# await sel_year.focus()
# i don't want to bother with age restrictions
print('filling in the "year" input field')
await sel_year.send_keys(str(random.randint(1980, 2005)))
await tab
# let's handle the cookie nag as well
cookie_bar_accept = await tab.find("accept all", best_match=True)
if cookie_bar_accept:
await cookie_bar_accept.click()
await tab.sleep(1)
next_btn = await tab.find(text="next", best_match=True)
# for btn in reversed(next_btns):
await next_btn.mouse_click()
print("sleeping 2 seconds")
await tab.sleep(2) # visually see what part we're actually in
print('finding "next" button')
next_btn = await tab.find(text="next", best_match=True)
print('clicking "next" button')
await next_btn.mouse_click()
# just wait for some button, before we continue
await tab.select("[role=button]")
print('finding "sign up" button')
sign_up_btn = await tab.find("Sign up", best_match=True)
# we need the second one
print('clicking "sign up" button')
await sign_up_btn.click()
print('the rest of the "implementation" is out of scope')
# further implementation outside of scope
await tab.sleep(10)
driver.stop()
# verification code per mail
if __name__ == "__main__":
# since asyncio.run never worked (for me)
# i use
uc.loop().run_until_complete(main())