
Python-Scraper basierend auf KI
English | 中文 | 日本語 | 한국어 | Русский | Türkçe | Deutsch | Español | français | Português | Italiano
ScrapeGraphAI ist eine Web-Scraping-Python-Bibliothek, die LLM und direkte Graphlogik verwendet, um Scraping-Pipelines für Websites und lokale Dokumente (XML, HTML, JSON, Markdown usw.) zu erstellen. Sage einfach, welche Informationen du extrahieren möchtest, und die Bibliothek erledigt das für dich!
ScrapeGraphAI bietet eine nahtlose Integration mit gängigen Frameworks und Tools, um deine Scraping-Fähigkeiten zu erweitern. Egal, ob du mit Python oder Node.js entwickelst, LLM-Frameworks verwendest oder mit No-Code-Plattformen arbeitest – wir haben mit unseren umfassenden Integrationsmöglichkeiten alles für dich abgedeckt.
Du findest weitere Informationen unter folgendem Link
Integrationen:
Die Referenzseite für Scrapegraph-ai ist auf der offiziellen PyPI-Seite verfügbar: pypi.
pip install scrapegraphai
# IMPORTANT (for fetching websites content)
playwright install
Hinweis: Es wird empfohlen, die Bibliothek in einer virtuellen Umgebung zu installieren, um Konflikte mit anderen Bibliotheken zu vermeiden 🐱
Es gibt mehrere standardmäßige Scraping-Pipelines, die zum Extrahieren von Informationen aus einer Website (oder einer lokalen Datei) verwendet werden können. Die häufigste ist SmartScraperGraph, die Informationen von einer einzelnen Seite extrahiert, basierend auf einem Benutzer-Prompt und einer Quell-URL.
from scrapegraphai.graphs import SmartScraperGraph
# Define the configuration for the scraping pipeline
graph_config = {
"llm": {
"model": "ollama/llama3.2",
"model_tokens": 8192,
"format": "json",
},
"verbose": True,
"headless": False,
}
# Create the SmartScraperGraph instance
smart_scraper_graph = SmartScraperGraph(
prompt="Extract useful information from the webpage, including a description of what the company does, founders and social media links",
source="https://scrapegraphai.com/",
config=graph_config
)
# Run the pipeline
result = smart_scraper_graph.run()
import json
print(json.dumps(result, indent=4))
[!NOTE] Für OpenAI und andere Modelle musst du nur die llm-Konfiguration ändern!
graph_config = { "llm": { "api_key": "YOUR_OPENAI_API_KEY", "model": "openai/gpt-4o-mini", }, "verbose": True, "headless": False, }
Die Ausgabe wird ein Dictionary wie das folgende sein:
{
"description": "ScrapeGraphAI transforms websites into clean, organized data for AI agents and data analytics. It offers an AI-powered API for effortless and cost-effective data extraction.",
"founders": [
{
"name": "",
"role": "Founder & Technical Lead",
"linkedin": "https://www.linkedin.com/in/perinim/"
},
{
"name": "Marco Vinciguerra",
"role": "Founder & Software Engineer",
"linkedin": "https://www.linkedin.com/in/marco-vinciguerra-7ba365242/"
},
{
"name": "Lorenzo Padoan",
"role": "Founder & Product Engineer",
"linkedin": "https://www.linkedin.com/in/lorenzo-padoan-4521a2154/"
}
],
"social_media_links": {
"linkedin": "https://www.linkedin.com/company/101881123",
"twitter": "https://x.com/scrapegraphai",
"github": "https://github.com/ScrapeGraphAI/Scrapegraph-ai"
}
}
Es gibt weitere Pipelines, die verwendet werden können, um Informationen von mehreren Seiten zu extrahieren, Python-Skripte zu generieren oder sogar Audiodateien zu erzeugen.