
순수 Rust로 작성된 libxml2 재구현
libxml2의 순수 Rust 재구현 — 오픈 소스 세계에서 사실상 표준인 XML/HTML 파싱 라이브러리입니다.
libxml2는 2025년 12월에 알려진 보안 문제로 인해 공식적으로 유지보수가 중단되었습니다. xmloxide는 동일한 적합성 테스트 스위트를 통과하는 메모리 안전하고 고성능인 대체제를 목표로 합니다.
unsafe가 전혀 없는 아레나 기반 트리html5::sax)css::select)으로 요소 쿼리, 결합자, 의사 클래스 및 빠른 #id 조회 포함matches(), replace(), tokenize(), upper-case(), lower-case(), abs(), min(), max() 등)를 포함한 완전한 표현식 파서 및 평가기serde 기능tokio::io::AsyncRead 소스에서 파싱을 위한 선택적 async 기능xmllint CLI — XML 파싱, 검증 및 쿼리를 위한 명령줄 도구Document는 독립적이며 Send + Syncinclude/xmloxide.h)encoding_rs만 사용 (라이브러리에는 다른 의존성이 없으며, clap은 CLI 전용)use xmloxide::Document;
let doc = Document::parse_str("<root><child>Hello</child></root>").unwrap();
let root = doc.root_element().unwrap();
assert_eq!(doc.node_name(root), Some("root"));
assert_eq!(doc.text_content(root), "Hello");
use xmloxide::Document;
use xmloxide::serial::serialize;
let doc = Document::parse_str("<root><child>Hello</child></root>").unwrap();
let xml = serialize(&doc);
assert_eq!(xml, "<root><child>Hello</child></root>");
use xmloxide::Document;
use xmloxide::xpath::{evaluate, XPathValue};
let doc = Document::parse_str("<library><book><title>Rust</title></book></library>").unwrap();
let root = doc.root_element().unwrap();
let result = evaluate(&doc, root, "count(book)").unwrap();
assert_eq!(result.to_number(), 1.0);
use xmloxide::sax::{parse_sax, SaxHandler, DefaultHandler};
use xmloxide::parser::ParseOptions;
struct MyHandler;
impl SaxHandler for MyHandler {
fn start_element(&mut self, name: &str, _: Option<&str>, _: Option<&str>,
_: &[(String, String, Option<String>, Option<String>)]) {
println!("Element: {name}");
}
}
parse_sax("<root><child/></root>", &ParseOptions::default(), &mut MyHandler).unwrap();
use xmloxide::html::parse_html;
let doc = parse_html("<p>Hello <br> World").unwrap();
let root = doc.root_element().unwrap();
assert_eq!(doc.node_name(root), Some("html"));
use xmloxide::css::select;
use xmloxide::Document;
let doc = Document::parse_str(r#"<div><p class="intro">Hello</p><p>World</p></div>"#).unwrap();
let root = doc.root_element().unwrap();
let intros = select(&doc, root, "p.intro").unwrap();
assert_eq!(intros.len(), 1);
assert_eq!(doc.text_content(intros[0]), "Hello");
use xmloxide::html5::parse_html5;
let doc = parse_html5("<p>Hello <b>world</b>").unwrap();
let root = doc.root_element().unwrap();
assert_eq!(doc.node_name(root), Some("html"));
Fragment parsing (the algorithm behind innerHTML) is also supported:
use xmloxide::html5::{parse_html5_with_options, Html5ParseOptions};
let opts = Html5ParseOptions {
scripting: false,
fragment_context: Some("body".to_string()),
};
let doc = parse_html5_with_options("<p>fragment</p>", &opts).unwrap();
use xmloxide::html5::sax::{Html5SaxHandler, parse_html5_sax};
struct LinkExtractor { hrefs: Vec<String> }
impl Html5SaxHandler for LinkExtractor {
fn start_element(&mut self, name: &str, attrs: &[(String, String)], _sc: bool) {
if name == "a" {
if let Some((_, href)) = attrs.iter().find(|(n, _)| n == "href") {
self.hrefs.push(href.clone());
}
}
}
}
let mut handler = LinkExtractor { hrefs: Vec::new() };
parse_html5_sax(r#"<a href="https://github.com/jonwiggins/xmloxide/blob/main/page">Link</a>"#, &mut handler);
assert_eq!(handler.hrefs, vec!["/page"]);
use xmloxide::parser::{parse_str_with_options, ParseOptions};
let opts = ParseOptions::default().recover(true);
let doc = parse_str_with_options("<root><unclosed>", &opts).unwrap();
for diag in &doc.diagnostics {
eprintln!("{}", diag);
}
# Parse and pretty-print
xmllint --format document.xml
# Validate against a schema
xmllint --schema schema.xsd document.xml
xmllint --relaxng schema.rng document.xml
xmllint --schematron schema.sch document.xml
xmllint --dtdvalid schema.dtd document.xml
# XPath query
xmllint --xpath "//title" document.xml
# Canonical XML
xmllint --c14n document.xml
# Parse HTML
xmllint --html page.html
| Module | Description |
|---|---|
tree | 아레나 기반 DOM 트리 (Document, NodeId, NodeKind) |
parser | 오류 복구를 지원하는 XML 1.0 재귀 하향 파서 |
parser::push | 청크 입력을 위한 푸시/증분 파서 |
html | 오류 허용 HTML 4.01 파서 |
html5 | WHATWG HTML Living Standard 파서 (토크나이저 + 트리 빌더) |
html5::sax | HTML5용 스트리밍 SAX 유사 API (DOM 트리 구축 없음) |
css | 문서 트리 쿼리를 위한 CSS 선택자 엔진 |
sax | SAX2 스트리밍 이벤트 기반 파서 |
reader | XmlReader 풀 기반 파싱 API |
serial | XML, HTML 및 HTML5 직렬화기, Canonical XML (C14N) 포함 |
xpath | XPath 1.0+ 표현식 파서 및 평가기 |
validation::dtd | DTD 파싱 및 검증 |
validation::relaxng | RelaxNG 스키마 검증 |
validation::xsd | XML Schema (XSD) 검증 |
validation::schematron | ISO Schematron 규칙 기반 검증 |
serde_xml | Serde XML (역)직렬화 (선택적 serde 기능) |
async_xml | tokio::io::AsyncRead를 통한 비동기 파싱 (선택적 async 기능) |
xinclude | XInclude 1.0 문서 포함 |
catalog | URI 해석을 위한 OASIS XML 카탈로그 |
encoding | 문자 인코딩 감지 및 트랜스코딩 |
ffi | C/C++ FFI 바인딩 (include/xmloxide.h) |
파싱 처리량은 libxml2와 경쟁력이 있습니다 — 대부분의 문서에서 3-4% 이내, SVG의 경우 12% 더 빠릅니다. 직렬화는 아레나 기반 트리 설계 덕분에 1.5-2.4배 더 빠릅니다. XPath는 모든 벤치마크에서 1.1-2.7배 더 빠릅니다.
Parsing:
| Document | Size | xmloxide | libxml2 | Result |
|---|---|---|---|---|
| Atom 피드 | 4.9 KB | 26.7 µs (176 MiB/s) | 25.5 µs (184 MiB/s) | 약 4% 느림 |
| SVG 도면 | 6.3 KB | 58.5 µs (103 MiB/s) | 65.6 µs (92 MiB/s) | 12% 더 빠름 |
| Maven POM | 11.5 KB | 76.9 µs (142 MiB/s) | 74.2 µs (148 MiB/s) | 약 4% 느림 |
| XHTML 페이지 | 10.2 KB | 69.5 µs (139 MiB/s) | 61.5 µs (157 MiB/s) | 약 13% 느림 |
| 대용량 (374 KB) | 374 KB | 2.15 ms (169 MiB/s) | 2.08 ms (175 MiB/s) | 약 3% 느림 |
Serialization:
| Document | Size | xmloxide | libxml2 | Result |
|---|---|---|---|---|
| Atom 피드 | 4.9 KB | 11.3 µs | 17.5 µs | 1.5배 더 빠름 |
| Maven POM | 11.5 KB | 20.1 µs | 47.5 µs | 2.4배 더 빠름 |
| 대용량 (374 KB) | 374 KB | 614 µs | 1397 µs | 2.3배 더 빠름 |
XPath: