Skip to content
KitploitKITPLOIT
도구블로그
제출
도구블로그
제출

해킹, 침투 테스트 및 사이버 보안 도구를 당신의 보안 무기고에!

Kitploit은 해킹, 사이버 보안 및 침투 테스트 도구 디렉토리입니다. 최신 프로젝트 업데이트를 발견하여 취약점을 찾고, 시스템을 분석하고, 테스트를 자동화하고, 보안을 강화하세요.

··피드·문의·개인정보·© 2026 Kitploit

도구 디렉토리

카테고리

모든 카테고리 보기
Loading categories
loglizer — 로그 기반 이상 탐지용 머신러닝 툴킷 [ISSRE'16] | Kitploit
도구/GitHubGitHub/logpai/loglizer
Machine LearningAnomaly DetectionLog Analysis
GitHublogpai/loglizer

loglizer

로그 기반 이상 탐지용 머신러닝 툴킷 [ISSRE'16]

저장소 보기
1.4k4372년 전Kitploit 검토 완료

인기

모두 보기 →

커뮤니티에서 가장 많이 사용되는 도구를 찾아보세요.

모든 도구 탐색

도구 컬렉션을 둘러보세요

모든 도구 보기 →
공유

loglizer

Loglizer는 자동 이상 탐지를 위한 머신러닝 기반 로그 분석 툴킷입니다.

Loglizer는 AI 기반 로그 빅데이터 분석 도구로, 자동 이상 탐지, 지능형 장애 진단 등의 시나리오에 사용할 수 있습니다.

로그는 많은 소프트웨어 시스템의 개발 및 유지보수 과정에서 필수적입니다. 로그는 시스템 운영 중의 상세한 런타임 정보를 기록하여 개발자와 지원 엔지니어가 시스템을 모니터링하고 비정상적인 동작과 오류를 추적할 수 있게 해 줍니다. Loglizer는 자동 이상 탐지를 위한 다양한 머신러닝 기반 로그 분석 기법을 구현한 툴킷을 제공합니다.

🔭 연구 논문 게재를 위해 loglizer를 사용하신다면 다음 논문을 인용해 주시기 바랍니다.

  • Shilin He, Jieming Zhu, Pinjia He, Michael R. Lyu. Experience Report: System Log Analysis for Anomaly Detection, IEEE International Symposium on Software Reliability Engineering (ISSRE), 2016. [Bibtex][중국어 버전] (ISSRE 최고 영향력 논문)

프레임워크

Framework of Anomaly Detection

이상 탐지를 위한 로그 분석 프레임워크는 일반적으로 다음 구성 요소로 이루어져 있습니다:

  1. 로그 수집: 로그는 런타임에 생성되어 Flume, Kafka와 같은 데이터 스트리밍 파이프라인을 통해 중앙 집중식 장소로 집계됩니다.
  2. 로그 파싱: 로그 파싱의 목표는 비정형 로그 메시지를 구조화된 이벤트 맵으로 변환하는 것이며, 이를 기반으로 정교한 머신러닝 모델을 적용할 수 있습니다. 로그 파싱에 대한 자세한 내용은 logparser 프로젝트에서 확인할 수 있습니다.
  3. 특징 추출: 구조화된 로그는 구간 윈도우, 슬라이딩 윈도우, 세션 윈도우를 통해 짧은 로그 시퀀스로 나눌 수 있습니다. 그런 다음 각 로그 시퀀스를 벡터화하기 위해 특징 추출을 수행합니다(예: 이벤트 카운팅 벡터 사용).
  4. 이상 탐지: 주어진 특징 벡터가 이상인지 여부를 확인하도록 이상 탐지 모델을 학습합니다.

모델

현재 사용 가능한 이상 탐지 모델:

로그 데이터

연구 목적으로 loghub에 레이블이 지정된 로그 데이터셋 모음을 수집했습니다. 데이터셋에 관심이 있으시면 링크를 따라 액세스 요청을 제출해 주세요.

설치

root@kitploit:~
git clone https://github.com/logpai/loglizer.git
cd loglizer
pip install -r requirements.txt

API 사용법

root@kitploit:~
# Load HDFS dataset. If you would like to try your own log, you need to rewrite the load function.
(x_train, y_train), (x_test, y_test) = dataloader.load_HDFS(...)

# Feature extraction and transformation
feature_extractor = preprocessing.FeatureExtractor()
feature_extractor.fit_transform(...) 

# Model training
model = PCA()
model.fit(...)

# Feature transform after fitting
x_test = feature_extractor.transform(...)
# Model evaluation with labeled data
model.evaluate(...)

# Anomaly prediction
x_test = feature_extractor.transform(...)
model.predict(...) # predict anomalies on given data

자세한 내용은 docs의 데모를 따라 시작하세요. 모든 ML 모델은 마법이 아니라는 점에 유의하세요. 자신의 데이터에 적용하려면 파라미터를 튜닝하는 방법을 직접 파악해야 합니다.

벤치마킹 결과

다음 결과를 재현하려면 전체 HDFS 데이터셋에서 benchmarks/HDFS_bechmark.py를 실행하세요(HDFS100k는 데모 전용입니다).

기여자

  • Shilin He, 홍콩 중문대학
  • Jieming Zhu, 홍콩 중문대학, 현재 Huawei Noah's Ark Lab 재직 중
  • Pinjia He, 홍콩 중문대학, 현재 ETH Zurich 재직 중

피드백

질문이나 피드백이 있으시면 이슈 페이지에 게시해 주세요.

변경 이력

  • May 14, 2016: 최초 커밋
  • Sep 21, 2017: 코드 및 readme 업데이트
  • Mar 21, 2018: 대부분의 코드 재작성 및 상세 주석 추가
  • Feb 18, 2019: 실습 데모와 함께 저장소 재구성
도구 다운로드
ModelPaper reference
지도 학습 모델
LR[EuroSys'10] Fingerprinting the Datacenter: Automated Classification of Performance Crises, by Peter Bodík, Moises Goldszmidt, Armando Fox, Hans Andersen. [Microsoft]
Decision Tree[ICAC'04] Failure Diagnosis Using Decision Trees, by Mike Chen, Alice X. Zheng, Jim Lloyd, Michael I. Jordan, Eric Brewer. [eBay]
SVM[ICDM'07] Failure Prediction in IBM BlueGene/L Event Logs, by Yinglung Liang, Yanyong Zhang, Hui Xiong, Ramendra Sahoo. [IBM]
비지도 학습 모델
LOF[SIGMOD'00] LOF: Identifying Density-Based Local Outliers, by Markus M. Breunig, Hans-Peter Kriegel, Raymond T. Ng, Jörg Sander.
One-Class SVM[Neural Computation'01] Estimating the Support of a High-Dimensional Distribution, by John Platt, Bernhard Schölkopf, John Shawe-Taylor, Alex J. Smola, Robert C. Williamson.
Isolation Forest[ICDM'08] Isolation Forest, by Fei Tony Liu, Kai Ming Ting, Zhi-Hua Zhou.
PCA[SOSP'09] Large-Scale System Problems Detection by Mining Console Logs, by Wei Xu, Ling Huang, Armando Fox, David Patterson, Michael I. Jordan. [Intel]
Invariants Mining[ATC'10] Mining Invariants from Console Logs for System Problem Detection, by Jian-Guang Lou, Qiang Fu, Shengqi Yang, Ye Xu, Jiang Li. [Microsoft]
Clustering[ICSE'16] Log Clustering based Problem Identification for Online Service Systems, by Qingwei Lin, Hongyu Zhang, Jian-Guang Lou, Yu Zhang, Xuewei Chen. [Microsoft]
DeepLog (예정)[CCS'17] DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning, by Min Du, Feifei Li, Guineng Zheng, Vivek Srikumar.
AutoEncoder (예정)[Arxiv'18] Anomaly Detection using Autoencoders in High Performance Computing Systems, by Andrea Borghesi, Andrea Bartolini, Michele Lombardi, Michela Milano, Luca Benini.
HDFS
모델정밀도재현율F1
LR0.9550.9110.933
Decision Tree0.9980.9980.998
SVM0.9590.9700.965
LOF0.9670.5610.710
One-Class SVM0.9950.2220.363
Isolation Forest0.8300.7760.802
PCA0.9750.6350.769
Invariants Mining0.8880.9450.915
Clustering1.0000.7200.837