
딥러닝을 활용한 비밀번호 헌팅
문서에서 비밀번호 후보를 분석하는 Docker 기반 애플리케이션입니다.
블로그 게시글 "DeepPass — Finding Passwords With Deep Learning" 에 모델의 접근 방식과 개발에 대한 자세한 내용이 나와 있습니다.
실행 방법: docker-compose up
이렇게 하면 문서를 업로드할 수 있는 http://localhost:5000이 노출됩니다.
API는 http://localhost:5000/api/passwords 에서 수동으로 사용할 수 있습니다:
C:\Users\harmj0y\Documents\GitHub\DeepPass>curl -F "file=@test_doc.docx" http://localhost:5000/api/passwords
[{"file_name": "test_doc.docx", "model_password_candidates": [{"left_context": ["for", "the", "production", "server", "is:"], "password": "P@ssword123!", "right_context": ["Please", "dont", "tell", "anyone", "on"]}, {"left_context": ["that", "the", "other", "password", "is"], "password": "LiverPool1", "right_context": [".", "This", "is", "our", "backup."]}], "regex_password_candidates": [{"left_context": ["for", "the", "production", "server", "is:"], "password": "P@ssword123!", "right_context": ["Please", "dont", "tell", "anyone", "on"]}], "custom_regex_matches": null}]
Apache Tika 는 다양한 문서 형식 에서 데이터를 추출하는 데 사용됩니다. Tensorflow Serving 은 모델 서빙에 사용됩니다.
신경망은 양방향 LSTM(Bidirectional LSTM)입니다:
embedding_dimension = 20
dropout = 0.5
cells = 200
model = Sequential()
model.add(Embedding(total_chars, embedding_dimension, input_length=32, mask_zero=True))
model.add(Bidirectional(LSTM(cells)))
model.add(Dropout(dropout))
model.add(Dense(1, activation='sigmoid'))
이 모델은 이 유출된 비밀번호 목록 에서 무작위로 선택된 2,000,000개의 비밀번호와 다양한 Google dorked 문서에서 추출한 2,000,000개의 용어로 훈련되었습니다. .1 테스트 세트에 대한 통계는 다음과 같습니다:
------------------
loss : 0.04804224148392677
tn : 199446.0
fp : 731.0
fn : 3281.0
tp : 196542.0
------------------
accuracy : 0.9899700284004211
precision : 0.9962944984436035
recall : 0.983580470085144
------------------
F1 score. : 0.9898966618590025
------------------
모델의 훈련 노트북은 ./notebooks/password_model_bilstm.ipynb 에 있습니다.