
A Java Burp Plugin that performs text clustering on responses to identify outliers/groups based on the actual content of the server responses, say from an Intruder run.
A Burp Suite extension for clustering HTTP responses to find outliers.
This is a plugin I've been thinking about building for many years, simply because I'm moderately annoyed by how we analyze results of fuzzing attacks in Burp intruder. We're typically looking for changes in server responses in our fuzzing attacks. And we look for those differences by looking at status codes, response time, and response sizes. All of these are indirect measures of the content of the response. Of course as pentesters we don't have time to read the content of thousands of server to look for differences in the response content. But certainly we can have algorithms do this for us.
That's where this plugin comes in. It uses text clustering techniques to analyze the content of the responses and put them into clusters based on their similarity. It gives you another way to view the results of your intruder fuzzing attacks, and quickly identifying those server responses that are different.
Colonel Clustered provides two different algorithms to perform the response clustering. The default algorithm is relatively fast, with an optional deeper analysis algorithm that excels at spotting outliers, but doesn't scale well. Well, neither scales well, I would hesitate to throw 50k responses at this plugin. The faster/default algorithm is O(n^2) in complexity, whereas the deep analysis algorithm is O(n^3). So keep that in mind as you send intruder results to it to process.
Content-Aware Tokenization: The extension first inspects the Content-Type header of each response to apply the most intelligent tokenization strategy:
Pre-Grouping: To remain fast even with thousands of responses, the extension performs a single pass to group all perfectly identical responses. It calculates a hash of each response's token set and groups all items that share the same hash. This means the expensive clustering algorithm only has to run on the much smaller set of unique response bodies.
Dual Clustering Algorithms: Colonel Clustered offers two distinct clustering algorithms:
Fast Scan (Default): A high-performance DBSCAN-based algorithm runs automatically when you send multiple request/responses to the extension.
Deep Analysis (Manual Trigger): The original, more computationally intensive hierarchical clustering algorithm is available via a "Deep Analysis" button. This option is designed for scenarios requiring a more granular and potentially different clustering perspective, utilizing Average Linkage for improved cluster cohesion.
Outlier Consolidation: After clustering, any resulting group containing only a single unique member is considered an outlier. All such outliers are then consolidated into a single, convenient "Outliers" group in the UI.
These two algorithms should provide an easy way to automatically identify server responses that differ in content, even when the response size is an unreliable measure of response differences.
Load the Extension:
Send Responses for Analysis:
Perform Deep Analysis (Optional):
Analyze the Results in the Quad-Pane UI:
Request/Response Pair: The original index of the item.Status Code: The HTTP response status code.Length: The length of the response body in bytes.Content-Type: The Content-Type header of the response.
This project uses Gradle. You need JDK version 17 installed to build the plugin.
git clone <repository-url>
cd ColonelClustered
./gradlew build
build/libs/ColonelClustered.jar.Drew Kirkpatrick
@hoodoer
[email protected]
You can find me over at Blackthorne Consulting.
This project is released into the public domain under the Unlicense. See the LICENSE file for details.