Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
collisions — Hash collisions and their exploitations | Kitploit
Tools/GitHubGitHub/decalage2/collisions
ExploitationHash AnalysisCryptographyBinary AnalysisLearning & Education
GitHubdecalage2/collisions

collisions

Hash collisions and their exploitations

View Repository
91304 years agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share

TL;DR getting an MD5 collision of these two images is now(*) trivial and instant.

MD5 page on Wikipedia ⟷ <a href=http://gunshowcomic.com/648>

Don't play with fire, don't rely on MD5.

(*) Colliding any pair of files has been possible for many years, but it takes several hours each time, with no shortcut. This page provide tricks specific to file formats and precomputed collision prefixes to make collision instant. git clone. Run Script. Done.

Hash collisions and exploitations

By Ange Albertini and Marc Stevens.

  • Introduction
  • Status
  • Attacks
    • Identical prefix
      • FastColl (MD5)
      • UniColl (MD5)
      • Shattered (SHA1)
    • Chosen-prefix collisions
      • HashClash (MD5)
      • Shambles (SHA1)
    • Attacks summary
  • Exploitations
    • Standard strategy
      • JPG
        • custom scans
      • PNG
        • incompatibility
      • GIF
      • GZIP
      • Portable Executable
      • MP4 and others
        • JPEG2000
      • PDF
        • JPG in PDF
      • ZIP
        • Zip-based formats
    • Uncommon strategies
      • MultiColls: multiple collisions chain
      • Validity
      • PolyColls: collisions of different file types
        • PE - JPG
        • PDF - PE
        • PDF - PNG
      • PileUps (multi-collision)
        • PE - PNG - MP4 - PDF
    • Use cases
      • Gotta collide 'em all!
      • Incriminating files
    • Failures
      • ELF
      • Mach-O
      • Java Class
      • TAR
    • Exploitations summary
    • Test files
  • References
  • Credits
  • Conclusion

Introduction

The goal is to explore extensively existing attacks - and show on the way how weak MD5 is (instant collisions of any JPG, PNG, PDF, MP4, PE...) - and also explore in detail common file formats to determine how they can be exploited with present or with future attacks.

Indeed, the same file format trick can be used on several hashes (the same JPG tricks were used for MD5, malicious SHA-1 and SHA1), as long as the collisions follow the same byte patterns.

This document is not about new attacks (the most recent one was documented in 2012), but about new forms of exploitations of existing attacks.

Status

Current status - as of December 2018 - of known attacks:

  • get a file to get another file's hash or a given hash: impossible

    • it's still even not practical with MD2.
    • works for simpler hashes(*)
  • get two different files with the same MD5: instant

    • examples: 1 ⟷ 2
  • make two arbitrary files get the same MD5: a few hours (72 hours.core)

    • examples: 1 ⟷ 2
  • make two arbitrary files of specific file formats (PNG, JPG, PE...) get the same MD5: instant

    • read below
  • get two different files with the same SHA1: 6500 years.core

    • get two different PDFs with the same SHA-1 to show a different picture: instant (the prefixes are already computed)

(*) example with crypt - thanks Sven!

>>> import crypt
>>> crypt.crypt("5dUD&66", "br")
'brokenOz4KxMc'
>>> crypt.crypt("O!>',%$", "br")
'brokenOz4KxMc'

Attacks

MD5 and SHA1 work with blocks of 64 bytes.

If two contents A & B have the same hash, then appending the same contents C to both will keep the same hash.

hash(A) = hash(B) -> hash(A + C) = hash(B + C)

Collisions work by inserting at a block boundary a number of computed collision blocks that depends on what came before in the file. These collision blocks are very random-looking with some minor differences (that follow a specific pattern for each attack) and they will introduce tiny differences while eventually getting hashes the same value after these blocks.

These differences are abused to craft valid files with specific properties.

File formats also work top-down, and most of them work by byte-level chunks.

Some 'comment' chunks can be inserted to align file chunks to block boundaries, to align specific structures to collision blocks differences, to hide the rest of the collision blocks randomness from the file parsers, and to hide otherwise valid content from the parser (so that it will see another content).

These 'comment' chunks are often not officially real comments: they are just used as data containers that are ignored by the parser (for example, PNG chunks with a lowercase-starting ID are ancillary, not critical).

Most of the time, a difference in the collision blocks is used to modify the length of a comment chunk, which is typically declared just before the data of this chunk: in the gap between the smaller and the longer version of this chunk, another comment chunk is declared to jump over one file's content A. After this file content A, just append another file content B.

Since file formats usually define a terminator that will make parsers stop after it, A will terminate parsing, which will make the appended content B ignored.

So typically at least two comments are needed - often three:

  1. alignment
  2. hide collision blocks
  3. hide one file content (for re-usable collisions)

These common properties of file formats make it possible - they are not typically seen as weaknesses, but they can be detected or normalized out:

  • dummy chunks - used as comments
  • more than one comment
  • huge comments (lengths: 64b for MP4, 32b for PNG -> trivial collisions. 16b for JPG, 8b for GIF -> no generic collision for GIF, limited for JPG)
  • store any data in a comment (ASCII or UTF8 could be enforced)
  • store anything after the terminator (usually used only for malicious purposes) - can be avoided by using two comments finishing at the same offsets.
  • no integrity check. CRC32 in PNG are usually ignored. However they can be all correct since the collision blocks declare chunks of different lengths - so even if the chunk's data starts differently, the chunk lengths are different
  • flat structure: ASN.1 defines parent structure with the length of all the enclosed substructures, which prevents these constructs: you'd need to abuse a length, but also the length of the parent.
  • put a comment before the header - this makes generic re-usable collisions possible.

Identical prefix

Download Tool