
Hash collisions and their exploitations
TL;DR getting an MD5 collision of these two images is now(*) trivial and instant.
⟷
<a href=http://gunshowcomic.com/648>
Don't play with fire, don't rely on MD5.
(*) Colliding any pair of files has been possible for many years, but it takes several hours each time, with no shortcut.
This page provide tricks specific to file formats and precomputed collision prefixes to make collision instant.
git clone. Run Script. Done.
By Ange Albertini and Marc Stevens.
The goal is to explore extensively existing attacks - and show on the way how weak MD5 is (instant collisions of any JPG, PNG, PDF, MP4, PE...) - and also explore in detail common file formats to determine how they can be exploited with present or with future attacks.
Indeed, the same file format trick can be used on several hashes (the same JPG tricks were used for MD5, malicious SHA-1 and SHA1), as long as the collisions follow the same byte patterns.
This document is not about new attacks (the most recent one was documented in 2012), but about new forms of exploitations of existing attacks.
Current status - as of December 2018 - of known attacks:
get a file to get another file's hash or a given hash: impossible
get two different files with the same MD5: instant
make two arbitrary files get the same MD5: a few hours (72 hours.core)
make two arbitrary files of specific file formats (PNG, JPG, PE...) get the same MD5: instant
get two different files with the same SHA1: 6500 years.core
(*) example with crypt - thanks Sven!
>>> import crypt
>>> crypt.crypt("5dUD&66", "br")
'brokenOz4KxMc'
>>> crypt.crypt("O!>',%$", "br")
'brokenOz4KxMc'
MD5 and SHA1 work with blocks of 64 bytes.
If two contents A & B have the same hash, then appending the same contents C to both will keep the same hash.
hash(A) = hash(B) -> hash(A + C) = hash(B + C)
Collisions work by inserting at a block boundary a number of computed collision blocks that depends on what came before in the file. These collision blocks are very random-looking with some minor differences (that follow a specific pattern for each attack) and they will introduce tiny differences while eventually getting hashes the same value after these blocks.
These differences are abused to craft valid files with specific properties.
File formats also work top-down, and most of them work by byte-level chunks.
Some 'comment' chunks can be inserted to align file chunks to block boundaries, to align specific structures to collision blocks differences, to hide the rest of the collision blocks randomness from the file parsers, and to hide otherwise valid content from the parser (so that it will see another content).
These 'comment' chunks are often not officially real comments: they are just used as data containers that are ignored by the parser (for example, PNG chunks with a lowercase-starting ID are ancillary, not critical).
Most of the time, a difference in the collision blocks is used to modify the length of a comment chunk,
which is typically declared just before the data of this chunk:
in the gap between the smaller and the longer version of this chunk,
another comment chunk is declared to jump over one file's content A.
After this file content A, just append another file content B.

Since file formats usually define a terminator that will make parsers stop after it,
A will terminate parsing, which will make the appended content B ignored.
So typically at least two comments are needed - often three:
These common properties of file formats make it possible - they are not typically seen as weaknesses, but they can be detected or normalized out: