
Collisions de hachage et leurs exploitations
TL;DR obtenir une collision MD5 de ces deux images est désormais(*) trivial et instantané.
⟷
<a href=http://gunshowcomic.com/648>
Ne jouez pas avec le feu, ne vous fiez pas à MD5.
(*) Collisionner n'importe quelle paire de fichiers est possible depuis de nombreuses années, mais cela prend plusieurs heures à chaque fois, sans raccourci.
Cette page fournit des astuces spécifiques aux formats de fichiers et des préfixes de collision précalculés pour rendre la collision instantanée.
git clone. Exécutez le script. Terminé.
Par Ange Albertini et Marc Stevens.
L'objectif est d'explorer en profondeur les attaques existantes - et de montrer au passage à quel point MD5 est faible (collisions instantanées de n'importe quel JPG, PNG, PDF, MP4, PE...) - et aussi d'explorer en détail les formats de fichiers courants pour déterminer comment ils peuvent être exploités avec des attaques présentes ou futures.
En effet, la même astuce de format de fichier peut être utilisée sur plusieurs hashs (les mêmes astuces JPG ont été utilisées pour MD5, SHA-1 malveillant et SHA1), tant que les collisions suivent les mêmes motifs d'octets.
Ce document ne traite pas de nouvelles attaques (la plus récente a été documentée en 2012), mais de nouvelles formes d'exploitation des attaques existantes.
État actuel - en décembre 2018 - des attaques connues :
obtenir un fichier qui a le hash d'un autre fichier ou un hash donné : impossible
obtenir deux fichiers différents avec le même MD5 : instantané
faire en sorte que deux fichiers arbitraires aient le même MD5 : quelques heures (72 heures.core)
faire en sorte que deux fichiers arbitraires de formats spécifiques (PNG, JPG, PE...) aient le même MD5 : instantané
obtenir deux fichiers différents avec le même SHA1 : 6500 ans.core
(*) exemple avec crypt - merci Sven !```
import crypt crypt.crypt("5dUD&66", "br") 'brokenOz4KxMc' crypt.crypt("O!>',%$", "br") 'brokenOz4KxMc'
# Attaques
MD5 et SHA1 fonctionnent avec des blocs de 64 octets.
Si deux contenus A et B ont le même hachage, alors ajouter le même contenu C aux deux conservera le même hachage.``` text
hash(A) = hash(B) -> hash(A + C) = hash(B + C)
Collisions work by inserting at a block boundary a number of computed collision blocks that depends on what came before in the file. These collision blocks are very random-looking with some minor differences (that follow a specific pattern for each attack) and they will introduce tiny differences while eventually getting hashes the same value after these blocks.
These differences are abused to craft valid files with specific properties.
File formats also work top-down, and most of them work by byte-level chunks.
Some 'comment' chunks can be inserted to align file chunks to block boundaries, to align specific structures to collision blocks differences, to hide the rest of the collision blocks randomness from the file parsers, and to hide otherwise valid content from the parser (so that it will see another content).
These 'comment' chunks are often not officially real comments: they are just used as data containers that are ignored by the parser (for example, PNG chunks with a lowercase-starting ID are ancillary, not critical).
Most of the time, a difference in the collision blocks is used to modify the length of a comment chunk,
which is typically declared just before the data of this chunk:
in the gap between the smaller and the longer version of this chunk,
another comment chunk is declared to jump over one file's content A.
After this file content A, just append another file content B.

Since file formats usually define a terminator that will make parsers stop after it,
A will terminate parsing, which will make the appended content B ignored.
So typically at least two comments are needed - often three: