Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

··Feeds·Contact·Privacy·© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
fnprint — match functions in binaries by what they do, not what their bytes look like. behavioral function fingerprinting via microexecution. | Kitploit
Tools/GitHubGitHub/1rhino2/fnprint
Vulnerability AnalysisDynamic Code Analysis (DAST)Reverse EngineeringMalware AnalysisBinary AnalysisFirmware Analysis
GitHub1rhino2/fnprint

fnprint

match functions in binaries by what they do, not what their bytes look like. behavioral function fingerprinting via microexecution.

View Repository
5345311 days agoNot yet reviewed
Website

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share

fnprint

fnprint matches functions in binaries by what they do, not by what their bytes or control-flow graphs look like. it runs each function in a tiny emulator with made-up inputs, records the side effects it produces, and hashes that behavior into a fingerprint. two functions that behave the same get similar fingerprints, even if they were built by a different compiler, at a different optimization level, or for a different CPU.

the point of doing it this way: byte signatures (FLIRT, FunctionID) break the moment code is recompiled, and CFG matchers (BinDiff, Diaphora) get shaky across -O0 vs -O3 and give up across architectures. behavior survives all of that a lot better. no training data, no model, one static binary.

three things it does that the others dont:

  • it is safe to point at malware. the emulator runs in a privilege-separated worker under a seccomp jail that forbids exec, sockets, file opens, and executable memory outright (unicorn is built as a no-JIT interpreter so it never needs one). a crafted binary that pops the emulator gets a process that can compute and nothing else. see SECURITY.md. BinDiff, Diaphora, BSim, FLIRT and the ML matchers all parse hostile bytes in-process.
  • it matches across architectures. index the x86-64 build you have symbols for, name the stripped aarch64 firmware. the effect model never sees the ISA. zlib x86-64 -O0 vs aarch64 -O0: 97.7% of functions named right.
  • triage gives you a verdict, not a diff. two corpora (known-vulnerable, patched), a margin, and a review queue of the functions that lean vulnerable. thats the n-day workflow as one command.

x86-64 and aarch64 ELF. see limits before you trust it.

show me

fnprint naming functions in a stripped, differently-compiled binary

point it at a stripped binary and a corpus of things you already have names for:

$ strip --strip-all mystery.so
$ nm mystery.so
nm: mystery.so: no symbols

$ fnprint index libz.so -o corpus.db          # a build you have symbols for
$ fnprint query mystery.so --corpus corpus.db --threshold 0.6
named 27 function(s):
  addr        sim    graph  name
  0x000022f9  100.0%     0%  adler32_z  (libz.so)
  0x00002a8d  100.0%     0%  compress2  (libz.so)
  0x00004500  100.0%     0%  deflateStateCheck  (libz.so)
  0x0000ad5d  100.0%     0%  inflateBackEnd  (libz.so)
  0x0000b69c  100.0%     0%  inflateStateCheck  (libz.so)
  0x00012f8c  100.0%     0%  _tr_tally  (libz.so)
  ...

that run is a fully stripped -O0 build named from an -O2 corpus. different optimization level, zero symbols left, 27 names come back and all 27 are right. it names what it is confident about and stays quiet about the rest. (the gif above is from 0.5, which found 9; 0.6 also recovers the static functions a stripped .so hides behind its exports.)

the same thing across CPUs. corpus from the x86-64 build, target is the aarch64 build of the same source, symbols gone:

$ fnprint index libz_x86_64.so -o corpus.db
$ fnprint query libz_arm64_stripped.so --corpus corpus.db
named 42 function(s):
  addr        sim    graph  name
  0x00001af4  100.0%     0%  adler32_z  (libz_x86_64.so)
  0x0000263c  100.0%     0%  compress2  (libz_x86_64.so)
  0x00002958  100.0%   100%  x2nmodp  (libz_x86_64.so)
  0x00002a98  100.0%     0%  crc32_z  (libz_x86_64.so)
  0x000034c4  100.0%   100%  crc32_combine64  (libz_x86_64.so)
  ...

42 named, 42 right, from a corpus built for a different CPU.

the other thing it does is diff two builds and tell you which functions changed behavior, which is handy when a vendor ships a new firmware and you want to know what actually moved. it works with or without symbol names: with them it aligns by name, without them it aligns the two builds on behavior plus the call graph, so two fully stripped firmware images still diff:

$ fnprint match old.so new.so
aligned 38 function pairs by behavior + call graph (no symbol names needed)
  unchanged:  31
  changed:    7
  low-signal: 57 (too small to judge)

changed behavior (lowest similarity first):
   30.5%  compress_block
   43.0%  deflate_rle -> deflate_huff
   68.8%  deflate_slow

for actual n-day work there is triage. build one corpus from the known vulnerable version of a function and one from the patched version, then rank an unknown build against both. a function close to the vulnerable side and clearly separated from the patched side is what you want in front of a human, not just a single match score you have to interpret:

$ fnprint index vuln.so    -o vuln.db
$ fnprint index patched.so -o patched.db
$ fnprint triage mystery.so --vuln vuln.db --patched patched.db
43 functions triaged: 1 look vulnerable, 0 patched, 42 inconclusive

review queue (vuln-leaning, strongest first):
   addr         vuln%  patched%  margin  matches
   0x000022f9  100.0     27.3   +72.7  adler32_z vs crc32_z

the 42 functions that are identical in both versions come back inconclusive on purpose, they can't be pinned to either side and shouldn't be flagged. --margin and --min-sim control how hard the two sides have to separate before it commits.

how it works

for each function:

  • map the binary and jump to the function with junk in the argument registers.
  • any read from memory we did not set up returns a deterministic value and the page gets mapped on the fly. wild pointers never crash the run, and the same input always gives the same trace. this is Godefroid's microexecution trick.
  • calls to other functions get stubbed (recorded, then skipped) so we never dive into libc and the run stays about this function.
  • we log an arch-neutral stream of effects: which argument buffers and struct fields it reads and writes, what value classes it writes (a copy of an input, a small constant, a pointer), calls it makes, branches it takes, what it returns. absolute addresses are thrown away, only offsets and shapes are kept.
  • that stream gets turned into shingles and a minhash signature. similarity is the fraction of matching minhash slots, which estimates how much two functions' behavior overlaps. an LSH band index keeps queries from comparing everything against everything.
  • calls are named from the import table only (plt stubs, got slots), because that is the part that survives strip. internal calls are anonymous in the print and become call-graph edges instead, from a static sweep of the body.
  • naming and diffing are a global 1:1 alignment, not a best-hit-per-function lookup: a corpus function can be claimed once, best pair first, then the assignment is rescored twice with a call-graph term (how many of a pair's paired callers and callees are paired with each other). that is what tells two behavioral twins apart, and what stops 27 -O2 functions that all inline the same state check from all claiming it.

no training data, no model. the same idea shows up in the literature as Blanket Execution (Egele et al, USENIX Security 2014); fnprint is a practical, maintained take on it with a CLI you can actually use. the effect model is arch-neutral by construction, which is why x86-64 and aarch64 prints of the same source line up without anything learned.

install

needs a rust toolchain, cmake, and a C toolchain. unicorn (the no-JIT fork) and capstone are built from source by cargo, nothing to apt install.

cargo install --path cli
# or just
cargo build --release   # binary at target/release/fnprint

usage

Download Tool