
A guide on how to write fast and memory friendly YARA rules
Writing efficient YARA rules is essential for maintaining fast and accurate scanning performance. This guide provides key principles and best practices to help you optimize your rules, reduce unnecessary computation, and avoid common pitfalls. It incorporates insights from industry experts, including Victor M. Alvarez, WXS, and contributions from the YARA community.
This section provides a concise summary of YARA performance best practices. For detailed explanations and examples, refer to the full guide below.
"Think of YARA as a two-step process: first, searching for all patterns listed in the strings, and second, evaluating the conditions. You can’t use well-formed conditions to make up for poorly chosen strings."
— Wesley Shields
YARA follows four main steps when scanning a file:
YARA searches for strings first, making string selection the single most important factor for rule efficiency.
✅ Best Practices for Strings:
\x00\x00\x00\x00 appears too frequently.nocase carefully – It generates exponentially more search variations.YARA evaluates conditions sequentially and stops at the first failure.
✅ Best Practices for Conditions:
filesize < X) before expensive conditions.for all i in (1..filesize) is inefficient).@) instead of regex for sequence checks.⚠ Note: Regex conditions do not short-circuit and are always evaluated last.
Modules like pe, elf, or magic must parse the entire file before evaluation, increasing scan time.
✅ Alternatives:
pe.is_pe, use uint16(0) == 0x5A4D to identify PE files.Excessive matches slow down scanning and may trigger "too many matches" errors.
✅ Fixing Inefficient Matches:
.*, .+, or {x,} without an upper bound.@herrcore has created a helpful video tutorial covering the topics discussed in this performance guide.
Introduction Into YARA - Writing Efficient YARA Rules
To get a better grip on what and where YARA performance can be optimized, it's useful to understand the scanning process. It's basically separated into 4 steps which will be explained very simplified using this examples rule:
import "math"
rule example_php_webshell_rule
{
meta:
description = "Just an example php webshell rule"
date = "2021/02/16"
strings:
$php_tag = "<?php"
$input1 = "GET"
$input2 = "POST"
$payload = /assert[\t ]{0,100}\(/
condition:
filesize < 20KB and
$php_tag and
$payload and
any of ( $input* ) and
math.entropy(500, filesize-500) >= 5
}
This step happens before the actual scan. YARA will look for so called atoms in the search strings to feed the Aho-Corasick automaton. The details are explained in the chapter atom but for now it's enough to know, that they're maximum 4 bytes longs and YARA picks them quite cleverly to avoid too many matches. In our example YARA might pick the following 4 atoms:
<?phGETPOSTsser (out of assert)Here the scan has started. Steps 2.-4. will be executed on all files. YARA will look in each file for the 4 atoms defined above with prefix tree called Aho-Corasick automaton. Any matches are handed over to the bytecode engine.
If there's e.g. a match on sser, YARA will check if it was prefixed by an a and continues with a t. If that is true, it will follow on with the regex [\t ]{0,100}\(. With this clever approach YARA avoids going with a slow regex engine over the complete files and just picks certain parts to look closer.
After all pattern matching is done, the conditions are checked.
YARA has another optimization mechanism to only do the CPU intense math.entropy check from our example rule, if the 4 conditions before it are satisfied. Explained in more details in the chapter Conditions and Short-Circuit Evaluation
If the conditions are satisfied, a match is reported. The scan continues with the next file in step 2.
YARA extracts from the strings short substrings up to 4 bytes long that are called "atoms". Those atoms can be extracted from any place within the string, and YARA searches for those atoms while scanning the file, if it finds one of the atoms then it verifies that the string actually matches.
For example, consider this strings:
/abc.*cde/
=> possible atoms are abc and cde, either one or the other can be used The abc atom is currently preferred because they have the same quality and it is the first of the two.
/(one|two)three/
=> possible atoms are one, two, thre and hree, we can search for thre (or hree) alone, or for both one and two. Atom thre is preferred because it will lead to less potential matches then one and two (these are shorter) and it does not contain double e (more unique letter the better).
YARA does its best effort to select the best atoms from each string, for example:
{ 00 00 00 00 [1-4] 01 02 03 04 }
=> here YARA uses the atom 01 02 03 04, because 00 00 00 00 is too common
{ 01 02 [1-4] 01 02 03 04 }
=> 01 02 03 04 is preferred over 01 02 because it's longer
So, the important point is that strings should contain good atoms. These are bad strings because they contain either too short or too common atoms:
{00 00 00 00 [1-2] FF FF [1-2] 00 00 00 00}
{AB [1-2] 03 21 [1-2] 01 02}
/a.*b/
/a(c|d)/
The worst strings are those that don't contain any atoms at all, like:
/\w.*\d/
/[0-9]+\n/
This regular expression don't contain any fixed substring that can be used as atom, so it must be evaluated at every offset of the file to see if it matches there.
Another good import recommendation is to avoid for loops with too many iterations, specially of the statement within the loop is too complex, for example:
strings:
$a = {00 00}
condition:
for all i in (1..#a) : (@a[i] < 10000)
This rule has two problems. The first is that the string $a is too common, the second one is that because $a is too common #a can be too high and can be evaluated thousands of times.
This other condition is also inefficient because the number of iterations depends on filesize, which can be also very high: