
Optimized HQC post-quantum KEM implementations for AVX2 and AVX-512 with GFNI, plus a unified benchmarking and NIST KAT suite for x86 IoT gateways.
This repository contains the optimized HQC implementations and benchmarking code accompanying the paper "Accelerating HQC for Post-Quantum TLS 1.3 on x86 IoT Gateways". It is a fork of the official HQC repository, extended with AVX-512 + GFNI implementations and a unified measurement suite that builds every compared implementation from the same code base, compiler flags, and benchmark protocol.
All implementations pass the NIST known-answer tests (KAT) for HQC-1, HQC-3, and HQC-5, and keep the input-independent control flow and memory-access pattern of the reference implementation.
Each variant under src/x86_64/ is selected at configure time with
-DHQC_X86_IMPL=<variant> and builds into its own library and test binaries.
| Variant | Paper label | ISA | Description |
|---|---|---|---|
avx256 | HQC reference | AVX2 | official optimized implementation (baseline) |
avx256_TC_Jang | Jang et al. | AVX2 | GFmulOpt Toom-Cook/Karatsuba multiplication |
avx256_TC_GFNI | Ours (AVX2) | AVX2+GFNI | our GFNI Reed-Solomon / Reed-Muller decoders on the AVX2 build |
avx256_FAFFT_GFNI_chen | Chen et al. | AVX2+GFNI | FAFFT multiplication, original non-cached flow |
avx256_FAFFT_GFNI_chen_x2 | Chen et al. (dagger) | AVX2+GFNI | FAFFT with the shared-operand ring_mul_x2 flow |
avx256_FAFFT_GFNI_ours | Ours FAFFT (AVX2) | AVX2+GFNI | FAFFT + zero-tail basis conversion + our decoders |
avx512_Robert | Robert–Véron | AVX-512 | Robert–Véron AVX-512 polynomial multiplication |
avx512_Cabral | Cabral et al. | AVX-512+GFNI | source code from Cabral et al. (includes their upstream HQC-3 hash_j fix of 2026-07-21) |
avx512_TC_GFNI | Ours TC/Karat (AVX-512) | AVX-512+GFNI | AVX-512 Toom-Cook/Karatsuba, GFNI decoders, AVX-512 Keccak |
avx512_FAFFT_GFNI | Ours FAFFT (AVX-512) | AVX-512+GFNI | AVX-512 port of Chen et al.'s AVX2 implementation |
The two Ours builds of each ISA differ only in the polynomial multiplier, so
their end-to-end comparison isolates the multiplier choice.
Requirements: GCC >= 12, CMake >= 3.21, an x86-64 CPU. The AVX-512 variants require AVX-512F/BW/DQ/VL/VBMI/VBMI2/IFMA/BITALG/VPOPCNTDQ, VPCLMULQDQ, and GFNI (Ice Lake / Tiger Lake / Rocket Lake / Zen 4 or newer).
# one variant
mkdir build-avx512_TC_GFNI && cd build-avx512_TC_GFNI
cmake -DHQC_ARCH=x86_64 -DHQC_X86_IMPL=avx512_TC_GFNI ..
cmake --build . -j
# or every variant at once (creates build-<variant>/ directories)
./build_all.sh
# from the repository root, after building
bash scripts/run_kat.sh # all variants
bash scripts/run_kat.sh avx512_TC_GFNI # one variant
Note: test_kat_hqc_* prints nothing on success and lines containing bad
on a mismatch; the script greps the output accordingly.
The paper's protocol: Turbo Boost and SMT disabled, the process pinned to one
physical core, and every reported value the median of 1,000 samples of 100
iterations after 10,000 warm-up iterations. The scripts pin to core 0 with
taskset.
bash scripts/measure_build.sh <variant> [...] # end-to-end KEM benchmark of the given variants
Raw logs are written to benchmark_results/raw/; the benchmark reports
min/median/mean, and the median is the value quoted in the paper.
src/common/ parameter-set independent HQC code (KEM layer, coding)
src/ref/ reference implementation
src/x86_64/common/ AVX2-era shared code
src/x86_64/<variant> one directory per compared implementation (see table)
lib/fips202/ Keccak (scalar PQClean + optional XKCP AVX-512 permutation)
tests/ KAT, API, unit tests, benchmarks
kats/ known-answer files used by the KAT tests
scripts/ KAT and measurement helper scripts
The base repository and our additions are released into the public domain
(see LICENSE). This repository additionally vendors third-party
implementations for measurement purposes; see
THIRD_PARTY_NOTICES.md for their publications,
origins, and licenses.