
Semantic search over videos using Gemini Embedding 2 or Qwen3-VL.
Semantic search over video footage. Type what you're looking for, get a trimmed clip back.
[!IMPORTANT] Official source: github.com/ssrajadh/sentrysearch is the only official home of SentrySearch. Other sites republishing or mirroring this project are not affiliated with or endorsed by the maintainer, always download from this repository.
Languages: English · 简体中文
New: MLX backend for Apple Silicon: runs the local 2B model twice as fast in under half the memory, at the same accuracy
SentrySearch splits your videos into overlapping chunks, embeds each chunk as video using Google's Gemini Embedding API, Alibaba DashScope (qwen-cloud), or a local Qwen3-VL model, and stores the vectors in a local ChromaDB database. When you search, your text query (or image, see search by image) is embedded into the same vector space and matched against the stored video embeddings. The top match is automatically trimmed from the original file and saved as a clip.
macOS/Linux:
curl -LsSf https://astral.sh/uv/install.sh | sh
Windows:
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
git clone https://github.com/ssrajadh/sentrysearch.git
cd sentrysearch
uv tool install .
Requires Python 3.11 or 3.12 (PyTorch wheels don't yet support 3.13+). If your default Python is newer, install a managed 3.12 and pin the tool install:
uv python install 3.12 uv tool install --python 3.12 .
--backend local, --backend mlx, or --backend qwen-cloud with DASHSCOPE_API_KEY in .env.sentrysearch init
This prompts for your Gemini API key, writes it to .env, and validates it with a test embedding.
sentrysearch index /path/to/footage
sentrysearch search "red truck running a stop sign"
ffmpeg is required for video chunking and trimming. If you don't have it system-wide, the bundled imageio-ffmpeg is used automatically.
Manual setup: If you prefer not to use
sentrysearch init, you can copy.env.exampleto.envand add your key from aistudio.google.com/apikey manually.
$ sentrysearch init
Enter your Gemini API key (get one at https://aistudio.google.com/apikey): ****
Validating API key...
Setup complete. You're ready to go — run `sentrysearch index <directory>` to get started.
If a key is already configured, you'll be asked whether to overwrite it.
Tip: Set a spending limit at aistudio.google.com/billing to prevent accidental overspending.
$ sentrysearch index /path/to/video/footage
Indexing file 1/3: front_2024-01-15_14-30.mp4 [chunk 1/4]
Indexing file 1/3: front_2024-01-15_14-30.mp4 [chunk 2/4]
...
Indexed 12 new chunks from 3 files. Total: 12 chunks from 3 files.
Options:
--chunk-duration 30 — seconds per chunk--overlap 5 — overlap between chunks--no-preprocess — skip downscaling/frame rate reduction (send raw chunks)--target-resolution 480 — target height in pixels for preprocessing--target-fps 5 — target frame rate for preprocessing--no-skip-still — embed all chunks, even ones with no visual change--rpm 10 — cap requests per minute to the cloud API (details below)--backend local — use a local model instead of Gemini (details below)--backend mlx — use the local model through MLX on Apple Silicon, faster and lighter than local (details below)$ sentrysearch search "red truck running a stop sign"
#1 [0.87] front_2024-01-15_14-30.mp4 @ 02:15-02:45
#2 [0.74] left_2024-01-15_14-30.mp4 @ 02:10-02:40
#3 [0.61] front_2024-01-20_09-15.mp4 @ 00:30-01:00
Saved clip: ./match_front_2024-01-15_14-30_02m15s-02m45s.mp4
If the best result's similarity score is below the confidence threshold (default 0.41, or 0.35 on --backend mlx), you'll be prompted before trimming:
No confident match found (best score: 0.28). Show results anyway? [y/N]:
With --no-trim, low-confidence results are shown with a note instead of a prompt.
Options: --results N, --output-dir DIR, --no-trim to skip auto-trimming, --threshold 0.5 to adjust the confidence cutoff, --save-top N to save the top N clips instead of just the best match, --dedupe to set how similar a result can be to a higher-ranked pick before it's dropped (on by default, so overlapping chunks of the same moment don't fill the list), and --rerank to ask a VLM to re-rank the returned candidates before trimming. Backend and model are auto-detected from the index — pass --backend or --model only to override.
# Save top 5 clips, dropping only near-identical chunks
sentrysearch search "red truck" --save-top 5 --dedupe 0.95
# Re-rank the top 10 embedding matches with a VLM before trimming
sentrysearch search "pedestrian crossing behind the car" --rerank --results 10
The --dedupe value is a cosine similarity ceiling (0–1). Any result whose similarity to an already-kept higher-ranked result exceeds this value is dropped. Lower values are stricter: 0.8 requires results to be very distinct, 0.95 only removes near-identical chunks. The default is 0.9; pass --dedupe 1 to keep every result. Search fetches extra candidates when deduping, so you still get the number of results you asked for.
--rerank extracts each returned candidate clip, sends it to a VLM with the query, and sorts likely visual matches ahead of embedding-only results. Gemini and qwen-cloud searches use Gemini 2.5 Flash for reranking; local searches use a local Qwen3-VL Instruct reranker. If reranking cannot run or a candidate cannot be scored, SentrySearch keeps the embedding-ranked results instead of failing the search.