
Benchmark and evaluation harness testing whether LLM agents resist malicious instructions hidden in multimodal skill images, with 108 cases across five risk categories and ASR/TSR scoring.
An agent skill for Claude Code, Codex, and similar coding-agent environments. It generates Xiaohongshu / Rednote carousel images, Live Photo motion cards, and WeChat 21:9 + 1:1 cover pairs from articles, copy, screenshots, product notes, subtitles, photos, or user-supplied videos.
Two visual systems share one workflow:
Sister project to guizang-ppt-skill. Shared visual language, separate maintenance. PPT solves "horizontal swipe talks"; this one solves "static feed images."

npx skills add https://github.com/op7418/guizang-social-card-skill --skill guizang-social-card-skill
Or paste this to an AI agent with shell access:
Install guizang-social-card-skill for me. Clone https://github.com/op7418/guizang-social-card-skill into ~/.claude/skills/guizang-social-card-skill, then verify that SKILL.md, assets/, and references/ exist.
If you already installed it, update with:
Update guizang-social-card-skill for me. Go to ~/.claude/skills/guizang-social-card-skill, run git pull, then tell me the latest commit.
Then ask your agent:
Make me a Swiss-style Xiaohongshu carousel from this article, 5 cards, IKB blue.
Other useful prompts:
Make me a 3:4 Xiaohongshu set from this product review, with editorial-style titles.
Turn this article into a WeChat cover pair: 21:9 hero + 1:1 share card, visually consistent.
I have 3 camping photos — make me an image-led Xiaohongshu carousel.
Turn this game guide copy into a Xiaohongshu set; pull some game art from Wallhaven.
I have a coffee video — make it a 5-second Xiaohongshu Live Photo card with text in a quiet zone.
Turn these three game clips into a triple Live Photo collage with Swiss-style guide copy.
.poster.xhs 1080×1440 (Xiaohongshu 3:4), .poster.wide 2100×900 (WeChat 21:9), .poster.square 1080×1080 (WeChat 1:1)5s for Xiaohongshu, 3s for WeChat Official Account articlesM01-M16, including Image-Led Cover, Pipeline, Before/After) + 12 Swiss (S01-S12, including KPI Tower, H-Bar Chart, Matrix + Hero)SOURCES.md.frame-shot / .device-browser / .device-phone utilitiesvalidate-social-deck.mjs auto-detects overflow, type cap violations, 4-band density gaps, and footer collisionsnode render.mjs outputs PNG directlyIn this skill, Live Photo means "put the user's video into a social-card layout." It is not random stock-video sourcing and not long-form video editing. First make the first frame work as a still card; then let 3s/5s of motion add evidence.
| Layout | Effect | Best for |
|---|---|---|
| Single-video motion card | One video cropped to full-canvas 3:4; when text is needed, use one short headline only and follow M16 Image-Led Cover / image-overlay rules for subject safety, quiet zones, type, and localized tint | Coffee, travel, fitness moves, product states, game moments |
| Two-grid / three-grid / four-grid puzzle | Multiple video wells inside one Live Photo, usually with no added text so strong footage leads | Travel scenery, fashion picks, room details, food process, workout actions |
| Triple Live Photo collage | Three short clips in parallel, increasing information density inside the short duration; add at most one real scene headline when needed | Guide steps, before/after, multi-angle demos, model test results |
| Long-video diagnosis | Sparse frames / contact sheet for long source videos, then recommend trimming, speed-up, splitting, or asking the user for an exact range | 1-3 minute user videos without a chosen moment |
| Publishing package | Debug JPG + MOV plus an iPhone-friendly .pvt package | Xiaohongshu and WeChat Official Account article Live Photos |
Judge the information budget first: WeChat 3s is best for one action point or one state change; Xiaohongshu 5s can hold one compact process; triple collages are for three parallel results, not a long sequential story.