AI/ML
Generative Media Skills
Multi-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI). High-quality image, video, and audio generation powered by muapi.ai.
npx skills add SamurAIGPT/Generative-Media-SkillsSkill Details
🎭 Generative Media Skills for AI Agents
The Ultimate Multimodal Toolset for Claude Code, Cursor, Gemini CLI, and OpenCode. A high-performance, schema-driven architecture for AI agents to generate, edit, and display professional-grade images, videos, and audio — powered by the muapi-cli.
🚀 Get Started | 🎬 Recipe Pack | 🎨 Expert Library | ⚙️ Core Primitives | 🤖 MCP Server | 📖 Reference
<p align="center"><a href="https://www.youtube.com/watch?v=SOXsxqnQGlc"><img src="https://i.ytimg.com/vi/SOXsxqnQGlc/maxresdefault.jpg" width="720"></a></p> <p align="center"><a href="https://www.youtube.com/watch?v=SOXsxqnQGlc"><b>▶ Watch: Best AI Video Generator (API) in 2026 (Quality, Price, Uncensored, Editing)</b></a></p>
Related Projects
- minimax-music-3-api — Python SDK for MiniMax Music 3.0 text-to-music generation on Muapi.
- awesome-minimax-music-3-prompts — Curated song prompts and lyrics-formatting guide for MiniMax Music 3.0.
- MiniMax-H3-API — Python SDK for MiniMax H3 video-generation workflows on Muapi.
- awesome-minimax-h3-prompts — Prompt gallery and runnable examples for the MiniMax H3 skills.
- Wan-3.0-API — Python SDK and MCP server for Wan 3.0-compatible video-generation workflows.
- Wan-3.0-Prime-API — Python SDK and MCP server for the higher-fidelity Wan 3.0 Prime tier.
- Open-Generative-AI — Free self-hosted AI media studio — GUI alternative to these skills for the same model set
- Awesome-GPT-Image-2-API-Prompts — Curated GPT-Image-2 prompts to use with these skills
- Awesome-Gemini-Omni-API-Prompts — Curated Gemini Omni prompts for video generation
- Gemini-Omni-1.1-Flash-API — Python SDK and MCP server for Google's newly announced Gemini Omni 1.1 Flash update
- AI-Voice-Agent — Self-hosted AI voice agent for real-time voice conversations, sales calls, and customer support
- awesome-ai-image-models — compare AI image models by API, price & quality
- flux-3-video-api — Python wrapper focused on FLUX 3 Text-to-Video and Image-to-Video
- ai-creator-academy — free curriculum teaching creators to monetize generative AI, built on these same skills
- Flux-3-Dev-API — Python wrapper for Black Forest Labs' FLUX 3 (Dev variant) — text-to-image, image-to-image, text-to-video, image-to-video
- Grok-Imagine-Image-2-API — Python SDK and MCP server for Grok Imagine Image 2.0 generation and editing through MuAPI
- midjourney-api — Python SDK for Midjourney V7, V8, and Niji image generation through MuAPI
- suno-api — Python SDK for Suno music, audio, and voice workflows through MuAPI
- awesome-flux-3-api-prompts — FLUX 3 API guide, prompts, and parameters
- seedance-2.5-mcp — MCP server for generating Seedance 2.5 Preview videos through MuAPI.
- seedance-2-mcp — MCP server for generating Seedance 2 videos through MuAPI.
- Text-to-Speech-API — narration and dialogue API examples for media workflows.
- Speech-to-Text-API — transcription and audio-understanding API examples.
- Voice-Cloning-API — consent-aware speaking and singing voice workflows.
- Image-Enhancement-API — image enhancement examples for creative pipelines.
- Video-Utilities-API — video upscaling and sound-generation utility examples.
- AI-3D-Model-API — 3D asset generation comparison and examples.
✨ Key Features
- 🤖 Agent-Native Design — CLI-powered scripts with structured JSON outputs, semantic exit codes, and
--jqfiltering for seamless agentic pipelines. - 🧠 Expert Knowledge Layer — Domain-specific skills that bake in professional cinematography, atomic design, and branding logic.
- ⚡ CLI-Powered Core — All primitives delegate to
muapi-cli— no curl, no JSON parsing, no boilerplate. - 🖼️ Direct Media Display — Use the
--viewflag to automatically download and open generated media in your system viewer. - 📁 Local File Support — Auto-upload images, videos, faces, and audio from your local machine to the CDN for processing.
- 🌈 100+ AI Models — One-click access to Midjourney v7, Flux Kontext, Seedance 2.0, Kling 3.0, Veo3, and more.
- 🔌 MCP Server — Run
muapi mcp serveto expose all 19 tools directly to Claude Desktop, Cursor, or any MCP-compatible agent.
🏗️ Scalable Architecture
This repository uses a Core/Library split to ensure efficiency and high-signal discovery for LLMs:
⚙️ Core Primitives (/core)
Thin wrappers around muapi-cli for raw API access.
core/media/— File uploadcore/edit/— Image editing (prompt-based)core/platform/— Setup, auth & result polling
📚 Expert Library (/library)
High-value skills that translate creative intent into technical directives.
- Cinema Director (
/library/motion/cinema-director/) — Technical film direction & cinematography. - Nano-Banana (
/library/visual/nano-banana/) — Reasoning-driven image generation (Gemini 3 Style). - UI Designer (
/library/visual/ui-design/) — High-fidelity mobile/web mockups (Atomic Design). - Logo Creator (
/library/visual/logo-creator/) — Minimalist vector branding (Geometric Primitives). - Seedance 2 (Doubao Video) (
/library/motion/seedance-2/) — Director-level cinematic video generation with text-to-video, image-to-video, and video extension with native audio-video sync. - AI Clipping (
/library/edit/ai-clipping/) — Long video → ranked vertical short clips in one managed API call. Server-side transcription, virality ranking, dedupe, and face-tracked auto-crop — no local Whisper or LLM. - YouTube Shorts (
/library/social/youtube-shorts/) — Platform-aware preset over AI Clipping (Shorts / TikTok / Reels / Feed defaults).
Plus 41 ready-to-run workflow recipes organized by output type — see 🎬 Recipe Pack below.
🎬 Recipe Pack
Forty-one LLM-orchestrated workflow recipes that combine multiple muapi-cli calls into named end-to-end pipelines (e.g. photo of person → 3D action figure, product photo → cinematic 10s ad). Each skill is a SKILL.md the agent reads and follows; bring your own consuming agent (Claude Code, Cursor, MCP) — these are recipes, not bash wrappers.
Motion / Video (16)
| Skill | Description |
|---|---|
| 3D Logo Animation | Transform a 2D logo into a premium 3D version and animate it with professional cinematic effects |
| AI Fight Scene Generator | High-cut-density action / fight scene — 16-cell storyboard image drives Seedance 2.0 i2v for shot-by-shot choreography |
| Animal Vlogger Video | Hilarious, ultra-realistic anthropomorphic-animal vlogger acting like a human in a real-world setting |
| Cartoon Dance Animation | Convert a photo into a Pixar-style 3D cartoon, then animate using a reference dance/motion video |
| Character Story Video | Multi-part animated story video — establish a consistent character then animate sequential scenes |
| Drone-Style Video | Aerial drone-perspective footage — bird's-eye sweeps, orbit shots, and flyover sequences |
| Giant Product Showcase | Dramatic giant-scale product visual (building-sized object next to a person), optionally animated |
| Jewelry Product Video | Luxury jewelry ad with high-end commercial cinematography and detailed macro animation |
| Music Video | Short music video from a song theme — keyframes, animation per beat, matching music track |
| One-Shot Video | Single continuous cinematic shot — no cuts, one seamless flowing scene |
| Cinematic Product Ad | Cinematic 5–10s product ad from a product photo + brand brief |
| Product Showcase Video | Dynamic product showcase with explosive ingredient arrangement + realistic motion animation |
| Product Video Ad Maker | High-end cinematic product video ad starting from a simple product photo |
| Talking Baby Video | Viral-style talking-baby video with custom costumes and scripts |
| UGC Lifestyle Try-On | UGC-style lifestyle photos & video of a person using your product — authentic, social-native |
| UGC Video Factory | Person photo + product photo + script → 10s vertical 9:16 UGC video ad with native dialogue (Nano-Banana Pro Edit → Seedance 2.0 VIP i2v) |
Social (5)
| Skill | Description |
|---|---|
| Instagram Post | Polished on-brand Instagram post — hero image + caption + hashtags |
| Product Campaign Pack | Full multi-channel campaign — hero visuals, social assets, short ad video, platform crops |
| RedNote Cover | Xiaohongshu (小红书) cover image — vibrant lifestyle aesthetic with typography overlay |
| Social Media Pack | Re-render a hero image into Instagram / TikTok / Shorts / X aspect ratios |
| UGC Ads Workflow | UGC video ad pipeline — combine selfie + product image, write script, animate |
Visual / Images & Design (21)
| Skill | Description |
|---|---|
| Action Figure Generator | Convert a photo of a person into a custom 3D action figure with collectible toy packaging |
| Ad Creative Set | High-converting ad set — hero image, copy variations, platform crops for Meta / Google / LinkedIn |
| Amazon Product Listing Pack | Full Amazon listing image set — hero, lifestyle, infographic, comparison/detail closeups |
| Blog Header | Professional 1200×628 blog header image with optional title composition guidance |
| Brand Kit | Cohesive brand visual kit — logo concept, color palette, typography pairings |
| Brochure Designer | Multi-page brochure — cover, inner spread, back — for business, real estate, events, launches |
| Couple Grid Creator | Stylized 6-box grid of a couple in romantic poses, each pose framed inside cardboard packaging |
| Brand Design Guide | Comprehensive design guide — palette, typography, UI components, visual identity rules |
| Fashion Try-On | Virtually try outfits by combining a person's photo + clothing item, optional fashion model video |
| Floor Plan Rendering | Design a 2D floor plan and convert into a realistic 3D architectural rendering |
| Interior Design | Pro interior design visualizations — redesign rooms, generate concepts, visualize furniture styles |
| Interior Design Visualizer | Generate an empty room and fill it with stylish furniture / decor; or redesign an existing room |
| Keyboard Art Maker | Artistic top-down photos of keyboard keycaps arranged to spell custom messages |
| Logo + Branding Package | Logo + full branding package — variations (dark/light/icon), palette, mockups |
| Logo Generator | Quick single-shot polished logo — fast, clean vector aesthetic with accurate brand-name text |
| Multi-Angle Reshoot | Re-render a subject from dramatic camera angles (fish-eye, bird's-eye, low, macro) — identity preserved |
| Multi-Angle Shots | Full multi-angle product shot set — front, side, back, top-down, 45° |
| Selfie with Celebrities | Realistic behind-the-scenes selfie of the user with a celebrity; optional cinematic long-take |
| Storyboard Generator | Generate N keyframes for a short story or scene sequence (image only, no video) |
| URL to Design | Analyze a website URL and generate a redesigned, improved UI with modern aesthetics |
| YouTube Thumbnail | High-CTR YouTube thumbnail — striking imagery, bold text placement, emotional face/subject |
Each recipe declares its inputs and a Steps body. Pass the inputs and let your agent execute the steps via muapi CLI calls (or raw API for endpoints that don't yet have a CLI alias — see the per-skill Notes for the Executing Agent footer).
🚀 Quick Start
1. Install the muapi CLI
The core scripts require muapi-cli. Install it once:
# via npm (recommended — no Python required)
npm install -g muapi-cli
# via pip
pip install muapi-cli
# or run without installing
npx muapi-cli --help
2. Configure Your API Key
# Interactive setup
muapi auth configure
# Or pass directly
muapi auth configure --api-key "YOUR_MUAPI_KEY"
# Get your key at https://muapi.ai/dashboard?utm_source=github&utm_medium=readme&utm_campaign=generative-media-skills
3. Install the Skills
# Install all skills to your AI agent
npx skills add SamurAIGPT/Generative-Media-Skills --all
# Or install a specific skill
npx skills add SamurAIGPT/Generative-Media-Skills --skill muapi-media-generation
# Install to specific agents
npx skills add SamurAIGPT/Generative-Media-Skills --all -a claude-code -a cursor
4. Generate Your First Image
muapi image generate "a cyberpunk city at night" --model flux-dev
# Download the result automatically
muapi image generate "a sunset over mountains" --model hidream-fast --download ./outputs
# Extract just the URL (agent-friendly)
muapi image generate "product on white bg" --model flux-schnell --output-json --jq '.outputs[0]'
5. Run an Expert Skill
# Use Nano-Banana reasoning to generate a 2K masterpiece
bash library/visual/nano-banana/scripts/generate-nano-art.sh \
--file ./my-source-image.jpg \
--subject "a glass hummingbird" \
--style "macro photography" \
--resolution "2k" \
--view
6. Direct a Cinematic Scene
cd library/motion/cinema-director
# Create a 10-second epic reveal
bash scripts/generate-film.sh \
--subject "a cybernetic dragon over Tokyo" \
--intent "epic" \
--model "kling-v3.0-pro" \
--duration 10 \
--view
# Animate a reference image into video
bash library/motion/seedance-2/scripts/generate-seedance.sh \
--mode i2v \
--file ./concept.jpg \
--subject "camera slowly pulls back to reveal the full landscape" \
--intent "reveal" \
--view
# Extend an existing video
bash library/motion/seedance-2/scripts/generate-seedance.sh \
--mode extend \
--request-id "YOUR_REQUEST_ID" \
--subject "camera continues pulling back to reveal the vast city" \
--duration 10
OpenCode
# Clone the repo and set the MUAPI_API_KEY env var
git clone https://github.com/SamurAIGPT/Generative-Media-Skills
export MUAPI_API_KEY=your_key_here
# Skills auto-load from .opencode/skills/ when you run opencode in this directory
opencode
🤖 MCP Server
Run muapi as a Model Context Protocol server so Claude Desktop, Cursor, or any MCP-compatible agent can call generation tools directly — no shell scripts needed.
muapi mcp serve
Claude Desktop config (~/Library/Application Support/Claude/claude_desktop_config.json):
{
"mcpServers": {
"muapi": {
"command": "muapi",
"args": ["mcp", "serve"],
"env": { "MUAPI_API_KEY": "your-key-here" }
}
}
}
This exposes 19 structured tools with full JSON Schema input/output definitions:
| Tool | Description |
|---|---|
muapi_image_generate | Text-to-image (14 models) |
muapi_image_edit | Image-to-image editing (11 models) |
muapi_video_generate | Text-to-video (13 models) |
muapi_video_from_image | Image-to-video (16 models) |
muapi_audio_create | Music generation (Suno) |
muapi_audio_from_text | Sound effects (MMAudio) |
muapi_enhance_upscale | AI upscaling |
muapi_enhance_bg_remove | Background removal |
muapi_enhance_face_swap | Face swap image/video |
muapi_enhance_ghibli | Ghibli style transfer |
muapi_edit_lipsync | Lip sync to audio |
muapi_edit_clipping | AI highlight extraction |
muapi_predict_result | Poll prediction status |
muapi_upload_file | Upload local file → URL |
muapi_keys_list | List API keys |
muapi_keys_create | Create API key |
muapi_keys_delete | Delete API key |
muapi_account_balance | Get credit balance |
muapi_account_topup | Add credits (Stripe checkout) |
⚡ Agentic Pipeline Examples
# Submit async, capture request_id, poll when ready
REQUEST_ID=$(muapi video generate "a dog running on a beach" \
--model kling-master --no-wait --output-json --jq '.request_id' | tr -d '"')
# ... do other work ...
muapi predict wait "$REQUEST_ID" --download ./outputs
# Pipe a prompt from another command
generate_prompt | muapi image generate - --model flux-dev
# Chain: upload → edit → download
URL=$(muapi upload file ./photo.jpg --output-json --jq '.url' | tr -d '"')
muapi image edit "make it look like a painting" --image "$URL" \
--model flux-kontext-pro --download ./outputs
📖 Schema Reference
This repository includes a streamlined schema_data.json that core scripts use at runtime to:
- Validate Model IDs: Ensures the requested model exists.
- Resolve Endpoints: Automatically maps model names to API endpoints.
- Check Parameters: Validates supported
aspect_ratio,resolution, anddurationvalues.
Discover all available models via the CLI:
muapi models list
muapi models list --category video --output-json
🔧 Compatibility
Optimized for the next generation of AI development environments:
- Claude Code — Direct terminal execution via tools + MCP server mode.
- Gemini CLI / Cursor / Windsurf — Seamless integration as local scripts.
- MCP — Full Model Context Protocol server with typed input/output schemas.
- CI/CD —
--output-json,--jq, semantic exit codes for scripting.
📄 License
MIT © 2026
muapi-3d-logo-animation
slug: muapi-3d-logo-animation name: muapi-3d-logo-animation version: "1.0.0" description: Transform a 2D logo into a premium 3D version and animate it with professional cinematic effects. acceptLicenseTerms: true
3D Logo Animation
Transform a 2D logo into a premium 3D version and animate it with professional cinematic effects.
Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
logo_image | image_url | yes | — | A clear 2D image of the logo to be converted to 3D. |
material_style | text | no | glossy glass and chrome | The material style for the 3D logo (e.g., gold, matte plastic, holographic). |
Steps
Phase A — 3D Logo Transformation
If {{logo_image}} is not provided, ask the user to upload their logo.
Once the logo is available, submit the plan with ONE step to convert it to 3D:
- 3D Logo Generation —
muapi image edit(model=nano-banana-2-edit):- Reference Image:
{{logo_image}} - Prompt:
Transform this 2D logo into a premium, high-quality 3D version. The logo should have depth and be made of {{material_style}}. Smooth edges, realistic reflections, and professional studio lighting. The logo is centered on a clean, minimal, out-of-focus background. High-end graphic design aesthetic, 8k resolution. - Aspect ratio: 1:1 or 4:3
- Reference Image:
Present the 3D logo to the user for approval.
Phase B — Cinematic Logo Animation
Once the 3D logo is ready, submit the plan to animate it:
- Logo Animation —
muapi video from-image(model=veo3.1-fast-image-to-video):- Reference Image: The 3D logo from Phase A.
- Prompt:
A professional cinematic logo reveal animation. The 3D logo rotates slowly with dynamic light sweeps reflecting off its {{material_style}} surface. Subtle camera movement, particle effects in the background, high-quality motion graphics style. - Aspect ratio: 16:9 or 1:1
After generation, present the final 3D logo animation to the user.
Trigger Keywords
3d logo, logo animation, 2d to 3d logo, animated logo, cinematic logo
Notes for the Executing Agent
- This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call
muapiCLI commands. Usemuapi auth configurefirst ifMUAPI_API_KEYis unset. - For model IDs without a CLI alias yet, fall back to the raw endpoint via
curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}'and poll withmuapi predict wait <request_id>. - Substitute
{{input_name}}placeholders with the user's actual inputs before issuing each call.
muapi-ai-fight-scene
slug: muapi-ai-fight-scene name: muapi-ai-fight-scene version: "1.0.0" description: Generate a high-cut-density action / fight scene by first composing a 16-cell storyboard image, then driving Seedance 2.0 image-to-video off that storyboard. Stacks GPT-Image-2 (character sheet + storyboard), Nano-Banana-2 (environment concept), and Seedance 2.0 i2v. acceptLicenseTerms: true
AI Fight Scene Generator
Generate a high-cut-density action / fight scene by first composing a 16-cell storyboard image, then driving Seedance 2.0 image-to-video off that storyboard.
The core idea: action tension comes from cut density, not single-shot quality. Forcing the video model to follow a pre-drawn 4×4 storyboard grid gives you 16 distinct shots in a 15-second clip — landing punches, reverse angles, ECUs, whip-pans — that no t2v prompt could choreograph on its own.
Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
character_description | text | yes | — | Full physical description of the fighter(s). Asymmetric details (eye colour, scar side, holster on left hip) help the model preserve identity across panels. |
environment_description | text | yes | — | The scene setting — e.g. "cyberpunk wet back-alley, neon kanji signage, Stray-game aesthetic, rain on chrome." |
action_script | text | yes | — | The action beat — prose or numbered beats. E.g. "Hero is cornered → blocks first punch → counter-elbow → throw opponent into trash cans → finisher." |
style_direction | text | no | cinematic action film, anamorphic lens, high contrast, motion blur on hits | Aesthetic / look tags applied to every frame. |
duration | int | no | 15 | Final video length in seconds. The storyboard's 16 cells map roughly 1 shot per second at default. |
aspect_ratio | text | no | 16:9 | Output aspect — 16:9 cinematic, 9:16 vertical, 1:1 square. |
Steps
Phase A — Character Sheet
Generate a clean turnaround-style character sheet using muapi image generate (model=gpt-image-2-text-to-image):
- Prompt:
Character reference sheet of {{character_description}}. Three views — front, 3/4, profile — on a neutral grey backdrop. Studio lighting, full body, no text overlays, photoreal. Asymmetric identifying details preserved on the correct side. {{style_direction}}. - Aspect ratio:
3:2
Present the character sheet and confirm identity details look right before proceeding. This image becomes reference #1 for later phases.
Phase B — Environment Concept
Use muapi image generate (model=nano-banana-2) to design the scene/world:
- Prompt:
Wide establishing shot of {{environment_description}}. No characters in frame — environment only. Strong perspective lines, depth, atmospheric haze. {{style_direction}}. Production-design concept art. - Aspect ratio:
{{aspect_ratio}}
Nano-Banana-2 is chosen here for its reasoning-driven composition — it's better than text-to-image-only models at producing locations with believable spatial logic (chokepoints, cover, sightlines) that an action scene can use. Present for approval. This becomes reference #2.
Phase C — 16-Cell Storyboard
Compose the action onto a single 4×4 storyboard image using muapi image edit (model=gpt-image-2-image-to-image):
- Reference Images: the character sheet from Phase A and the environment plate from Phase B.
- Prompt:
Compose a 4×4 storyboard grid (16 numbered cells) for the following action sequence: {{action_script}} CHARACTER (use reference image 1 identity throughout, asymmetric details preserved): {{character_description}} LOCATION (use reference image 2 spatial layout): {{environment_description}} Each cell labels: SHOT # (1–16) · SIZE (WIDE / MS / CU / ECU) · CAMERA-MOVE arrow (push, pull, whip, dolly, crash-zoom, handheld) · 1-word RHYTHM note (BEAT / IMPACT / RECOVERY / RESET). Vary shot size aggressively — never two WIDEs in a row. Land every IMPACT on a CU or ECU. Hand-drawn comic-book ink-and-wash style, monochrome with selective red accents on hits. Numbered cells, clear gutters between panels. Aesthetic: {{style_direction}}. - Aspect ratio:
1:1(square works best for a 4×4 grid)
Present the storyboard to the user. Confirm:
- The 16 shots read clearly
- Identity stays consistent cell-to-cell
- Cut density / shot-size variation looks aggressive enough
If a panel reads poorly, regenerate just the storyboard with that cell's note bolded ("CELL 7 must be an ECU on the right fist").
Phase D — Storyboard → Video (Seedance 2.0)
Hand the storyboard to muapi video from-image (model=seedance-v2.0-i2v):
- Reference Image: the 16-cell storyboard from Phase C.
- Prompt:
Generate a {{duration}}-second action sequence that strictly follows the 16-cell storyboard reference image, cell-by-cell, top-left to bottom-right. - Honour each cell's labelled SHOT SIZE and CAMERA-MOVE — match cuts to the storyboard's rhythm notes. - Strong cinematic feel and shot language. Exaggerated dynamics. Hits land hard with motion blur and impact frames. - Camera language: anamorphic, handheld where the storyboard calls for it, locked-off where it doesn't. - Native audio: impact sfx on every IMPACT cell, footsteps, fabric/Foley, restrained low score under the action. Action being rendered: {{action_script}}. Aesthetic: {{style_direction}}. - Duration:
{{duration}}(default 15) - Aspect ratio:
{{aspect_ratio}}
After generation, present the final video. If the cut density feels too low or shots don't match the storyboard, regenerate Phase D first (cheaper than rebuilding the storyboard) with the prompt emphasising "strict cell-by-cell adherence" more aggressively.
Notes
- Why the storyboard image and not a text storyboard? Seedance 2.0 i2v anchors its motion plan to the visual reference. A grid of 16 drawn cells gives it 16 visual targets to hit — text descriptions of shots get averaged into mush.
- Asymmetric character details matter. Without something like "scar over the right eyebrow" or "leather glove on the left hand only", identity drift between cells is the #1 failure mode.
- Use
seedance-2.0-i2v-480pto draft. Cheaper preview pass before committing to the full-resseedance-v2.0-i2vrun. - For longer fights, chain two runs: first run uses storyboard A (cells 1–16, beats 1–15s); second run uses storyboard B (cells 17–32, beats 15–30s) with the last cell of A as a continuity anchor in B's first cell.
- Language: Both English and Chinese prompts work in all four models, so the storyboard cell labels can be in either language.
Trigger Keywords
fight scene, action sequence, storyboard to video, cut density, cinematic action, combat choreography, seedance 2 storyboard
Pipeline at a Glance
character_description ──► [GPT-Image-2 t2i] ─► character sheet ──┐
│
environment_description ─► [Nano-Banana-2 t2i] ─► environment plate ┼─► [GPT-Image-2 i2i] ─► 16-cell storyboard ─► [Seedance 2.0 i2v] ─► 15s action video
│
action_script + style_direction ───────────────────────────────────►┘
Notes for the Executing Agent
- This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call
muapiCLI commands. Usemuapi auth configurefirst ifMUAPI_API_KEYis unset. - For model IDs without a CLI alias yet, fall back to the raw endpoint via
curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}'and poll withmuapi predict wait <request_id>. - Phase C uses TWO reference images (character sheet + environment plate). When calling
gpt-image-2-image-to-image, pass them as a list underimages_list(or the model's documented multi-ref field). - Substitute
{{input_name}}placeholders with the user's actual inputs before issuing each call.
muapi-animal-video-generator
slug: muapi-animal-video-generator name: muapi-animal-video-generator version: "1.0.0" description: Create a hilarious and ultra-realistic video of an anthropomorphic animal acting like a human vlogger in a real-world setting. acceptLicenseTerms: true
Animal Vlogger Video
Create a hilarious and ultra-realistic video of an anthropomorphic animal acting like a human vlogger in a real-world setting.
Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
animal_type | text | no | monkey | The type of animal vlogging (e.g., monkey, dog, cat, bear). |
location | text | no | busy streets of Noida, India | The setting for the vlog. |
clothing | text | no | a bright red t-shirt | What the animal is wearing. |
script | text | no | क्या यार, कहाँ फंस गया नॉएडा में आके? इससे तो बैंगलोर ही अच्छा था. यहाँ तो बड़ी गर्मी है भाई. पसीने ही छूट गए मेरे तो. | The dialogue for lip-sync. |
Steps
Phase A — Generate Vlogger Image
Submit the plan with ONE step to create the base image of the animal vlogger:
- Image Generation —
muapi image generate(model=nano-bananaornano-banana-pro):- Prompt:
Ultra-realistic, cinematic portrait of an expressive {{animal_type}} wearing {{clothing}}, holding a selfie stick on the {{location}}. The {{animal_type}} is anthropomorphic, walking upright, and looks like a believable vlogger. Sweating slightly in the heat, highly detailed, photorealistic. Background shows a bustling street, people eating at stalls, vibrant urban details. - Aspect ratio: 9:16 (for vertical vlog style)
- Prompt:
After generating the image, present it to the user.
Phase B — Video Generation
Once the image is approved, submit a second the plan with ONE step to animate it with lip-sync:
- Video Generation —
muapi video generateormuapi video from-image(model=veo3.1-fast-image-to-video):- Reference Image: The generated image from Phase A.
- Prompt:
Create an ultra-realistic cinematic video featuring a lifelike {{animal_type}} vlogging with a selfie stick on the {{location}}. The film starts in selfie mode, camera slightly wide-angle. The {{animal_type}} is expressive, looks mildly disgruntled, and speaks naturally. Occasional quick zoom-in on the disappointed face. Background features busy street life, vibrant colors. Smooth gimbal-like camera motion. - Dialogue for Lip-Sync (if tool supports audio/lipsync):
{{script}} - Aspect ratio: 9:16
After generation, present the final funny video.
Trigger Keywords
animal video, funny monkey video, animal vlogger, vlog video, monkey in noida
Notes for the Executing Agent
- This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call
muapiCLI commands. Usemuapi auth configurefirst ifMUAPI_API_KEYis unset. - For model IDs without a CLI alias yet, fall back to the raw endpoint via
curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}'and poll withmuapi predict wait <request_id>. - Substitute
{{input_name}}placeholders with the user's actual inputs before issuing each call.
muapi-cartoon-dance-animation
slug: muapi-cartoon-dance-animation name: muapi-cartoon-dance-animation version: "1.0.0" description: Convert a photo of a person into a Pixar-style 3D cartoon character, then animate it using a reference dance or motion video. acceptLicenseTerms: true
Cartoon Dance Animation
Convert a photo of a person into a Pixar-style 3D cartoon character, then animate it using a reference dance or motion video.
Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
user_image | image_url | yes | — | A clear full-body or medium-shot photo of the person to be cartoonified. |
reference_video | video_url | no | — | A video containing the specific dance or motion to apply to the character. |
Steps
Phase A — Cartoon Character Generation
If {{user_image}} is not provided, ask the user to upload their photo.
Once the photo is available, submit the plan with ONE step to cartoonify the image:
- Image Generation —
muapi image edit(model=nano-banana-2-edit):- Reference Image:
{{user_image}} - Prompt:
Use the uploaded input photo as the exact same person in the final render. Preserve identity accurately: same face shape, eyes, nose, lips, jawline, skin tone, hairstyle, hairline, expression, age, and overall vibe. Do NOT change the person into a different face. Keep it clearly recognizable as the same person. Create one full-size ultra-high-quality 3D stylized character illustration, Pixar-inspired but original, based on the input person. Smooth plastic-like skin, soft rounded facial features, big expressive eyes, small nose, subtle blush (very minimal), cozy wholesome aesthetic. High-end character sculpting with stylized proportions while maintaining the real person’s likeness. 👕 Outfit / Costume (MUST MATCH INPUT) Keep the costume/outfit EXACTLY the same as the input image. Do not change colors, fabric type, accessories, layers, patterns, logos, or fit. No added glasses, no headphones, no new jacket, no new styling. 💇 Hair (Exact Match) Hair must remain the same as the input image: same hairstyle, same length, same hairline, only converted into clean stylized 3D hair shapes. 🎨 Render Quality Premium character sculpting, soft studio lighting, global illumination, subsurface scattering, soft shadows, cinematic depth of field, crisp edges. Octane/Arnold render look, ultra-clean, high-quality shading, 8K detail. 🎯 Composition Single full-size image (NOT a grid). Full-body or medium shot matching the input pose and vibe. Minimal clean studio background (solid color), no clutter. - Negative Prompt:
No outfit change, no costume change, no new clothes, no extra accessories, no glasses, no headphones, no makeup, no cosmetics, no lipstick, no eyeliner, no facial redesign, no different face, no extra limbs, no deformed hands, no scary look, no photoreal skin pores, no wrinkles, no blur, no noise, no watermark, no logo, no text. - Aspect ratio: Maintain the aspect ratio of the input image or default to 9:16.
- Reference Image:
Present the generated cartoon character to the user for approval.
Phase B — Motion Control Animation
After the character is approved, ask the user to upload a reference_video (if not already provided) containing the dance or movement they want the character to perform.
Once the video is provided, submit the plan with ONE step:
- Motion Control Video Generation —
muapi video from-imageoredit_video(model=kling-v2.6-std-motion-control):- Reference Image: The cartoon image generated in Phase A.
- Reference Video:
{{reference_video}} - Prompt:
Smooth, fluid 3D character animation. The 3D character perfectly replicates the movements and dance from the reference video. High frame rate, dynamic motion, consistent character details, Pixar animation quality.
After generation, present the final animated dance video to the user.
Trigger Keywords
cartoon dance, 3d animation, pixar character, animate my photo, motion control video, dance video, cartoonify and animate
Notes for the Executing Agent
- This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call
muapiCLI commands. Usemuapi auth configurefirst ifMUAPI_API_KEYis unset. - For model IDs without a CLI alias yet, fall back to the raw endpoint via
curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}'and poll withmuapi predict wait <request_id>. - Substitute
{{input_name}}placeholders with the user's actual inputs before issuing each call.
muapi-character-story-video
slug: muapi-character-story-video name: muapi-character-story-video version: "1.0.0" description: Create a multi-part animated story video by first establishing a consistent character and then generating sequential scenes and animating them. acceptLicenseTerms: true
Character Story Video
Create a multi-part animated story video by first establishing a consistent character and then generating sequential scenes and animating them.
Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
character_description | text | yes | — | Description of the main character (e.g. "a cute piglet wearing a leather aviator jacket and goggles"). |
story_premise | text | yes | — | The overall story arc (e.g. "building a jetpack and flying to space"). |
reference_image | image_url | no | — | Optional starting image of the character to maintain consistency. |
Steps
This skill involves multiple phases to build a cohesive narrative.
Phase A — Character Establishment
If {{reference_image}} is NOT provided, submit the plan with ONE step to create the character:
- Character Creation —
muapi image generate(model=nano-banana-pro):- Prompt:
{{character_description}}, introducing the main character, cinematic lighting, highly detailed, Pixar 3D animation style. - Aspect ratio: 4:5 or 1:1
- Prompt:
If {{reference_image}} IS provided, use it as the established character and proceed to Phase B.
After generation, ask the user to confirm the character design before proceeding.
Phase B — Sequential Scene Generation
Once the character is established, create the story beats (e.g., Scene 1, Scene 2, Scene 3).
Submit the plan using muapi image edit (model=nano-banana-2-edit or flux-kontext-pro-i2i) to maintain character consistency. Use the established character image as the reference for ALL these steps.
- Scene 1 (Beginning)
- Reference: Character Image
- Prompt:
The character ({{character_description}}) in the first scene of the story: [Describe the beginning of {{story_premise}}]. Cinematic lighting, Pixar 3D animation style, storybook illustration.
- Scene 2 (Middle)
- Reference: Character Image
- Prompt:
The character ({{character_description}}) in the second scene: [Describe the climax or middle action of {{story_premise}}]. Cinematic lighting, Pixar 3D animation style, storybook illustration.
- Scene 3 (End)
- Reference: Character Image
- Prompt:
The character ({{character_description}}) in the final scene: [Describe the resolution of {{story_premise}}]. Cinematic lighting, Pixar 3D animation style, storybook illustration.
Note: All scenes should be generated in parallel or sequentially depending on the story flow.
After generating the scenes, present them to the user and ask if they are ready to animate the story.
Phase C — Animation (Sequel Part 1, Part 2, Part 3)
Submit the plan to animate the generated scenes using an image-to-video model (e.g., kling-v3.0-pro-image-to-video or veo3.1-image-to-video).
- Part 1 Video
- Input: Scene 1 Image
- Prompt:
Cinematic animation of the scene, character comes to life, subtle natural movements, high quality 3D animation.
- Part 2 Video
- Input: Scene 2 Image
- Prompt:
Cinematic animation of the scene, character comes to life, dynamic action, high quality 3D animation.
- Part 3 Video
- Input: Scene 3 Image
- Prompt:
Cinematic animation of the scene, character comes to life, triumphant resolution, high quality 3D animation.
After generating the videos, present them to the user as a multi-part story sequence. You may also suggest using the muapi predict result + ffmpeg concat tool to merge them into a single movie if requested.
Trigger Keywords
character story, story video, animated story, sequel video, multi part video, sequential story
Notes for the Executing Agent
- This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call
muapiCLI commands. Usemuapi auth configurefirst ifMUAPI_API_KEYis unset. - For model IDs without a CLI alias yet, fall back to the raw endpoint via
curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}'and poll withmuapi predict wait <request_id>. - Substitute
{{input_name}}placeholders with the user's actual inputs before issuing each call.
muapi-drone-style-video
slug: muapi-drone-style-video name: muapi-drone-style-video version: "1.0.0" description: Generate aerial drone-perspective footage — sweeping bird's-eye views, orbit shots, and flyover sequences for landscapes, architecture, and events. acceptLicenseTerms: true
Drone-Style Video
Generate aerial drone-perspective footage — sweeping bird's-eye views, orbit shots, and flyover sequences for landscapes, architecture, and events.
Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
location_or_subject | text | yes | — | What to shoot from above (e.g. "mountain valley at sunrise", "luxury villa by the ocean", "crowded city intersection"). |
shot_type | text | no | reveal | Camera movement style — 'reveal' (ascend & reveal), 'orbit' (circle subject), 'flyover' (pass over), 'top-down' (bird's eye static). |
style | text | no | golden hour, cinematic, 4K, ultra-detailed | Visual atmosphere (e.g. "dramatic storm clouds", "misty morning", "blue hour city lights"). |
aspect_ratio | text | no | 16:9 | Output aspect ratio. |
reference_image | image_url | no | — | Optional aerial/location reference image. |
Steps
Phase A — Generate Drone Footage
Submit the plan with ONE step:
- Aerial video — If
{{reference_image}}is provided, usemuapi video generate(model=veo3.1-image-to-video); otherwise usemuapi video generate(model=veo3.1-text-to-video).- Build prompt based on
{{shot_type}}:- reveal:
Drone camera starts low, slowly ascends and reveals {{location_or_subject}}, sweeping wide aerial perspective, {{style}} - orbit:
Drone camera orbits {{location_or_subject}} in a smooth circular arc, 360-degree aerial rotation, {{style}} - flyover:
Drone camera flies low and fast over {{location_or_subject}}, tracking forward momentum, depth of field, {{style}} - top-down:
Perfect overhead bird's eye view of {{location_or_subject}}, drone looking straight down, minimal distortion, {{style}}
- reveal:
- Append to all prompts:
DJI-quality drone footage, stabilized gimbal, no shake, cinematic color grade, photorealistic - Aspect ratio:
{{aspect_ratio}}
- Build prompt based on
After generation, offer:
- A different shot type variation
- Adding wind/ambient audio via
mmaudio-v2-video-to-video - Upscaling via
ai-video-upscaler-pro
Notes
- For architecture, emphasize "slow orbit to reveal full building facade".
- For landscapes, use "magic hour lighting" for the best results.
veo3.1-text-to-videoproduces the best physics and camera motion for aerial scenes.
Trigger Keywords
drone, aerial, bird's eye, flyover, aerial shot, drone footage, top down, overhead video
Notes for the Executing Agent
- This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call
muapiCLI commands. Usemuapi auth configurefirst ifMUAPI_API_KEYis unset. - For model IDs without a CLI alias yet, fall back to the raw endpoint via
curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}'and poll withmuapi predict wait <request_id>. - Substitute
{{input_name}}placeholders with the user's actual inputs before issuing each call.
muapi-giant-product-showcase
slug: muapi-giant-product-showcase name: muapi-giant-product-showcase version: "1.0.0" description: Create a dramatic "Giant Product" visual where a regular item is showcased as a massive, building-sized object next to a person, then optionally animate the scene. acceptLicenseTerms: true
Giant Product Showcase
Create a dramatic "Giant Product" visual where a regular item is showcased as a massive, building-sized object next to a person, then optionally animate the scene.
Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
product_image | image_url | yes | — | A clear image of the product to be made giant. |
person_description | text | no | a stylishly dressed man | Description of the person standing next to the giant product. |
Steps
Phase A — Giant Product Visualization
If {{product_image}} is not provided, ask the user to upload a photo of the product.
Once the photo is available, submit the plan with ONE step to create the giant product scene:
- Scene Generation —
muapi image edit(model=nano-banana-2-edit):- Reference Image:
{{product_image}} - Prompt:
A professional commercial photograph featuring a massive, giant-sized version of the product from the reference image. The product is the size of a person and is standing on a clean, modern floor. Next to the giant product, {{person_description}} is leaning against it or standing nearby, highlighting the enormous scale. High-end product photography, soft studio lighting, realistic reflections, 8k resolution. - Aspect ratio: 3:4 or 4:5
- Reference Image:
Present the generated giant product image to the user for approval.
Phase B — Animation (Optional)
After the image is generated, ask the user if they would like to animate the scene into a cinematic showcase video.
If requested, submit the plan with ONE step:
- Video Generation —
muapi video from-image(model=veo3.1-fast-image-to-video):- Reference Image: The giant product image from Phase A.
- Prompt:
Cinematic slow-motion camera movement around the giant product. The person next to it moves naturally, looking at the camera or adjusting their pose. Dynamic lighting, high-quality textures, professional commercial vibe. - Aspect ratio: 9:16 or 4:5
After generation, present the final product showcase video.
Trigger Keywords
giant product, massive object, product showcase, scale comparison, product animation
Notes for the Executing Agent
- This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call
muapiCLI commands. Usemuapi auth configurefirst ifMUAPI_API_KEYis unset. - For model IDs without a CLI alias yet, fall back to the raw endpoint via
curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}'and poll withmuapi predict wait <request_id>. - Substitute
{{input_name}}placeholders with the user's actual inputs before issuing each call.
muapi-jewelry-product-video
slug: muapi-jewelry-product-video name: muapi-jewelry-product-video version: "1.0.0" description: Create a luxury jewelry advertisement with high-end commercial cinematography and detailed macro animation. acceptLicenseTerms: true
Jewelry Product Video
Create a luxury jewelry advertisement with high-end commercial cinematography and detailed macro animation.
Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
jewelry_description | text | no | a delicate rose gold ring with a lotus design and a sparkling diamond | Detailed description of the jewelry item. |
surface_description | text | no | a beige surface | The surface the jewelry is resting on. |
Steps
Phase A — High-End Jewelry Rendering
Submit the plan with ONE step to create the base luxury image:
- Luxury Image Generation —
muapi image generate(model=nano-banana-2-edit):- Prompt:
Style: Luxury product ad, high-end commercial feel. Scene: {{jewelry_description}} resting on {{surface_description}}. A soft, warm light highlights the diamond, creating subtle highlights on the metal. 100mm macro lens photography, shallow DOF, incredible detail, elegant and minimal composition. - Aspect ratio: 1:1 or 4:5
- Prompt:
Present the luxury image to the user for approval.
Phase B — Cinematic Animation
Once the image is approved, submit the plan with TWO sequential video steps to build the commercial:
-
Macro Rotation —
muapi video from-image(model=grok-imagine-image-to-video):- Reference Image: The luxury image from Phase A.
- Prompt:
[00:00–00:02] Close-up shot, 100mm macro lens, shallow DOF. A soft, warm light highlights the diamond, creating subtle highlights on the rose gold. Slight 1-second camera rotation around the ring. Smooth, elegant movement.
-
Facet Gliding —
muapi video from-imageormuapi video from-image(model=grok-imagine-image-to-video):- Reference Image: The luxury image from Phase A.
- Prompt:
[00:02–00:05] Extreme close-up on the diamond, 200mm macro lens, razor-thin DOF. A focused LED light illuminates the diamond, catching every facet. The camera glides slowly over the diamond, showcasing its brilliance. Ethereal, sparkling highlights.
Note: You can use the muapi predict result + ffmpeg concat tool to merge these shots into a final 5-second commercial.
After generation, present the final jewelry commercial video to the user.
Trigger Keywords
jewelry video, luxury ad, diamond animation, ring commercial, high-end jewelry showcase
Notes for the Executing Agent
- This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call
muapiCLI commands. Usemuapi auth configurefirst ifMUAPI_API_KEYis unset. - For model IDs without a CLI alias yet, fall back to the raw endpoint via
curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}'and poll withmuapi predict wait <request_id>. - Substitute
{{input_name}}placeholders with the user's actual inputs before issuing each call.
Related Skills
- SkillsPublic repository for Agent SkillsAI/MLView Details
- Agent SkillsProduction-grade engineering skills for AI coding agents.AI/MLView Details
- Awesome Claude SkillsA curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflowsAI/MLView Details
- Claude Code Best Practicefrom vibe coding to agentic engineering - practice makes claude perfectAI/MLView Details