AI/ML

Generative Media Skills

作者 SamurAIGPT4,205

Multi-modal Generative Media Skills for AI Agents (Claude Code, Cursor, Gemini CLI). High-quality image, video, and audio generation powered by muapi.ai.

agent-skillsagent-toolsai-agentsai-videoclaude-codeclaude-code-skills
安装命令
npx skills add SamurAIGPT/Generative-Media-Skills
在 GitHub 打开
支持的客户端
Claude CodeCursorVS Code CopilotWindsurf

Skill 详情

🎭 Generative Media Skills for AI Agents

Powered by MuAPI

The Ultimate Multimodal Toolset for Claude Code, Cursor, Gemini CLI, and OpenCode. A high-performance, schema-driven architecture for AI agents to generate, edit, and display professional-grade images, videos, and audio — powered by the muapi-cli.

🚀 Get Started | 🎬 Recipe Pack | 🎨 Expert Library | ⚙️ Core Primitives | 🤖 MCP Server | 📖 Reference


<p align="center"><a href="https://www.youtube.com/watch?v=SOXsxqnQGlc"><img src="https://i.ytimg.com/vi/SOXsxqnQGlc/maxresdefault.jpg" width="720"></a></p> <p align="center"><a href="https://www.youtube.com/watch?v=SOXsxqnQGlc"><b>▶ Watch: Best AI Video Generator (API) in 2026 (Quality, Price, Uncensored, Editing)</b></a></p>

Related Projects

✨ Key Features

  • 🤖 Agent-Native Design — CLI-powered scripts with structured JSON outputs, semantic exit codes, and --jq filtering for seamless agentic pipelines.
  • 🧠 Expert Knowledge Layer — Domain-specific skills that bake in professional cinematography, atomic design, and branding logic.
  • ⚡ CLI-Powered Core — All primitives delegate to muapi-cli — no curl, no JSON parsing, no boilerplate.
  • 🖼️ Direct Media Display — Use the --view flag to automatically download and open generated media in your system viewer.
  • 📁 Local File Support — Auto-upload images, videos, faces, and audio from your local machine to the CDN for processing.
  • 🌈 100+ AI Models — One-click access to Midjourney v7, Flux Kontext, Seedance 2.0, Kling 3.0, Veo3, and more.
  • 🔌 MCP Server — Run muapi mcp serve to expose all 19 tools directly to Claude Desktop, Cursor, or any MCP-compatible agent.

🏗️ Scalable Architecture

This repository uses a Core/Library split to ensure efficiency and high-signal discovery for LLMs:

⚙️ Core Primitives (/core)

Thin wrappers around muapi-cli for raw API access.

  • core/media/ — File upload
  • core/edit/ — Image editing (prompt-based)
  • core/platform/ — Setup, auth & result polling

📚 Expert Library (/library)

High-value skills that translate creative intent into technical directives.

  • Cinema Director (/library/motion/cinema-director/) — Technical film direction & cinematography.
  • Nano-Banana (/library/visual/nano-banana/) — Reasoning-driven image generation (Gemini 3 Style).
  • UI Designer (/library/visual/ui-design/) — High-fidelity mobile/web mockups (Atomic Design).
  • Logo Creator (/library/visual/logo-creator/) — Minimalist vector branding (Geometric Primitives).
  • Seedance 2 (Doubao Video) (/library/motion/seedance-2/) — Director-level cinematic video generation with text-to-video, image-to-video, and video extension with native audio-video sync.
  • AI Clipping (/library/edit/ai-clipping/) — Long video → ranked vertical short clips in one managed API call. Server-side transcription, virality ranking, dedupe, and face-tracked auto-crop — no local Whisper or LLM.
  • YouTube Shorts (/library/social/youtube-shorts/) — Platform-aware preset over AI Clipping (Shorts / TikTok / Reels / Feed defaults).

Plus 41 ready-to-run workflow recipes organized by output type — see 🎬 Recipe Pack below.


🎬 Recipe Pack

Forty-one LLM-orchestrated workflow recipes that combine multiple muapi-cli calls into named end-to-end pipelines (e.g. photo of person → 3D action figure, product photo → cinematic 10s ad). Each skill is a SKILL.md the agent reads and follows; bring your own consuming agent (Claude Code, Cursor, MCP) — these are recipes, not bash wrappers.

Motion / Video (16)

SkillDescription
3D Logo AnimationTransform a 2D logo into a premium 3D version and animate it with professional cinematic effects
AI Fight Scene GeneratorHigh-cut-density action / fight scene — 16-cell storyboard image drives Seedance 2.0 i2v for shot-by-shot choreography
Animal Vlogger VideoHilarious, ultra-realistic anthropomorphic-animal vlogger acting like a human in a real-world setting
Cartoon Dance AnimationConvert a photo into a Pixar-style 3D cartoon, then animate using a reference dance/motion video
Character Story VideoMulti-part animated story video — establish a consistent character then animate sequential scenes
Drone-Style VideoAerial drone-perspective footage — bird's-eye sweeps, orbit shots, and flyover sequences
Giant Product ShowcaseDramatic giant-scale product visual (building-sized object next to a person), optionally animated
Jewelry Product VideoLuxury jewelry ad with high-end commercial cinematography and detailed macro animation
Music VideoShort music video from a song theme — keyframes, animation per beat, matching music track
One-Shot VideoSingle continuous cinematic shot — no cuts, one seamless flowing scene
Cinematic Product AdCinematic 5–10s product ad from a product photo + brand brief
Product Showcase VideoDynamic product showcase with explosive ingredient arrangement + realistic motion animation
Product Video Ad MakerHigh-end cinematic product video ad starting from a simple product photo
Talking Baby VideoViral-style talking-baby video with custom costumes and scripts
UGC Lifestyle Try-OnUGC-style lifestyle photos & video of a person using your product — authentic, social-native
UGC Video FactoryPerson photo + product photo + script → 10s vertical 9:16 UGC video ad with native dialogue (Nano-Banana Pro Edit → Seedance 2.0 VIP i2v)

Social (5)

SkillDescription
Instagram PostPolished on-brand Instagram post — hero image + caption + hashtags
Product Campaign PackFull multi-channel campaign — hero visuals, social assets, short ad video, platform crops
RedNote CoverXiaohongshu (小红书) cover image — vibrant lifestyle aesthetic with typography overlay
Social Media PackRe-render a hero image into Instagram / TikTok / Shorts / X aspect ratios
UGC Ads WorkflowUGC video ad pipeline — combine selfie + product image, write script, animate

Visual / Images & Design (21)

SkillDescription
Action Figure GeneratorConvert a photo of a person into a custom 3D action figure with collectible toy packaging
Ad Creative SetHigh-converting ad set — hero image, copy variations, platform crops for Meta / Google / LinkedIn
Amazon Product Listing PackFull Amazon listing image set — hero, lifestyle, infographic, comparison/detail closeups
Blog HeaderProfessional 1200×628 blog header image with optional title composition guidance
Brand KitCohesive brand visual kit — logo concept, color palette, typography pairings
Brochure DesignerMulti-page brochure — cover, inner spread, back — for business, real estate, events, launches
Couple Grid CreatorStylized 6-box grid of a couple in romantic poses, each pose framed inside cardboard packaging
Brand Design GuideComprehensive design guide — palette, typography, UI components, visual identity rules
Fashion Try-OnVirtually try outfits by combining a person's photo + clothing item, optional fashion model video
Floor Plan RenderingDesign a 2D floor plan and convert into a realistic 3D architectural rendering
Interior DesignPro interior design visualizations — redesign rooms, generate concepts, visualize furniture styles
Interior Design VisualizerGenerate an empty room and fill it with stylish furniture / decor; or redesign an existing room
Keyboard Art MakerArtistic top-down photos of keyboard keycaps arranged to spell custom messages
Logo + Branding PackageLogo + full branding package — variations (dark/light/icon), palette, mockups
Logo GeneratorQuick single-shot polished logo — fast, clean vector aesthetic with accurate brand-name text
Multi-Angle ReshootRe-render a subject from dramatic camera angles (fish-eye, bird's-eye, low, macro) — identity preserved
Multi-Angle ShotsFull multi-angle product shot set — front, side, back, top-down, 45°
Selfie with CelebritiesRealistic behind-the-scenes selfie of the user with a celebrity; optional cinematic long-take
Storyboard GeneratorGenerate N keyframes for a short story or scene sequence (image only, no video)
URL to DesignAnalyze a website URL and generate a redesigned, improved UI with modern aesthetics
YouTube ThumbnailHigh-CTR YouTube thumbnail — striking imagery, bold text placement, emotional face/subject

Each recipe declares its inputs and a Steps body. Pass the inputs and let your agent execute the steps via muapi CLI calls (or raw API for endpoints that don't yet have a CLI alias — see the per-skill Notes for the Executing Agent footer).


🚀 Quick Start

1. Install the muapi CLI

The core scripts require muapi-cli. Install it once:

# via npm (recommended — no Python required)
npm install -g muapi-cli

# via pip
pip install muapi-cli

# or run without installing
npx muapi-cli --help

2. Configure Your API Key

# Interactive setup
muapi auth configure

# Or pass directly
muapi auth configure --api-key "YOUR_MUAPI_KEY"

# Get your key at https://muapi.ai/dashboard?utm_source=github&utm_medium=readme&utm_campaign=generative-media-skills

3. Install the Skills

# Install all skills to your AI agent
npx skills add SamurAIGPT/Generative-Media-Skills --all

# Or install a specific skill
npx skills add SamurAIGPT/Generative-Media-Skills --skill muapi-media-generation

# Install to specific agents
npx skills add SamurAIGPT/Generative-Media-Skills --all -a claude-code -a cursor

4. Generate Your First Image

muapi image generate "a cyberpunk city at night" --model flux-dev

# Download the result automatically
muapi image generate "a sunset over mountains" --model hidream-fast --download ./outputs

# Extract just the URL (agent-friendly)
muapi image generate "product on white bg" --model flux-schnell --output-json --jq '.outputs[0]'

5. Run an Expert Skill

# Use Nano-Banana reasoning to generate a 2K masterpiece
bash library/visual/nano-banana/scripts/generate-nano-art.sh \
  --file ./my-source-image.jpg \
  --subject "a glass hummingbird" \
  --style "macro photography" \
  --resolution "2k" \
  --view

6. Direct a Cinematic Scene

cd library/motion/cinema-director

# Create a 10-second epic reveal
bash scripts/generate-film.sh \
  --subject "a cybernetic dragon over Tokyo" \
  --intent "epic" \
  --model "kling-v3.0-pro" \
  --duration 10 \
  --view

# Animate a reference image into video
bash library/motion/seedance-2/scripts/generate-seedance.sh \
  --mode i2v \
  --file ./concept.jpg \
  --subject "camera slowly pulls back to reveal the full landscape" \
  --intent "reveal" \
  --view

# Extend an existing video
bash library/motion/seedance-2/scripts/generate-seedance.sh \
  --mode extend \
  --request-id "YOUR_REQUEST_ID" \
  --subject "camera continues pulling back to reveal the vast city" \
  --duration 10

OpenCode

# Clone the repo and set the MUAPI_API_KEY env var
git clone https://github.com/SamurAIGPT/Generative-Media-Skills
export MUAPI_API_KEY=your_key_here

# Skills auto-load from .opencode/skills/ when you run opencode in this directory
opencode

🤖 MCP Server

Run muapi as a Model Context Protocol server so Claude Desktop, Cursor, or any MCP-compatible agent can call generation tools directly — no shell scripts needed.

muapi mcp serve

Claude Desktop config (~/Library/Application Support/Claude/claude_desktop_config.json):

{
  "mcpServers": {
    "muapi": {
      "command": "muapi",
      "args": ["mcp", "serve"],
      "env": { "MUAPI_API_KEY": "your-key-here" }
    }
  }
}

This exposes 19 structured tools with full JSON Schema input/output definitions:

ToolDescription
muapi_image_generateText-to-image (14 models)
muapi_image_editImage-to-image editing (11 models)
muapi_video_generateText-to-video (13 models)
muapi_video_from_imageImage-to-video (16 models)
muapi_audio_createMusic generation (Suno)
muapi_audio_from_textSound effects (MMAudio)
muapi_enhance_upscaleAI upscaling
muapi_enhance_bg_removeBackground removal
muapi_enhance_face_swapFace swap image/video
muapi_enhance_ghibliGhibli style transfer
muapi_edit_lipsyncLip sync to audio
muapi_edit_clippingAI highlight extraction
muapi_predict_resultPoll prediction status
muapi_upload_fileUpload local file → URL
muapi_keys_listList API keys
muapi_keys_createCreate API key
muapi_keys_deleteDelete API key
muapi_account_balanceGet credit balance
muapi_account_topupAdd credits (Stripe checkout)

⚡ Agentic Pipeline Examples

# Submit async, capture request_id, poll when ready
REQUEST_ID=$(muapi video generate "a dog running on a beach" \
  --model kling-master --no-wait --output-json --jq '.request_id' | tr -d '"')

# ... do other work ...

muapi predict wait "$REQUEST_ID" --download ./outputs

# Pipe a prompt from another command
generate_prompt | muapi image generate - --model flux-dev

# Chain: upload → edit → download
URL=$(muapi upload file ./photo.jpg --output-json --jq '.url' | tr -d '"')
muapi image edit "make it look like a painting" --image "$URL" \
  --model flux-kontext-pro --download ./outputs

📖 Schema Reference

This repository includes a streamlined schema_data.json that core scripts use at runtime to:

  • Validate Model IDs: Ensures the requested model exists.
  • Resolve Endpoints: Automatically maps model names to API endpoints.
  • Check Parameters: Validates supported aspect_ratio, resolution, and duration values.

Discover all available models via the CLI:

muapi models list
muapi models list --category video --output-json

🔧 Compatibility

Optimized for the next generation of AI development environments:

  • Claude Code — Direct terminal execution via tools + MCP server mode.
  • Gemini CLI / Cursor / Windsurf — Seamless integration as local scripts.
  • MCP — Full Model Context Protocol server with typed input/output schemas.
  • CI/CD--output-json, --jq, semantic exit codes for scripting.

📄 License

MIT © 2026


muapi-3d-logo-animation


slug: muapi-3d-logo-animation name: muapi-3d-logo-animation version: "1.0.0" description: Transform a 2D logo into a premium 3D version and animate it with professional cinematic effects. acceptLicenseTerms: true

3D Logo Animation

Transform a 2D logo into a premium 3D version and animate it with professional cinematic effects.

Inputs

NameTypeRequiredDefaultDescription
logo_imageimage_urlyesA clear 2D image of the logo to be converted to 3D.
material_styletextnoglossy glass and chromeThe material style for the 3D logo (e.g., gold, matte plastic, holographic).

Steps

Phase A — 3D Logo Transformation

If {{logo_image}} is not provided, ask the user to upload their logo.

Once the logo is available, submit the plan with ONE step to convert it to 3D:

  1. 3D Logo Generationmuapi image edit (model=nano-banana-2-edit):
    • Reference Image: {{logo_image}}
    • Prompt: Transform this 2D logo into a premium, high-quality 3D version. The logo should have depth and be made of {{material_style}}. Smooth edges, realistic reflections, and professional studio lighting. The logo is centered on a clean, minimal, out-of-focus background. High-end graphic design aesthetic, 8k resolution.
    • Aspect ratio: 1:1 or 4:3

Present the 3D logo to the user for approval.

Phase B — Cinematic Logo Animation

Once the 3D logo is ready, submit the plan to animate it:

  1. Logo Animationmuapi video from-image (model=veo3.1-fast-image-to-video):
    • Reference Image: The 3D logo from Phase A.
    • Prompt: A professional cinematic logo reveal animation. The 3D logo rotates slowly with dynamic light sweeps reflecting off its {{material_style}} surface. Subtle camera movement, particle effects in the background, high-quality motion graphics style.
    • Aspect ratio: 16:9 or 1:1

After generation, present the final 3D logo animation to the user.

Trigger Keywords

3d logo, logo animation, 2d to 3d logo, animated logo, cinematic logo


Notes for the Executing Agent

  • This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call muapi CLI commands. Use muapi auth configure first if MUAPI_API_KEY is unset.
  • For model IDs without a CLI alias yet, fall back to the raw endpoint via curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}' and poll with muapi predict wait <request_id>.
  • Substitute {{input_name}} placeholders with the user's actual inputs before issuing each call.

muapi-ai-fight-scene


slug: muapi-ai-fight-scene name: muapi-ai-fight-scene version: "1.0.0" description: Generate a high-cut-density action / fight scene by first composing a 16-cell storyboard image, then driving Seedance 2.0 image-to-video off that storyboard. Stacks GPT-Image-2 (character sheet + storyboard), Nano-Banana-2 (environment concept), and Seedance 2.0 i2v. acceptLicenseTerms: true

AI Fight Scene Generator

Generate a high-cut-density action / fight scene by first composing a 16-cell storyboard image, then driving Seedance 2.0 image-to-video off that storyboard.

The core idea: action tension comes from cut density, not single-shot quality. Forcing the video model to follow a pre-drawn 4×4 storyboard grid gives you 16 distinct shots in a 15-second clip — landing punches, reverse angles, ECUs, whip-pans — that no t2v prompt could choreograph on its own.

Inputs

NameTypeRequiredDefaultDescription
character_descriptiontextyesFull physical description of the fighter(s). Asymmetric details (eye colour, scar side, holster on left hip) help the model preserve identity across panels.
environment_descriptiontextyesThe scene setting — e.g. "cyberpunk wet back-alley, neon kanji signage, Stray-game aesthetic, rain on chrome."
action_scripttextyesThe action beat — prose or numbered beats. E.g. "Hero is cornered → blocks first punch → counter-elbow → throw opponent into trash cans → finisher."
style_directiontextnocinematic action film, anamorphic lens, high contrast, motion blur on hitsAesthetic / look tags applied to every frame.
durationintno15Final video length in seconds. The storyboard's 16 cells map roughly 1 shot per second at default.
aspect_ratiotextno16:9Output aspect — 16:9 cinematic, 9:16 vertical, 1:1 square.

Steps

Phase A — Character Sheet

Generate a clean turnaround-style character sheet using muapi image generate (model=gpt-image-2-text-to-image):

  • Prompt: Character reference sheet of {{character_description}}. Three views — front, 3/4, profile — on a neutral grey backdrop. Studio lighting, full body, no text overlays, photoreal. Asymmetric identifying details preserved on the correct side. {{style_direction}}.
  • Aspect ratio: 3:2

Present the character sheet and confirm identity details look right before proceeding. This image becomes reference #1 for later phases.

Phase B — Environment Concept

Use muapi image generate (model=nano-banana-2) to design the scene/world:

  • Prompt: Wide establishing shot of {{environment_description}}. No characters in frame — environment only. Strong perspective lines, depth, atmospheric haze. {{style_direction}}. Production-design concept art.
  • Aspect ratio: {{aspect_ratio}}

Nano-Banana-2 is chosen here for its reasoning-driven composition — it's better than text-to-image-only models at producing locations with believable spatial logic (chokepoints, cover, sightlines) that an action scene can use. Present for approval. This becomes reference #2.

Phase C — 16-Cell Storyboard

Compose the action onto a single 4×4 storyboard image using muapi image edit (model=gpt-image-2-image-to-image):

  • Reference Images: the character sheet from Phase A and the environment plate from Phase B.
  • Prompt:
    Compose a 4×4 storyboard grid (16 numbered cells) for the following action sequence:
    {{action_script}}
    
    CHARACTER (use reference image 1 identity throughout, asymmetric details preserved):
    {{character_description}}
    
    LOCATION (use reference image 2 spatial layout):
    {{environment_description}}
    
    Each cell labels: SHOT # (1–16) · SIZE (WIDE / MS / CU / ECU) · CAMERA-MOVE arrow (push, pull, whip, dolly, crash-zoom, handheld) · 1-word RHYTHM note (BEAT / IMPACT / RECOVERY / RESET).
    
    Vary shot size aggressively — never two WIDEs in a row. Land every IMPACT on a CU or ECU.
    Hand-drawn comic-book ink-and-wash style, monochrome with selective red accents on hits.
    Numbered cells, clear gutters between panels.
    
    Aesthetic: {{style_direction}}.
    
  • Aspect ratio: 1:1 (square works best for a 4×4 grid)

Present the storyboard to the user. Confirm:

  • The 16 shots read clearly
  • Identity stays consistent cell-to-cell
  • Cut density / shot-size variation looks aggressive enough

If a panel reads poorly, regenerate just the storyboard with that cell's note bolded ("CELL 7 must be an ECU on the right fist").

Phase D — Storyboard → Video (Seedance 2.0)

Hand the storyboard to muapi video from-image (model=seedance-v2.0-i2v):

  • Reference Image: the 16-cell storyboard from Phase C.
  • Prompt:
    Generate a {{duration}}-second action sequence that strictly follows the 16-cell storyboard reference image, cell-by-cell, top-left to bottom-right.
    
    - Honour each cell's labelled SHOT SIZE and CAMERA-MOVE — match cuts to the storyboard's rhythm notes.
    - Strong cinematic feel and shot language. Exaggerated dynamics. Hits land hard with motion blur and impact frames.
    - Camera language: anamorphic, handheld where the storyboard calls for it, locked-off where it doesn't.
    - Native audio: impact sfx on every IMPACT cell, footsteps, fabric/Foley, restrained low score under the action.
    
    Action being rendered: {{action_script}}.
    Aesthetic: {{style_direction}}.
    
  • Duration: {{duration}} (default 15)
  • Aspect ratio: {{aspect_ratio}}

After generation, present the final video. If the cut density feels too low or shots don't match the storyboard, regenerate Phase D first (cheaper than rebuilding the storyboard) with the prompt emphasising "strict cell-by-cell adherence" more aggressively.

Notes

  • Why the storyboard image and not a text storyboard? Seedance 2.0 i2v anchors its motion plan to the visual reference. A grid of 16 drawn cells gives it 16 visual targets to hit — text descriptions of shots get averaged into mush.
  • Asymmetric character details matter. Without something like "scar over the right eyebrow" or "leather glove on the left hand only", identity drift between cells is the #1 failure mode.
  • Use seedance-2.0-i2v-480p to draft. Cheaper preview pass before committing to the full-res seedance-v2.0-i2v run.
  • For longer fights, chain two runs: first run uses storyboard A (cells 1–16, beats 1–15s); second run uses storyboard B (cells 17–32, beats 15–30s) with the last cell of A as a continuity anchor in B's first cell.
  • Language: Both English and Chinese prompts work in all four models, so the storyboard cell labels can be in either language.

Trigger Keywords

fight scene, action sequence, storyboard to video, cut density, cinematic action, combat choreography, seedance 2 storyboard

Pipeline at a Glance

character_description ──► [GPT-Image-2 t2i]   ─► character sheet ──┐
                                                                    │
environment_description ─► [Nano-Banana-2 t2i] ─► environment plate ┼─► [GPT-Image-2 i2i] ─► 16-cell storyboard ─► [Seedance 2.0 i2v] ─► 15s action video
                                                                    │
action_script + style_direction ───────────────────────────────────►┘

Notes for the Executing Agent

  • This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call muapi CLI commands. Use muapi auth configure first if MUAPI_API_KEY is unset.
  • For model IDs without a CLI alias yet, fall back to the raw endpoint via curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}' and poll with muapi predict wait <request_id>.
  • Phase C uses TWO reference images (character sheet + environment plate). When calling gpt-image-2-image-to-image, pass them as a list under images_list (or the model's documented multi-ref field).
  • Substitute {{input_name}} placeholders with the user's actual inputs before issuing each call.

muapi-animal-video-generator


slug: muapi-animal-video-generator name: muapi-animal-video-generator version: "1.0.0" description: Create a hilarious and ultra-realistic video of an anthropomorphic animal acting like a human vlogger in a real-world setting. acceptLicenseTerms: true

Animal Vlogger Video

Create a hilarious and ultra-realistic video of an anthropomorphic animal acting like a human vlogger in a real-world setting.

Inputs

NameTypeRequiredDefaultDescription
animal_typetextnomonkeyThe type of animal vlogging (e.g., monkey, dog, cat, bear).
locationtextnobusy streets of Noida, IndiaThe setting for the vlog.
clothingtextnoa bright red t-shirtWhat the animal is wearing.
scripttextnoक्या यार, कहाँ फंस गया नॉएडा में आके? इससे तो बैंगलोर ही अच्छा था. यहाँ तो बड़ी गर्मी है भाई. पसीने ही छूट गए मेरे तो.The dialogue for lip-sync.

Steps

Phase A — Generate Vlogger Image

Submit the plan with ONE step to create the base image of the animal vlogger:

  1. Image Generationmuapi image generate (model=nano-banana or nano-banana-pro):
    • Prompt: Ultra-realistic, cinematic portrait of an expressive {{animal_type}} wearing {{clothing}}, holding a selfie stick on the {{location}}. The {{animal_type}} is anthropomorphic, walking upright, and looks like a believable vlogger. Sweating slightly in the heat, highly detailed, photorealistic. Background shows a bustling street, people eating at stalls, vibrant urban details.
    • Aspect ratio: 9:16 (for vertical vlog style)

After generating the image, present it to the user.

Phase B — Video Generation

Once the image is approved, submit a second the plan with ONE step to animate it with lip-sync:

  1. Video Generationmuapi video generate or muapi video from-image (model=veo3.1-fast-image-to-video):
    • Reference Image: The generated image from Phase A.
    • Prompt: Create an ultra-realistic cinematic video featuring a lifelike {{animal_type}} vlogging with a selfie stick on the {{location}}. The film starts in selfie mode, camera slightly wide-angle. The {{animal_type}} is expressive, looks mildly disgruntled, and speaks naturally. Occasional quick zoom-in on the disappointed face. Background features busy street life, vibrant colors. Smooth gimbal-like camera motion.
    • Dialogue for Lip-Sync (if tool supports audio/lipsync): {{script}}
    • Aspect ratio: 9:16

After generation, present the final funny video.

Trigger Keywords

animal video, funny monkey video, animal vlogger, vlog video, monkey in noida


Notes for the Executing Agent

  • This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call muapi CLI commands. Use muapi auth configure first if MUAPI_API_KEY is unset.
  • For model IDs without a CLI alias yet, fall back to the raw endpoint via curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}' and poll with muapi predict wait <request_id>.
  • Substitute {{input_name}} placeholders with the user's actual inputs before issuing each call.

muapi-cartoon-dance-animation


slug: muapi-cartoon-dance-animation name: muapi-cartoon-dance-animation version: "1.0.0" description: Convert a photo of a person into a Pixar-style 3D cartoon character, then animate it using a reference dance or motion video. acceptLicenseTerms: true

Cartoon Dance Animation

Convert a photo of a person into a Pixar-style 3D cartoon character, then animate it using a reference dance or motion video.

Inputs

NameTypeRequiredDefaultDescription
user_imageimage_urlyesA clear full-body or medium-shot photo of the person to be cartoonified.
reference_videovideo_urlnoA video containing the specific dance or motion to apply to the character.

Steps

Phase A — Cartoon Character Generation

If {{user_image}} is not provided, ask the user to upload their photo.

Once the photo is available, submit the plan with ONE step to cartoonify the image:

  1. Image Generationmuapi image edit (model=nano-banana-2-edit):
    • Reference Image: {{user_image}}
    • Prompt: Use the uploaded input photo as the exact same person in the final render. Preserve identity accurately: same face shape, eyes, nose, lips, jawline, skin tone, hairstyle, hairline, expression, age, and overall vibe. Do NOT change the person into a different face. Keep it clearly recognizable as the same person. Create one full-size ultra-high-quality 3D stylized character illustration, Pixar-inspired but original, based on the input person. Smooth plastic-like skin, soft rounded facial features, big expressive eyes, small nose, subtle blush (very minimal), cozy wholesome aesthetic. High-end character sculpting with stylized proportions while maintaining the real person’s likeness. 👕 Outfit / Costume (MUST MATCH INPUT) Keep the costume/outfit EXACTLY the same as the input image. Do not change colors, fabric type, accessories, layers, patterns, logos, or fit. No added glasses, no headphones, no new jacket, no new styling. 💇 Hair (Exact Match) Hair must remain the same as the input image: same hairstyle, same length, same hairline, only converted into clean stylized 3D hair shapes. 🎨 Render Quality Premium character sculpting, soft studio lighting, global illumination, subsurface scattering, soft shadows, cinematic depth of field, crisp edges. Octane/Arnold render look, ultra-clean, high-quality shading, 8K detail. 🎯 Composition Single full-size image (NOT a grid). Full-body or medium shot matching the input pose and vibe. Minimal clean studio background (solid color), no clutter.
    • Negative Prompt: No outfit change, no costume change, no new clothes, no extra accessories, no glasses, no headphones, no makeup, no cosmetics, no lipstick, no eyeliner, no facial redesign, no different face, no extra limbs, no deformed hands, no scary look, no photoreal skin pores, no wrinkles, no blur, no noise, no watermark, no logo, no text.
    • Aspect ratio: Maintain the aspect ratio of the input image or default to 9:16.

Present the generated cartoon character to the user for approval.

Phase B — Motion Control Animation

After the character is approved, ask the user to upload a reference_video (if not already provided) containing the dance or movement they want the character to perform.

Once the video is provided, submit the plan with ONE step:

  1. Motion Control Video Generationmuapi video from-image or edit_video (model=kling-v2.6-std-motion-control):
    • Reference Image: The cartoon image generated in Phase A.
    • Reference Video: {{reference_video}}
    • Prompt: Smooth, fluid 3D character animation. The 3D character perfectly replicates the movements and dance from the reference video. High frame rate, dynamic motion, consistent character details, Pixar animation quality.

After generation, present the final animated dance video to the user.

Trigger Keywords

cartoon dance, 3d animation, pixar character, animate my photo, motion control video, dance video, cartoonify and animate


Notes for the Executing Agent

  • This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call muapi CLI commands. Use muapi auth configure first if MUAPI_API_KEY is unset.
  • For model IDs without a CLI alias yet, fall back to the raw endpoint via curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}' and poll with muapi predict wait <request_id>.
  • Substitute {{input_name}} placeholders with the user's actual inputs before issuing each call.

muapi-character-story-video


slug: muapi-character-story-video name: muapi-character-story-video version: "1.0.0" description: Create a multi-part animated story video by first establishing a consistent character and then generating sequential scenes and animating them. acceptLicenseTerms: true

Character Story Video

Create a multi-part animated story video by first establishing a consistent character and then generating sequential scenes and animating them.

Inputs

NameTypeRequiredDefaultDescription
character_descriptiontextyesDescription of the main character (e.g. "a cute piglet wearing a leather aviator jacket and goggles").
story_premisetextyesThe overall story arc (e.g. "building a jetpack and flying to space").
reference_imageimage_urlnoOptional starting image of the character to maintain consistency.

Steps

This skill involves multiple phases to build a cohesive narrative.

Phase A — Character Establishment

If {{reference_image}} is NOT provided, submit the plan with ONE step to create the character:

  1. Character Creationmuapi image generate (model=nano-banana-pro):
    • Prompt: {{character_description}}, introducing the main character, cinematic lighting, highly detailed, Pixar 3D animation style.
    • Aspect ratio: 4:5 or 1:1

If {{reference_image}} IS provided, use it as the established character and proceed to Phase B.

After generation, ask the user to confirm the character design before proceeding.

Phase B — Sequential Scene Generation

Once the character is established, create the story beats (e.g., Scene 1, Scene 2, Scene 3). Submit the plan using muapi image edit (model=nano-banana-2-edit or flux-kontext-pro-i2i) to maintain character consistency. Use the established character image as the reference for ALL these steps.

  1. Scene 1 (Beginning)
    • Reference: Character Image
    • Prompt: The character ({{character_description}}) in the first scene of the story: [Describe the beginning of {{story_premise}}]. Cinematic lighting, Pixar 3D animation style, storybook illustration.
  2. Scene 2 (Middle)
    • Reference: Character Image
    • Prompt: The character ({{character_description}}) in the second scene: [Describe the climax or middle action of {{story_premise}}]. Cinematic lighting, Pixar 3D animation style, storybook illustration.
  3. Scene 3 (End)
    • Reference: Character Image
    • Prompt: The character ({{character_description}}) in the final scene: [Describe the resolution of {{story_premise}}]. Cinematic lighting, Pixar 3D animation style, storybook illustration.

Note: All scenes should be generated in parallel or sequentially depending on the story flow.

After generating the scenes, present them to the user and ask if they are ready to animate the story.

Phase C — Animation (Sequel Part 1, Part 2, Part 3)

Submit the plan to animate the generated scenes using an image-to-video model (e.g., kling-v3.0-pro-image-to-video or veo3.1-image-to-video).

  1. Part 1 Video
    • Input: Scene 1 Image
    • Prompt: Cinematic animation of the scene, character comes to life, subtle natural movements, high quality 3D animation.
  2. Part 2 Video
    • Input: Scene 2 Image
    • Prompt: Cinematic animation of the scene, character comes to life, dynamic action, high quality 3D animation.
  3. Part 3 Video
    • Input: Scene 3 Image
    • Prompt: Cinematic animation of the scene, character comes to life, triumphant resolution, high quality 3D animation.

After generating the videos, present them to the user as a multi-part story sequence. You may also suggest using the muapi predict result + ffmpeg concat tool to merge them into a single movie if requested.

Trigger Keywords

character story, story video, animated story, sequel video, multi part video, sequential story


Notes for the Executing Agent

  • This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call muapi CLI commands. Use muapi auth configure first if MUAPI_API_KEY is unset.
  • For model IDs without a CLI alias yet, fall back to the raw endpoint via curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}' and poll with muapi predict wait <request_id>.
  • Substitute {{input_name}} placeholders with the user's actual inputs before issuing each call.

muapi-drone-style-video


slug: muapi-drone-style-video name: muapi-drone-style-video version: "1.0.0" description: Generate aerial drone-perspective footage — sweeping bird's-eye views, orbit shots, and flyover sequences for landscapes, architecture, and events. acceptLicenseTerms: true

Drone-Style Video

Generate aerial drone-perspective footage — sweeping bird's-eye views, orbit shots, and flyover sequences for landscapes, architecture, and events.

Inputs

NameTypeRequiredDefaultDescription
location_or_subjecttextyesWhat to shoot from above (e.g. "mountain valley at sunrise", "luxury villa by the ocean", "crowded city intersection").
shot_typetextnorevealCamera movement style — 'reveal' (ascend & reveal), 'orbit' (circle subject), 'flyover' (pass over), 'top-down' (bird's eye static).
styletextnogolden hour, cinematic, 4K, ultra-detailedVisual atmosphere (e.g. "dramatic storm clouds", "misty morning", "blue hour city lights").
aspect_ratiotextno16:9Output aspect ratio.
reference_imageimage_urlnoOptional aerial/location reference image.

Steps

Phase A — Generate Drone Footage

Submit the plan with ONE step:

  1. Aerial video — If {{reference_image}} is provided, use muapi video generate (model=veo3.1-image-to-video); otherwise use muapi video generate (model=veo3.1-text-to-video).
    • Build prompt based on {{shot_type}}:
      • reveal: Drone camera starts low, slowly ascends and reveals {{location_or_subject}}, sweeping wide aerial perspective, {{style}}
      • orbit: Drone camera orbits {{location_or_subject}} in a smooth circular arc, 360-degree aerial rotation, {{style}}
      • flyover: Drone camera flies low and fast over {{location_or_subject}}, tracking forward momentum, depth of field, {{style}}
      • top-down: Perfect overhead bird's eye view of {{location_or_subject}}, drone looking straight down, minimal distortion, {{style}}
    • Append to all prompts: DJI-quality drone footage, stabilized gimbal, no shake, cinematic color grade, photorealistic
    • Aspect ratio: {{aspect_ratio}}

After generation, offer:

  • A different shot type variation
  • Adding wind/ambient audio via mmaudio-v2-video-to-video
  • Upscaling via ai-video-upscaler-pro

Notes

  • For architecture, emphasize "slow orbit to reveal full building facade".
  • For landscapes, use "magic hour lighting" for the best results.
  • veo3.1-text-to-video produces the best physics and camera motion for aerial scenes.

Trigger Keywords

drone, aerial, bird's eye, flyover, aerial shot, drone footage, top down, overhead video


Notes for the Executing Agent

  • This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call muapi CLI commands. Use muapi auth configure first if MUAPI_API_KEY is unset.
  • For model IDs without a CLI alias yet, fall back to the raw endpoint via curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}' and poll with muapi predict wait <request_id>.
  • Substitute {{input_name}} placeholders with the user's actual inputs before issuing each call.

muapi-giant-product-showcase


slug: muapi-giant-product-showcase name: muapi-giant-product-showcase version: "1.0.0" description: Create a dramatic "Giant Product" visual where a regular item is showcased as a massive, building-sized object next to a person, then optionally animate the scene. acceptLicenseTerms: true

Giant Product Showcase

Create a dramatic "Giant Product" visual where a regular item is showcased as a massive, building-sized object next to a person, then optionally animate the scene.

Inputs

NameTypeRequiredDefaultDescription
product_imageimage_urlyesA clear image of the product to be made giant.
person_descriptiontextnoa stylishly dressed manDescription of the person standing next to the giant product.

Steps

Phase A — Giant Product Visualization

If {{product_image}} is not provided, ask the user to upload a photo of the product.

Once the photo is available, submit the plan with ONE step to create the giant product scene:

  1. Scene Generationmuapi image edit (model=nano-banana-2-edit):
    • Reference Image: {{product_image}}
    • Prompt: A professional commercial photograph featuring a massive, giant-sized version of the product from the reference image. The product is the size of a person and is standing on a clean, modern floor. Next to the giant product, {{person_description}} is leaning against it or standing nearby, highlighting the enormous scale. High-end product photography, soft studio lighting, realistic reflections, 8k resolution.
    • Aspect ratio: 3:4 or 4:5

Present the generated giant product image to the user for approval.

Phase B — Animation (Optional)

After the image is generated, ask the user if they would like to animate the scene into a cinematic showcase video.

If requested, submit the plan with ONE step:

  1. Video Generationmuapi video from-image (model=veo3.1-fast-image-to-video):
    • Reference Image: The giant product image from Phase A.
    • Prompt: Cinematic slow-motion camera movement around the giant product. The person next to it moves naturally, looking at the camera or adjusting their pose. Dynamic lighting, high-quality textures, professional commercial vibe.
    • Aspect ratio: 9:16 or 4:5

After generation, present the final product showcase video.

Trigger Keywords

giant product, massive object, product showcase, scale comparison, product animation


Notes for the Executing Agent

  • This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call muapi CLI commands. Use muapi auth configure first if MUAPI_API_KEY is unset.
  • For model IDs without a CLI alias yet, fall back to the raw endpoint via curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}' and poll with muapi predict wait <request_id>.
  • Substitute {{input_name}} placeholders with the user's actual inputs before issuing each call.

muapi-jewelry-product-video


slug: muapi-jewelry-product-video name: muapi-jewelry-product-video version: "1.0.0" description: Create a luxury jewelry advertisement with high-end commercial cinematography and detailed macro animation. acceptLicenseTerms: true

Jewelry Product Video

Create a luxury jewelry advertisement with high-end commercial cinematography and detailed macro animation.

Inputs

NameTypeRequiredDefaultDescription
jewelry_descriptiontextnoa delicate rose gold ring with a lotus design and a sparkling diamondDetailed description of the jewelry item.
surface_descriptiontextnoa beige surfaceThe surface the jewelry is resting on.

Steps

Phase A — High-End Jewelry Rendering

Submit the plan with ONE step to create the base luxury image:

  1. Luxury Image Generationmuapi image generate (model=nano-banana-2-edit):
    • Prompt: Style: Luxury product ad, high-end commercial feel. Scene: {{jewelry_description}} resting on {{surface_description}}. A soft, warm light highlights the diamond, creating subtle highlights on the metal. 100mm macro lens photography, shallow DOF, incredible detail, elegant and minimal composition.
    • Aspect ratio: 1:1 or 4:5

Present the luxury image to the user for approval.

Phase B — Cinematic Animation

Once the image is approved, submit the plan with TWO sequential video steps to build the commercial:

  1. Macro Rotationmuapi video from-image (model=grok-imagine-image-to-video):

    • Reference Image: The luxury image from Phase A.
    • Prompt: [00:00–00:02] Close-up shot, 100mm macro lens, shallow DOF. A soft, warm light highlights the diamond, creating subtle highlights on the rose gold. Slight 1-second camera rotation around the ring. Smooth, elegant movement.
  2. Facet Glidingmuapi video from-image or muapi video from-image (model=grok-imagine-image-to-video):

    • Reference Image: The luxury image from Phase A.
    • Prompt: [00:02–00:05] Extreme close-up on the diamond, 200mm macro lens, razor-thin DOF. A focused LED light illuminates the diamond, catching every facet. The camera glides slowly over the diamond, showcasing its brilliance. Ethereal, sparkling highlights.

Note: You can use the muapi predict result + ffmpeg concat tool to merge these shots into a final 5-second commercial.

After generation, present the final jewelry commercial video to the user.

Trigger Keywords

jewelry video, luxury ad, diamond animation, ring commercial, high-end jewelry showcase


Notes for the Executing Agent

  • This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call muapi CLI commands. Use muapi auth configure first if MUAPI_API_KEY is unset.
  • For model IDs without a CLI alias yet, fall back to the raw endpoint via curl -X POST https://api.muapi.ai/api/v1/<endpoint> -H "x-api-key: $MUAPI_API_KEY" -H 'content-type: application/json' -d '{...}' and poll with muapi predict wait <request_id>.
  • Substitute {{input_name}} placeholders with the user's actual inputs before issuing each call.