Skip to content
Masko logomasko
Docs
Documentation

Talking animations

A talking animation is your mascot saying a line in its own voice. It moves while it talks, following the movements you write into the line, and the video carries the sound. The MP4 and the transparent WebM and HEVC files all keep the voice. You also get the voice alone as an MP3 and a transcript with the time of every word, ready for captions.

Talking with a voice is open to every account. Adding a voice (ideas, adjusting, samples and keeping one) unlocks once the account has spent $50 on Masko: credit purchases, Marketplace purchases and plan payments made from 30 September 2026 count, and a team mascot also counts the team's purchases. Before that, these calls return 403 with details.reason: "voice_locked", paid_cents and required_cents.

1. Give the mascot a voice

A mascot speaks with one voice. A variant speaks with the Original's voice until you give it its own. Designing a voice takes three calls: get ideas, hear three samples, keep one.

Masko writes voice ideas from the mascot's references and description. Pass refresh=true for other ideas.

curl https://api.masko.ai/v1/mascots/MASCOT_ID/voice/suggestions \
  -H "Authorization: Bearer $MASKO_API_KEY"
{
  "data": {
    "ideas": [
      {
        "title": "Cozy and warm",
        "description": "A round, cozy pond creature with a big belly. Male voice, warm and round, medium-low, smiling while he talks, relaxed and kind, easy natural pace. American accent."
      },
      {
        "title": "Excited little sidekick",
        "description": "An excited little cartoon sidekick. Young male voice, fast and bouncy, high energy, can barely contain his excitement, slightly breathless and giggly. American accent."
      }
    ]
  }
}

To move a description in one direction, such as "Older", "Slower" or "A little grumpy", send it with a direction. The response has the rewritten text, what changed, and directions that fit the new text. Without direction, it returns only the directions.

{
  "data": {
    "description": "A round, cozy pond creature with a big belly. Male voice, warm and round, medium-high, smiling while he talks, relaxed and kind, easy natural pace. American accent.",
    "changes": [{ "from": "medium-low", "to": "medium-high" }],
    "directions": ["Younger", "More excited", "Slower", "Drier", "Sleepier"]
  }
}

Three voices designed from one description, each reading the same line, for 1 credit. sample_line is optional.

curl -X POST https://api.masko.ai/v1/mascots/MASCOT_ID/voice/samples \
  -H "Authorization: Bearer $MASKO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "description": "An excited little cartoon sidekick. Young male voice, fast and bouncy, high energy, slightly breathless and giggly. American accent."
  }'
{
  "data": {
    "samples": [
      { "id": "5b1c2d8e-8a0f-4c55-9f7e-1d2a3b4c5d61", "url": "https://storage.googleapis.com/...", "duration": 9.6 },
      { "id": "6c2d3e9f-9b10-4d66-a08f-2e3b4c5d6e72", "url": "https://storage.googleapis.com/...", "duration": 8.3 },
      { "id": "7d3e4f0a-ac21-4e77-b190-3f4c5d6e7f83", "url": "https://storage.googleapis.com/...", "duration": 9.3 }
    ],
    "sample_line": "Hi! I'm Gubby. I live in your app, and I love showing people around.",
    "cost_credits": 1,
    "expires_at": "2026-09-30T09:12:00Z"
  }
}

Keep one sample within 24 hours. It replaces the voice the mascot had. An optional name renames it.

curl -X PUT https://api.masko.ai/v1/mascots/MASCOT_ID/voice \
  -H "Authorization: Bearer $MASKO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "sample_id": "7d3e4f0a-ac21-4e77-b190-3f4c5d6e7f83" }'
{
  "data": {
    "object": "voice",
    "name": "Gubby's voice",
    "description": "An excited little cartoon sidekick. Young male voice, fast and bouncy, high energy, slightly breathless and giggly. American accent.",
    "sample_url": "https://storage.googleapis.com/...",
    "source": "original",
    "created_at": "2026-09-29T09:14:00Z"
  }
}

GET /v1/mascots/{id}/voice returns the voice, or null. source is original for the mascot's own voice, variant for a variant's own voice, and creator for a Marketplace mascot that speaks with its creator's voice. Keeping a new voice for the Original also changes it for every variant without its own.

For a variant, use the same calls under /v1/mascots/{id}/variants/{variantId}/voice. DELETE on the variant's voice makes it speak with the Original's voice again.

2. Write the line

Send the line as speech.script, with movements in square brackets placed where each one starts. Words in brackets are never spoken. A feeling after a comma also steers the voice.

[waves hello, excited] Hi! I'm Gubby. [points to the right] The docs are right here!

You can also send only the words as speech.text, with an optional speech.direction such as "excited, points at the docs at the end". Masko writes the movements and returns the script it used.

The voice speaks every language its voice model supports. Set speech.language to an ISO 639-1 code, such as fr, or leave it out to detect the language from the words.

A line holds up to about 14 seconds of speech. The video is 5 to 15 seconds and follows the length of the speech.

3. Estimate for free

Add dry_run: true to get the estimate without charging or generating anything.

curl -X POST https://api.masko.ai/v1/mascots/MASCOT_ID/generate \
  -H "Authorization: Bearer $MASKO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "type": "animation",
    "item_id": "ITEM_ID",
    "speech": { "text": "Hi! I'\''m Gubby. Here'\''s how the Masko API works.", "direction": "excited about what he does" },
    "dry_run": true
  }'
{
  "data": {
    "dry_run": true,
    "script": "[waves hello, excited] Hi! I'm Gubby. [opens an arm to the right] Here's how the Masko API works.",
    "estimate": { "speech_seconds": 4.3, "duration": 6, "max_duration": 6, "credits": 18, "too_long": false }
  }
}

4. Generate

A talking animation is an animation with speech. It starts from item_id or source_image_asset_id and ends on end_image_asset_id, or on its start image when you leave that out. variant_id works as for any animation and picks that variant's voice.

curl -X POST https://api.masko.ai/v1/mascots/MASCOT_ID/generate \
  -H "Authorization: Bearer $MASKO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "type": "animation",
    "item_id": "ITEM_ID",
    "end_image_asset_id": "END_IMAGE_ID",
    "speech": {
      "script": "[waves hello, excited] Hi! I'\''m Gubby. [opens an arm to the right] Here'\''s how the Masko API works."
    }
  }'
{
  "data": {
    "job_id": "3c4d5e6f-7a8b-4c9d-8e0f-1a2b3c4d5e6f",
    "status": "pending",
    "type": "animation",
    "animation_model": "standard",
    "item_id": "4b5c6d7e-8f90-4a1b-9c2d-3e4f5a6b7c8d",
    "item_name": "Hi, I'm Gubby",
    "estimated_cost": 18,
    "estimate": { "speech_seconds": 4.3, "duration": 6, "max_duration": 6, "credits": 18, "too_long": false },
    "script": "[waves hello, excited] Hi! I'm Gubby. [opens an arm to the right] Here's how the Masko API works.",
    "asset_ids": {
      "video": "5d6e7f80-91a2-4b3c-8d4e-5f6a7b8c9d0e",
      "webm": "6e7f8091-a2b3-4c4d-9e5f-6a7b8c9d0e1f",
      "hevc": "7f8091a2-b3c4-4d5e-8f6a-7b8c9d0e1f20",
      "audio": "8091a2b3-c4d5-4e6f-9a7b-8c9d0e1f2031",
      "transcript": "91a2b3c4-d5e6-4f7a-8b8c-9d0e1f203142"
    },
    "poll_url": "/api/v1/jobs/3c4d5e6f-7a8b-4c9d-8e0f-1a2b3c4d5e6f"
  }
}

A talking animation costs 3 credits per second of video. The length is only known once the voice is recorded, so the request charges the longest the line could need (estimate.max_duration). The unused seconds are refunded as soon as the voice is recorded, and the job's cost_credits shows the settled price.

Poll the job as for any animation. Send the same Idempotency-Key to retry a request safely.

The transcript

The transcript asset is a JSON file. Read it with GET /v1/assets/{id} and fetch its file_url. Times are seconds from the first frame of the video, so they include the short silence before the first word. Lines end at ., ! or ?; use them for captions.

{
  "text": "Hi! I'm Gubby. Here's how the Masko API works.",
  "language": "en",
  "duration": 6,
  "words": [
    { "text": "Hi!", "start": 0.26, "end": 0.6 },
    { "text": "I'm", "start": 0.92, "end": 1.04 },
    { "text": "Gubby.", "start": 1.16, "end": 1.5 },
    { "text": "Here's", "start": 2.04, "end": 2.24 },
    { "text": "how", "start": 2.26, "end": 2.38 },
    { "text": "the", "start": 2.42, "end": 2.5 },
    { "text": "Masko", "start": 2.54, "end": 2.9 },
    { "text": "API", "start": 3.0, "end": 3.46 },
    { "text": "works.", "start": 3.56, "end": 3.9 }
  ],
  "lines": [
    { "text": "Hi! I'm Gubby.", "start": 0.26, "end": 1.5 },
    { "text": "Here's how the Masko API works.", "start": 2.04, "end": 3.9 }
  ],
  "movements": [
    { "text": "waves hello, excited", "start": 0.26, "end": 2.04 },
    { "text": "opens an arm to the right", "start": 2.04, "end": 3.9 }
  ]
}

The audio asset is an MP3 of the voice alone. The video assets report has_audio: true in their metadata.

Limits

  • speech needs a start image: item_id or source_image_asset_id.
  • The line decides duration and loop, so a request with speech refuses duration, loop, reverse, auto_reverse, sizes, animation_prompt, image_prompt and the image crops. Talking animations use the standard model; animation_model: "premium" is refused.
  • A mascot without a voice returns 409 with details.reason: "voice_required", before any charge.
  • A line longer than about 14 seconds of speech returns 400; split it into two animations.
  • Talking animations cannot be reversed or edited, because the voice would not follow. To get a new take, send the same script again.
  • generate-batch does not accept speech; send each talking animation to /generate.