# Seed Audio 1.0 All-in-one audio generation: dialogue, music, ambience, and sound effects in one pass, with voice cloning from up to 3 reference clips (under 30s each). The model field picks the variant: multilingual (default, 20 auto-detected languages, in-prompt timestamp control) or base (English/Chinese with the preset voice library). - Model ID: `bytedance/seed-audio-1.0/text-to-audio` - Provider: ByteDance - Modality: text → audio - Type: Text to Audio - Status: stable - Delivery: async - Output format: mp3 ## Authentication Send your API key in the `X-API-Key` header. The same key reads the org balance at `GET https://api.dev.pika.art/billing/balance`, which is free to call and confirms the request can be paid for before you submit it. ## Endpoint ``` POST https://api.dev.pika.art/v1/media/bytedance/seed-audio-1.0/text-to-audio ``` ## Reference uploads To use a local file (image, audio, or video) as a model input, first request a presigned upload URL: `POST /v1/media/uploads` with the `X-API-Key` header and a JSON body `{ "content_type": "image/png", "size_bytes": 12345 }`, where `size_bytes` is the file's exact byte length. The response returns `upload_url` (temporary, valid for 5 minutes), `headers`, and `url` (the permanent Pika URL). PUT the file bytes to `upload_url` sending the returned `headers` verbatim — both `Content-Type` and `Content-Length` are signed, and storage returns a bare 403 if either is missing or differs. Then pass `url` into fields such as `image`, `image_url`, `end_image_url`, `image_urls`, `first_frame_image`, or `mask_image_url`. ## Generate `POST /v1/media/bytedance/seed-audio-1.0/text-to-audio` ```bash curl -X POST https://api.dev.pika.art/v1/media/bytedance/seed-audio-1.0/text-to-audio \ -H "X-API-Key: YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "prompt": "No speech. A thunderstorm approaching: distant rumble building to a close crack, heavy rain, wind through trees, 20 seconds.", "output_format": "mp3", "speed": 1, "volume": 1, "pitch": 0, "enable_subtitles": false }' ``` ### Request body | Param | Type | Required | Default | Description | | --- | --- | --- | --- | --- | | `model` | enum | no | - | Model variant. Leave unset to infer: `seed-audio-1.0` when a preset `voice` is given, else `seed-audio-1.0-multilingual`. Set explicitly to override. `seed-audio-1.0-multilingual`: 20 languages auto-detected, voice via cloning references only. `seed-audio-1.0`: English/Chinese with the preset voice library (the `voice` field). | | `prompt` | string | yes | - | What to generate — dialogue, music, ambience, and sound effects in one pass (multilingual variant: any of 20 languages, auto-detected). Reference supplied clips as @Audio1...@Audio3. Output up to ~120 seconds. | | `voice` | enum | no | - | Preset speaker (base variant only — requires model=seed-audio-1.0; the multilingual variant's speaker map differs). Counts toward the 3 voice/audio reference cap. | | `audio_urls` | string[] | no | - | Up to 3 voice-cloning reference clips, EACH UNDER 30 SECONDS; refer to them in the prompt as @Audio1...@Audio3. Cannot be combined with image_url. | | `image_url` | string | no | - | A single image reference (the generated voice/scene matches it). Cannot be combined with voice or audio_urls. | | `output_format` | enum | no | mp3 | Output Format | | `sample_rate` | enum | no | - | Sample Rate | | `speed` | number | no | - | Speed | | `volume` | number | no | - | Volume | | `pitch` | integer | no | - | Pitch | | `enable_subtitles` | boolean | no | false | Return per-sentence and per-word timestamps for the generated speech. When true, the completed job's audio output carries a `subtitle` asset — a JSON URL with {text, sentences[], words[]} (start_time/end_time in ms). | ### Accepted values | Param | Values | Default | | --- | --- | --- | | `model` | `seed-audio-1.0-multilingual`, `seed-audio-1.0` | — | | `output_format` | `mp3`, `wav`, `pcm`, `ogg_opus` | `mp3` | | `sample_rate` | `8000`, `16000`, `24000`, `32000`, `44100`, `48000` | — | Returns a job object, not the final output — normally with status `queued`. Store the `id` and poll until a terminal state. A rejected submit (insufficient balance, rate limit, unpriceable input) still returns the job object, with status `failed` and `error` set. An `Idempotency-Key` replay returns the existing job in its current state — including `failed`, so retry a failure with a fresh key, never the same one. ### Response **200 OK · queued** ```json { "id": "media_8f3a2c91-5b7d-4e0a-9c26-31d4f2a8e6b0", "status": "queued" } ``` ## Poll status `GET /v1/media/jobs/{request_id}` ```bash curl https://api.dev.pika.art/v1/media/jobs/{request_id} \ -H "X-API-Key: YOUR_API_KEY" ``` Poll the job by id until it reaches a terminal state: `completed` or `failed`. ### The job object - `id` (string): Unique identifier for the job. - `status` (enum): One of `queued`, `running`, `completed`, `failed`. - `output` (object): Present once the job completes. - `media_type` (enum): `audio`. - `output.audio.url` (string): URL of the generated audio. - `error` (object): Present if the job failed. `error.code` is the stable machine-readable value to branch on — one of `invalid_input`, `content_moderation`, `provider_error`, `provider_timeout`, `provider_unavailable`, `rate_limited`, `insufficient_balance`, `membership_required`, `cycle_limit_exceeded`, `admission_suspended`, `timed_out`, `internal`. Unrecognized values normalize to `provider_error`. `error.message` is diagnostic text and varies. ### Response **200 OK · running** ```json { "id": "media_8f3a2c91-5b7d-4e0a-9c26-31d4f2a8e6b0", "status": "running" } ``` **200 OK · completed** ```json { "id": "media_8f3a2c91-5b7d-4e0a-9c26-31d4f2a8e6b0", "status": "completed", "output": { "media_type": "audio", "audio": { "url": "https://api.dev.pika.art/v1/files/audio_8f3a2c91.mp3", "content_type": "audio/mpeg" } } } ``` **200 OK · failed** ```json { "id": "media_8f3a2c91-5b7d-4e0a-9c26-31d4f2a8e6b0", "status": "failed", "error": { "code": "provider_error", "message": "upstream generation failed" } } ``` ## Get result `GET /v1/media/jobs/{request_id}/content` ```bash curl https://api.dev.pika.art/v1/media/jobs/{request_id}/content \ -H "X-API-Key: YOUR_API_KEY" ``` Once the job completes, fetch a download URL for the generated media. ### Response **200 OK · content URL** ```json { "url": "https://api.dev.pika.art/v1/files/audio_8f3a2c91.mp3" } ``` ## Errors Errors use conventional HTTP status codes and always return JSON, in one of two shapes. Request-level errors (auth, unknown path, schema validation, idempotency conflict) are `{"message": "..."}`. Submit rejections that occur after the job row is created (insufficient balance, rate limit, unpriceable input, dispatch unavailable) return the full failed job envelope — branch on `error.code`, and retry with a fresh `Idempotency-Key`, since a failed job replays on its key. - 401 Unauthorized — The API key is missing or invalid. Body: `{"message":"Invalid API key"}` - 403 Forbidden — An inactive key is rejected with a `{"message"}` body. A submit that the org balance or postpaid cycle limit cannot cover is rejected after the job row exists and returns the failed job envelope. `GET /billing/balance` separates the two: a `200` means the key is active, so the rejection was the balance or the cycle limit. Body: `{"id":"media_8f3a2c91-5b7d-4e0a-9c26-31d4f2a8e6b0","status":"failed","error":{"code":"insufficient_balance","message":"Insufficient org balance"}}` - 404 Not Found — No job with that id exists in your org, or the media path names an unknown vendor/model/function. Body: `{"message":"media job not found"}` - 409 Conflict — The result was requested before the job completed, or an Idempotency-Key header was reused with a different body. Body: `{"message":"media job is not ready"}` - 422 Unprocessable Entity — The request body failed validation: an invalid enum value, a missing required field, a wrong type, or malformed JSON. The JSON message names the offending field. A schema-valid parameter combination that cannot be priced instead fails after job creation and returns the failed job envelope with error code `invalid_input`. Body: `{"message":"duration: Input should be less than or equal to 15"}` - 429 Too Many Requests — Org limit reached: requests per minute or day, or concurrent jobs. Returned as the failed job envelope; check the Retry-After header and retry with a fresh Idempotency-Key. Body: `{"id":"media_8f3a2c91-5b7d-4e0a-9c26-31d4f2a8e6b0","status":"failed","error":{"code":"rate_limited","message":"rate limit exceeded: rpm"}}` - 503 Service Unavailable — The model rail or a backend dependency is temporarily unavailable. Returned as the failed job envelope; retry with backoff and a fresh Idempotency-Key. Body: `{"id":"media_8f3a2c91-5b7d-4e0a-9c26-31d4f2a8e6b0","status":"failed","error":{"code":"provider_unavailable","message":"media dispatch unavailable"}}` ## Pricing - default: $0.15 per minute Only successful generations are charged. Check the balance against this price before submitting: `GET https://api.dev.pika.art/billing/balance` returns `balance_micro_usd`, which is US dollars multiplied by 1,000,000, and `postpaid` when the org bills by invoice instead. For an active postpaid org, compare the price against `postpaid.cycle.remaining_micro_usd` rather than the prepaid balance. Report a shortfall to the user rather than submitting a request the balance cannot cover. --- Model page: https://dev.pika.art/models/bytedance/seed-audio-1.0/text-to-audio API index: https://dev.pika.art/llms.txt