Video
Text to video
Which text-to-video model fits the shot you have in mind, prompts that actually work, and the full API run from create to finished file.
One sentence in, one MP4 out. That is the whole promise of text-to-video, and it is close enough to true that a two-person team can ship a product teaser, a looping hero background or forty ad variants in an afternoon.
What nobody tells you is that the eleven text-to-video models here are not interchangeable. One of them renders 4K. One writes its own soundtrack. One runs thirty seconds but never goes above 720p. Pick by the shot you have in mind, not by the leaderboard — and since all of them run on one key and one balance, you can put the same prompt through three of them and keep the winner.
What people build with it
| What you are making | What actually matters | Start with |
|---|---|---|
| Product teaser, hero loop | resolution, clean camera motion | veo3.1-quality |
| Vertical social ad | cost per variant, speed | veo3.1-fast, pixverse-v6 |
| A beat with dialogue or ambience | audio in the same render | doubao-seedance-2.0 |
| B-roll and establishing shots | length, believable camera | sora2, sora-2-official |
| Ambient background behind a headline | duration over sharpness | grok-imagine-1.5-video |
| Testing forty hooks before picking one | price per clip | veo3.1-fast, doubao-seedance-2.0-fast |
One habit is worth more than any model choice: draft on a fast model until the prompt is right, then run the final prompt once on the expensive one. The prompt transfers. The money spent on twelve throwaway 4K renders does not.
Which model, and when
veo3.1-quality — the resolution ceiling. It and veo3.1-fast are the only
models that accept resolution: "4k". Use it when the clip will be seen
full-bleed on a desktop hero or cut into real footage. veo3.1-fast is the
same family at draft speed: find the prompt there, then upgrade.
sora2 / sora2-pro — the long take. 10, 15 or 25 seconds, with the most
coherent camera work of the set: dolly-ins, orbits and handheld follows tend to
survive to the end of the clip instead of dissolving halfway through.
sora2-pro is the same model with more compute behind it.
sora-2-official — the duration ladder. 4, 8, 12, 16 or 20 seconds, priced
per second of what you actually requested. When the clip has to drop into an
exact slot on a timeline, this is the one that lets you buy exactly that many
seconds instead of trimming.
doubao-seedance-2.0 — the one that speaks. generateAudio defaults to
true, so the render arrives with room tone, footsteps and, if the prompt asks
for it, a line of dialogue already mixed in. That removes a whole sound pass.
doubao-seedance-2.0-fast trades some fidelity for turnaround, and
doubao-seedance-1-5-pro is the previous generation — still the cheapest of
the three.
pixverse-v6 — the knob-heavy one. negativePrompt, seed, motionMode,
watermark, and an audio switch that defaults to false. seed is the real
reason to care: fix it, change one word in the prompt, and you see the effect
of that one word instead of an unrelated re-roll. That is how a shot gets
tuned rather than gambled on.
grok-imagine-1.5-video — the long one. Up to 30 seconds, capped at 720p.
For a muted loop behind a headline nobody will ever notice the resolution, and
you get three times the runtime.
wan2.7-video — the literal one. 5, 10 or 15 seconds at 720p or 1080p.
Worth trying when a prompt keeps coming back over-stylised elsewhere; it stays
closer to plain photographic realism.
Writing a prompt that renders
A prompt that works reads like a shot list, not a wish. Five slots, in this order, carry almost all of the result:
subject → what it does → how the camera moves → light and time of day → look.
"A lighthouse" gives you a lottery ticket. "A lighthouse on a basalt cliff at dawn, slow push-in, fog rolling over the water, shot on 35 mm" gives you a shot. Three more things worth knowing before you spend anything:
- Motion belongs in the prompt, not in your hopes. Name the move — slow push-in, orbit left, static locked-off, handheld follow. Models that are not told how the camera behaves invent a drift, and it is usually the wrong one.
- Negatives only work where they exist. Only
pixverse-v6takes anegativePrompt. Everywhere else, saying "no text, no watermark" in the prompt can just as easily summon text. Describe what you want instead. - One idea per clip. Two actions in ten seconds is where models start morphing objects mid-shot. Cut it into two generations and edit them together.
Recipes
Each of these is a full prompt you can paste, with the model it was written for and what tends to come back.
One shot, described in five parts — veo3.1-quality,
aspectRatio: "16:9", resolution: "720p":
Close up shot (composition) of melting icicles (subject) on a frozen rock wall(context) with cool blue tones (ambiance), zoomed in (camera motion) maintainingclose-up detail of water drips (action).
The prompt names composition, subject, context, ambience, camera motion and action without padding. The result keeps all six instructions in one coherent macro shot.
Source: Google AI for Developers, used under CC BY 4.0; format and dimensions adapted.
Vertical ad hook, spoken line — doubao-seedance-2.0, size: "9:16",
resolution: "1080p", duration: 8, generateAudio: true:
A barista slides a paper cup across the counter and says "your usual, right?",warm indoor light, handheld, slight rack focus onto the cup, cafe ambience
Comes back with the ambience and the line already in the audio track. Lip sync is decent at conversational pace and falls apart on long sentences — keep spoken lines under about six words.
Detail adds control — veo3.1-quality, aspectRatio: "16:9",
resolution: "720p":
A close-up cinematic shot follows a desperate man in a weathered green trenchcoat as he dials a rotary phone mounted on a gritty brick wall, bathed in theeerie glow of a green neon sign. The camera dollies in, revealing the tensionin his jaw and the desperation etched on his face as he struggles to make thecall. The shallow depth of field focuses on his furrowed brow and the blackrotary phone, blurring the background into a sea of neon colors and indistinctshadows, creating a sense of urgency and isolation.
Every detail points in the same direction: the dolly-in, shallow focus, green neon and facial tension all sell urgency. Long prompts work when their clauses reinforce one shot instead of asking for more events.
Source: Google AI for Developers, used under CC BY 4.0; format and dimensions adapted.
Subject and context only — veo3.1-quality, aspectRatio: "16:9",
resolution: "720p":
A satellite floating through outer space with the moon and some stars in thebackground.
This is enough to produce a coherent clip, but framing and motion are now the model's decision. Compare it with the recipes above: brevity buys speed, while camera language buys control.
Source: Google AI for Developers, used under CC BY 4.0; format and dimensions adapted.
Tuning a shot instead of re-rolling it — pixverse-v6, size: "16:9",
duration: 5, seed: 42, negativePrompt: "text, watermark, distorted hands":
A cyclist crests a hill at dusk, silhouette against an orange sky, cameratracking alongside, dust in the air
Keep the seed, change dusk to overcast noon, re-run: same composition, same rider, different light. That is the loop that turns generation into art direction.
About these samples. Ready-made examples are copied to our own storage only when their source permits reuse. The exact prompt appears above each result, and the original source and license are linked below the media.
Pick a model
Every model has its own input DTO, so the parameter names differ between families. The table below covers the text-to-video capable public models.
| Model | Duration (s) | Resolution | Ratio field | Audio switch |
|---|---|---|---|---|
veo3.1-fast | provider default | 720p 1080p 4k | aspectRatio | none |
veo3.1-quality | provider default | 720p 1080p 4k | aspectRatio | none |
sora2 | 10 15 25 | provider default | aspectRatio | none |
sora2-pro | 10 15 25 | provider default | aspectRatio | none |
sora-2-official | 4 8 12 16 20 | provider default | aspectRatio | none |
doubao-seedance-2.0 | 4–15 | 480p 720p 1080p | size | generateAudio |
doubao-seedance-2.0-fast | 4–15 | 480p 720p 1080p | size | generateAudio |
doubao-seedance-1-5-pro | 4–15 | 480p 720p 1080p | size | generateAudio |
pixverse-v6 | 1–15 | 360p 540p 720p 1080p | size | audio |
grok-imagine-1.5-video | 6 10 15 20 25 30 | 480p 720p | size | none |
wan2.7-video | 5 10 15 | 720p 1080p | size | none |
Full parameter tables live in Video models.
Aspect ratio, resolution and duration
Three fields carry most of the framing, and two of them are spelled differently depending on the model.
aspectRatio is used by sora2, sora2-pro, sora-2-official,
veo3.1-fast and veo3.1-quality. It accepts exactly two values:
| Value | Shape |
|---|---|
16:9 | landscape (the default) |
9:16 | portrait |
size is used by everything else and the accepted set is wider — Seedance
takes 16:9, 9:16, 1:1, 4:3, 3:4 and 21:9; Pixverse adds 2:3 and
3:2; wan2.7-video takes 16:9, 9:16, 1:1, 4:3 and 3:4.
aspectRatioandsizeare not interchangeable. Sendingsizetoveo3.1-fast— oraspectRatiotopixverse-v6— is an unknown field for that DTO and the request fails validation with 400invalid_input. There is no silent fallback.
resolution is a string enum whose casing follows the model. Veo uses 720p,
1080p and 4k (lower-case k); Seedance, Pixverse, Grok and Wan use only
p tiers. duration is always an integer number of seconds, and for most
models it is a fixed enum rather than a free range — see the table above.
Submit the request
The request body is always the same three keys: model, input and an
optional webhook.
The example below asks veo3.1-quality for a 4K landscape shot.
curl -X POST https://api.apihubs.ru/api/v1/generation/create \-H "Authorization: Bearer sk-your-key" \-H "Content-Type: application/json" \-d '{"model": "veo3.1-quality","input": {"prompt": "A lighthouse on a basalt cliff at dawn, slow push-in, fog rolling over the water","aspectRatio": "16:9","resolution": "4k"}}'
The response comes back before any provider work has started.
{"code": 200,"data": {"taskId": "019ca881-9503-7270-a560-f00fc2b15785","status": "not_started","createdAt": "2026-02-28T11:22:33.000Z"}}
The balance is debited at this point, not on completion. If the balance
will not cover the request the call returns 402 insufficient_balance and no
task is created. A job that ends in failed is refunded automatically.
The same job on doubao-seedance-2.0, which uses size and can render its own
soundtrack:
{"model": "doubao-seedance-2.0","input": {"prompt": "A lighthouse on a basalt cliff at dawn, fog rolling over the water","size": "16:9","resolution": "1080p","duration": 8,"generateAudio": true,"cameraFixed": false}}
Poll for the result
GET /api/v1/task/status/{taskId} returns the same envelope at every stage.
The status walks not_started → processing → finished or failed.
curl https://api.apihubs.ru/api/v1/task/status/019ca881-9503-7270-a560-f00fc2b15785 \-H "Authorization: Bearer sk-your-key"
While the provider is working:
{"code": 200,"data": {"taskId": "019ca881-9503-7270-a560-f00fc2b15785","status": "processing","createdTime": "2026-02-28T11:22:33.000Z"}}
When it lands:
{"code": 200,"data": {"taskId": "019ca881-9503-7270-a560-f00fc2b15785","status": "finished","files": [{"fileUrl": "https://storage.apihubs.ru/generated/video-abc123.mp4","fileType": "video"}],"createdTime": "2026-02-28T11:22:33.000Z"}}
files is only populated when status is finished. Treat its absence as
"not ready", never as "no output". Each fileUrl points at storage and is
valid for 24 hours — download it, do not hot-link it. See
Files & storage.
A complete run
The loop below submits, polls with a fixed interval and a wall-clock deadline, and distinguishes the three ways a job can end: success, provider failure and your own timeout.
const BASE = "https://api.apihubs.ru/api/v1";const KEY = process.env.API_STOCK_KEY!;type TaskFile = { fileUrl: string; fileType: "image" | "video" | "music" };type TaskData = {taskId: string;status: "not_started" | "processing" | "finished" | "failed";files?: TaskFile[];output?: Record<string, unknown>;createdTime: string;errorMessage?: string;};async function createVideo(prompt: string): Promise<string> {const res = await fetch(`${BASE}/generation/create`, {method: "POST",headers: {Authorization: `Bearer ${KEY}`,"Content-Type": "application/json",},body: JSON.stringify({model: "veo3.1-quality",input: { prompt, aspectRatio: "16:9", resolution: "4k" },}),});const body = await res.json();if (!res.ok) {// { code, error: { message, type, code } }throw new Error(`${body.error.code}: ${body.error.message}`);}return body.data.taskId; // → "019ca881-9503-7270-a560-f00fc2b15785"}async function waitForTask(taskId: string,{ intervalMs = 10_000, timeoutMs = 20 * 60_000 } = {},): Promise<TaskData> {const deadline = Date.now() + timeoutMs;while (Date.now() < deadline) {const res = await fetch(`${BASE}/task/status/${taskId}`, {headers: { Authorization: `Bearer ${KEY}` },});const body = await res.json().catch(() => null);if (!res.ok) {const message = body?.error?.message ?? `${res.status} ${res.statusText}`;throw new Error(message);}const data = body.data as TaskData;if (data.status === "finished") return data;if (data.status === "failed") {throw new Error(data.errorMessage ?? "generation failed");}await new Promise((r) => setTimeout(r, intervalMs));}// The task is still alive server-side — resume polling later with the id.throw new Error(`timed out waiting for ${taskId}`);}const taskId = await createVideo("A lighthouse on a basalt cliff at dawn, slow push-in, fog rolling over the water",);const task = await waitForTask(taskId);console.log(task.files?.[0].fileUrl); // → https://storage.apihubs.ru/...mp4
Video renders run in minutes, not seconds. A 10-second poll interval is reasonable; anything under a second only burns rate-limit budget. See Rate limits for the 120 requests per 60 seconds ceiling.
Let the result come to you
Polling is optional. Pass a webhook URL on create and the finished task is
POSTed to it, byte-identical to the task-status envelope.
{"model": "sora2","input": {"prompt": "Neon-lit rain on an empty parking garage, handheld","aspectRatio": "9:16","duration": 15},"webhook": "https://example.com/hooks/api-stock/8f2c1e9a-secret"}
There is no signature header on the delivery, so the URL itself has to carry the secret. Delivery is retried 13 times with exponential backoff starting at 10 seconds — roughly 22 hours of attempts. Full details in Webhooks.
When it fails
Two categories of failure, and they arrive through different channels.
Request-time errors are HTTP errors on the create call, with the standard envelope:
{"code": 400,"error": {"message": "Input parameters do not match the requirements for model veo3.1-quality","type": "BadRequest","code": "invalid_input"}}
Match on error.code, never on error.message — the code is stable, the
message is prose. invalid_input means the body will never validate, so
retrying it unchanged is pointless.
Generation-time errors land on the task, with HTTP 200 and
status: "failed":
{"code": 200,"data": {"taskId": "019ca881-9503-7270-a560-f00fc2b15785","status": "failed","createdTime": "2026-02-28T11:22:33.000Z","errorMessage": "This generation may violate our content policy. Please try a different prompt."}}
By the time you see failed, the platform has already exhausted its own
retries — the same provider up to its attempt limit, then any reserve provider
capable of the request — and refunded the debit. A content-policy
rejection is terminal and never retried; changing the prompt is the only fix.
Next steps
- Image to video — animate a still instead of starting from text.
- Video models — the exhaustive parameter tables.
- Polling task status and Webhooks.
- Errors & error codes — the full code list.
- Pricing & balance — how a render is priced.
Run this with your own key
Every model in this guide is live in the API Hubs catalog — one API key, one prepaid balance, no separate signup per provider.