A retry policy for an async video API, error by error
Updated 2026-10-02
Retry logic for a video API has one extra constraint compared with most APIs: creation is billed, in full, once, at the moment the job is created. A naive "retry on any failure" wrapper can create a second paid job while the first is still running. The goal is to retry only where it is safe, and to make the unsafe cases impossible to hit by accident.
The error envelope
Errors use an OpenAI-style body: {"error": {"message": ..., "type": ..., "code": ...}}. Branch on the HTTP status first and code second.
| Status / code | Meaning | Policy |
|---|---|---|
| 400 | Unknown model, or fields the model rejects (for example start_image_url on a model without image input) | Never retry. Fix the request. |
| 401 | Missing, revoked or expired key | Never retry. Alert. |
402 spend_cap_exceeded | This key's monthly cap is reached | Never retry. Raise cap or wait. |
402 insufficient_credits | Organisation balance is empty | Never retry. Top up. |
403 model_not_allowed | Model is outside the key's allow-list | Never retry. Fix config. |
| 429 | Per-key rate limit | Sleep Retry-After seconds, then retry. |
| 500 / 502 / 503 / 504 | Every candidate in the fallback chain failed | Retry with backoff. Not billed. |
Rate-limit buckets are smooth rather than per-minute cliffs, so honouring Retry-After exactly is better than guessing a fixed sleep.
Retrying job creation
import random, time, requests
BASE = "https://videorouter.sh/api/v1"
RETRY_STATUS = {429, 500, 502, 503, 504}
def create_job(payload, key, max_attempts=5):
for attempt in range(max_attempts):
try:
r = requests.post(f"{BASE}/videos", json=payload, timeout=(5, 30),
headers={"Authorization": f"Bearer {key}"})
except requests.ConnectionError:
# No response at all. The job MAY have been created. See below.
raise AmbiguousCreate()
if r.status_code == 202 or r.ok:
return r.json()
if r.status_code not in RETRY_STATUS:
raise ApiError(r.status_code, r.json().get("error"))
wait = float(r.headers.get("Retry-After", 0)) or min(60, 2 ** attempt)
time.sleep(wait + random.uniform(0, 1)) # jitter
raise ApiError(503, "exhausted retries")
Notice what is not retried: connection errors and read timeouts on creation. If the connection drops after the server accepted the job, you will not know whether a job exists. That is the duplicate-cost hazard.
Idempotency and duplicate jobs
The documentation does not describe an idempotency-key header for video creation, so assume none and build the protection yourself:
- Write your own job row first with a client-generated id, status
pending_create, and the full request payload. - Create the upstream job, then store the returned
idon that row. - On an ambiguous failure, mark the row
unknownand have a human or a conservative policy decide. Do not blindly re-submit; one extra clip is a full extra charge. - Deduplicate in your own layer: a hash of (model, prompt, inputs, duration) with a short window prevents double-clicks and queue redeliveries from creating twins.
Two platform behaviours add to this. A job that ends failed because every upstream host failed is not billed, so retrying that is safe. And the opt-in hedge failover.on_timeout_sec deliberately resubmits to the next host without cancelling the first, billing both if both land. Do not enable it unless you want that trade, and never combine it with your own timeout-resubmit logic.
Polling with a deadline
def wait_for(job_id, key, deadline_s=600, interval_s=5):
start = time.monotonic()
while True:
r = requests.get(f"{BASE}/videos/{job_id}", timeout=(5, 30),
headers={"Authorization": f"Bearer {key}"})
if r.status_code in RETRY_STATUS: # transient on the status call
time.sleep(interval_s); continue # safe: reading is free and idempotent
r.raise_for_status()
job = r.json()
if job["status"] == "completed":
return job["data"][0]["url"]
if job["status"] == "failed":
raise JobFailed(job["error"])
if time.monotonic() - start > deadline_s:
raise PollTimeout(job_id) # job may still finish: do NOT resubmit
time.sleep(interval_s)
Polling is free and idempotent, so retrying the status call aggressively is fine. The distinction between a safe read retry and an unsafe create retry is the main idea of this page. Slow is not failed: on PollTimeout, keep the job id, resume polling later from a worker, and only resubmit after the job reports failed.
Timeouts that match the call
- Create: short connect timeout, modest read timeout (the call returns a job id, not a video).
- Poll: short per-request timeout, long overall deadline sized to the model. Fast models and premium ones differ widely; measure your own p95 rather than copying a number.
- Download: separate, larger timeout, and copy the file to your own storage promptly.
Testing the failure paths
Retry code that has never run is a liability. Wrap the HTTP call behind a small interface and test your policy with a fake that returns each status in the table: a 429 with a Retry-After, a run of 503s followed by success, a 402 with each code, and a connection error on create. Assert two things: the 400 family is never retried, and the ambiguous create never produces a second POST. You can also provoke real errors cheaply: a bogus model id returns a 400 and an invalid key a 401, and neither costs anything.
Classify failures into three buckets
Permanent (400/401/402/403): surface to the developer, no retry. Transient (429, 5xx): backoff with jitter, bounded attempts. Ambiguous (connection reset on create): park for reconciliation. Most production incidents with paid async APIs come from treating the third bucket as the second.
Keys and caps that turn 402s into a safety net are in authentication and spend controls; the worker architecture that stores these job rows is in the pipeline article. See also async jobs and the quickstart.
Frequently asked questions
Which errors should I retry on a video API?
Retry 429 (after Retry-After) and 5xx with exponential backoff and jitter. Never retry 400, 401, 402 or 403; they need a request, key or budget fix.
Can a retry create a duplicate paid job?
Yes. Creation is billed once at creation, and no idempotency key is documented, so a blind re-submit after an ambiguous failure can be charged twice. Track your own job rows and reconcile unknown states.
What should I do when polling times out?
Keep the job id and resume polling later. A slow job is not a failed job; resubmit only after the status is failed.
Are failed jobs billed?
A job that fails because every upstream host failed is not billed. A request that is accepted and then produces an unusable clip is still a normal billed clip.
Keep reading
- Async Video Jobs Explained — Polling, Timeouts and Retries
- Video API Authentication and Spend Controls: Keys, Caps, Limits
- Production Video Generation Pipeline: Queue, Workers, Python
- Python Video API Client: A Typed Wrapper With Retries
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →