Home › Guides › Production pipeline

Architecture for an async video API in production

Updated 2026-10-02

The tutorial version of a video API call is a script that submits, sleeps in a loop and downloads. That works for ten clips. It breaks when the process restarts mid-poll, when two workers submit the same request, or when a finished video's URL is gone by the time someone looks. This page describes the smallest architecture that avoids those failures.

Principle: your database is the source of truth

The API gives you a job id and a status. Treat both as a cache of external state. Your own table records what you intended, what you submitted, and what you stored. Every worker is then a stateless loop over that table, and a crash loses nothing.

CREATE TABLE video_jobs (
  id            uuid PRIMARY KEY,           -- your id, created before any API call
  state         text NOT NULL,              -- pending | submitted | polling | stored | failed | unknown
  request       jsonb NOT NULL,             -- model, prompt, inputs, duration_secs
  upstream_id   text,                       -- the id returned by POST /videos
  attempts      int  NOT NULL DEFAULT 0,
  next_poll_at  timestamptz,
  output_key    text,                       -- your storage path, not the provider URL
  error         text,
  created_at    timestamptz DEFAULT now(),
  updated_at    timestamptz DEFAULT now()
);

The three workers

1. Submitter. Claims rows in pending (using SELECT ... FOR UPDATE SKIP LOCKED so two workers never take the same row), calls POST /videos, and writes upstream_id and state submitted in the same transaction boundary as closely as you can manage. Creation is billed once, so this is the step you must not repeat. If the call outcome is ambiguous (a dropped connection), set unknown instead of retrying; see the error cookbook.

2. Poller. Selects rows whose next_poll_at has passed, calls GET /videos/{id}, and reschedules. Polling is free, so the cost of being generous is only your own load; a few seconds between polls with a little jitter is plenty. Use a per-job deadline, and let a job that exceeds it move to unknown for review rather than resubmitting.

3. Archiver. When status is completed, stream the file from the returned URL into your own object storage and record output_key. Do this immediately. Provider URLs are delivery links, not storage, and their lifetime is not something to build on. Only mark the row stored after the copy succeeds and its size is non-zero.

Skeleton

import time, random, requests

BASE = "https://videorouter.sh/api/v1"
H = {"Authorization": "Bearer " + KEY}

def submit_one(db):
    row = db.claim("pending")                    # FOR UPDATE SKIP LOCKED
    if not row: return
    try:
        r = requests.post(BASE + "/videos", json=row.request, headers=H, timeout=(5, 30))
    except requests.RequestException:
        db.set(row.id, state="unknown", error="ambiguous create")
        return
    if r.status_code in (429, 500, 502, 503, 504):
        wait = float(r.headers.get("Retry-After", 0)) or 10
        db.requeue(row.id, delay=wait)           # safe: no job was created
    elif r.ok:
        db.set(row.id, state="polling", upstream_id=r.json()["id"],
               next_poll_at=db.now_plus(5))
    else:
        db.set(row.id, state="failed", error=r.text)   # 4xx: do not retry

def poll_one(db):
    row = db.claim_due("polling")
    if not row: return
    j = requests.get(BASE + "/videos/" + row.upstream_id, headers=H, timeout=(5, 30)).json()
    if j["status"] == "completed":
        db.set(row.id, state="archive", error=None)
        archive(row, j["data"][0]["url"])
    elif j["status"] == "failed":
        db.set(row.id, state="failed", error=str(j["error"]))
    else:
        db.set(row.id, next_poll_at=db.now_plus(5 + random.random() * 2))

Keep each function small and the loop outside it, so you can run them as separate processes, cron jobs or queue consumers; the database table is the only coordination you need.

Concurrency and backpressure

Retries at the pipeline level

A failed job caused by upstream failure is not billed, so an automatic re-enqueue with a new row and a small attempt cap is reasonable. A job that completed but looks bad is a billed clip; a re-render is a new, deliberate spend and belongs behind a review decision, not an automatic loop.

What to measure

MetricWhy
Time from create to completed, per modelSizes your deadlines and tells you which models suit interactive use
Failure rate by error classSeparates your bugs (400s) from upstream trouble (5xx)
Cost per accepted clipThe only cost number that includes retries and rejects
Jobs in unknownEach one is potential double-billing; keep the count near zero
Archive lagCompleted-but-not-stored age; alert before a link could expire

Log the job id, model, and the provider reported in the response with every state change, so a billing question can be traced to a row.

Deployment shape

You do not need a heavyweight orchestrator. One process running the three loops on threads or asyncio is a fine start; split them into separate services when one needs to scale independently, typically the archiver, which moves the most bytes. Make every loop safe to run twice concurrently, which the row-claiming query already gives you, and make every state transition conditional on the previous state so a late duplicate event cannot move a row backwards. When you outgrow polling for notification purposes, add a webhook receiver that simply sets next_poll_at to now, and keep the poller as the fallback, so a missed notification never strands a job.

Choosing models and hosts in the pipeline

Put the model id in config, not code. An unpinned model id routes to the cheapest healthy host and falls back across hosts for submission failures; use a provider object only where policy requires a specific host. Switching a model for a workload then becomes a config change you can A/B on cost per accepted clip.

The state machine behind the API is in async jobs explained, and the request itself in the quickstart. Get a key, then start with a single worker and the table above.

Frequently asked questions

Why keep my own job table instead of just polling the API?

It survives restarts, prevents duplicate submissions across workers, and records where each output is stored. The API's job id is external state; your row is the source of truth.

Should I store the provider's video URL?

No. Copy the file into your own object storage as soon as the job completes and store your own key. Provider URLs are delivery links, not durable storage.

How often should the poller call the API?

Polling is free, so a few seconds between polls with jitter and a per-job deadline is fine. Avoid holding connections open.

Can failed jobs be retried automatically?

A job that failed because every upstream host failed is not billed, so a bounded automatic retry is reasonable. Re-rendering a completed but unwanted clip is a new charge and should be a deliberate decision.

Keep reading

Using AI video is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →