API keys and spend limits for an AI video API
Updated 2026-10-02
Video is the expensive end of generative AI: one request can cost real money, and a loop with a bug can submit hundreds. Authentication is therefore not just "who is calling" but "how much damage can this credential do". This page covers the controls VideoRouter keys carry and a setup that keeps a mistake cheap.
The key and how it is sent
Keys look like llmr_sk_live_<random> and go in a standard bearer header on every request to https://videorouter.sh/api/v1:
Authorization: Bearer llmr_sk_live_...
The plaintext secret is shown once, at creation. Only a hash plus a prefix and last four characters are stored, so you cannot recover a lost key; you replace it. A missing, malformed, revoked or expired key returns 401. New accounts must verify their email before a key can be created.
Per-key policy: where the real protection lives
Every key carries its own policy, set at creation and editable later. There is no separate budget system to wire up. The fields that matter for video:
| Field | Default | What it does for you |
|---|---|---|
monthly_spend_cap_usd | uncapped | Hard ceiling on the key's spend for the month. Hitting it returns 402 with code spend_cap_exceeded. |
model_allowlist | all models | Restricts the key to named model slugs. A request for anything else gets 403 model_not_allowed. |
rpm_limit | unlimited | Requests per minute, enforced per key as a token bucket. Exceeding it gives 429 with a Retry-After header. |
expires_at | never | Set an expiry in days at creation; the key returns 401 after that date. |
One caveat to verify for your own setup: the documented scopes field lists chat and embeddings endpoints, and the documentation does not spell out a video scope. Do not rely on scopes to fence off video; use the model allow-list and spend cap, which are model- and money-based and apply regardless. Test the behaviour with a throwaway key before you depend on it.
A key layout that bounds blast radius
- One key per environment. Separate keys for local development, CI, staging and production. A leaked dev key should not be able to spend the production budget.
- Tight caps in non-production. A development key with a small monthly cap and an allow-list of one cheap model turns a runaway loop into a
402instead of an invoice. - Allow-list the models you actually ship. If production uses two models, list two. This also stops a typo or a prompt-injected model name from selecting a premium model.
- One key per service or tenant where you need attribution. Usage is attributable per key, which is far easier than reverse-engineering a shared one.
- Short-lived keys for experiments. Use
expires_atfor contractors and hack-week keys.
The two independent budget checks
A key's monthly cap and your organisation's prepaid credit balance are checked separately. A 402 from either looks similar, so branch on the code: spend_cap_exceeded means this key hit its own ceiling (raise it or wait for the month), insufficient_credits means the balance is empty (top up). Your alerting should distinguish them, because the owners differ.
import requests
def create_job(payload: dict, key: str) -> dict:
r = requests.post(
"https://videorouter.sh/api/v1/videos",
headers={"Authorization": f"Bearer {key}"},
json=payload,
timeout=30,
)
if r.status_code == 402:
code = r.json()["error"].get("code")
raise BudgetError(code) # page the right owner: cap vs credits
r.raise_for_status()
return r.json()
Secrets hygiene
- Load keys from the environment or a secrets manager. Never commit them, never bake them into a front-end bundle or mobile app: anything the client can read, a user can read.
- Call the API from your server and expose your own narrow endpoint to the browser, where you can apply per-user quotas the key policy cannot express.
- Keep keys out of logs. Log the prefix and last four characters, which is what the dashboard shows.
- Rotate with the dashboard's roll action: it issues a new secret that inherits the same policy and revokes the old one in the same transaction, so there is never a window where both work. Deploy the new secret promptly, because the old one stops immediately.
- Revoke, do not just delete from code, when a key may have leaked.
Minting keys per customer
If you build a product on top and want a key per end customer, the platform documents a separate provisioning-kind key that manages standard keys over /api/v1/keys. It cannot call inference endpoints itself and can only see keys that spend against your own organisation's balance. It is issued on request rather than self-serve.
Auditing and offboarding
Treat key management as a recurring task, not a one-time setup. List your active keys every quarter, match each one to an owner and a running service, and revoke any you cannot account for. When an engineer leaves or a contractor finishes, rolling the keys they could see is cheaper than wondering whether a copy survives on a laptop. Because a rolled key inherits its policy, rotation does not require re-deriving caps and allow-lists, which removes the usual excuse for postponing it. Pair rotation with a deploy step that reads the new secret from your secrets manager at startup, so a roll does not need a code change.
What the platform will not do for you
Caps bound money, not abuse inside your own product. If end users can trigger generations, add your own per-user rate limit and cost estimate before calling the API. And remember that a request is billed once at creation from the requested duration, so a retry that creates a second job is a second charge; the error-handling cookbook covers how to avoid that.
Next: build the worker side in a production pipeline, or see the job lifecycle. Create an account and set a cap on your first key before sending traffic.
Frequently asked questions
How do I limit how much a single API key can spend?
Set monthly_spend_cap_usd on the key. When the ceiling is reached, requests return a 402 with the code spend_cap_exceeded.
Can I restrict a key to specific video models?
Yes. The key's model_allowlist limits which model slugs it can call; others return 403 with model_not_allowed.
What happens to the old key when I roll it?
A new secret is issued with the same policy and the old one is revoked in the same transaction, so it stops working immediately.
Should I use the same key in development and production?
No. Use separate keys per environment, with a small cap and a narrow model allow-list on development keys.
Keep reading
- Async Video Jobs Explained — Polling, Timeouts and Retries
- Video API Error Handling Cookbook: Retries, Backoff, Duplicates
- Production Video Generation Pipeline: Queue, Workers, Python
- Python Video API Client: A Typed Wrapper With Retries
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →