Soft host preference vs hard pin: choosing how tightly to route
Updated 2026-10-02
By default a gateway picks the host for you, and for most traffic that is the right answer. But sometimes you need to say "this host, or none". VideoRouter gives you two different tools for that, they behave very differently when something goes wrong, and mixing them up is a common source of either surprise outages or surprise bills. This page lays out both, with the failure modes of each. The authoritative reference is the provider selection docs.
The default: no routing hints at all
Send just model and prompt. The request goes to the cheapest healthy host for that checkpoint. If that host rejects the submission, it retries once on the same host, then walks every other host confirmed to serve the exact checkpoint, cheapest first, until one accepts. This costs you no configuration, which is why everything below is optional.
Tool one: the soft preference, model/host
Append a host segment to the model id:
{ "model": "minimax/h3/fal", "prompt": "a paper airplane gliding over a city" }
The documentation is explicit that this is a preference, not a pin. The platform tries that host first, and if it errors, still falls back across the other hosts serving the same checkpoint. Use it when you want a particular host most of the time but would rather get a result from somewhere else than an error. Typical reasons: you have measured that host to be consistent on your prompts, or you want repeat runs to land on one host for easier comparison.
Tool two: the hard pin
A hard pin uses the provider object:
{
"model": "alibaba/wan-3.0",
"prompt": "a paper airplane gliding over a city",
"provider": { "only": ["fal"], "allow_fallbacks": false }
}
only is an allow-list of provider slugs and can name several. allow_fallbacks: false then disables retries within that model's candidates entirely. Per the documentation, only the cheapest surviving candidate from the allow-list is ever attempted, and a rejected submission comes back as an error immediately: no same-host retry, and none of the other hosts in your list tried either. It is exactly one attempt. The same page notes that a models[] fallback list is a separate mechanism and may still be consulted.
Side by side
Soft: model/host | Hard: only + allow_fallbacks: false | |
|---|---|---|
| Host tried first | The named one | Cheapest in your allow-list |
| On rejection | Retry, then other hosts | Error returned immediately |
| Can run on a host you did not name | Yes | No |
| Availability | Higher | Lower, you own the failure |
| Predictability of where data goes | Lower | Higher |
Middle options you may actually want
Two settings sit between "anywhere" and "exactly one attempt". provider.only without allow_fallbacks: false narrows the eligible set but keeps the normal retry walk over that set. provider.ignore does the reverse: exclude hosts you do not want and let everything else compete. And provider.sort or provider.policy reorders the same walk by price, latency, reliability or queue length instead of price alone. For most compliance or consistency needs, an allow-list of two or three trusted hosts with fallbacks left on is a better trade than a hard single-host pin.
When to hard pin
- Data handling. You have approved specific hosts for user content and a request must never go elsewhere.
- Controlled comparisons. You are benchmarking one host, and a silent fallback would contaminate the result.
- Account-specific reasons. A host you want exercised for contractual or reconciliation purposes.
- Cost certainty. You want to be sure a request is never billed at another host's rate.
When none of those apply, stay with the default or the soft preference.
Failure modes of a hard pin
- One host's bad hour is your bad hour. With no fallback, a transient rejection becomes an error to your user. Your client needs its own retry policy, and since a rejected submission is not a created job, retrying a 5xx is safe under the usual rules in the error-handling cookbook.
- Typos in slugs. A misspelled or unsupported host slug leaves no candidates. Test each pin once with a short cheap request before shipping it.
- Models change hosts. A pinned host may stop serving a model. Alert on the specific error rather than discovering it from user complaints.
- Over-restrictive hints may be ignored. The documentation says an over-restrictive
onlyor similar filter degrades to "ignore this filter" in some cases rather than failing, so confirm behaviour on your own request instead of assuming a pin held. Check which host actually served the job. - Hedging does not mix. The opt-in
failover.on_timeout_sechedge resubmits to the next-cheapest host without cancelling the first and bills both if both finish, which contradicts a pin's intent.
A sensible escalation path
- Start with no hints.
- If you need consistency, add a soft
model/hostpreference. - If you need a boundary, use
onlywith several hosts and fallbacks on. - Only then use
allow_fallbacks: false, and pair it with client-side retries and monitoring.
For the surrounding client code, see the Python wrapper, which passes a provider dict straight through. The quickstart covers the basics, and you can create a key to try a pin against a short clip.
Frequently asked questions
What is the difference between model/host and provider.only?
A model/host suffix is a soft preference that tries that host first but still falls back to others. provider.only with allow_fallbacks false is a hard pin that makes exactly one attempt within your allow-list and returns an error if it is rejected.
Can provider.only list more than one host?
Yes. It is an allow-list, so you can name several hosts; with allow_fallbacks false only the cheapest surviving candidate from that list is attempted.
Is a pinned job that fails billed?
A rejected submission does not create a job. In general, a job that fails because every upstream host failed is not billed.
Should I hard pin by default?
No. The default routes to the cheapest healthy host with automatic fallback. Pin hard only for data-handling, controlled comparison or contractual reasons.
Keep reading
- Async Video Jobs Explained — Polling, Timeouts and Retries
- Video API Authentication and Spend Controls: Keys, Caps, Limits
- Video API Error Handling Cookbook: Retries, Backoff, Duplicates
- Production Video Generation Pipeline: Queue, Workers, Python
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →