Routing Best Practices#
Use these patterns with Routing Rules after providers are connected. Create each rule from Add a routing rule and verify the result in Traces.
Target and fallback providers must be available to the requesting user. Routing does not override configured provider access.
Request types#
OptScale AI distinguishes two chat request types in routing conditions:
- Chat Completion — Non-streaming chat completion requests.
- Chat Completion Stream — Streaming chat completion requests, including OptScale AI Chat.
Use:
request_type == "chat_completion_stream"for Chat-only traffic.request_type == "chat_completion" || request_type == "chat_completion_stream"for both streaming Chat and non-streaming API clients.
A condition of only Request type = Chat Completion matches non-streaming API requests only. It does not match OptScale Chat.
Rule Builder patterns#
Match Chat traffic#
request_type == "chat_completion_stream"
Match Chat and API traffic#
request_type == "chat_completion" || request_type == "chat_completion_stream"
Combine conditions#
Use AND when all conditions must match and OR when any condition may match.
For example:
(request_type == "chat_completion" || request_type == "chat_completion_stream")
&& model == "gpt-4o"
Use nested groups only when the routing condition cannot be expressed clearly with a flat set of rules.
For configuration steps, see Add a routing rule.
Routing patterns#
Primary provider with fallback#
Use this pattern when one provider should handle matching requests and another should take over if the primary target cannot process the request.
Example:
- Scope — Organization
- Condition — Chat and API traffic
- Target — Provider A /
gpt-4o - Fallback — Provider B /
gpt-4o-mini
request_type == "chat_completion" || request_type == "chat_completion_stream"
The target and fallback providers must both be available to the requesting user.
To test fallback, temporarily disable the primary target provider or model, send another matching request, and confirm Traces shows the expected fallback destination.
Weighted traffic split#
Use this pattern to distribute matching requests across multiple provider/model targets.
Example:
- Target 1 — Provider A /
gpt-4o/ weight0.7 - Target 2 — Provider B /
gpt-4o/ weight0.3
Target weights must sum to 1.
Routing is selected independently for each request; the split is not session-sticky.
Use a fallback separately when failed requests should be retried through another destination. A fallback is not an additional weight bucket.
Prefer a lower-cost model#
Use this pattern when users select an expensive model and a cheaper model on the same provider should handle matching requests.
Example:
- Condition — Chat and API traffic for
gpt-4o - Target — Provider A /
gpt-4o-mini - Fallback — Provider A /
gpt-4o
Fall back to the original model if the cheaper path fails.
Scope-specific overrides#
Use more specific scopes when a team or individual user requires routing different from the organization default.
Routing scope precedence is:
Employee → Team → Organization
Within the same scope, lower Priority values are evaluated first.
Typical patterns:
- Team override — Apply a different target to a pilot or specialized team.
- Employee exception — Route one user to a specific provider for testing or troubleshooting.
Keep overrides narrow. Prefer Disable over Delete when pausing an override you may reuse.
User-specific exception#
Use Scope = Employee when one user must always reach a specific provider, for example during troubleshooting.
Keep the exception narrow and disable it when it is no longer required.
Multi-provider failover#
Use this pattern for ordered recovery across providers.
Example:
- Target — Provider A /
gpt-4o - Fallbacks — Provider B /
gpt-4o, then Provider C
The gateway tries fallbacks in list order until a request succeeds or no fallbacks remain.
Note
Use Fallbacks for failure recovery. Use weighted Targets when traffic should be distributed across providers during normal operation.
Migrate traffic between providers#
Use this pattern when clients still send requests to Provider A, but matching traffic should be handled by Provider B without a client-side configuration change.
Example:
- Condition — Chat and API traffic where Provider is Provider A
- Target — Provider B
- Fallback — Provider A
(request_type == "chat_completion" || request_type == "chat_completion_stream")
&& provider == "Provider A"
If Provider B cannot process a request, the gateway retries it using Provider A.
Match a model family#
Use regex only when multiple model identifiers must be matched by a shared pattern.
Example:
model.matches("^gpt-4o")
Prefer an exact model match when possible.
Canary rollout#
Use a team-scoped weighted rule to test a new provider with limited exposure.
Example:
- Scope — Team
- Target 1 — Existing provider / weight
0.9 - Target 2 — New provider / weight
0.1 - Fallback — Existing provider
Validate the rollout in Traces before expanding the rule to the organization scope.
Verify routing#
- Confirm the rule is Enabled and that the requesting user can access every target and fallback provider.
- Send matching Chat or API requests.
- In Traces, confirm the applied routing rule, provider, and model.
- For fallback, temporarily disable the primary target and confirm the fallback destination in Traces.
For the full verification procedure, see Verify routing.
See also#
- Routing Rules — Configure rule scope, conditions, targets, and fallbacks.
- Allowed providers — Confirm that users can access target and fallback providers.
- Traces — Confirm which provider and model handled a request.
- Architecture Overview — Providers and routers — Review when a routing rule is required.