HERO RUN
Live routing

Every model, routed to the fastest and cheapest gateway.

Hero Run runs one catalog across many inference gateways. When several serve the same model, we route to the cheapest and fail over to the next on error. Hero goes further: it reads your prompt and picks a right-sized model. These numbers come from real traffic and update continuously.

0
Models
one catalog
0
Gateways
routed live
0
With failover
2+ gateways
0
Hero routes
prompts routed

Gateway speed

Throughput and time-to-first-token (TTFT), measured from live streamed runs.

GatewayRunstok/sTTFT
Loading…

What Hero routes

Hero scores each prompt and sends it to a right-sized tier: simple work to a fast cheap model, hard reasoning to a frontier model, at one flat price.

0
Cheap
simple prompts
0
Mid
everyday work
0
Frontier
hard reasoning
Try Hero Mint an API key

TTFT is time to the first streamed token, measured on real runs; tok/s is effective throughput including latency. Sample sizes vary by gateway (see Runs). Numbers reflect the models actually routed through each gateway, not a fixed benchmark model.