Jev Router Wins on Steerability but Costs 38% More Than Direct Calls

Arena.ai put TypeSafe's Jev Router through 4,700+ agentic sessions and found smart routing choices that still don't beat calling DeepSeek directly.

·
·
·
Jev Router Wins on Steerability but Costs 38% More Than Direct Calls
Read4 min
TypeNews
  • Arena.ai evaluated TypeSafe's Jev Router on Agent Arena across 4,700+ real agent sessions.
  • Jev posts +3.6% net improvement at $0.155 median cost per task, below DeepSeek V4.1 Flash.
  • Costs 38% more than DeepSeek V4.1 Flash (Max) with 1.7x higher median model latency.
  • Steerability score of +10.0% nearly matches Claude Opus 5.5 (High) at +10.48%.
  • Routes most calls to DeepSeek V4.1 Flash (28.1%), GPT-6 Astra (23.5%), GPT-6.1 Sol (22.8%).
  • P90 end-to-end latency hits 30.59s versus 14.26s for DeepSeek direct.

Jev Router gains steerability at a cost

Model routers aim to reduce inference costs by sending straightforward requests to cheaper models and difficult requests to stronger ones. A router must choose well enough to offset its own latency, fee, and effect on downstream token usage.

An Agent Arena benchmark tested Jev Router across more than 4,700 real-world agent sessions. The results show sensible model selection and strong responses to user corrections, alongside higher cost and latency than a direct call to the router’s most-used model.

Direct calls win the baseline

Treated as an Agent Arena leaderboard entry, Jev recorded a +3.6% net-improvement score at a median cost of $0.155 per task. DeepSeek V4.1 Flash ranked higher. At comparable end-task performance, the router cost 38% more than calling DeepSeek V4.1 Flash (Max) directly, while median request latency was 1.7 times higher.

Latency percentile Jev Router DeepSeek V4.1 Flash (Max)
Median 6.18 seconds 3.64 seconds
P90 30.59 seconds 14.26 seconds

The tail-latency gap exceeded the median gap: Jev’s 90th-percentile response time was more than twice DeepSeek’s. Each routed request adds a serial decision before generation begins, while the selected model can also change downstream latency and token usage. The aggregate results do not attribute the cost difference to individual components.

Course corrections favor Jev

Jev’s clearest advantage appeared in steerability, which measures how well an agent responds when a user pushes back or changes direction during a session. Its +10.0% score trailed Claude Opus 5.5 (High) by 0.48 percentage points.

Metric Jev Router DeepSeek V4.1 Flash
Steerability +10.0% -0.2%
Praise vs. Complaint +6.3% +2.1%
Confirmed Success +2.2% +8.3%
Bash Recovery -0.7% +7.8%
Tool Hallucination +0.2% +0.4%

Agent Arena reports these as signed benchmark scores, with higher values indicating better results. Jev led the user-feedback-oriented measures, while DeepSeek performed better on confirmed completion, shell-error recovery, and the tool-hallucination metric. Correction-heavy interactive products may value Jev’s profile more than batch or autonomous systems do.

The route mix favors efficient models

The routing log also shows which models Jev selected during production-style agent loops. The four largest shares accounted for 85.5% of all calls:

  • DeepSeek V4.1 Flash: 28.1%
  • GPT-6 Astra: 23.5%
  • GPT-6.1 Sol: 22.8%
  • GPT-6 Luna: 11.1%
  • Other models: 14.5%

The three named OpenAI models received 57.4% of calls, and no Anthropic model appeared among the four largest routes. Agent Arena also reported that four of the five most-used models sat on the cost-performance Pareto frontier, where no available alternative was both cheaper and better.

That distribution documents an economically plausible policy. Measuring the value of each routing decision would require request-level comparisons showing whether a cheaper model could have produced the same result.

A cheap decision still adds a hop

TypeSafe’s technical overview describes Jev as a decision model that returns structured choices and calibrated confidence scores in one forward pass, meaning one model evaluation. Its published token price is about $0.042 per million input tokens, with no output-token charge.

Each routed request therefore involves two sequential operations: Jev selects a model, then the chosen model handles the task. Jev’s direct fee is small at the published rate. Total task cost can still rise through the selected model mix, longer outputs, repeated tool calls, or other downstream behavior.

A production routing policy must balance four measurements together:

  1. Task completion: whether the agent produces a verified result.
  2. Steerability: whether it incorporates corrections during the session.
  3. Total cost: routing, generation, retries, and tool use.
  4. End-to-end latency: especially P90 and other tail percentiles.

An independent RouterArena experiment found a related limitation. A Jev-based router matched or exceeded the best single model on accuracy, while an ablation that removed Jev and retained the retrieval evidence achieved the same gain. In that setup, retrieval accounted for the measured improvement.

Test the system against one strong model

Teams evaluating a router should establish a direct-call baseline with the strongest low-cost model in the candidate pool, then replay representative production traces through both configurations. The comparison should include:

  • verified completion and failure rates;
  • performance after user corrections;
  • tool errors, shell recovery, and hallucinated calls;
  • median, P90, and timeout latency;
  • total cost per completed task, including retries; and
  • results segmented by interactive, batch, and autonomous workloads.

In Agent Arena’s test, direct DeepSeek V4.1 Flash delivered comparable aggregate performance with lower cost and roughly half the tail latency. Jev’s strongest case comes from interactive workloads where users frequently redirect the agent and the steerability gain outweighs the additional request hop.

Trending
  • No trending articles

Comments

avatar

Next Reads