A task-aware model router reduced the median cost of its Open SWE coding agent from $2.61 to $0.94 per thread. Instead of sending every request to the most capable model, the router chose among fast, balanced and performance tiers after reading the first user message.
The company tested the approach across 973 live threads. Routed tasks produced a merged pull request 29.2 percent of the time, compared with 27.3 percent for requests always sent to the strongest model. That difference was not statistically significant. Median cost fell 64 percent, while mean cost dropped 42 percent and the 90th-percentile cost dropped 37 percent.
Most work avoided the top tier: 56 percent went to the balanced model, 34 percent to the fast model and 10 percent to the performance model. LangChain argues that routing belongs inside an agent harness because it can use task, tool and domain context that a generic model gateway may not have.
The current router makes one choice at the beginning and keeps that model for the full thread. It does not yet adapt when a conversation becomes harder, and the production test relied mainly on merge rates and sparse user feedback. LangChain plans controlled benchmark testing and routing for subagents before treating the design as settled.