If you’re running Claude Code, Codex, or Cursor at any real scale, your AI token bill is probably the fastest-growing line item in your engineering budget. Vilnius-based nexos.ai just launched a smart LLM router designed to fix that without touching developer workflows.
The Problem It Solves
Coding agents like Claude Code autonomously decide which model to call and when. That makes cost optimization hard. You can’t just swap to a cheaper model across the board: cheap models fail on complex planning tasks, and expensive frontier models get used for trivial edits that don’t need them.
As nexos.ai head of product Žilvinas Girėnas put it:
“Point it at one cheap model and quality goes down on hard tasks. Point it at one expensive model and you end up paying Opus prices for requests the agent itself considered Haiku-grade work. The answer is knowing what the agent is actually trying to do at any given moment.”
What the Production Numbers Show
In production testing, nexos.ai found that coding agent traffic splits cleanly into two buckets:
- 16% of requests (Planning): Routed to frontier reasoning models like Claude Opus to decide what to build and how.
- 84% of requests (Editing): Routed to cost-efficient open-weight models like Kimi or GLM to apply diffs, write files, and execute the existing plan.
The cost impact was material. One workload that would have cost over $9,200 at frontier-model prices came in at roughly $3,800 after routing, a 59.2% reduction. A second test hit 60.4%. Quality held because the expensive model still handled every decision that determined the outcome.
Mirror Benchmarking
The routing decisions are powered by what nexos.ai calls Mirror benchmarking: continuous evaluation of live production traffic at the session level, rather than static public benchmarks. The router tracks how sessions grow, where costs land, and how real-world failures occur to build a picture of actual customer demand patterns. It reads each request without modifying it and switches models only at natural session breakpoints to keep the cache intact and avoid costly rebuilds.
Company Background
nexos.ai was founded in 2024 by Tomas Okmanas and Eimantas Sabaliauskas, who also co-founded Nord Security (a $3B cybersecurity unicorn) and Oxylabs. The company raised an $8M seed round in early 2025 from Index Ventures, Creandum, Dig Ventures, and angel investors. It is based in Lithuania and originated within the Tesonet tech accelerator ecosystem.
