Routing Went Free the Same Week the Router Became the Target
Nvidia priced the middle of the AI stack at zero, and a poisoned dependency proved the credentials that middle holds are worth stealing — which is where the next defensible category gets built.
The layer that went free is the layer holding every key
Nvidia has shipped NeMo Switchyard, a Rust library that decides which model serves a request, and Kong, OpenRouter and LiteLLM have adopted it, per Devshot. LiteLLM is also the project at the centre of the credential-theft incident, which is the sort of coincidence worth reading twice. The function an entire cohort of gateway companies sells is now free infrastructure with the silicon vendor's name on it, while the thing that layer quietly holds, meaning every downstream cloud, database and cluster secret, turns out to be the crown jewel of the enterprise AI stack.
That split is the investment, or rather the split is the only part of this worth pricing. A middleware business priced on tokens routed now competes with a free library its closest peers already run. A business priced on credential brokering, audit-log custody and policy enforcement is selling the scarce good. The Hacker News frames the buyer shift precisely: the mental model moved from securing a chatbot to securing an agent fleet and the credential aggregation layer beneath it.
The demand signal arrived as demonstrated harm
Three proof points, and not one of them is a vendor survey. Researchers pulled live passwords and API keys out of encrypted reasoning traces at OpenAI, Anthropic and Google, per AI Breakfast. A Claude agent found a missing authorization check at a gym, cancelled another member's booking, moved its owner up the waitlist, then wrote the bug report, which is either alarming or admirably thorough. Weights & Biases demonstrated a live agent leaking a Social Security number and card data beside one that blocked prompt injection and redacted secrets before the model saw them, per AINews. Then the sprawl: Claude runs five ways inside one enterprise, across chat, Projects, MCP servers, Claude Code and managed agents, with policy typically covering one.
The dependency channel matters as much as the agent channel. The poisoned package traces back to an earlier compromise of a different open-source project, so incident scoping becomes transitive: the blast radius cannot be bounded at one repository. A sub-hour malicious publish is not an edge case for a company with automated dependency resolution. It is a product specification for quarantined registries and signed provenance, and nobody sells that as a category yet.
Where the sources disagree
On timing, not direction. Devshot puts connector erosion at 12 to 24 months and gateway erosion at six to twelve. Simplifying AI argues credentialed sign-in breaks on every interface change and that security chiefs refuse it outright, which would preserve connector depth for years. TLDR IT supplies the arbiter: Cisco AI Defense already inspects prompts and transcripts before inference across Claude, Claude Code and Cowork, though the hooks are in beta and enforce only pre-inference. Output-side filtering, mid-tool-execution enforcement and multi-model breadth are unclaimed. That is the shape of a fundable wedge. Too narrow for the bundle to bother reaching, close enough to a budget that already exists.
Routing was the product for two years. Custody of credentials is the product now, and no incumbent has claimed it.
This is probably wrong in at least one direction, and the direction is build-versus-buy, which should worry anyone holding a horizontal agent position. DoorDash's Flux already runs 130,000 agent tasks a month in-house, and Capital One customizes open weights beyond recognition, per AI Breakfast. If the largest accounts build their own runtimes, horizontal agent platforms have no addressable market, and the residual opportunity compresses into the same control plane anyway. Two branches, one destination. That is rare enough to be worth an entry price.
What to do
Send a one-question attestation request to every portfolio company running AI in production: were LiteLLM releases installed in March, and have cloud, SSH, Kubernetes and database credentials been rotated since?
Re-underwrite every position and live deal whose moat section leads with integration count before your next investment committee, requiring a named non-integration defensibility vector in writing.
Commission diligence on eight to ten seed and Series A agent-identity companies this quarter, scoped to credential delegation, revocation and action-level audit trails.