The Ship-Ready Output Bar — 500 Bankers, 22% Jury Skepticism, and the $1B Proof That Owning the Loop Is Everything
The Quality Gap Nobody Measured
A relationship manager at a bank opens a chat window, types a prompt about a client portfolio, gets a paragraph back, and rewrites it before sending. She does this four times a day. She is not a skeptic. She is doing the work the model did not finish. A survey of 500 banking professionals this week found AI outputs consistently unusable for client-facing communication. Jury selection in the Musk v. OpenAI trial surfaced the same pattern from the other end of the labor market: 22% of non-tech workers in San Francisco — nurses, caretakers, painters — say AI makes their work slower because they have to double-check output. Platform analytics are still counting sessions. The metric that matters is whether the output shipped without a human rewrite.
The bottleneck in AI product adoption has shifted from model capability to the distance between model output and a send-ready artifact.
Who Solved It and Who Didn't
Replit went from $2.8M to ~$1B ARR in 18 months with 300% net revenue retention. A user types a prompt, sees a working prototype, edits the parts that are wrong, and deploys. All of that happens on one surface. Cursor, pitched into the same category, reportedly runs -23% gross margins and is exploring a $60B exit to SpaceX. The model is not the difference. The loop is. Replit owns compute, deployment, collaboration, and the AI layer. Cursor is an AI panel over VS Code. When the product is a thin interface on someone else's foundation model, inference cost is the P&L.
The platform giants reached the same conclusion. Google Gemini now produces documents, spreadsheets, and presentations inside the chat interface. Microsoft embedded AI contract agents into Word that surface clauses, risks, and changes at a glance. Mistral launched Workflows to connect models to business processes. Standalone AI chat is a dead-end UX. The teams that are winning are shipping finished artifacts inside the tools where work already happens.
The Reddit Counter-Example
Reddit posted $663M revenue, beat estimates by $53.2M, and reported a 30% YoY jump in weekly search users driven by Reddit Answers. The feature summarizes human forum posts. This works because the output is the artifact. The user reads the summary and leaves. There is no draft to rewrite. Search DAUs, WAUs, and queries are up meaningfully, which is what "ship-ready AI output" looks like when the artifact is small enough to fit in the answer.
The Diagnostic
Draw a 2x2. One axis: does the AI output ship to the next person without a human rewriting it. Other axis: does the user come back next week. The cell with revenue in it is "ships without rewrite, user returns." A product sitting in "needs rewriting, user returns" is a tolerated workflow, not a moat, and it collapses the week a competitor crosses into the top-right. The 500 bankers are not a verdict on AI in finance. They are a verdict on which teams decided the quality bar was someone else's problem.
What to do
Audit every AI feature that requires copy-paste, reformat, or human rewrite before the output is usable. Tag the top 3 by usage volume by end of this sprint.
Redesign the highest-usage AI feature to produce an inline, editable artifact rather than chat output. Ship a prototype within two sprints.
Replace 'sessions per user' and 'prompts per week' with 'time-to-usable-output' and 'output acceptance rate' as primary AI feature metrics. Instrument by end of June.