Security & Threat Intelligence

The Watch

The Signal

Anthropic shipped Claude Computer Use this week

Simultaneously, ByteDance's DeerFlow 2.0 (bash terminal, persistent memory, autonomous sub-agent spawning) hit #1 on GitHub Trending.

In Play

  1. AI Agents Gain Physical Desktop Control — With Known Prompt Injection Risk

    Anthropic's Claude Computer Use controls macOS screens, cursors, and apps — including Slack and Workspace — with acknowledged prompt injection risk. ByteDance's DeerFlow 2.0 runs bash terminals in Docker with persistent memory. Anthropic's Dispatch creates a phone-to-desktop remote C2 channel. EDR/XDR stacks have zero detection for AI UI automation patterns.

    Ask Clarity
  2. AI Vendor Governance Crisis: Safety Deprioritized, Political Risk Emerges

    OpenAI's CEO stepped back from safety oversight during the 'Spud' model launch push. A federal judge reversed the Trump administration's designation of Anthropic as a 'supply chain risk.' Both Anthropic ($60B IPO, Q4 2026) and OpenAI face governance transitions. AI vendor safety is becoming a variable cost cut under competitive pressure.

    Ask Clarity
  3. Shadow AI Agent Adoption Hits Measurable Enterprise Scale

    Microsoft reports 62% of UK businesses are already running AI agents; 84% of security leaders flag unauthorized 'shadow agents.' Pinterest published a production MCP governance architecture (registry approval, dual-identity auth, centralized discovery) — the first credible reference model. Business thought leaders are explicitly framing IT security as the bottleneck to circumvent.

    Ask Clarity
  4. Capital Markets Migrating to Blockchain Rails — Largest New Financial Attack Surface

    DTCC ($3.7Q annual volume), NYSE, Tradeweb, and Nasdaq are actively deploying on-chain settlement with H1 2026 targets. Tradeweb already executed Saturday Treasury financing on-chain with BofA, Citadel, and Virtu. Atomic settlement eliminates the T+1 fraud detection window. Middleware layer is explicitly unbuilt.

    Ask Clarity

Deep Dives

Claude Computer Use ships prompt injection to your desktop — and three more agent attack surfaces landed this week

The Week AI Agents Got Physical

Four distinct AI agent capabilities shipped this week that your existing security controls were not built to detect or contain. This isn't a theoretical escalation of previous shadow AI concerns — these are specific products with specific attack surfaces that need immediate assessment.

1. Claude Computer Use: Prompt Injection → Desktop Takeover

Anthropic's Computer Use gives Claude direct control of macOS screens, cursors, and applications — including Slack and Google Workspace. It first tries API-level connectors, then falls back to raw UI automation: clicking, typing, reading screen content. It operates in the background via Claude Cowork and Claude Code, and supports scheduled recurring tasks.

The critical detail: Anthropic explicitly warns about prompt injection, advising users not to let Claude access sensitive data during this research preview. This means a malicious instruction embedded in any document, email, webpage, or Slack message that Claude processes could redirect the agent to exfiltrate data, send messages, or modify files — all under the user's legitimate identity and session.

Anthropic shipped a desktop automation agent, then published a warning that its own product can be hijacked by adversarial content. Your employees will adopt it anyway.

2. Anthropic Dispatch: Phone-to-Desktop C2

The new Dispatch tool lets users text tasks from their phone to have Claude execute them on their desktop. From a security perspective, this is a remote command-and-control channel to workstations that bypasses your VPN, MDM, and conditional access policies. A compromised phone → Dispatch → desktop automation chain gives an attacker hands-on-keyboard-equivalent access through a legitimate product.

3. DeerFlow 2.0: ByteDance's Agent Framework With Bash Access

ByteDance's open-source autonomous agent framework hit #1 on GitHub Trending. It runs 100% locally in isolated Docker sandboxes with persistent filesystem and bash terminal, spawns parallel sub-agents autonomously, and maintains persistent cross-session memory. The layered risks: supply chain exposure from a ByteDance-originated repository, container escape risk from Docker sandboxes with bash access, memory poisoning across sessions, and autonomous sub-agent spawning that amplifies any single compromise.

Detection Gap Analysis

Attack VectorProductDetection Gap
Prompt Injection → Desktop ActionsClaude Computer UseEDR/XDR not tuned for AI UI automation
Remote C2 via Legitimate ProductAnthropic DispatchMDM/EDR lacks Dispatch visibility
Container Escape + Memory PoisoningDeerFlow 2.0Standard dependency scanning misses agent-specific risks

MITRE ATT&CK mapping: T1059 (Command Interpreter via UI automation), T1071 (Application Layer Protocol via Slack/Workspace), T1041 (Exfiltration over C2 via Dispatch), T1204 (User Execution — user grants agent access).


The Pinterest MCP Governance Model

Pinterest published the most detailed enterprise reference architecture for securing AI agent tool access: registry-based approval, layered authentication (user JWTs + service identities), centralized discovery, and full audit logging. The dual-identity model is essential — agents act on behalf of users but with service-level access. Without this separation, you cannot distinguish between a user's legitimate request and an agent's autonomous lateral movement. If you're deploying MCP infrastructure, this is your benchmark.

What to do

  1. Inventory Claude Pro/Max subscriptions across your macOS fleet and issue guidance restricting Computer Use from accessing corporate apps until prompt injection mitigations are validated

  2. Assess whether your EDR/XDR detects AI agent UI automation patterns (programmatic cursor movement, rapid app switching, screen reading) and whether DLP covers the Dispatch mobile-to-desktop pipeline

  3. If developers are using DeerFlow 2.0, audit Docker sandbox configurations, dependency chain, and persistent memory storage mechanisms against your open-source governance policy

  4. Adopt Pinterest's MCP governance pattern for any planned agent infrastructure: registry-based approval, dual-identity auth, centralized discovery, audit logging

AI vendor safety is now a variable cost — OpenAI drops oversight, Anthropic faces political targeting, products vanish overnight

Three Governance Shocks in One Week

Three distinct events this week converge on a single conclusion: your implicit trust in AI vendor safety governance is no longer defensible. Each one independently warrants a third-party risk register update; together, they demand a structural reassessment of how you model AI vendor risk.

1. OpenAI's CEO Disengages From Safety

Sam Altman announced he's stepping back from direct oversight of safety and security teams to focus on infrastructure, capital raising, and the imminent launch of a frontier model codenamed 'Spud'. The combination of maximum capability push plus minimum CEO-level safety attention is the governance equivalent of running with scissors during a sprint.

This is not an isolated event — it's part of a pattern where safety is subordinated to capability during competitive pressure. The difference now is that it's happening during what OpenAI itself describes as an AGI race reaching 'fever pitch.'

2. Anthropic Designated — Then Un-Designated — as Supply Chain Risk

The Trump administration attempted to designate Anthropic as a 'supply chain risk,' ordering federal agencies to cut ties. A federal judge reversed it on free speech grounds. But the precedent stands: an AI vendor can be politically targeted with an overnight designation that disrupts enterprise access. If your production workflows depend on Claude APIs, you now carry political concentration risk that didn't exist 12 months ago.

3. Products Vanish Under Compute Pressure

OpenAI's Sora shutdown — killing a $1B Disney partnership struck just three months prior — was already covered, but the pattern now has a second data point: OpenAI is consolidating ChatGPT, Codex, and Atlas browser into a single desktop superapp, with product lines being killed to reallocate GPUs for Spud. If OpenAI will break a billion-dollar deal to reallocate compute, your API dependency is not safe from the same calculus.

When your AI vendor's CEO publicly steps away from safety oversight during the biggest capability push in the company's history, that's not a press release — it's a third-party risk event.

Cross-Source Pattern: Safety as Variable Cost

Multiple sources this week surface the same structural dynamic from different angles. Anthropic's explosive growth — $1B to $20B ARR in ~14 months — is driven primarily by agentic tool capabilities, not safety features. Both OpenAI and Anthropic are heading toward IPOs (OpenAI in 2026, Anthropic targeting $60B in Q4 2026). As the race intensifies toward IPOs, safety is becoming a cost center that frontier labs cut when competitive pressure peaks.

For security teams, this means the implicit trust model — 'the AI vendor has a safety team, so model outputs are reasonably safe' — is actively weakening. Your defense-in-depth strategy for AI-integrated applications needs to assume vendor-side safety is unreliable and build independent guardrails.

What to do

  1. Trigger a third-party risk reassessment of OpenAI and document the safety governance change as a material vendor risk event for your next risk committee briefing

  2. Add political/regulatory risk as an explicit scored dimension in your AI vendor risk framework, with tested contingency for sudden access disruption within 48 hours

  3. Ensure application-layer guardrails (input validation, output filtering, rate limiting) work independently of vendor model safety for all AI-integrated production systems

  4. Prepare red team scenarios for OpenAI's 'Spud' model release within 48 hours of availability — focus on prompt injection bypass, guardrail regression, and novel capability abuse

$3.7 quadrillion in settlement volume is moving to smart contract rails — your fraud playbooks assume reversible transactions

The Migration Is Already Executing

This is not a future threat — the migration is underway. DTCC received an SEC No-Action Letter in December 2025 to tokenize real-world assets on approved blockchains. NYSE announced 24/7 on-chain trading for U.S. equities and ETFs with stablecoin funding and instant settlement. Tradeweb already executed real-time on-chain U.S. Treasury financing against USDC — on a Saturday — with Bank of America, Citadel Securities, DTCC, and Virtu Financial. Nasdaq filed its own rule change with the SEC.

DTCC processed $3.7 quadrillion in transactions in 2024. That volume is heading to smart contract rails with production targets as early as H1 2026.

What Changes in Your Threat Model

DimensionTraditional SettlementOn-Chain Settlement
Settlement FinalityT+1 with reversal capabilityAtomic, instant, irreversible
Operating HoursMarket hours (9:30–4:00 ET)24/7/365
Key ControlsIdentity-based (KYC, account auth)Cryptographic key possession
Fraud ResponseChargebacks, freezes, court ordersAppend-only; contain and prevent only
Code RiskProprietary, audited infrastructureSmart contracts, often open-source

The T+1 settlement window wasn't just an efficiency constraint — it was an implicit fraud detection layer. Instant, irreversible settlement means your fraud detection must be pre-transaction, not post-settlement. Every existing playbook that assumes transactions can be reversed, frozen, or charged back breaks on atomic settlement rails.

The largest financial infrastructure migration in 30 years is happening on smart contract rails, and most security teams are still running risk assessments designed for a world where settlements take a day and transactions can be reversed.

The Middleware Gap Is a Security Gap

The institutions building on-chain rails (DTCC, NYSE, Tradeweb) are not building the middleware. Investment firms are explicitly calling for startups to fill the compliance, tooling, and cross-border distribution layers. This means a wave of early-stage companies with immature security programs will become critical infrastructure in the financial supply chain — handling irreversible transactions at institutional scale.

Known DeFi attack vectors — smart contract exploits, bridge hacks, oracle manipulation, flash loan attacks, governance takeovers — have collectively caused billions in losses in the crypto ecosystem. These same vectors now apply to infrastructure handling quadrillions in annual volume. The regulatory frameworks (CLARITY Act, Genius Act) are still in formation, and SOC 2 and SOX don't meaningfully address smart contract risk.

What to do

  1. Audit financial counterparty and vendor relationships to identify which are migrating to blockchain settlement, and update third-party risk questionnaires to include smart contract audit practices and key management

  2. Brief your SOC on the operational differences of atomic settlement: 24/7 attack windows, no reversal capability, and pre-transaction fraud detection requirements

  3. Establish vendor vetting criteria for tokenization middleware startups before business stakeholders bring integration requests

The bottom line

AI agents crossed from 'access your data' to 'control your desktop' this week — Anthropic shipped Claude Computer Use with acknowledged prompt injection risk while OpenAI's CEO walked away from safety oversight, and Microsoft data confirms 62% of UK businesses already have agents running that security teams never scoped. Your security architecture was built for a world where software reads data and humans take actions; that boundary dissolved this week, and every prompt injection is now functionally closer to remote code execution.