AI Governance & Cybersecurity  ·  27 August 2026

Is Your Company Spoilt for Agents?

Ollama, Google and Anthropic have all shipped the same design: flexible model routing, sold as a feature. None of them shipped a way to ask, afterwards, which model actually answered, on whose account, or under whose data-handling terms.
By Alan Wright  ·  The Haunted Lighthouse Limited  ·  Peel, Isle of Man

There is a version of this conversation happening in a finance meeting somewhere right now.

“We've got a ten grand bill from Ollama over the weekend.”

“Who's using Ollama?”

“No idea.”

Nobody in that room did anything wrong, exactly. They just discovered, the expensive way, that “which AI answered this prompt” and “who's paying for it” have quietly become two separate questions, and most companies have only ever budgeted for the first one.


The toggle nobody reads

Start with the specific case, because it just became very concrete. Ollama shipped v0.33.0 on 25 August 2026, and folded the whole thing straight into the menu bar: turn individual Ollama models on or off for use inside Claude Desktop, pick from your local models without leaving Claude, and cloud models appear the moment you're signed in. Ollama's own launch line is “one toggle, cloud and local models just work.” Flip it on, and your prompts can go to a locally-hosted model, or to Ollama Cloud, depending on which slot is mapped and switched on.

Ollama states it runs a zero-data-retention policy across both its cloud and local offerings, telemetry off by default. Worth noting, because it heads off the wrong objection: the risk here was never that Ollama secretly trains on your prompts. It's what happens after the toggle is flipped and everyone forgets which model is actually listening.

The instinct is to assume the danger here is a blended bill: forget the toggle's on, and your Anthropic subscription silently racks up API-rate charges alongside it. That's not actually how it works. Anthropic's own gateway documentation is clear that the two paths don't share a meter; when the toggle's on, billing goes entirely to the third-party provider's account, and the claude.ai subscription “isn't used or charged” at all. So there's no state where forgetting the setting quietly inflates an Anthropic invoice.

That's the reassuring finding. It's also not the point, because it doesn't make the underlying problem go away, it just relocates it.


What actually goes wrong

Two failure modes survive the correction, and one of them is worse than a bill.

The first is the straightforward one: if the mapped slot points at a cloud-hosted model rather than a local one, and whoever set it up assumed “on Ollama” meant “on my own hardware, therefore free,” that's a genuine surprise bill, just on Ollama's meter instead of Anthropic's. And it's a nastier shape of surprise than most finance teams are used to, because a subscription hitting its usage cap just throttles you; an unmonitored metered slot has no ceiling at all. It will happily answer agentic loops all weekend and hand you the invoice on Monday.

The second failure mode doesn't show up on an invoice at all, which is exactly why it's more dangerous. Picture a firm running a contract review, believing it's getting frontier-grade Claude, when the routing slot is actually mapped to a local model a fraction of the size. Nothing in the interface flags which engine actually answered. The output comes back fluent and confident either way. For a law firm, or anyone whose work carries professional liability, a wrong answer delivered with total conviction is a materially worse outcome than an inflated bill, and there is currently no native mechanism that tells the user which model just did the work.

Scale that up. Right now this is a Mac problem, because that's where the gateway toggle currently lives. It's not staying that way. Wider install base, same setting nobody reads, times ten, and it starts looking exactly like shadow IT to a SOC analyst running the traffic logs: unexplained outbound connections to Ollama Cloud from a laptop that's supposed to be running one vendor's assistant and nothing else.


It was never really about Ollama

Widen the lens and Ollama stops being the story. It's this month's proof case for a pattern that's showing up across the entire agentic-tooling landscape: vendors are shipping “choice of model” and “accountability for that choice” as two separate features, and the second one keeps arriving late, or not at all. It's the same shape as the gap in Windows Recall's TPM-backed vault, built on an operating system that never gated who could read the screen in the first place: elaborate protection for the part that was never the actual weak point.

Google's own Gemini stack makes the same point from a different angle, and arguably makes it harder. Google AI Studio and the free tier of the Gemini Developer API train on your inputs and outputs by default, with no toggle available inside the free tier itself; the only way out is to enable billing and move to paid. Workspace and Vertex AI, by contrast, carry a contractual no-training guarantee. Same company, same model family, opposite data handling, and nothing in the interface tells a developer copy-pasting a client's document into AI Studio which regime they've just landed in.

Google adds a wrinkle that goes further than a training toggle. Even with Gemini's consumer “Keep Activity” setting switched off, any conversation a human reviewer has already read is retained for up to three years, disconnected from the account, surviving deletion entirely. Turning the setting off stops the harvesting going forward. It does nothing about what's already happened.

Checking any of this for yourself takes a couple of minutes, once you know where to look, and that qualifier is most of the problem. The training toggle isn't sitting on Google's main settings screen; it's under Gemini Apps Activity in your Google Account, or reachable directly at myactivity.google.com/product/gemini, several taps deep from wherever you'd naturally land looking for it. I already run mine with Keep Activity switched off, which stops future conversations being used for training or handed to a human reviewer; it does nothing about anything Google's already looked at, and finding the toggle in the first place meant going looking for it, not stumbling across it. If someone who audits this stuff for a living has to hunt for the setting, the employee who got handed a Google AI Pro seat by IT and never asked what it does isn't finding it either, and that's before anyone gets to the separate, less-discoverable distinction between AI Studio and Vertex.

Claude's own consumer-versus-enterprise split runs on the identical logic, just without that particular tail: personal Free, Pro, and Max accounts train on conversations by default unless a user opts out, via a single toggle under Settings, Privacy, “Help Improve Claude”; Team, Enterprise, and API access carry a contractual no-training carve-out that was never subject to that toggle in the first place. Which means, in practice, that IP exposure at most companies isn't a story about “the cloud” being unsafe in the abstract. It's a story about someone using their personal account instead of the one the company actually sanctioned and paid for, because nothing stopped them and nothing told anyone afterward.

My own workflow already splits along that exact line, though not because I set out to prove a point. Editorial back-and-forth runs through the consumer interface, where that toggle is one setting away from mattering. Anything that counts as serious development work goes through the API instead, on Opus 5 or Fable 5, and doesn't touch the consumer product at all; the no-training terms there aren't a preference I selected, they're what API access carries by default, contractually, regardless of what any toggle is set to. The lesson isn't “remember to check your settings.” It's that for anything that actually matters, the work shouldn't be routed anywhere near a setting that needs checking.

Put a face on that risk and it stops being abstract. Samsung found out the hard way in 2023, when semiconductor engineers pasted proprietary source code and defect-detection algorithms into ChatGPT's free tier on three separate occasions in twenty days, each one apparently reasonable in isolation: fix this bug, optimise this test sequence, summarise this meeting. None of them realised the input became training material the moment they hit enter. Samsung banned generative AI tools company-wide within a month, but the code was already gone; nobody has ever managed to claw proprietary information back out of a model once it's baked into the weights. Three years on, the same mistake is one keystroke away in a biotech lab rather than a semiconductor fab: a researcher pastes an unpublished compound structure, a novel synthesis route, or early trial data into a free AI Studio session, or a personal AI account, to get a faster literature comparison, on the entirely reasonable assumption that it's just a smarter search engine. Trade secret protection depends on the information staying secret, and a patent application built on that same research now has to survive a novelty challenge from data nobody outside the company was ever supposed to see. You can turn a training toggle off going forward. You cannot un-train a model on someone else's potential medical breakthrough.


No regulation is coming to save you from this

Before assuming Brussels will eventually force the attribution logging this piece is asking for, check the calendar. The EU AI Act's high-risk obligations, the ones requiring record-keeping and human oversight for AI systems in the more sensitive Annex III categories, were due to bite on 2 August 2026. The Digital Omnibus on AI, Regulation (EU) 2026/1744, entered into force on 27 July 2026, six days before that deadline, and pushed standalone high-risk compliance to 2 December 2027. Sixteen months of runway just appeared for anyone racing toward the original date, and the routing problem in this piece sits outside that framework regardless: general commercial contract review doesn't fall inside any of Annex III's eight listed categories, which cover things like employment decisions, credit scoring, and essential services, not a law firm reviewing its own client's paperwork. The “administration of justice” category, which sounds like it might catch legal work, is scoped to systems assisting judicial authorities, not private practice.

What is live, and has been since 2 August 2025, are the transparency obligations on providers of general-purpose AI models: Anthropic, Google, and the rest have to publish documentation of what their models can do. That's a paper trail about the model. Nothing in it obliges anyone to log, inside your own stack, which model actually answered a given prompt, and no provider-level transparency duty reaches down to fill that gap for you.

GDPR, or the Isle of Man's equivalent regime, has the sharper edge here, and it isn't waiting on any 2027 deadline. The accountability principle at Article 5(2) requires you to demonstrate, not merely assert, how personal data was processed; Article 30 requires your records of processing to name your actual processors; Article 28 requires authorisation before a sub-processor is engaged at all. A DPIA carried out under Article 35 on the assumption that a named vendor handles the processing is wrong the moment routing silently swaps that vendor for one with different data-handling terms, and nobody revises a DPIA because a menu-bar toggle changed state. That's not a future compliance risk with a deadline attached. It's a live gap in paperwork you may already have signed off as accurate.


The Theatre Pulldown version of events

Line the vendor examples up against that regulatory timeline and the institutional claim and the operational reality come apart in exactly the same place every time. The claim is flexibility, or breathing room: look how much model choice we've handed users, look how much extra time Brussels has just bought industry to get this right. The reality is that the choice and the accountability for the choice were shipped as unrelated features, the extra compliance time was never aimed at this particular gap in the first place, and the accountability half is either an afterthought, a paid-tier upsell, or simply absent.

None of this requires a hostile actor, a bad model, or a broken feature. Ollama's toggle works exactly as designed. Google's tier split works exactly as designed. Claude's Enterprise carve-out works exactly as designed. The Digital Omnibus's deferral works exactly as legislated. Every one of these systems is functioning correctly, and every one of them still leaves the CFO, and the DPO, with no way to answer a one-line question after the fact: which model, on which account, under which policy, actually did last night's work?

That's the audit a board should be running before procurement rather than after the invoice. Not “which AI tools have we approved,” because that question is already being asked everywhere. The one nobody's asking is whether anything in the stack can tell you, with a straight face, which engine answered a given prompt, on whose account, and under whose data-handling terms, without someone having to reconstruct it from memory after the fact.

If the honest answer is “we'd have to ask around,” the flexibility you've been sold isn't a feature. It's an audit gap wearing a better name.


Sources


Questions about this analysis, or interested in working with The Haunted Lighthouse?
consultancy@haunted.lighthouse.co.im

The Sovereign Auditor covers digital sovereignty, cybersecurity governance, and data protection policy—with particular focus on Isle of Man jurisdiction and Crown Dependency issues.

Support independent analysis. Subscribe directly—or scan on your phone.

Payments via PayPal. Credentials delivered by email. No Substack. No Stripe. No middlemen.