🏢 Advanced business · Guide 2 of 2
What Leaves Your Network: A Data Policy for Agent Tools
9 min read · Last reviewed 23 Sep 2026

Someone in your organisation is going to ask where the code goes when an engineer runs an agent. It is a fair question and it does not have one answer. An agent session can touch a model provider, the tool vendor's own backend, two telemetry pipelines, three MCP servers and a cloud runner that clones your repository into a virtual machine you do not operate — and the vendor's privacy page only describes the first of those.
This guide is about mapping that surface and deciding what you will accept. It is not legal advice. If you work in health, finance, public sector or anywhere personal data is regulated, your counsel signs this off, not a directory page.
Five places your code can go
The model provider. Prompts, file contents and tool output are sent to a model for inference. This is the flow every vendor documents, and it is the one people worry about most. It is often not the riskiest.
The tool vendor's own backend. Separate from the model. Editors and agents may index your repository for retrieval, store embeddings, assemble prompts server-side, sync rules and settings, or keep session history for resume. A vendor can honestly say "we do not train on your code" while still storing a copy of it. These are different questions and you have to ask both.
Telemetry and error reporting. Usage metrics and crash reports, often sent to third-party logging and error-tracking services rather than the vendor's own infrastructure. Good vendors document exactly what these contain, redact known secret patterns, and give you environment variables to switch each one off. Defaults differ by plan and by which model provider you are routed through, so check the defaults for your configuration rather than the headline.
Every MCP server the agent can call. Covered below, because this is where most of the real exposure sits.
Any cloud agent runner. Background agents and web sessions clone your repository into a machine the vendor operates. Credentials, network egress and retention there are a separate contract from the one covering the desktop client.
There is also a local surface that nobody asks about. Agents keep session transcripts on the laptop, frequently in plain text, for weeks by default, and those files are inside your engineers' backups. Anthropic documents this for Claude Code explicitly, down to the directory and the retention setting. That is more transparency than most vendors offer, and it applies to all of them in some form.
The questions to ask each vendor, and where the answers live
Vendors publish this material and update it more often than any guide can. Ask five questions and read the answers at the source.
| Question | Why it matters |
|---|---|
| Do you train on our inputs and outputs, by default? | Defaults differ between consumer subscriptions and commercial plans at the same vendor |
| How long do you retain prompts and responses? | Abuse-monitoring and safety logs are usually retained even when training is off |
| Who are your sub-processors? | Your customers may have contractual rights over this list |
| Where is data processed and stored? | Residency and encryption controls are often plan-gated |
| Is zero-retention available, and what does it disable? | It is normally an approval, not a checkbox, and it costs you features |
Anthropic's position for Claude Code is set out at data usage and zero data retention, with the underlying policies in the Privacy Center. OpenAI documents API retention, abuse-monitoring logs and zero-retention eligibility in data controls. Cursor publishes privacy and data governance alongside its security page. GitHub keeps Copilot's commitments in its Trust Center. Read those, not a summary of them.
Two things are worth knowing before you read. First, zero data retention is generally granted to qualifying accounts after review rather than enabled from a settings page, and it switches off features that need server-side storage — cloud sessions, shared artefacts, feedback submission. Second, and more important for what follows: a zero-retention agreement covers the model provider's handling of your prompts. It says nothing about the third-party tools and MCP servers your agent calls. Anthropic states this plainly in its own documentation.
An MCP server is an outbound channel holding your credentials
This is the part that is genuinely under-documented, so be precise about the mechanics. You install a server. You give it a token — for your issue tracker, your error monitor, your database, your cloud account. From then on, the model decides when to call it, with what arguments, based on text it has read. Two consequences follow.
The first is exfiltration. Any server that can send data outward is a channel out of your network that bypasses the controls you built for browsers and email. A server approved for reading issues may also be able to create them, and an issue body is a fine place to put whatever the agent has in context.
The second is injection. Whatever a server returns becomes context, and context behaves like instructions. A poisoned issue description, a doctored error payload or a hostile web page fetched mid-session can steer the next tool call. The MCP specification's own security document treats this as a first-class concern and argues for minimal, progressively elevated scopes rather than broad ones.
GitHub MCP Server 🔌 MCP serverFree
GitHub's official MCP server: issues, pull requests, code, Actions and security alerts
Approve MCP servers the way you approve integrations, not the way you approve editor plugins. Who wrote it, what it can write as well as read, what credential it holds, and what happens if it is wrong.
A local server is not automatically safer
"It runs on the developer's machine" gets treated as an answer. It is not one.
A local server runs with the privileges of the person who started it, which on an engineer's laptop means their SSH keys, cloud credentials and everything they can reach on your network. Local says nothing about network egress: plenty of local servers exist precisely to call remote APIs. And the launch command is a supply chain — a server started by fetching a package at run time gets whatever that package contains today, from whoever published it. The MCP specification documents this attack class directly, including malicious startup commands in shared configuration.
Directory listings are not audits, and the vendors say so. Anthropic's own documentation notes that connectors in its directory are reviewed against listing criteria but not security-audited, and that it does not manage any MCP server. Our MCP server listings are the same thing: a map of what exists and what it connects to, not a security clearance.
MCP Reference Servers 🔌 MCP serverFree
The official reference MCP servers: fetch, filesystem, git, memory and more
Secrets hygiene, in the order that actually helps
- Credentials in environment variables or a secrets manager, never in the config file. Config files get committed, shared in a setup guide and pasted into a support thread.
- One scoped token per server, read-only wherever the job allows. A server that reads issues does not need write access to your repositories.
- Separate machine identities. Tokens issued to a shared bot account with its own permissions can be revoked and audited. Tokens minted from a senior engineer's personal access are indistinguishable from that engineer in every log you own.
- Short lifetimes and a rotation you have practised. The question is not whether a token leaks but how long the window is.
- Treat a secret pasted into a prompt as a rotated secret. Prompts may be logged by the provider, by your own gateway, or by a crash report. Rotate it; do not reason about it.
- Understand what file exclusions are. Telling an agent not to read
.envis a useful default and not a boundary: an agent that can run shell commands can read anything the user can. If a file must be unreachable, it should not be on the machine.
Self-hosting and open source as genuine controls
Two structural moves reduce the number of parties involved rather than adding promises from them.
An open-source client with your own API key removes the tool vendor from the data path entirely. Your code goes to the model provider you chose, under the commercial terms you signed, and there is no second backend holding an index of your repository.
OpenCode Open source
Open-source terminal coding agent that works with any model provider
Running the agent runtime on your own infrastructure does the same for execution: the repository, the working directory and the shell stay on machines you operate.
OpenHands Open source
Open-source platform for autonomous coding agents, local or in the cloud
A third option suits larger teams: route every request through one gateway or through your existing cloud account. Claude Code, for example, supports a custom base URL, corporate gateways and inference through Amazon Bedrock, Google Cloud or Microsoft Foundry, which keeps traffic inside a cloud contract you already have. You get one place to log, one place to apply limits, and one place to revoke.
Be honest about what this buys. Self-hosting moves risk rather than deleting it: you now own the patching, the logs and the incident response. The model still reads your code unless you run a local model, and local models remain meaningfully weaker at agentic work. And a gateway that logs everything is itself a new store of prompts and code, which needs the same retention answer you demanded from the vendor.
A checklist you can run in an afternoon
Before approving any agent tool or MCP server:
- Which company receives the inference request, and under which terms — consumer or commercial?
- Does the tool have a backend of its own, and what does it store there?
- What is the retention window for prompts and responses, and for safety or abuse logs?
- What telemetry is on by default, and what switches it off?
- Is there a cloud or background mode? Is it on by default? What credentials does it hold?
- Which MCP servers will this agent be able to reach, and what token does each hold?
- For each server: who maintains it, can it write, and how does it get updated?
- Where are session transcripts written locally, for how long, and are they in laptop backups?
- Is the sub-processor list published, and does it conflict with anything you promised a customer?
- Who at your organisation revokes all of this in one afternoon if you need to?
If you cannot answer 1, 6 and 10, you are not ready to approve it. The rest can follow.
The risk you cannot remove
Prompt injection through content the agent reads is unsolved. An agent with shell access is a general-purpose program driven by text from sources you do not control. Contractual data-handling promises are enforceable in court and are not technical controls. Defaults change between releases, so a configuration you audited in March may behave differently in September.
The rational response is a small approved surface, scoped and rotatable credentials, and a habit of re-checking after upgrades. Pair this with the rollout mechanics in rolling out AI coding to a team, and send the engineers who implement it to the developer guides. Smaller operations working alone will find the proportionate version in the solo builder guides, and the ecosystem map shows which servers and skills work with which app before you start approving anything.
Read the official docs for…
- MCP security — security best practices in the specification: confused deputy, token passthrough, local server compromise, scope minimisation.
- Claude Code data handling — data usage for training, retention, telemetry and local transcripts, and zero data retention for what ZDR does and does not cover.
- Restricting MCP servers centrally — managed MCP for allow-lists, deny-lists and fixed server sets.
- OpenAI data controls — developers.openai.com for retention, abuse-monitoring logs and zero-retention eligibility.
- Cursor privacy and governance — privacy and data governance for privacy mode at team level, residency and encryption keys.
- GitHub Copilot commitments — the Copilot Trust Center for enterprise data use, and enterprise policies for the controls an owner can set.
Mentioned in this guide
GitHub MCP Server 🔌 MCP serverFree
GitHub's official MCP server: issues, pull requests, code, Actions and security alerts
OpenCode Open source
Open-source terminal coding agent that works with any model provider
MCP Reference Servers 🔌 MCP serverFree
The official reference MCP servers: fetch, filesystem, git, memory and more
OpenHands Open source
Open-source platform for autonomous coding agents, local or in the cloud