Knowledge Graph · Companion essay
Who has your data?
Your data is already tracked, sold, and resold by companies you cannot easily leave. OpenClaw is just the next door — every key you paste adds another company to the count.

The world you already live in
Before we talk about any agent, count the companies that already hold a piece of you. You give them data every day, usually for free:
- WhatsApp → Meta. Personal messages and calls are end-to-end encrypted, so Meta cannot read their content. WhatsApp still operates on Meta infrastructure and can share account and related information with other Meta products; when you add WhatsApp to Accounts Center, Meta’s own help pages say that information can be used to personalize products and show ads across Meta apps.1 noyb has documented Meta’s push to serve ads on WhatsApp using personal data from Instagram and Facebook under that model.2 You are not Meta’s customer. You are the inventory.
- Instagram, same house. Who you look at, how long you look, what you save, what you almost buy — Meta’s ad systems are built on behavioral signals across its apps. Linked Accounts Center profiles let those signals combine.
- Google holds a profile that spans Search, YouTube, Maps, and Android. Google stopped scanning consumer Gmail content to personalize ads in 2017; ads still appear in Gmail, targeted with other Google account activity such as Search and YouTube.3 When Google Assistant activates, the device sends the request to Google servers to fulfill it; human reviewers can process de-identified query text for quality.4 Location from Maps and Android is a continuous input to the same account graph.
- YouTube knows what you almost believe. Starts, finishes, rewinds, and early skips feed recommendation systems. That watch graph is one of the richest continuous portraits of attention any consumer product holds — and it sits inside the same Google account used for ads and Search.
- Netflix watches you watch. Plays, early abandons, and which artwork you click are logged. Netflix has publicly described A/B testing and personalizing thumbnails so the image you see is chosen to match your history — the same data that drives recommendations and content decisions.5
- Spotify keeps context on a timeline. Listening history, playlists, and session context power recommendations. Spotify’s own ad products market contextual and mood-adjacent targeting — workout vs. wind-down playlists as signals of receptivity for advertisers.6
- Apple has listened through Siri grading. In 2019, The Guardian reported that Apple contractors graded a sample of Siri recordings and regularly heard private situations, including medical conversations; Apple then suspended human grading globally and later made participation opt-in.7
- Amazon is listening in the cloud. After the wake word, Echo devices stream audio to Amazon’s cloud to process the request. Voice recordings are saved by default until you change retention or delete them in Alexa Privacy settings.8
- Location is a market and a government tool. The Supreme Court held in Carpenter v. United States (2018) that accessing historical cell-site location information from carriers is a Fourth Amendment search that generally requires a warrant.9 Separately, commercial data brokers sold precise location that could place people at medical clinics, places of worship, and other sensitive sites; the FTC has ordered firms including X-Mode/Outlogic and InMarket to stop selling or sharing that class of precise and sensitive location data.1011
None of these companies asked you to build a profile. You built it for them — one tap, one play, one search at a time — and they kept what the product design allowed them to keep.

The data does not always stay in one silo. A secondary market of data brokers and aggregators — firms such as Acxiom, Experian, Oracle, LiveRamp, and Epsilon, among many others — buys, blends, and resells consumer attributes keyed to identifiers. WhatsApp contacts you chose to sync, Gmail-adjacent Google activity, Spotify session context, and Netflix watches need not remain only where you generated them. Matched profiles can be sold onward to advertisers, and in some cases to insurers, landlords, employers, and campaigns. Location data is the clearest recent example of that secondary market colliding with regulation: the FTC’s 2024 orders against X-Mode/Outlogic and InMarket were aimed at brokers who collected and sold precise location without adequate consent, including trails that could reveal visits to sensitive places.
The ad market is what makes much of the consumer internet run. Platforms watch what you do, model what you are likely to want, and put it in front of you when you are most likely to act. YouTube queues the next video most likely to keep you watching. Netflix autoplays and rewrites thumbnails to genres you click. Spotify on free tier inserts ads into sessions it can describe contextually. Instagram learns which framing holds your thumb an extra half-second. Meta, Google, Amazon, and the brokers behind them feed those signals into auctions that follow you across the web — the shoes you glanced at once reappear days later because a model still marks you warm. The system was not built primarily to inform you. It was built to predict you. The more of your data it holds, the harder it is to live outside the prediction.
So when the next product asks you to paste one more key, that is the world it asks you to add to. You are not starting from zero. You are adding a line to a ledger that was already long.
Where OpenClaw fits in the ledger
OpenClaw is a clean place to watch the same pattern arrive in agent form. It started as Clawdbot — a weekend build that could read files, run shell commands, and call one model, Claude, for anything harder than ls. One key, one company, one connection. The pitch was simple: the agent runs on your machine; the only traffic leaving the laptop is the prompt you typed and the reply that came back.
That is not what setup looks like now. The wizard still only requires a model key to start, then offers optional steps for web search and messaging. Say yes and the prompts appear. Beyond the wizard, a full install can ask for a memory key to index files, a web-fetch key to read pages, a speech key for voice, a vision key for images. Messaging alone can mean ten channels — Slack, Discord, Telegram, WhatsApp, iMessage, Teams, Feishu, QQ Bot, Signal, Google Chat — each wanting its own bot token or login. The single door became a hallway, and onboarding treats every door as a normal install step.
Each key is a contract. Not only “let the agent call the API,” but “let the provider see the traffic that flows through it.” Prompts, transcripts, images, schedules, friends’ numbers — whatever the key unlocks — cross a wire you do not control and land on infrastructure you do not own. That part rarely makes the landing page. It is worth reading key by key.
Count the keys

Privacy is not a switch. It is a ledger. The honest way to read a product’s posture is to count the keys it can ask for and, for each one, ask what flows across it and who is on the other end. Here is that ledger for a present-day OpenClaw-style install — the wizard forces only the first row; every row below is one settings screen away. The examples in the middle column are illustrative of the kind of data each capability exposes, not claims about any single user’s history:
| Key you paste | What that path can carry | Who is on the other end |
|---|---|---|
OPENAI_API_KEY / ANTHROPIC_API_KEY / XAI_API_KEY / GEMINI_API_KEY | Prompts, attached context (for example an email you asked to summarize), and the model’s replies | OpenAI, Anthropic, xAI, or Google — each a U.S.-based provider with its own retention and training policies |
OPENAI_API_KEY / VOYAGE_API_KEY (memory / embeddings) | Chunks of notes, files, and conversations indexed for retrieval | OpenAI or Voyage AI |
BRAVE_API_KEY / TAVILY_API_KEY / EXA_API_KEY | Search queries the agent runs on your behalf, and thus the intent behind them | Brave, Tavily, or Exa |
FIRECRAWL_API_KEY | URLs you ask it to fetch, and the page text returned | Firecrawl |
DEEPGRAM_API_KEY / ELEVENLABS_API_KEY | Voice audio and/or synthesized speech for the turn | Deepgram or ElevenLabs |
GEMINI_API_KEY / XAI_API_KEY / OPENAI_API_KEY (vision) | Images you ask the agent to interpret | Google, xAI, or OpenAI |
| Slack / Discord / Telegram / WhatsApp / iMessage / Teams bot token | Messages and channel traffic the bot is allowed to read or send | Slack (Salesforce), Discord, Telegram, Meta, Apple, or Microsoft |
Seven capability lines, each with its own key — and depending on providers, potentially a dozen companies holding pieces of the same life. The wizard forces only the first; the rest are optional but one settings screen away, and the product feels thinner without them. None of them alone is the whole story. The union of those datasets, held across houses you do not live in, is a fuller portrait than any single company could build — and each piece crosses a wire you opened yourself.
A second question sits inside this ledger: which country owns the model you talk to, and what that means for what you believe. That is its own piece: Who owns the answer?
It did not used to be this many doors. Clawdbot shipped with one key because it did one thing — talk to a model. Every key since was added so a feature would work: memory, search, voice, channels. The features are real. So is the cost. Each key that unlocks a capability also opens a live line to a new company — and the line stays open whether or not you use that feature today.
Why local-by-default inverts this

The Knowledge Graph stack starts from a different default: if a capability can run on your hardware, it does. Voice is recognized on your chip. Images are understood on your chip. Notes and memory are indexed on your chip. Files and history live on your disk. Nothing of that leaves the laptop, because there is no company on the other end waiting for it — the model is a file on disk, the store is a file you own.
The trade-off is real. Your laptop has limits, and the rest of the Knowledge Graph series is honest about them. Local costs memory, battery, and the discipline not to reach for a cloud API the moment a task gets hard. The payoff is simpler: your information stays yours. Prompts, files, voice, and photos do not cross a wire to a server you do not control. Privacy is not a promise in a policy PDF. It is what happens when nothing leaves the machine.
The honest position
This is not a religion. A laptop cannot do everything a cloud cluster can. When you hit a task it truly cannot do, sending that one job out can be a fair trade — eyes open, one thing at a time, never as the default path. The position worth defending is narrower:
- Default to local. If the model fits, it runs here. Your information stays on disk until you have a real reason to send it. The count of companies holding your data starts at zero — not because they refused it, but because you never sent it.
- One capability, one company — never a bundle. When a product needs a dozen parts of your life in a dozen houses to feel complete, it is a cloud product in a local costume — and the costume is thin.
- Count the connections, not the promises. Privacy is not what the privacy policy says. It is the set of wires your data actually flows through. Read the table, not the marketing.
The Knowledge Graph series is, among other things, an attempt to build a personal agent that does not require you to sign away the personal layer to make it work. Utility versus personal AI matters here: a utility agent on your shell is one exposure profile; a personal agent on your relationships is a heavier one — and which world you are in is decided by what leaves the machine.
Reclaiming your data is not a single act. It is the refusal, each time a settings screen appears, to open a door you do not need. The longer argument for that refusal is here.
Sources
- Meta Help — Adding your WhatsApp account to an Accounts Center (cross-product use and ads when linked)
- noyb — WhatsApp ads using Instagram and Facebook personal data (16 Jun 2025)
- New York Times — Google will no longer scan Gmail for ad targeting (23 Jun 2017)
- Google Safety Center — Protecting your Google Assistant privacy (activation sends request to Google servers; review of de-identified queries)
- Netflix Technology Blog — Artwork Personalization (personalized thumbnails from viewing behavior)
- Spotify Advertising — Contextual advertising (playlist/context signals for ad targeting)
- The Guardian — Apple contractors hear confidential details on Siri recordings (26 Jul 2019); Apple later suspended grading and moved to opt-in
- Amazon Help — Alexa, Echo devices, and your privacy (cloud processing; default retention until deleted)
- Oyez — Carpenter v. United States (2018; warrant generally required for historical CSLI)
- FTC — Order prohibiting X-Mode/Outlogic from selling sensitive location data (9 Jan 2024)
- FTC — Final order prohibiting InMarket from selling or sharing precise location data (1 May 2024)