Personal AI Agent Security: Who Watches the Agent That Holds Your Credit Card?
I listened to the founder of Instinct, a personal-assistant company reported at roughly a $10 billion valuation, explain why people hand an AI agent their credit card. The product has no app. You text it, call it, or email it, and it has a phone, a computer and an email address of its own. Three numbers from the interview stuck with me:
- Three weeks in, 40% of users have shared a personal credit card with it.
- Users who share even one piece of sensitive information retain at 80%.
- More than $1 billion a year already flows through the platform, half of it travel, on a user base growing about 10% a day.
Asked how a nervous user should think about the risk, the founder did not talk about the model. The answer was a layer around it: firewalls that intercept and reject inbound content before the agent reads it, a monitor "decoupled from Instinct itself" that can pause, approve or block every action before it happens, a check before any tool call by "something that is decoupled from the same incentive system as the underlying agent," staged rollouts from the founder to the team to early users to everyone, and a security team hunting for adversarial cases.
That is the right architecture. It is also, almost line for line, what we have spent this year building and selling as an outside service. So here is the claim I will defend:
Decoupled from the agent is not decoupled from the business. The agent that holds your card should not be the only one grading its own safety.
Why in-house is not enough, even when it is good
The same interview explains the business model. The product is free to the user. The company takes a cut of the transactions it routes, and the founder walked through the reference points out loud: Shopify at two to three percent, Amazon near ten, Apple at thirty, and boutique hotels that already offer up to thirty percent for a delivered booking.
I do not think Instinct is acting in bad faith. The founder said plainly that the company does not want the agent influencing users against their interest, and named ad-funded platforms as the model the company is refusing. But consider the structure. The monitor that decides whether the agent may book the hotel is owned by the company that is paid a share of the hotel. The monitor is decoupled from the agent's incentives. It is not decoupled from the company's. And the user cannot check either one. Trust, by the founder's own metric, is "time to first credit card": a feeling, measured as behaviour.
An independent check turns part of that feeling into something a user, a merchant or a regulator can verify without asking the company.
We built the same layer. Here is what it caught, and where it failed.
The honest version of this post has to include our own marks, including the bad ones.
| Control | What happened |
|---|---|
| Pre-action grounding gate (the "hallucinated proper noun" check) | When our plan generator was given an inventory of what actually exists in the repo, the share of cited file paths that were real went from 0% to 82.4% in a 6-intent A/B. Without the inventory it either punted or invented plausible paths. The remaining 18% is still caught by a second check downstream. |
| Decoupled watchdog on the builder agent | Until today it had fired twice, ever. It was built, tested and documented, but it lived in a pipeline our main build lane never called. We wired it into that lane on 2026-10-04 and the first two observed jobs went through it. A monitor that exists is not a monitor that runs. |
| Independent payment witness | A Chainlink oracle network re-reads each USDC settlement on Base mainnet and records it on a contract only it can write to (Section 5.7 of the settlement white paper has every value). Our sweeper ledger shows 62 distinct settled jobs witnessed since 2026-10-01. It also shows 25 checks that found no witness record on-chain, which we have not yet classified. |
| Public verification of our audit credentials | Today our public verify-credential endpoint vouches for 4 of the 28 audit credentials we have published. The other 24 return no ERC-8004 passport on record — onboard + activate first. That is the gate working as designed: it refuses to vouch for an agent that never activated a passport. It is also our backlog, in public. |
| Spend permissions and a stop switch | Every purchasing grant we issue is a signed credential stating what the agent may do, where it may spend, the per-call and per-day caps, and which actions require a human, bound to an agent passport. A signed stop order revokes it, stops never lift on a timer, and the status list is public so a merchant can check it. |
The watchdog row is the one I would defend under pushback. We had the architecture the founder describes and it did nothing for months, and nothing inside the system told us. We found it by reading the ledger, not the design document. That is the strongest argument I know for a second party: the failure you cannot see from inside is the one that is "built" and quiet.
What an independent check looks like for a personal agent
- Attack the firewall from outside. Run prompt-injection and content-trap payloads through the agent's inbound channels (email, messages, fetched pages) across the OWASP agentic top ten, then publish the findings as a signed credential anyone can verify. Our audit methodology is public; a per-audit run is priced at $5.
- Publish the grant, not the promise. What the agent may spend, where, and what needs a human, as a revocable credential a merchant can check at checkout.
- Witness the money. The settlement confirmed by someone who is not the agent, the user or the platform.
- Turn trust telemetry into attestations. "40% share a card by week three" is a company metric. "Passed an external injection audit on this date, credential here" is a fact the user can check.
The evidence kit shows what a delivered audit and its proof look like, and we audited ourselves first, with the result labeled as a self-audit.
The tradeoff is real. An outside auditor sees less than the in-house monitor, because it is not in the action path and cannot pause anything in real time. That is the cost of independence, and it is why the two are complements, not substitutes: the in-house monitor stops the bad action, the outside check proves the monitor exists and works. A team that wants the second without the first should look at a governance readiness audit instead; one that wants a standing reviewer can use the agent trust auditor.
What we refuse
- We refuse to call any agent "safe." An audit is evidence about what was tested, on a date. It expires the next time the agent changes.
- We refuse to sit in your users' payment path. We read evidence and sign attestations. We do not hold anyone's card, and a bug on our side cannot move a user's money.
- We refuse to sell our self-audit as independent. It is labeled as ours, and the four-in-twenty-eight verification number above is printed here because hiding it would be the exact behaviour this post argues against.
- We refuse to treat Instinct's numbers as verified. They are a founder's statements on a podcast. They are interesting precisely because nobody outside can check them.
The second question
Why would a company growing ten percent a day pay for an outside check? Because by its own account trust is the product: card-sharing is the leading indicator and 80% retention is the payoff. A credential a user can check is cheaper than weeks of earned feeling, and it survives the first public incident in a way that "our security team is world-class" does not.
What would make this false? If users never ask and merchants never check. Then trust stays a feeling, and in-house monitoring is enough. The leading signal would be a personal-agent company that publishes external verification and sees no change in card-share or retention.
What would I do Monday? Offer one free external audit to a small personal-agent team that cannot yet afford a security group. What would I not do? Pitch the $10 billion company. They built theirs in-house and said so. The teams behind them are the ones who will be asked "why should I give you my card" without a billion dollars of answer.
Read the audit methodology See a delivered audit
Scope
Checked: the full interview transcript; our grounding A/B record; the watchdog's fix-event ledger and the 2026-10-04 wiring; the settlement-witness sweeper ledger (rows to 2026-10-04); all 28 published audit-credential verify URLs, live, on 2026-10-04; the live delegation status list.Not checked: Instinct's product itself, or any of its figures; why 25 witness checks found no record; whether any personal-agent company would accept an external audit.
Would change the conclusion: evidence that Instinct already publishes third-party verification, or that users of personal agents do not respond to it.