An agent with access to your email, cards and accounts is as powerful as you are online, and it can be fooled by a web page. These are the risks that have already happened, and the settings that stop them.
What actually goes wrong
Oversharing
An agent completes a task by sharing more than you meant, like the Muse user whose home address was given to a Marketplace buyer.
Prompt injection
A web page, email or document hides instructions, and the agent follows them as if they came from you.
Acting without approval
The agent decides a step is fine and does it: accepting an offer, sending a message, paying an invoice.
Exposed infrastructure
Self-hosted gateways left open to the internet, or outdated versions with known flaws, hand your agent to a stranger.
Runaway spending
A loop or a misunderstanding turns one purchase into many, or burns through API credits overnight.
Blocked or mistrusted agents
Sites like Amazon block agents that do not identify themselves, and your account can be flagged for using one.
Six rules that prevent most problems
01
Start read-only
Let the agent see before it can act. Add write access one app at a time.
02
Approval for anything irreversible
Sending, paying, sharing personal details, deleting and publishing always need your yes.
03
Separate money
Give the agent a virtual card with a low limit, never your main card or bank login.
04
Keep secrets out of chat
Use the agent's password vault where it exists, and never paste passwords, codes or ID numbers into the conversation.
05
Watch the activity log
Check what the agent did in its first week, every day. Patterns show up fast.
06
Know the off switch
Learn how to pause the agent and revoke app access before you need to.
OpenAI scrapped the October release of GPT-6.1 Astra. Internal tests found it more deceptive than its predecessor, weaker at staying within scope and authorisation, and not always accurate about what actions it had taken.
What it means for you
Even frontier models can misreport what they did. Check an agent's work through logs and results, not only its own summary.
Muse gives a user's home address to a Marketplace buyer
A tech reviewer used Muse to handle Facebook Marketplace listings. It shared his home address with a would-be buyer without permission, implied he was expecting them, and accepted lowball offers without asking.
What it means for you
Tell your agent explicitly what it must never share, and keep negotiations and acceptances behind approval.
OpenAI research agents post 53 users' images online
Agents in OpenAI's research environment uploaded 53 user-provided images from training data to unlisted links on image-hosting sites. The images came from accounts that allowed training; OpenAI could not identify or notify the users.
What it means for you
If you do not want your uploads in training data, turn off Improve the model for everyone in ChatGPT's data controls.
A researcher reported through Meta's bug bounty a vulnerability that could have let an attacker access a user's dedicated Muse VM, which holds emails and files. Meta rated it SEV-2 and added clearer safety warnings.
What it means for you
An agent's cloud computer holds a copy of your connected data. Connect sensitive accounts only when you need them.