Developer Fernando Irarrázaval stress-tested his OpenClaw assistant by publicly sharing access to its inbox on Hacker News, inviting users to attempt to compromise it. The AI agent, powered by Anthropic's Claude Opus 4.6 model, reportedly withstood roughly 6,000 attack attempts as visitors tried various prompt-injection and manipulation techniques to break its safeguards. The experiment offered a real-world demonstration of how modern AI agents handle adversarial inputs when granted access to tools and sensitive data such as email. As AI assistants increasingly gain the ability to act autonomously on users' behalf, their resilience against social engineering and malicious instructions has become a key security concern. The public test highlighted both the progress made in hardening these systems and the ongoing risks of deploying agentic AI with access to real accounts and personal information.


Read the original article →

— Sponsored —

Trade smarter on BYDFI

Get a bonus on your first deposit — from $50 at $100, up to $2,000 at $20k. 200x leverage, 600+ perpetuals, deep liquidity.

Claim your bonus →