Will AI Agents Hack Everything?
Public 2025 threat reporting shows agents used for recon, malware mutation, and tool abuse. What changed when models got hands, why refusals are not enough, and how Agnostics release testing catches agent failures before launch.
Agents can browse, call APIs, and run scripts at machine speed. Public threat reports documented state-sponsored abuse and LLM-assisted malware. Your job is not to pretend that stops. Your job is measurable defense on the agent you ship, with attack packs, findings, and a Release Gate before users find the same breaks.
The launch question teams avoid
Every team shipping agents eventually hits the same quiet moment in a standup. Someone asks whether the product could be used to break into things. The room gets uncomfortable. Someone mentions guardrails. Someone else mentions responsible AI. The ticket moves to next sprint.
That dodge is expensive. Agents are not chatbots with ambition. They are loops that plan, fetch context, pick tools, and act. When those actions touch customer data, billing, internal admin APIs, or a browser session, you are shipping a new class of risk whether or not the slide deck says security.
The honest launch question is narrower and harder: can an adversary reach a harmful side effect through the agent path we configured for production? Not in theory. Not on a trimmed demo. On the target with the same tools, auth, retrieval, and confirmation rules you intend to ship.
Agnostics exists to make that question answerable before the launch button. Attack packs pressure your configured target. Findings show reproduction evidence. The Release Gate turns results into Ready, Monitor, Fix, or Blocked for this scan and this target.
What changed when agents got tool access
For years, LLM risk debates centered on words: toxic output, hallucinated facts, leaked secrets in a reply. That mattered. It still matters. Tool access changed the unit of harm from text to action.
An agent with a browser can enumerate pages you forgot to lock down. An agent with shell access can run recon scripts faster than a human operator typing. An agent wired to your ticket system can exfiltrate context that never appeared in the chat transcript users see.
Public threat reports documented this shift in 2025. Researchers reported state-sponsored actors using agent-style workflows for reconnaissance and intrusion support. Security teams also reported malware families that query LLMs to rewrite payloads, generate variants, and iterate on social engineering copy.
None of that means your shipping agent will join a nation-state campaign. It means the skill floor dropped and the speed ceiling rose. Defenders now compete with automation that never gets tired and never feels guilty about skipping confirmation.
Release testing has to follow the wiring you ship. When Agnostics scans agent targets, we treat tools, retrieval connectors, and memory as part of the attack surface, not accessories to the chat UI.
If your threat model still lists only prompt injection in chat, update it this week. Tool schemas, OAuth scopes, browser cookies, and shared agent memory belong on the same diagram as your API gateway.
Public reporting described patterns. Your obligation is to test whether those patterns reach your configured endpoint before customers do.
Everyone's a hacker now (the asymmetric skill floor)
Offense used to reward depth: years of protocol knowledge, custom exploit development, patient recon. LLM-assisted workflows compress parts of that stack into prompts and tool calls anyone can iterate.
You do not need to believe every product manager will turn malicious to see the problem. Curious users, angry customers, bored teenagers, and competitors running black box tests all probe the same surface. Many will phrase requests badly. Some will phrase them well.
Agents amplify curiosity into action. A user who would never open a terminal might still ask an agent to export all records, rotate credentials, or browse an internal URL because the UI made agency feel normal.
Asymmetric skill floor means your least technical adversary can attempt attacks that previously required a specialist. Your defense cannot assume attackers will miss the obvious path.
Quantitative release testing is how product teams reclaim asymmetry on the defense side. Defined attack packs, recorded findings, and retests after fixes let you improve faster than anecdotal bug reports from social media.
- Lower skill floor: more people can attempt tool abuse through natural language
- Higher speed ceiling: automation retries variants without burnout
- Broader surface: retrieval, memory, and sub-agents add injection points
- Same launch bar: harmful side effects still block ship when your policy says they should
Vibe coding and vibe hacking use the same capabilities
Vibe coding is the shorthand for building software by describing intent and letting models draft code, wire APIs, and iterate until something runs. It is fast. It is seductive. It is also how half-finished permission checks reach production.
Vibe hacking is the same loop pointed at offense. Describe the goal. Let the model draft a script, refine a payload, or chain tools until something works. Public reporting on LLM-querying malware illustrated the pattern: automation asks the model for the next variant when defenders adjust.
The capabilities are symmetric. The same assistant that scaffolds a CRUD app can scaffold a credential-stuffing helper if the operator has credentials and weak gates. Your product decision is whether the agent you ship makes that loop easy on your infrastructure.
Teams that vibe-code features without adversarial scans often vibe-ship vulnerabilities. Missing authorization on a new tool endpoint is invisible in a demo that only shows happy paths.
Release testing breaks the symmetry in your favor. You pressure the agent you built with the same creative iteration attackers will use, but on staging, with evidence, before launch.
Three AI attack roles: operator, builder, and enabler
Public threat reporting clusters agent abuse into three roles. You do not need to memorize vendor taxonomies. You need to map roles to tests on your target.
The operator role uses agents to run workflows: browse targets, summarize recon, draft phishing, queue follow-up actions. Impact depends on what your agent can reach. A read-only research agent leaks less than an agent with write tools on production data.
The builder role uses models to create or mutate artifacts: scripts, configs, droppers, obfuscation layers. Your app may not host malware generation, but internal copilots with code execution can still produce harmful outputs that become inputs elsewhere.
The enabler role uses models to remove friction: explain errors, suggest bypasses, translate security jargon into steps. That role shows up in support bots that over-explain internal APIs and in agents that coach users around guardrails you thought were fixed.
When Agnostics scans agent targets, attack packs target operator-style hijacks and enabler-style scope breaks on your configured tools. Builder-style risks on customer-facing agents often show up as unsafe tool actions or boundary bypass on code and exec paths you exposed.
- Operator: chained tool use toward recon or exfiltration on your endpoints
- Builder: generated code or configs that escape intended sandboxes
- Enabler: coaching that defeats confirmations, roles, or policy text
Mutating malware and machine-speed recon in public reporting
Researchers reported malware that queries LLMs to mutate signatures and refresh lures. Defenders see a moving target because automation asks for another variant instead of hand-rewriting.
Separate public reports described agent-assisted recon at machine speed: summarizing exposed assets, drafting exploit paths, and parallelizing grunt work that used to take days. The headline is not that models replaced elite operators. It is that iteration got cheaper.
Your agent product can accidentally become recon infrastructure if it browses broadly, caches sensitive pages, or returns internal URLs in answers. Even benign features leak architecture hints attackers reuse.
Traditional malware categories still apply: persistence, credential access, lateral movement. LLMs change how fast variants appear and how readable playbooks become. They do not repeal network segmentation or least privilege on your APIs.
Release testing should include scenarios where untrusted retrieval or user input steers browsing and tool plans toward paths you intended to block. Indirect injection through web pages is not a future threat. It is a current agent failure class.
You cannot outsource that class to a vendor refusal list. You can reproduce it on staging with attack packs, capture the tool plan that resulted, and decide whether your Release Gate stays Blocked until the pattern closes.
Why traditional detection assumptions break
Signature-based defenses assume attackers reuse bytes. LLM-assisted mutation generates fresh text and restyled code faster than hash blocklists update.
Rule-based content filters assume harm looks like known strings. Agents assemble benign-looking steps that become harmful only in sequence: fetch internal doc, extract token pattern, call admin tool.
Human SOC review assumes speed limits. Machine-speed loops generate more attempts per hour than analysts can read, especially when each attempt uses different phrasing.
Detection still matters at the network and identity layer. Zero trust, egress controls, and tool allowlists reduce blast radius. They do not replace testing whether your agent obeys those controls when a user applies social pressure.
Assume detection will miss novel phrasing on day one. Assume your agent will face prompts no rule writer predicted. Release testing records failures with reproduction steps so engineering fixes patterns instead of playing alert whack-a-mole after launch.
Detection is a safety net. Release testing is the proof your agent respects the net before you walk the wire.
The offense-defense blur nobody wants to schedule
Security teams live with uncomfortable dual use. Password auditing scripts, port scanners, and exploit tutorials are legitimate in authorized hands and illegal elsewhere. Agents collapse the distance between ask and execute.
A developer copilot asked to improve login performance might suggest query changes that weaken tenant isolation if tools can run SQL. A support agent asked to troubleshoot billing might read fields policy says should stay masked.
Pen testers use the same agent features product teams ship. The difference is scope, authorization, and intent. Your app cannot read intent reliably. It sees language and tool permissions.
Public discourse sometimes frames this as models going rogue. Most launch incidents look like authorized tools doing unauthorized work because instructions hijacked the plan. That is an application-layer failure you can test.
Schedule adversarial scans with the same seriousness you schedule pen tests. The output should feed a Release Gate, not a PDF that nobody links to the launch checklist.
Why model refusals alone fail for your app
Base model refusals are a provider-layer control. They reduce obvious abuse in generic chat. They are not a contract about your retrieval corpus, your tool wiring, or your multi-agent handoffs.
Indirect injection hides instructions where refusals never look: white text on a web page, metadata in a PDF, a prior ticket body loaded into context. The model may comply with the hidden instruction while refusing the visible one.
Tool schemas bypass refusal semantics entirely. The model does not need to say it will exfiltrate data. It calls export_user_records and returns a polite summary.
Fine-tuned or system-prompted agents can be more compliant with user goals, which helps product until it hurts security. Helpful tone is not authorization.
Your release bar must be runtime behavior on the configured target. When Agnostics scans agent targets, findings show tool calls, retrieval snippets, and model output together. A refusal in chat paired with a successful tool call is still a blocker under most policies.
What Agnostics tests on agent targets
Agnostics tests the agent as deployed: endpoint, auth mode, tool list, retrieval sources, memory behavior, and confirmation rules. We do not grade abstract model cards. We pressure the path users and integrators hit.
Scans run attack packs aligned to exposure. Prompt injection and boundary bypass test instruction hijacking across user input, retrieval, and tool results. Unsafe tool actions test whether side effects require the consent you think they require.
Sensitive context exposure packs probe whether session data, other users' fields, or internal runbooks leak through answers or tool payloads. Permission abuse packs attempt actions outside the role you configured.
Each finding includes severity, reproduction steps, and evidence: prompts used, retrieval hits, tool invocations, and model output. That transcript is what engineering needs, not a vague severity ticket from a generic scanner.
The Release Gate summarizes whether this scan met your policy for this target. Ready, Monitor, Fix, or Blocked is a product decision with audit trail, not a vibes check from the last demo.
Multi-agent topologies get the same treatment. If a router hands work to specialist agents, configure the target to mirror those handoffs. Hijacking often appears in intermediate messages users never see.
Attack pack mapping for agent risks
Attack packs group scenarios by failure class so coverage stays discussable in release reviews. You can add packs over time. You cannot wish away classes you never tested.
Our unsafe-tool-actions pack surfaces refunds, deletes, exports, and admin calls triggered without the confirmations your policy requires. Start here if your agent touches money, records, or infrastructure.
Prompt injection packs pressure instruction hijacking from user messages, retrieved pages, and poisoned tool output. Essential for any agent that reads untrusted text before acting.
Sensitive-context-exposure packs target cross-session leaks, oversharing from retrieval, and answers that repeat secrets adjacent to the conversation. Critical when agents sit on support or internal knowledge bases.
Boundary-bypass packs attack role, scope, and persona limits you advertised to customers. Permission-abuse packs attempt tool calls outside configured roles even when the user sounds polite.
- unsafe-tool-actions: harmful or irreversible side effects without proper gates
- prompt-injection: hostile instructions steering plans and tool selection
- sensitive-context-exposure: data spills through answers or tool payloads
- boundary-bypass: broken scope, role, or policy promises under pressure
- permission-abuse: successful calls outside least-privilege intent
Findings that should block ship
Not every finding is launch-critical. Agent products need explicit never events, the outcomes that block release if observed on staging under adversarial coverage.
Unconfirmed destructive tool success should block ship: deleting records, issuing refunds, rotating credentials, or invoking admin-only endpoints because the model treated a paraphrase as consent.
Cross-tenant or cross-user data access through tools or answers should block ship. Support agents that return another customer's fields fail the minimum bar for any multi-tenant product.
Instruction hijacking that routes browsing or retrieval toward internal assets you meant to exclude should block ship, even if no exploit chain completed yet. The plan leak is the finding.
Repeated confirmation bypass on financial or privacy workflows should block ship. Monitor state may be acceptable for cosmetic scope breaks with documented acceptance. It is not acceptable for wire transfers.
Severity in Agnostics ties to launch impact on this target, not abstract CVSS theater. Your release policy decides which severities map to Blocked versus Fix versus Monitor.
If you cannot name three never events for your agent, you do not have a release policy yet. You have optimism.
Release Gate for agent workflows
Agent launches fail when security findings live in one doc and the ship checklist lives in another. The Release Gate connects them on the same target record.
Ready means this scan with these packs met your policy bar. Monitor means ship with documented gaps stakeholders explicitly accept. Fix means address findings before launch. Blocked means critical items remain open.
Agent workflows need gate states per target, not per model vendor. You might ship a read-only research agent while a write-capable ops agent stays in Fix because tool abuse findings remain open.
Export release reports when compliance or leadership needs sign-off. Reports should reference findings and retests, not hand-wavy summaries about responsible AI.
Sample Demo Data demonstrates gate states without touching your endpoints. Use it to train the team on how Blocked differs from Monitor before you run the first live scan.
Retest after you narrow tools
The first agent scan often looks ugly because early prototypes grant wide tool access to unblock demos. Engineering responds by removing tools, adding confirmations, and tightening scopes. That is good.
Narrowing tools without retest is faith-based security. Attackers do not care that you commented out a function. They care whether the remaining path still reaches the same outcome through injection or permission gaps.
Retests run the same attack packs against the same target class after material fixes. Improved means verified under the same coverage, not assumed because someone merged a PR titled harden agent.
Watch for partial fixes that block the literal scenario but leave the failure class open. Changing one refund tool name does not stop permission abuse if another write tool still inherits the same credential.
When Agnostics scans agent targets after a narrowing pass, compare findings scan to scan. New Blocked items after a fix attempt mean the patch missed the pattern. Cleared criticals with stable coverage mean you can argue honestly for Ready.
Continuous scans on agent changes
Agents are living systems. New tools land mid-sprint. Retrieval corpora update daily. Model upgrades shift compliance and creativity. Memory formats change without a press release.
Continuous scanning means rerunning relevant attack packs after material changes and before major releases, not scanning once before Product Hunt and never again.
Treat production incidents as free test cases. When support forwards a weird transcript, reproduce it on staging, add coverage if packs missed the pattern, and retest.
Agent red teaming strategy should name owners: who updates the target config, who triages findings, who decides Monitor versus Blocked, who schedules rescans after corpus imports.
The goal is not zero findings forever. The goal is shrinking time from new risk introduction to measured result.
Pair rescans with change logs. When someone adds a browser tool or expands retrieval to a new domain, the security note should include which attack packs rerun and who owns the Release Gate update.
What to do this week: actionable checklist
You do not need a twenty-page threat brief to move. You need one configured target and one honest scan.
Pick the agent endpoint closest to production wiring, not the sanitized demo. Match auth, tools, retrieval, and confirmation behavior you intend to ship.
Write three never events. Example: unconfirmed refund, export of another user's records, successful admin API call from a standard session. Share them with engineering and security.
Run unsafe-tool-actions and prompt-injection packs first on that target. Read findings for patterns, not one weird reply. Hand engineering reproduction transcripts.
Schedule a thirty-minute Release Gate review with product and engineering. Decide Blocked, Fix, Monitor, or Ready with names attached. Book the retest date before anyone closes the meeting.
- Configure one production-faithful agent target in Agnostics
- Document three never events that block launch for that target
- Run unsafe-tool-actions and prompt-injection attack packs
- Triage findings with evidence and assign pattern fixes
- Retest after fixes and record Release Gate state before ship
What Agnostics does not claim
Agnostics did not discover state-sponsored campaigns or author third-party threat research. We synthesize public reporting so product teams know what to test, not to claim credit for intelligence work.
Agnostics does not guarantee your agent cannot be broken after launch. Novel phrasing, corpus drift, provider changes, and zero-day tooling remain ongoing risks.
Agnostics does not replace secure engineering, identity controls, network segmentation, or secret hygiene. Scans test whether your AI surface fails under adversarial input on the configured target.
Agnostics does not prove negative about malware your customers run on their laptops. We focus on the agent workflow you ship and the side effects it can reach.
Sample Demo Data in the interactive demo illustrates findings and Release Gate states without connecting to your endpoints or executing live attacks in demo mode.
A Ready Release Gate means this scan met your policy for this target at this time. It is not a permanent guarantee against agent abuse.
The uncomfortable answer
Will AI agents hack everything? No. Will they keep breaking things teams thought were safe? Yes, especially when tools outrun tests.
Public threat reporting will keep documenting operator, builder, and enabler abuse. Vibe coding will keep shipping half-baked permissions. Vibe hacking will keep iterating until something works on your staging clone.
Your job is not to win a philosophical debate about AGI risk on Twitter. Your job is measurable defense on the agent you ship this quarter.
Configure the target. Run the packs. Record findings with evidence. Fix patterns. Retest. Let the Release Gate say Blocked until never events stop reproducing.
That loop will not make headlines. It will keep you out of the incident postmortem that ends with we should have tested tool calls before launch.
Agents will keep breaking assumptions. Your advantage is a release process that treats breaks as expected input, measures them, and blocks ship until the never events stop reproducing under defined coverage.
Questions
Will AI agents hack everything?
No single technology hacks everything. Agents do lower the skill floor and raise iteration speed for abuse on configured surfaces. Public threat reports documented recon automation, malware mutation assistance, and tool-driven intrusion support. Your defense is release testing on the agent you ship, not denial.
What changed in the 2025 agent threat landscape?
Public reporting highlighted tool-using workflows used for recon, social engineering drafts, and malware variant generation via LLM queries. Harm shifted from text-only failures to chained actions through browsers, APIs, and scripts. Product teams feel this as more ways users and attackers reach side effects through natural language.
How is vibe hacking related to vibe coding?
Both use describe, generate, iterate loops with models and tools. Vibe coding builds features quickly. Vibe hacking applies the same iteration to offense. Symmetric capabilities mean you must adversarially test agents built fast, especially tool permissions and confirmation flows.
Why are model refusals not enough for agent security?
Refusals are provider-layer controls on generic abuse. They do not cover indirect injection in retrieval, tool calls that bypass chat semantics, or permission gaps in your APIs. Release testing must observe runtime behavior including tool payloads on your configured target.
Which Agnostics attack packs matter most for agents?
Start with unsafe-tool-actions and prompt-injection on any agent that can change data or call APIs. Add sensitive-context-exposure when agents handle multi-user or internal knowledge. Add boundary-bypass and permission-abuse when you advertise strict scope or role limits.
What findings should block an agent launch?
Block on unconfirmed destructive or financial tool success, cross-tenant data access, hijacked plans that reach internal assets you meant to exclude, and repeated confirmation bypass on privacy or money workflows. Your release policy maps severities to Blocked, Fix, Monitor, or Ready.
When should we retest an agent after fixes?
Retest after narrowing tools, adding confirmations, changing auth, updating retrieval sources, swapping models, or modifying memory behavior. Use the same attack packs to verify the failure class closed, not just the single scenario that embarrassed the demo.
Does Agnostics claim to stop nation-state agents?
No. Agnostics synthesizes public threat reporting into release testing guidance and runs attack packs on your configured targets. We do not claim to discover campaigns or guarantee post-launch safety. A Ready Release Gate means this scan met your policy at this time.