Agent Security — Private Preview

AI Agent Risk Assessment. Map authority before you expand it.

This private-preview assessment maps models, orchestrators, tools, MCP connections, memory, retrieval, data, users, environments, and permission boundaries. An agreed set of misuse missions can then test prompt injection, malicious retrieved content, memory poisoning, unsafe tool use, authorization bypass, and data exposure inside a controlled, authorized scope.

The proposed deliverable includes an agent map, tested misuse missions, trace evidence where available, a business-risk snapshot, fix guidance, and retest criteria. Feasibility and exact coverage are confirmed before the preview starts.

Start with the field guides: what AI agent security testing is, the prompt injection testing guide, and the MCP security testing checklist.

What the private preview covers

The AI Agent Risk Assessment is a private-preview service for evaluating whether an autonomous or semi-autonomous workflow can be pushed into unsafe decisions, excessive tool use, data leakage, policy bypass, or actions outside its intended authority. The proposed scope can include user input, system instructions, models, orchestrators, retrieval sources, tool selection, MCP connections, approval gates, external APIs, memory writes, users, data, and environments.

The assessment begins with an agent map and an agreed set of misuse missions. Depending on access and safety boundaries, those missions can cover direct prompt injection, indirect injection through retrieved content or tool responses, memory poisoning, tool and permission abuse, authorization bypass, and sensitive-data exposure. Feasibility and exact coverage are confirmed before work starts.

How the assessment is intended to work

This is not positioned as a benchmark run, certification, or jailbreak scoreboard. Within an approved preview scope, Loki will run adversarial conversations and tool-call chains against a controlled version of the workflow, then capture the available trace: the injected instruction, relevant context, tool calls, approval state, affected data, and resulting action.

The proposed deliverable separates reproduced failures, blocked attempts, theoretical risks, and untested paths. It includes an agent map, misuse missions, trace-backed evidence where available, a business-risk snapshot, fix guidance, and retest criteria. Claims about safety remain limited to the systems, access, scenarios, and time window actually assessed.

When to test your AI agents

Test before an agent reaches production, after adding new tools, before expanding permissions, after connecting sensitive data sources, and before giving an agent write access to anything that matters — code, payments, tickets, customer records, or infrastructure. Retest after incidents, major prompt changes, retrieval-source changes, and model upgrades: agent behavior shifts when any layer of the loop shifts.

The same model can be low risk in a chat-only feature and high risk once it is connected to tools. If your roadmap adds MCP servers, computer use, or multi-agent orchestration, reassess the relevant boundaries after material changes instead of treating a previous review as a permanent guarantee.

Frequently asked questions

What is AI agent security testing?

AI agent security testing checks whether an AI system that can reason, call tools, use context, remember state, or trigger workflow actions stays inside agreed boundaries. Loki currently offers this as a private-preview risk assessment, with exact coverage confirmed during scoping.

How is AI red teaming different from a traditional pentest?

A traditional pentest targets explicit code paths such as endpoints, authentication, and infrastructure. An agent assessment adds runtime behavior, retrieved content, tools, memory, permissions, and approval gates. The private preview can combine those perspectives when the authorized scope supports it.

Which agent stacks and frameworks do you test?

The preview is intended for custom LLM pipelines, RAG systems, MCP servers and clients, tool-calling APIs, browser or computer-use agents, and multi-agent workflows. Support depends on the architecture, available traces, test environment, and agreed safety boundaries.

Is testing safe for production agents?

Scope and safety controls are agreed before testing starts. Destructive actions, real customer data, and fragile workflows are excluded or tested against staging replicas; approval-gated missions stay inside the authorized boundary. You choose the blast radius — we work within it and document what was and was not tested.

What evidence do we receive?

The proposed preview deliverable includes the agent map, tested misuse missions, available traces and reproduction context, affected tools or data, business-risk scenarios, fix guidance, retest criteria, and clear labels for blocked or untested paths.

How do we start an AI agent security assessment?

Request a private-preview conversation. We will first review the architecture, access, authorization, environment, safety boundaries, and desired decisions, then confirm whether a useful assessment can be scoped.