What we do
We test the parts of an agent that sit between the model and the action. Whether a planted fact survives a reload. Whether one user can see another user’s memory. Whether a document someone uploads can change how the agent answers a later question. Whether an instruction hidden inside retrieved content gets followed.
Each of those is a measurement with a defined pass condition, not an opinion about risk posture. You get the result per scenario, the fixtures that produced it, and a re-runnable harness so you can check us.
Why memory is the weak joint
Prompt injection gets the attention, but a prompt injection ends with the conversation. Something written into memory does not. It persists, it gets retrieved into future contexts, and it arrives wearing the same clothes as anything the user said themselves.
The frameworks are improving, and we have measured several of them correcting real gaps after a private report. That is the point of a standing measurement rather than a one-off review: it catches both the hole and the fix.
How we work with the tools we test
We maintain an open conformance suite for agent memory integrity. The attacks are versioned, the results are published as data rather than prose, and every row can be reproduced from the repository.
Maintainers get findings privately first and are welcome to dispute a row with evidence. Several have, and the suite is better for it.
WHAT THIS RESTS ON
Open conformance suite
Versioned attacks, published fixtures, results re-runnable from the repo
Vendor fixes measured
At least one framework improved its defences after a private report and the change was recorded
In the working groups
Participating in the OWASP agentic security effort on agent memory
Questions we hear before engagements
Does this work on an agent we built ourselves?
Yes. The suite needs a way to write to memory and a way to query it. Anything with those two front doors can be measured, framework or not.
Is this a penetration test?
No. A penetration test looks for a way in. This measures whether specific, named failure modes occur, with a pass condition defined before the test runs, so the result is comparable over time.
What do we get at the end?
A result per scenario, the fixtures, the harness, and a written read on which gaps matter for your deployment rather than in the abstract.
Related capability
If your system has to be right, let’s talk.
Start the conversation →