Mikhail Gribov PRO
mihailgribov
AI & ML interests
Understanding LLMs from the inside - probing internals, and testing what survives when the model becomes an agent
Recent Activity
upvoted an article 1 day ago
Give Your Coding Agents a Memory You Own repliedto their post 1 day ago
How often can an email make your AI agent move money?
We gave the agent one job: log an incoming email. But the emails carried an indirect prompt injection - a second instruction, written for the agent rather than for a person: make a payment.
Across nine agentic models, the same injected emails produced payment orders in **0% to 42%** of cases. All nine ran under the same conditions - one agent, one set of tools, the same 395 emails - so the numbers compare directly.
And the average score hides the interesting part: different models fail on different kinds of injections.
Full experiment and results:
https://huggingface.co/blog/mihailgribov/agentic-models-measured-on-the-injections-that-mov
The bench is public too - run your own model through the same test:
https://github.com/mihail-gribov/quadrat-ipi-model-eval
https://huggingface.co/datasets/mihailgribov/quadrat-ipi
#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents posted an update 4 days ago
How often can an email make your AI agent move money?
We gave the agent one job: log an incoming email. But the emails carried an indirect prompt injection - a second instruction, written for the agent rather than for a person: make a payment.
Across nine agentic models, the same injected emails produced payment orders in **0% to 42%** of cases. All nine ran under the same conditions - one agent, one set of tools, the same 395 emails - so the numbers compare directly.
And the average score hides the interesting part: different models fail on different kinds of injections.
Full experiment and results:
https://huggingface.co/blog/mihailgribov/agentic-models-measured-on-the-injections-that-mov
The bench is public too - run your own model through the same test:
https://github.com/mihail-gribov/quadrat-ipi-model-eval
https://huggingface.co/datasets/mihailgribov/quadrat-ipi
#prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents