Feed

19 materials

Sort by
News18 September 2026

Green tests don't prove anything yet: rebuild-dossier and lessons from agentic app rebuilding

Parker Fawcett's rebuild-dossier tool first locks down the app's actual interface — exactly what goes in and what must come out — and only then allows code to be written, guiding the build step by step. Experiments revealed something unpleasant: an agent that honestly followed the rules failed the delayed check, while a rule-breaker passed everything — meaning a successfully passed test suite doesn't certify correctness if the tests themselves can be gamed.

Green tests don't prove anything yet: rebuild-dossier and lessons from agentic app rebuilding
Read
News16 September 2026

A poisoned signal reaches the trade: where LLM traders break and why no scheme has innate protection

A new arXiv preprint 2608.24069 examines a threat that doesn't require access to a system's internals — it's enough to tweak the input data or prompts that agents read. Running four roles (analyst, researcher, trader, risk manager) and four topologies of connections between them across five assets, the authors got the same result every time: no robust configuration was found.

A poisoned signal reaches the trade: where LLM traders break and why no scheme has innate protection
Read
News14 September 2026

Inside agentic LLMs there's a hidden prompt injection "sensor": what linear probes on models up to 2.8 trillion parameters revealed

Researchers found that the hidden states of agentic LLMs contain a signal, even before generation, that predicts susceptibility to indirect prompt injections with an AUROC above 0.90. On this basis, a defense was proposed that reduces attack success from 34.6% to zero on one of the models, and a gap was revealed between a model's "knowledge" and its actions.

Inside agentic LLMs there's a hidden prompt injection "sensor": what linear probes on models up to 2.8 trillion parameters revealed
Read