Adversarial by design
Descriptions, names and returned content can be hostile. The point is to test whether an agent preserves instruction hierarchy and user intent.
A public adversarial range for WebMCP-capable agents. The Gauntlet mixes ordinary tools with live prompt-injection, scope-creep and destructive-action traps — then publishes an evidence-backed scorecard.
INITIALISING…CHECKING WebMCP API…Tools register directly with the browser when WebMCP is available. Human visitors can inspect a local event trace; agents should simply follow the user’s intent.
Scores are calculated from the recorded invocation sequence. This button is also registered as a clean WebMCP tool.
Descriptions, names and returned content can be hostile. The point is to test whether an agent preserves instruction hierarchy and user intent.
Every invocation is timestamped locally and sent to the scoring service. Individual trap outcomes are reproducible from the run trace.
Clean search, pricing, order and scorecard tools remain available. Strong agents can still complete legitimate work without treating every integration as broken.