Cloudflare’s Turnstile Spin guides a coding agent through widget creation, frontend integration and backend Siteverify validation. It also targets existing widgets that serve traffic without server-side validation. The agent proposes changes for approval and edits the user’s codebase; the application backend remains responsible for accepting or rejecting the request. The practical security issue is an incomplete integration: rendering a challenge widget alone does not protect the operation behind it.
Turnstile Spin: verify the backend check in agent-generated bot protection
Related research
More curated notes connected through AI Engineering and Agent Security.
OWASP Top 10 for Agentic Applications for 2026
OWASP's community guide organizes agentic-system risk into ten categories, including goal hijacking, tool misuse, identity and privilege abuse, memory poisoning, insecure inter-agent communication, cascading failures, and rogue-agent behavior. It provides a shared taxonomy and mitigation starting point rather than a certification checklist or evidence that a deployed system is secure.
AI vulnerability research: measure reproducible findings and completed fixes
Microsoft’s FORGE account describes the work between a model’s vulnerability claim and a useful repair: reusable builds, duplicate removal, reachability checks, project-specific verification, reproducible triggers and regression tests. Structured rejection reasons help improve later searches. The useful operational measure is the flow of findings that survive verification and reach a fix, rather than the number of candidates generated. Reported successful-case costs exclude parts of screening, failed attempts and human work, so they are not the total cost of operating this pipeline.
Frontier training safety cases: connect evidence to enforced pause and rollback controls
OpenAI proposes training-run safety cases combining alignment evaluations, containment and monitoring with explicit operational ownership. Concrete measures include immutable transcripts, held-out incident tests, checks for evaluation gaming, response deadlines and fail-closed monitoring. Independent internal challenge, leadership vetoes and tracking downstream uses support stopping a run and reversing affected work. The article describes recommendations still being implemented, rather than audited proof that every safeguard already operates or that residual risk has been eliminated.