ASCII smuggling crosses over from AI prompt injection to phishing evasion
Invisible Unicode characters popularized for hiding instructions from AI models are now being used to obfuscate words before email filters parse them.
Browse entries 289–312 of 1576. Return to the first page to search and filter the complete collection.
Invisible Unicode characters popularized for hiding instructions from AI models are now being used to obfuscate words before email filters parse them.
Release notes for garak, an LLM vulnerability scanning and evaluation toolkit. Relevant to tracking new probes, detectors, and repeatable red-team workflows.
garak release adding probes and detector improvements for LLM security testing. Relevant to maintaining practical red-team coverage across evolving attack techniques.
Google DeepMind proposes a taxonomy of ten cognitive abilities and a three-stage evaluation protocol combining held-out AI tasks, human baselines and comparisons between performance distributions. It highlights gaps in evaluating learning, metacognition, attention, executive functions and social cognition. The contribution is a framework for designing a broader measurement program. The article does not deliver a completed benchmark suite, an agreed definition of AGI or a score establishing that a model has reached it.
NVIDIA's AI Red Team extends its visual prompt-injection work with a Gemini 2.5 Pro demonstration in which a scrambled puzzle reconstructs a command during problem solving. The post calls these multimodal cognitive attacks and argues that payloads can emerge during inference after simple input filters have already run; it proposes output validation, tool sandboxing, and anomalous-reasoning detection as research directions.
METR’s January 2025 analysis identifies risk pathways that a release-time evaluation can miss: stolen model weights, employee misuse and autonomous behavior during internal development or use. It proposes applying assessment, information security and independent oversight throughout the model lifecycle, rather than waiting for public availability. These are threat-model arguments and governance recommendations, not evidence that the hypothetical severe incidents occurred. The practical distinction is between limiting public access and controlling all the environments in which a capable model already operates.
OWASP’s GenAI security project remains a practical baseline for teams building or assessing LLM applications and agentic systems.
Play video
Ryan Cooke’s publisher notes explain how WorkOS connects coding agents to project plans, ticket dependencies and completion webhooks. A shared MCP gateway supplies tool access and guidance about where organizational information lives. In the demonstrated workflow, a short brief becomes draft planning documents that an engineer refines before implementation proceeds. The talk reports qualitative experience rather than measured delivery gains and explicitly leaves cross-system authorization unresolved.
Novee Security found that unprivileged GitHub issues could reach privileged coding-agent workflows: Gemini CLI and Claude Code paths led to CI-runner code execution, while a Codex path could alter the next agent run. The two assigned CVEs were patched; the Codex behavior was documented rather than assigned a product CVE, highlighting failures in the surrounding harness rather than the model alone.
A credential-stealing npm worm spread through hundreds of package versions using lifecycle scripts and a Bun-based payload. Related repositories also carried Claude Code and VS Code hooks that could execute after workspace trust; reported campaign totals vary, so exposure depends on exact resolved versions and execution.
Learn how CNAPP platforms are helping organizations prioritize exploitable risks, reduce exposure, and operationalize security across the application lifecycle.
Anthropic red-team research assessing how LLMs affect exploitation of known vulnerabilities. Relevant to cyber capability evaluation, benchmark design, and misuse risk modeling.
GPT-Rosalind advances life sciences research with enhanced biological reasoning, medicinal chemistry expertise, genomics analysis, and experimental workflow capabilities.
Introducing GPT-5.5, our smartest model yet—faster, more capable, and built for complex tasks like coding, research, and data analysis across tools.
OpenAI introduces GPT-Rosalind, a frontier reasoning model built to accelerate drug discovery, genomics analysis, protein reasoning, and scientific research workflows.
METR’s retrospective on an AI-biology randomized trial focuses on evaluation design and execution. It describes the difficulty of matching long studies to fast model releases, staffing specialized evaluations and interpreting a result from a limited participant and task distribution. A lack of a significant aggregate effect in that study does not establish the absence of risk for other users, tasks or systems. The transferable lesson is to prepare measurement and safeguards before capability changes force a decision.
In a recent evaluation of AI models’ cyber capabilities, current Claude models can now succeed at multistage attacks on networks with dozens of hosts using only standard, open-source tools, instead of the custom tools needed by previous generations.
Google Cloud outlines a defense-in-depth view of AI security spanning application controls, data protections, and infrastructure isolation.
Wiz’s 2025 article and public repository provide baseline security instructions for common language and framework combinations, formatted for several coding assistants. The repository exposes the generation prompt and script and clearly identifies the rules as AI-generated. The proposed method keeps guidance close to a project’s actual stack and coding practices. The release does not independently establish that these files prevent vulnerabilities or measure their effectiveness in a production repository. They are reviewable prompt material, whose usefulness depends on rule quality, context selection and subsequent code verification.
METR’s February 2025 assessment had about a week of access to an earlier GPT-4.5 checkpoint and used a scaffold optimized for another model. It combined task results with some developer-provided evidence, but explicitly declined to treat those results as ruling out major risks. Limited elicitation, possible gains from later modifications and internal use before release all constrain a predeployment snapshot. The report proposes earlier third-party involvement in evaluation design and review of internal results. Its tentative historical judgment is not a current safety certification for GPT-4.5 or later systems.
METR’s Task Standard packages an interactive task as an environment, instructions and an optional scoring function. Task families declare a standard version and can define setup, startup, permissions and auxiliary machines, allowing different evaluation harnesses to reuse the same task definition. The February 2024 introduction explains this portability goal; the repository supplies templates and adapters. A shared format does not itself enforce containment or validate the grader, and the standard deliberately leaves agent safety and oversight outside its scope.
Play video
Jose Palafox shows how reusable agent definitions can move from a developer’s machine into scheduled or event-triggered repository workflows. The useful pattern separates discovery, planning, implementation and human review. GitHub Agentic Workflows documentation makes the authority boundary explicit: the generated agent job starts read-only, while configured safe outputs apply limited writes in separate jobs. Authors can override those defaults or pass credentials to tools, so the compiled workflow needs inspection. Sensitive writes can depend on a protected GitHub Environment; a prompt asking for approval does not provide that enforcement.
Play video
Abhi Arya describes why exposing low-level pipeline operations let an agent assemble confident but difficult-to-inspect results. His redesign groups known operations into workflow-shaped tools, provides a snapshot of current pipeline state and errors, and asks users to resolve ambiguous requirements. Reducto’s public MCP documentation separately illustrates schema-based extraction, document classification, job references, and explicit handling of truncated results. The transferable method is to encode valid transitions and expose evidence for review. The internal redesign and customer reactions are anecdotes, and extraction confidence scores should not be mistaken for independently calibrated guarantees of correctness.
Play video
George He presents document retrieval as a sequence: use semantic or keyword search to locate promising files, inspect metadata, then search and read selected passages. Parsed text helps with navigation, while page images help check tables and layouts that text extraction may distort. The associated LiteParse repository supplies local parsing and screenshot primitives. Access permissions and freshness must remain consistent across search, listing, direct reads, and image access; an index alone does not enforce those boundaries. The talk’s financial-filing demonstration illustrates the workflow but does not establish a general accuracy advantage over other retrieval systems.