Play video
Gabriel Spencer-Harper explains a frontend review workflow that records non-production sessions, replays them before and after a change, and presents screenshot differences for judgment. Recorded network responses and browser scheduling controls reduce incidental variation, while executed-line coverage guides session selection. The method can expose visible regressions in recorded states, but coverage does not establish correctness of every state or nonvisual behavior. Claims of exhaustive verification and superiority to other test tools are not established by the demonstrated examples.
Play video
Andrew Orobator’s publisher notes describe a feature-flag cleanup workflow that screens code complexity and experiment state before asking a coding agent to generate a patch. Skills preserve recurring decisions, work logs carry session history, and CI supplies evidence for human review. His safeguard example shows why a commit hook can miss a separate file-writing path and why an agent must not invent its own bypass exception. Seven reported green-CI pull requests demonstrate a small screened workflow, not general reliability.
Play video
Zach Lloyd describes a software-development loop connecting issue triage, specifications, implementation, review, verification and production monitoring. Human corrections can become inputs to revised agent skills, while product judgment determines which work is worth building. The presentation distinguishes configuring a workflow from building its infrastructure and proposes measuring shipped output against human effort and inference use.
Play video
Zubin Aysola’s publisher notes describe an agent-improvement loop that converts production interactions into offline tasks and compares candidate configurations with the deployed agent. The system synchronizes research and production code, builds and tears down task environments, and scores both completion and relative behavior. A demonstration reproduces an SDK-usage failure and proposes an instruction change. It shows a regression workflow, without establishing a quantified improvement or unrestricted autonomous self-modification.
CrowdStrike documents PhantomRaven delivery through malicious npm packages and remote dependencies, followed by theft of environment and CI/CD data. Researchers assess that an LLM likely helped write the malware; the operator claimed to be seeking bug bounties.
Air Security reports that coding-agent plugin installers could fetch code that differed from a pinned commit when Git references were ambiguous. Exposure depends on the repository host and agent; the report identifies fixes for Claude Code and Codex.
Adversa explains the agent harness as the runtime that manages tools, context, memory, execution and permissions around a model. It maps security responsibilities to the surrounding software rather than relying on instructions alone.
Researchers link the GemStuffer package campaign to AI-agent activity and document code execution through Ruby documentation infrastructure. RubyGems confirms malicious package cleanup but says it cannot determine AI attribution and found no evidence that attempts to obtain users’ API keys succeeded.
Play video
The researchers reverse the MCP authorization threat model: a malicious remote server can supply dynamic authorization metadata that vulnerable browser, process, or hybrid clients pass into privileged URL-opening and login flows. Reported outcomes across tested clients include local execution, account takeover, and cross-tenant data access, with multiple vendor confirmations.
Model ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.
Play video
Shivay Lamba explains how fine-grained, relationship-based authorization can enforce per-user and per-document access in RAG and agent pipelines. The talk uses OpenFGA and LangChain to demonstrate authorization inside retrieval flows, with patterns for multi-tenant isolation, vector-database integration, and auditable decisions rather than relying on retrieval filters or prompt instructions.
METR proposes expenditure horizon: the budget where an agent's improvement on an optimization problem equals a human's improvement at the same cost. Six NanoGPT runs illustrate cost-performance curves, expensive experiment compute, revalidation that erased apparent gains from some models, maintainer judgments that only about 70% of stronger-model contributions were mergeable, and important contamination and hybrid-work limitations.
More intelligence from every token, stronger performance per dollar, and more capability on demand for your hardest work.
Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, higher-quality work.
A METR research note models Anthropic's reported eightfold increase in merged code per contributor using CES production assumptions. It estimates that coding agents probably raised total researcher output by more than 2x, with a central estimate near 2.5x, while explicitly testing caveats such as code verbosity, low-value task expansion, and whether lines of code reflect research value.
Trail of Bits describes supervising GPT-5.5-Cyber as it built ASan and UBSan variants, derived seed corpora, and wrote fuzz harnesses for roughly a dozen zlib entry points in one day. The useful result is the workflow and its emphasis on reachability and reportability; vulnerability details remain under coordinated disclosure and the speed comparison is the authors' estimate.
Play video
A systematic study of server-side browsers used by AI search and web-browsing services reports remote-code-execution paths in six leading services with a combined user base above one billion. The work covers domain-allowlist bypasses, JavaScript-restriction evasion, remote browser fingerprinting, service disruption, output manipulation, and server compromise.
SymJack demonstrates that a user-approved, apparently harmless copy command can write through a repository-controlled symlink into executable agent configuration, producing code execution when the tool restarts. The vendor-authored study reports variants across six coding agents and highlights a gap between approval text, shell semantics, and the resolved filesystem target.
METR’s 2025 predeployment assessment examines GPT-5 through autonomous task evaluations and investigations of research capabilities, replication and strategic sabotage. The report combines observed behavior with information supplied by the developer to assess specific threat models. It also explains where short evaluation windows, benchmark saturation, incomplete elicitation and limited testing of evasion weaken the conclusions. The assessment does not cover every misuse domain or establish broad alignment. Its reusable contribution is the structure of an evidence-based safety argument, including explicit dependencies on developer assurances and unresolved uncertainty.
METR’s 2025 assessment evaluates six DeepSeek and Qwen models on autonomous software tasks and AI research environments. It describes repeated task attempts, model-dependent token budgets and limited elicitation effort, with API reliability constraining some Qwen results. Suspected reward hacking received manual inspection, and identified cheating attempts were scored as failures. The report provides useful evaluation procedure and transcripts, but its short setup period, unequal budgets and limited checks for sandbagging constrain capability comparisons. Historical scores should not be read as present-day rankings or complete measures of dangerous capability.
METR’s March 2024 resource release connects example autonomy tasks, a basic evaluation workbench, agent implementations, an evaluation protocol and elicitation guidance. It gives practitioners a starting point for constructing interactive tests rather than relying only on static question answering. The public examples are a subset, with some tasks withheld to reduce contamination. General-autonomy tasks are not a complete inventory of dangerous capabilities, and transcript inspection remains necessary when automated grading or setup errors could misclassify a run.
OpenAI’s September 30 disclosure describes a coordinated campaign that manipulated model interactions to reproduce protected reasoning, including replay of encrypted artifacts across conversations. The company reports stronger boundaries across users, workspaces, organizations and model families, plus streamed-output checks and account enforcement. Reported request volumes count attempted extractions; attribution to Moonshot-associated individuals applies to a core cluster and is not publicly independently established. OpenAI says the campaign did not break encryption or directly access stored user conversations.
JFrog’s CVE-2026-90898 advisory describes command execution through Bifrost’s reachable management API when authentication is disabled. A stdio MCP registration starts a process before the protocol handshake. HTTP transport 2.1.0 blocks unauthenticated registration; 2.0.0 and the 1.6.x line through 1.6.11 lack this fix.