Full Archive · Page 13

Research archive, page 13

Browse entries 289–312 of 1576. Return to the first page to search and filter the complete collection.

Google DeepMind Blog March 17, 2026 news

Cognitive evaluations: cover missing abilities and compare human distributions

Google DeepMind proposes a taxonomy of ten cognitive abilities and a three-stage evaluation protocol combining held-out AI tasks, human baselines and comparisons between performance distributions. It highlights gaps in evaluating learning, metacognition, attention, executive functions and social cognition. The contribution is a framework for designing a broader measurement program. The article does not deliver a completed benchmark suite, an agreed definition of AGI or a score establishing that a model has reached it.

NVIDIA AI Red Team August 7, 2025 analysis

How Hackers Exploit AI’s Problem-Solving Instincts

NVIDIA's AI Red Team extends its visual prompt-injection work with a Gemini 2.5 Pro demonstration in which a scrambled puzzle reconstructs a command during problem solving. The post calls these multimodal cognitive attacks and argues that payloads can emerge during inference after simple input filters have already run; it proposes output validation, tool sandboxing, and anomalous-reasoning detection as research directions.

METR January 17, 2025 analysis

METR argues that safety evaluation must start before public deployment

METR’s January 2025 analysis identifies risk pathways that a release-time evaluation can miss: stolen model weights, employee misuse and autonomous behavior during internal development or use. It proposes applying assessment, information security and independent oversight throughout the model lifecycle, rather than waiting for public availability. These are threat-model arguments and governance recommendations, not evidence that the hypothetical severe incidents occurred. The practical distinction is between limiting public access and controlling all the environments in which a capable model already operates.

Event-driven coding agents: project handoffs still need an authorization model video thumbnail Play video
AI Engineer September 27, 2026 video

Event-driven coding agents: project handoffs still need an authorization model

Ryan Cooke’s publisher notes explain how WorkOS connects coding agents to project plans, ticket dependencies and completion webhooks. A shared MCP gateway supplies tool access and guidance about where organizational information lives. In the demonstrated workflow, a short brief becomes draft planning documents that an engineer refines before implementation proceeds. The talk reports qualitative experience rather than measured delivery gains and explicitly leaves cross-system authorization unresolved.

The Hacker News AI Security August 7, 2026 news

Claude Code and Gemini CLI Flaws Let a GitHub Issue Reach CI Workflow Secrets

Novee Security found that unprivileged GitHub issues could reach privileged coding-agent workflows: Gemini CLI and Claude Code paths led to CI-runner code execution, while a Codex path could alter the next agent run. The two assigned CVEs were patched; the Codex behavior was documented rather than assigned a product CVE, highlighting failures in the surrounding harness rather than the model alone.

The Hacker News AI Security August 4, 2026 news

Keyv-Linked npm Worm Poisons Hundreds of Packages, Plants Claude Code and VS Code Hooks

A credential-stealing npm worm spread through hundreds of package versions using lifecycle scripts and a Bun-based payload. Related repositories also carried Claude Code and VS Code hooks that could execute after workspace trust; reported campaign totals vary, so exposure depends on exact resolved versions and execution.

METR February 19, 2026 analysis

High-stakes uplift studies: design around release timing and limited generalization

METR’s retrospective on an AI-biology randomized trial focuses on evaluation design and execution. It describes the difficulty of matching long studies to fast model releases, staffing specialized evaluations and interpreting a result from a limited participant and task distribution. A lack of a significant aggregate effect in that study does not establish the absence of risk for other users, tasks or systems. The transferable lesson is to prepare measurement and safeguards before capability changes force a decision.

Wiz AI Security June 6, 2025 tool

Wiz publishes stack-specific instruction files for safer AI-assisted coding

Wiz’s 2025 article and public repository provide baseline security instructions for common language and framework combinations, formatted for several coding assistants. The repository exposes the generation prompt and script and clearly identifies the rules as AI-generated. The proposed method keeps guidance close to a project’s actual stack and coding practices. The release does not independently establish that these files prevent vulnerabilities or measure their effectiveness in a production repository. They are reviewable prompt material, whose usefulness depends on rule quality, context selection and subsequent code verification.

METR February 27, 2025 analysis

GPT-4.5’s short evaluation illustrates the limits of release-time safety checks

METR’s February 2025 assessment had about a week of access to an earlier GPT-4.5 checkpoint and used a scaffold optimized for another model. It combined task results with some developer-provided evidence, but explicitly declined to treat those results as ruling out major risks. Limited elicitation, possible gains from later modifications and internal use before release all constrain a predeployment snapshot. The report proposes earlier third-party involvement in evaluation design and review of internal results. Its tentative historical judgment is not a current safety certification for GPT-4.5 or later systems.

METR February 29, 2024 tool

Portable agent evaluation tasks need explicit environments, permissions and graders

METR’s Task Standard packages an interactive task as an environment, instructions and an optional scoring function. Task families declare a standard version and can define setup, startup, permissions and auxiliary machines, allowing different evaluation harnesses to reuse the same task definition. The February 2024 introduction explains this portability goal; the repository supplies templates and adapters. A shared format does not itself enforce containment or validate the grader, and the standard deliberately leaves agent safety and oversight outside its scope.

Agents in CI: inspect permissions and gate external writes video thumbnail Play video
AI Engineer October 10, 2026 video

Agents in CI: inspect permissions and gate external writes

Jose Palafox shows how reusable agent definitions can move from a developer’s machine into scheduled or event-triggered repository workflows. The useful pattern separates discovery, planning, implementation and human review. GitHub Agentic Workflows documentation makes the authority boundary explicit: the generated agent job starts read-only, while configured safe outputs apply limited writes in separate jobs. Authors can override those defaults or pass credentials to tools, so the compiled workflow needs inspection. Sensitive writes can depend on a protected GitHub Environment; a prompt asking for approval does not provide that enforcement.

Design MCP tools around valid workflows and visible uncertainty video thumbnail Play video
AI Engineer October 8, 2026 video

Design MCP tools around valid workflows and visible uncertainty

Abhi Arya describes why exposing low-level pipeline operations let an agent assemble confident but difficult-to-inspect results. His redesign groups known operations into workflow-shaped tools, provides a snapshot of current pipeline state and errors, and asks users to resolve ambiguous requirements. Reducto’s public MCP documentation separately illustrates schema-based extraction, document classification, job references, and explicit handling of truncated results. The transferable method is to encode valid transitions and expose evidence for review. The internal redesign and customer reactions are anecdotes, and extraction confidence scores should not be mistaken for independently calibrated guarantees of correctness.

Combine document search with bounded file reads and page-image checks video thumbnail Play video
AI Engineer October 7, 2026 video

Combine document search with bounded file reads and page-image checks

George He presents document retrieval as a sequence: use semantic or keyword search to locate promising files, inspect metadata, then search and read selected passages. Parsed text helps with navigation, while page images help check tables and layouts that text extraction may distort. The associated LiteParse repository supplies local parsing and screenshot primitives. Access permissions and freshness must remain consistent across search, listing, direct reads, and image access; an index alone does not enforce those boundaries. The talk’s financial-filing demonstration illustrates the workflow but does not establish a general accuracy advantage over other retrieval systems.