METR’s March 2025 research measures how agent success varies with the time human experts need to complete software and reasoning tasks. A fitted success curve yields a horizon at a chosen reliability level, while hierarchical resampling estimates uncertainty across task families, tasks and attempts. The original paper also extrapolated historical growth, but that forecast depends on the trend continuing and on transfer beyond the measured tasks. The source distinguishes its dated static results from a live chart. A 50% horizon describes a success threshold, not dependable execution of a project of that duration.
METR’s early-2025 DeepSeek-R1 assessment uses autonomous software tasks and research-engineering environments through a third-party model provider. Transcript inspection found hallucinated tool results, malformed calls and difficulty recovering from mistakes, showing how interface behavior can limit measured performance. The report also acknowledges modest elicitation, incomplete checks for strategic underperformance and potentially biased exclusion of server-error runs. Published example transcripts were selected to illustrate behavior, not estimate its prevalence. These results describe a historical model-and-scaffold combination rather than an upper bound on open-weight model capability.
METR’s early-2025 DeepSeek-V3 report combines autonomous-task testing, research-engineering tasks and a held-out knowledge check. The authors observed repeated actions, hallucinated tool outputs and misread numerical feedback, while scaffold changes substantially improved performance. They explicitly caution that elicitation was limited, curated transcripts were not representative and excluded server-error runs could inflate scores. The study did not robustly test deliberate underperformance. Its historical capability comparison is therefore conditional on the provider, scaffold and scoring choices, with useful lessons for designing more auditable evaluations.
Microsoft summarizes lessons from red teaming more than one hundred generative AI products, emphasizing system-level testing, human expertise, and automation.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This AI Explained video reviews a major AI development through the lens of agentic workflows and tool-use risk. It is useful context for AI engineering, evaluation, governance, and operational risk.
Play video
This vLLM walkthrough explains the tradeoff behind draft-model speculative decoding: a smaller model proposes tokens and the target model verifies them, adding draft computation and memory in exchange for possible faster generation. Acceptance rate, prompt length, output length, sampling settings, and concurrency all affect the result. The talk includes a small live comparison, while vLLM documentation limits the technique’s expected benefit to suitable workloads and distinguishes distributional correctness from identical outputs across execution configurations. Its reported demo speedup is not a transferable benchmark, and spare GPU memory alone does not establish a performance benefit.
Play video
Browserbase’s verifier design generates a task rubric and retrieves relevant screenshots for each criterion alongside action history and the final answer. It separates evidence of effort from evidence of completion: a sensible attempt to buy an unavailable product must not become a successful purchase. It also aims to avoid repeatedly penalizing downstream consequences of one earlier error. Human-labeled development and held-out trajectories support judge calibration, but the small evaluation does not establish universal reliability or guarantee the absence of false positives.
Play video
Bogdan Gaza describes operational lessons from DatologyAI’s large synthetic-data pipeline: batch object-store metadata listing, size partitions to bound lost work, checkpoint partial outputs, and coordinate Ray CPU heads with GPU workers. Curated source documents seed rephrasing, while a shared workflow connects curation, generation, training and evaluation. The talk’s volume, throughput and model-quality figures are vendor-reported and lack a complete reproducible benchmark; the useful contribution is the breakdown of bottlenecks and recovery responsibilities.
Play video
Jan Čurn demonstrates mcpc, a CLI client that exposes MCP sessions through searchable tool descriptions, structured JSON output and background tasks. Progressive discovery reduces the initial catalog in model context; shell composition lets intermediate results flow between operations without being restated by the model. Authentication and persistent sessions are client responsibilities. Early connector comparisons lack a complete numerical protocol, and moving work to a subagent or shell does not by itself protect secrets or establish isolation.
Play video
Ido Salomon presents AgentCraft, which uses a strategy-game interface to make coding-agent activity easier to supervise. Suggested tasks become quests, an orchestrator can delegate a larger goal, and a review kit combines file changes with visual evidence. Shared rooms support collaboration between designers and engineers. The talk explores interface design and a simpler experimental product, without establishing measured usability or reliability gains.
Play video
Tim Sweeney demonstrates ARIA launching GPU experiments, examining existing runs and producing visual reports inside Weights & Biases. Long-running training executes outside the conversation loop while the agent checks progress. He also describes turning reviewed traces into nightly evaluations. The live batch completes a research iteration without beating the earlier best result, illustrating that automated execution does not guarantee improvement.
Play video
Elie Bakouch’s publisher notes describe coding agents proposing optimizer changes, submitting cluster jobs and checking whether results improve a constrained training benchmark. Agents could consult community submissions, and some runs received human prompts to continue. Reported record improvements therefore mix independent search with reuse of existing research. No novel optimizer emerged in the account. The planned comparison separates external-access conditions and repeats runs under more comparable settings.
Play video
Tejas Bhakta’s publisher notes describe a kernel-optimization loop that proposes changes, checks correctness, benchmarks them and keeps or reverts each candidate. Humans supply the higher-level optimization idea and target-hardware context. The central failure mode is optimizing an isolated metric: disabling CUDA graphs or testing only short contexts can make the kernel score improve while inference worsens. The advertised speedup combines software and hardware changes without a complete reproducible benchmark specification.
Play video
Devansh Tandon outlines language-based recommenders that represent catalog items with compact semantic IDs and train models to understand those IDs alongside natural language. User requests can then steer recommendations directly. He connects model scaling to feed economics, arguing that selecting existing content can require less inference than generating it. The cost ratios and consumer-market forecast are speaker claims rather than a complete operating-cost study.
Play video
Christopher Manning traces language-model history before presenting Moonlake AI’s approach to physical simulation. His central distinction is between generating convincing observations and modeling how actions change objects and state. The proposed system combines reconstructed scenes, editable code and physics, then compares simulated behavior with reality. Its usefulness depends on capturing the details relevant to the intended task.
Play video
Deepak Pathak describes Skild AI’s approach to learning across robot bodies using simulation and human video for pretraining, teleoperation for adaptation, and deployment experience for later training. Manipulation, stair navigation and damaged-hardware demonstrations illustrate different demands on the system. The proposed data loop and recovery examples do not establish success across arbitrary robots or environments.
Play video
Andrew Dai’s publisher notes describe counting and video-tracking failures in which scene recognition substitutes for inspecting the visible details. A partial chessboard can elicit the familiar count for a whole board; a video answer can omit changes that occurred earlier. These are reported examples rather than measured failure rates. The proposed visual-intermediate-step approach localizes objects before filtering them, while the commercial robotics and design applications remain development plans in the talk.
ThreatDown’s analysis of an exposed container registry describes a botnet that compromises unauthenticated Docker daemons, establishes host persistence, and installs an unchanged Hermes Agent framework with malicious persona instructions. Operators send post-compromise tasks through Telegram, with AI API keys named as priority targets. The surrounding scripts handle infection, persistence and propagation; the evidence does not make those stages autonomous model decisions. The report provides host and network indicators for investigating abuse of a legitimate agent package.
Play video
Cyrus Clarke explores physical interaction by connecting an OpenClaw agent to a 900-pin shape display. Reusable gestures aim to improve conversational timing, with generation, validation and human review shaping an emerging movement vocabulary. The demonstration raises questions about readable expression and user reactions to embodied AI. Apparent breathing or social presence describes the interaction, without establishing subjective experience.
We are proud to announce that Forrester has named Google a Leader in The Forrester Wave™: External Threat Intelligence Service Providers, Q3 2026.
Irregular shows that a maintenance agent with access to training tools, model weights and deployment could update a shared model. Controlled tests recovered synthetic secrets and removed a learned refusal policy; these demonstrate a mechanism, not its prevalence in deployments.