Full Archive · Page 19

Research archive, page 19

Browse entries 433–456 of 1576. Return to the first page to search and filter the complete collection.

METR March 19, 2025 analysis

METR’s time-horizon method ties agent success to human task duration

METR’s March 2025 research measures how agent success varies with the time human experts need to complete software and reasoning tasks. A fitted success curve yields a horizon at a chosen reliability level, while hierarchical resampling estimates uncertainty across task families, tasks and attempts. The original paper also extrapolated historical growth, but that forecast depends on the trend continuing and on transfer beyond the measured tasks. The source distinguishes its dated static results from a live chart. A 50% horizon describes a success threshold, not dependable execution of a project of that duration.

METR March 5, 2025 analysis

DeepSeek-R1 evaluation highlights tool-interface failures and elicitation limits

METR’s early-2025 DeepSeek-R1 assessment uses autonomous software tasks and research-engineering environments through a third-party model provider. Transcript inspection found hallucinated tool results, malformed calls and difficulty recovering from mistakes, showing how interface behavior can limit measured performance. The report also acknowledges modest elicitation, incomplete checks for strategic underperformance and potentially biased exclusion of server-error runs. Published example transcripts were selected to illustrate behavior, not estimate its prevalence. These results describe a historical model-and-scaffold combination rather than an upper bound on open-weight model capability.

METR February 12, 2025 analysis

DeepSeek-V3 assessment shows how scaffold and error handling shape evaluation

METR’s early-2025 DeepSeek-V3 report combines autonomous-task testing, research-engineering tasks and a held-out knowledge check. The authors observed repeated actions, hallucinated tool outputs and misread numerical feedback, while scaffold changes substantially improved performance. They explicitly caution that elicitation was limited, curated transcripts were not representative and excluded server-error runs could inflate scores. The study did not robustly test deliberate underperformance. Its historical capability comparison is therefore conditional on the provider, scaffold and scoring choices, with useful lessons for designing more auditable evaluations.

Evaluate speculative decoding against the workloads you actually serve video thumbnail Play video
AI Engineer October 6, 2026 video

Evaluate speculative decoding against the workloads you actually serve

This vLLM walkthrough explains the tradeoff behind draft-model speculative decoding: a smaller model proposes tokens and the target model verifies them, adding draft computation and memory in exchange for possible faster generation. Acceptance rate, prompt length, output length, sampling settings, and concurrency all affect the result. The talk includes a small live comparison, while vLLM documentation limits the technique’s expected benefit to suitable workloads and distinguishes distributional correctness from identical outputs across execution configurations. Its reported demo speedup is not a transferable benchmark, and spare GPU memory alone does not establish a performance benefit.

Browser-agent judges: check completion against visible evidence video thumbnail Play video
AI Engineer October 5, 2026 video

Browser-agent judges: check completion against visible evidence

Browserbase’s verifier design generates a task rubric and retrieves relevant screenshots for each criterion alongside action history and the final answer. It separates evidence of effort from evidence of completion: a sensible attempt to buy an unavailable product must not become a successful purchase. It also aims to avoid repeatedly penalizing downstream consequences of one earlier error. Human-labeled development and held-out trajectories support judge calibration, but the small evaluation does not establish universal reliability or guarantee the absence of false positives.

Synthetic-data pipelines: make metadata discovery, retries and scheduling measurable video thumbnail Play video
AI Engineer October 2, 2026 video

Synthetic-data pipelines: make metadata discovery, retries and scheduling measurable

Bogdan Gaza describes operational lessons from DatologyAI’s large synthetic-data pipeline: batch object-store metadata listing, size partitions to bound lost work, checkpoint partial outputs, and coordinate Ray CPU heads with GPU workers. Curated source documents seed rephrasing, while a shared workflow connects curation, generation, training and evaluation. The talk’s volume, throughput and model-quality figures are vendor-reported and lack a complete reproducible benchmark; the useful contribution is the breakdown of bottlenecks and recovery responsibilities.

MCP tool discovery: compose operations without loading every definition and result video thumbnail Play video
AI Engineer October 2, 2026 video

MCP tool discovery: compose operations without loading every definition and result

Jan Čurn demonstrates mcpc, a CLI client that exposes MCP sessions through searchable tool descriptions, structured JSON output and background tasks. Progressive discovery reduces the initial catalog in model context; shell composition lets intermediate results flow between operations without being restated by the model. Authentication and persistent sessions are client responsibilities. Early connector comparisons lack a complete numerical protocol, and moving work to a subagent or shell does not by itself protect secrets or establish isolation.

I Turned Coding Agents Into a Strategy Game — Ido Salomon, AgentCraft video thumbnail Play video
AI Engineer September 27, 2026 video

I Turned Coding Agents Into a Strategy Game — Ido Salomon, AgentCraft

Ido Salomon presents AgentCraft, which uses a strategy-game interface to make coding-agent activity easier to supervise. Suggested tasks become quests, an orchestrator can delegate a larger goal, and a review kit combines file changes with visual evidence. Shared rooms support collaboration between designers and engineers. The talk explores interface design and a simpler experimental product, without establishing measured usability or reliability gains.

An AI Research Agent That Runs Your Experiments — Tim Sweeney, Weights & Biases video thumbnail Play video
AI Engineer September 26, 2026 video

An AI Research Agent That Runs Your Experiments — Tim Sweeney, Weights & Biases

Tim Sweeney demonstrates ARIA launching GPU experiments, examining existing runs and producing visual reports inside Weights & Biases. Long-running training executes outside the conversation loop while the agent checks progress. He also describes turning reviewed traces into nightly evaluations. The live batch completes a research iteration without beating the earlier best result, illustrating that automated execution does not guarantee improvement.

Research-agent speedruns: control access to prior work before claiming discovery video thumbnail Play video
AI Engineer September 26, 2026 video

Research-agent speedruns: control access to prior work before claiming discovery

Elie Bakouch’s publisher notes describe coding agents proposing optimizer changes, submitting cluster jobs and checking whether results improve a constrained training benchmark. Agents could consult community submissions, and some runs received human prompts to continue. Reported record improvements therefore mix independent search with reuse of existing research. No novel optimizer emerged in the account. The planned comparison separates external-access conditions and repeats runs under more comparable settings.

Agent-written GPU kernels: a local benchmark win can slow the full model video thumbnail Play video
AI Engineer September 26, 2026 video

Agent-written GPU kernels: a local benchmark win can slow the full model

Tejas Bhakta’s publisher notes describe a kernel-optimization loop that proposes changes, checks correctness, benchmarks them and keeps or reverts each candidate. Humans supply the higher-level optimization idea and target-hardware context. The central failure mode is optimizing an isolated metric: disabling CUDA graphs or testing only short contexts can make the kernel score improve while inference worsens. The advertised speedup combines software and hardware changes without a complete reproducible benchmark specification.

Why LLM Recommenders Will Be AI's Biggest Consumer App — Devansh Tandon, Meta video thumbnail Play video
AI Engineer September 25, 2026 video

Why LLM Recommenders Will Be AI's Biggest Consumer App — Devansh Tandon, Meta

Devansh Tandon outlines language-based recommenders that represent catalog items with compact semantic IDs and train models to understand those IDs alongside natural language. User requests can then steer recommendations directly. He connects model scaling to feed economics, arguing that selecting existing content can require less inference than generating it. The cost ratios and consumer-market forecast are speaker claims rather than a complete operating-cost study.

World Models Need Causality, Not Pretty Pixels — Christopher Manning, Moonlake AI video thumbnail Play video
AI Engineer September 24, 2026 video

World Models Need Causality, Not Pretty Pixels — Christopher Manning, Moonlake AI

Christopher Manning traces language-model history before presenting Moonlake AI’s approach to physical simulation. His central distinction is between generating convincing observations and modeling how actions change objects and state. The proposed system combines reconstructed scenes, editable code and physics, then compares simulated behavior with reality. Its usefulness depends on capturing the details relevant to the intended task.

Robotics Has Been Stuck for 70 Years — Deepak Pathak, Skild AI video thumbnail Play video
AI Engineer September 24, 2026 video

Robotics Has Been Stuck for 70 Years — Deepak Pathak, Skild AI

Deepak Pathak describes Skild AI’s approach to learning across robot bodies using simulation and human video for pretraining, teleoperation for adaptation, and deployment experience for later training. Manipulation, stair navigation and damaged-hardware demonstrations illustrate different demands on the system. The proposed data loop and recovery examples do not establish success across arbitrary robots or environments.

Visual reasoning evaluations: require evidence from the actual image video thumbnail Play video
AI Engineer September 23, 2026 video

Visual reasoning evaluations: require evidence from the actual image

Andrew Dai’s publisher notes describe counting and video-tracking failures in which scene recognition substitutes for inspecting the visible details. A partial chessboard can elicit the familiar count for a whole board; a video answer can omit changes that occurred earlier. These are reported examples rather than measured failure rates. The proposed visual-intermediate-step approach localizes objects before filtering them, while the commercial robotics and design applications remain development plans in the talk.

ThreatDown September 22, 2026 news

CARBONATO: exposed Docker APIs enable an agent-assisted botnet

ThreatDown’s analysis of an exposed container registry describes a botnet that compromises unauthenticated Docker daemons, establishes host persistence, and installs an unchanged Hermes Agent framework with malicious persona instructions. Operators send post-compromise tasks through Telegram, with AI API keys named as priority targets. The surrounding scripts handle infection, persistence and propagation; the evidence does not make those stages autonomous model decisions. The report provides host and network indicators for investigating abuse of a legitimate agent package.

I Gave an AI a Body — Cyrus Clarke, MIT Media Lab video thumbnail Play video
AI Engineer September 22, 2026 video

I Gave an AI a Body — Cyrus Clarke, MIT Media Lab

Cyrus Clarke explores physical interaction by connecting an OpenClaw agent to a 900-pin shape display. Reusable gestures aim to improve conversational timing, with generation, validation and human review shaping an emerging movement vocabulary. The demonstration raises questions about readable expression and user reactions to embodied AI. Apparent breathing or social presence describes the interaction, without establishing subjective experience.