Full Archive · Page 16

Research archive, page 16

Browse entries 361–384 of 1576. Return to the first page to search and filter the complete collection.

Training-data pipelines: profile fetch, preprocessing and transfer separately video thumbnail Play video
AI Engineer October 10, 2026 video

Training-data pipelines: profile fetch, preprocessing and transfer separately

Tarun Sunkaraneni traces a multimodal training pipeline from serial image fetches through concurrent preprocessing, prefetch queues and object references. Each change exposes another bottleneck: waiting for data, copying large arrays, then network pressure as workers scale. Prefetching complicates checkpoint recovery, and spreading workers can trade locality for network capacity. Ray’s documentation limits zero-copy NumPy reads to workers on the same node; references do not eliminate cross-node transfer. The talk’s utilization figures describe its experiment, not an expected gain for other workloads.

Turn recurring agent mistakes into checks that cannot pass by deleting evidence video thumbnail Play video
AI Engineer October 9, 2026 video

Turn recurring agent mistakes into checks that cannot pass by deleting evidence

Greg Pstrucha describes moving repeated coding corrections into types, linters, and tests that run outside the model. His Sentry examples connect API response types to generated schemas and validate code examples in agent instructions. The talk also exposes a failure mode: an agent can satisfy a check by removing an example or changing its formatting, leaving the intended requirement unmet. Sentry’s public endpoint-documentation workflow provides a concrete companion example of comparing declared response types with actual behavior. These checks improve the feedback available to a coding agent, but they do not replace judgment about design or correctness.

JFrog Security Research October 7, 2026 analysis

LMCache advisory: restrict the multiprocess ZeroMQ listener

JFrog’s October 7 advisory describes CVE-2026-105192 in LMCache’s multiprocess transport: an unauthenticated message can reach Python pickle deserialization before request-handler checks and execute code with the service’s privileges. Remote exposure depends on making the ZeroMQ listener reachable; it binds to localhost by default, and in-process use does not expose this transport. The advisory lists versions from 0.3.9 through 0.5.5 and 0.5.6 release-candidate/development builds as affected, with no fixed release identified at publication.

Lumen Black Lotus Labs October 7, 2026 analysis

PoeLLM: exposed AI services become mining and scanning infrastructure

Lumen’s Black Lotus Labs describes a campaign targeting exposed AI and development services, including LiteLLM, Ollama, Gotenberg and Gitea. Infected hosts run cryptocurrency miners and can become scanners or exploit servers. The malware derives command-and-control addresses from words in a GitHub-hosted poem, letting the operator redirect bots by changing that document. This is a parsing and lookup mechanism, not evidence that a live language model controls each bot. The report provides infrastructure indicators and sample analysis, while some proposed initial-access paths remain inferred.

Agent-facing interfaces: expose tool prerequisites and keep bulk data outside context video thumbnail Play video
AI Engineer October 4, 2026 video

Agent-facing interfaces: expose tool prerequisites and keep bulk data outside context

Sarah Simionescu describes how cross-application agents need more than an MCP connection: they must discover relevant operations, resolve prerequisites and move intermediate data between services. A debugging mock-up returns tools together with a plan; an analytics demonstration stores a PostHog cohort outside model context while using database schema and samples to construct a Metabase query. These examples illustrate interface design, not an independently established speed or accuracy advantage; the talk provides no reproducible comparative benchmark.

Choose prompts, retrieval and fine-tuning according to the failure you need to fix video thumbnail Play video
AI Engineer October 4, 2026 video

Choose prompts, retrieval and fine-tuning according to the failure you need to fix

Anant Srivastava separates stable behavioral instructions, changing factual knowledge and learned task behavior. His examples show how training on historical support tickets can preserve obsolete product facts even after a prompt update, while fine-tuning on runbooks leaves missing-document retrieval unresolved. He proposes code-aware chunking and permission metadata for retrieval, and fine-tuning only after human judgments converge on a stable task. This is an architectural diagnostic, not a measured comparison proving one storage choice always wins.

Training recovery: coordinate node replacement, job requeueing and checkpoints video thumbnail Play video
AI Engineer October 3, 2026 video

Training recovery: coordinate node replacement, job requeueing and checkpoints

Crusoe describes a Slurm-on-Kubernetes recovery path that marks a failed GPU node down, signals and requeues its jobs, replaces unhealthy capacity, and restarts training from an application checkpoint. The demonstration injects an XID 79 failure signal; it does not physically break a GPU. Recovery depends on replacement policy, spare nodes and checkpoint save/load code. Its reported timing belongs to that demonstration, and a shutdown grace period cannot guarantee a fresh checkpoint after hardware failure.

Local agent memory: separate persistence, retrieval and permission to share video thumbnail Play video
AI Engineer October 2, 2026 video

Local agent memory: separate persistence, retrieval and permission to share

Dylan Couzon demonstrates an offline object-memory application using local detection, embeddings and Qdrant Edge storage. Retrieved sightings include images and timing metadata, showing how a persistent index supplies continuity beyond the current context window. Portability across reasoners or devices assumes the same embedding model; changing that representation needs a separate migration plan. The demo’s latency and footprint are workload-specific, and proposed whole-day assistants and shared memories are extensions rather than validated outcomes.

Ground coding-agent dependency reviews in current upstream evidence video thumbnail Play video
AI Engineer October 2, 2026 video

Ground coding-agent dependency reviews in current upstream evidence

Jakub Hojsan uses an API migration to show why a plausible diff can be misread when a coding agent relies on stale model knowledge. A version-change rule triggers retrieval of upstream documentation or changelogs, while query-specific excerpts and retrieval traces make the evidence inspectable without loading entire pages. The method separates tool availability from actually invoking verification. The talk does not establish a fixed knowledge-age gap for every model or the vendor’s claimed search-cost advantage.

Measure AI development impact beyond usage and pull-request throughput video thumbnail Play video
AI Engineer September 30, 2026 video

Measure AI development impact beyond usage and pull-request throughput

Justin Reock separates utilization, impact and cost when evaluating AI-assisted development. His pipeline example shows how rapid code generation can increase batching when builds and reviews remain slow, while the suggested measurements combine PR size and cycle time with failures, review corrections and developer experience. The talk’s organizational observations are associations and self-reported outcomes, not a causal estimate of AI’s effect. Faster releases or higher token consumption alone do not demonstrate greater business value.

Cleafy September 28, 2026 news

RATHat: model-assisted targeting and UI recovery rely on an existing Android foothold

Cleafy’s analysis distinguishes two AI uses in RATHat: the operator panel estimates victim value from stolen SMS messages, while device-side Gemini calls help locate controls when fixed UI automation fails. The malware first requires an Accessibility grant and a successful wireless-debugging pairing path; an operator can then deploy a separate shell-level service. That service can survive app removal until reboot. The analyzed samples do not show an LLM performing fraudulent transfers, and some native capture tools fail on Android 14 and later.

GLM-5.2: Open Weights, Near-Frontier Intelligence — Zixuan Li, Z.ai video thumbnail Play video
AI Engineer September 27, 2026 video

GLM-5.2: Open Weights, Near-Frontier Intelligence — Zixuan Li, Z.ai

Zixuan Li introduces GLM-5.2 through its coding and agentic capabilities, adjustable thinking budget and open-weight deployment options. He separates the model from Z Code, a coding harness that also accepts other models. The talk explains the roles of local inference, domain fine-tuning and ecosystem tooling; its benchmark placements are Z.ai’s reported comparisons with incomplete evaluation conditions.

Improving agent skills and memory through reviewed changes and task evaluations video thumbnail Play video
AI Engineer September 27, 2026 video

Improving agent skills and memory through reviewed changes and task evaluations

Suraj Gupta’s publisher notes distinguish agents doing recurring work from agents proposing improvements to that work. Warp’s triage example turns human feedback into a skill-change pull request, while persistent memory retains investigation findings with editing and provenance controls. A separate evaluation loop compares models on recurring task classes. The demonstration does not quantify memory savings, and customer-facing routing evaluations were planned rather than available in the account.

Get Out of the Model's Way — Kevin Hou, Google Antigravity video thumbnail Play video
AI Engineer September 27, 2026 video

Get Out of the Model's Way — Kevin Hou, Google Antigravity

Kevin Hou describes Antigravity’s move toward model-led teams that create specialist subagents, listen for events and generate task-specific interfaces. Examples include an operating-system kernel demonstration and parallel investigations of evaluation differences. The reported costs and completion times belong to particular demonstrations; generated hypotheses and working applications still require independent review.

Agent improvement loops: version the whole configuration and test user outcomes video thumbnail Play video
AI Engineer September 26, 2026 video

Agent improvement loops: version the whole configuration and test user outcomes

Roland Gavrilescu’s publisher notes propose preserving an agent’s prompts, skills, evaluations, tools and environment choices as a versioned configuration. Failures become tests, repeated procedures become skills, and human feedback defines what counts as useful work. Candidate changes should then face production experiments before promotion. The examples explain an improvement process, but do not provide controlled outcome measurements or establish that a higher-level agent can replace human judgment.

Long-running agent memory: test provenance, uncertainty and privacy together video thumbnail Play video
AI Engineer September 26, 2026 video

Long-running agent memory: test provenance, uncertainty and privacy together

Erina Karati’s publisher notes use a simulated game village to expose memory failures: agents can retain a topic while losing its source, certainty or implications for later plans. The proposed evaluation records observations, memory writes, retrievals and belief changes across complete scenarios. It freezes the harness and evaluator while searching a limited policy space. A corrected rumor example is illustrative; the talk does not claim repeated evidence of general improvement.

AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile video thumbnail Play video
AI Engineer September 25, 2026 video

AI-Generated Code Is Already Competing With Human Code — Daksh Gupta, Greptile

Daksh Gupta compares likely agent-written and human-written pull requests using Greptile’s enterprise review data. Authorship is inferred from metadata, while reverts, flagged issues and review rounds serve as quality proxies. He reports broadly similar aggregate results with different failure patterns, then describes validating related code and running applications in sandboxes. The observational comparisons do not control all differences in task selection or difficulty.

Robot reliability: evaluate recovery and new-site transfer separately video thumbnail Play video
AI Engineer September 24, 2026 video

Robot reliability: evaluate recovery and new-site transfer separately

Jason Ma’s publisher notes describe using a video-based progress model to find robot failures, then collecting human demonstrations of recovery and fine-tuning the policy. Napkin folding exposes both bad grasps and quality failures that a nominal task-completion measure can miss. The reported 24-hour success rate belongs to a particular folding evaluation, with no sample size supplied in the notes. Transfer to a new site and recovery from unfamiliar mistakes require separate evidence.

Document ingestion: correct OCR without rewriting the source video thumbnail Play video
AI Engineer September 23, 2026 video

Document ingestion: correct OCR without rewriting the source

Adit Abraham’s publisher notes explain a document pipeline that combines layout detection, selective vision-model processing and targeted OCR corrections. The key failure mode is a model silently fixing the document itself, such as replacing a printed but incorrect total. Separate representations can support retrieval and structured reasoning, while iterative chart reconstruction can expose extraction mistakes. Reported benchmark improvements are incomplete comparisons; the notes provide no numerical tolerance establishing exact chart recovery.

From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI video thumbnail Play video
AI Engineer September 23, 2026 video

From VLM/VLA's to Embodied Agents — Armen Aghajanyan, Perceptron AI

Armen Aghajanyan describes Perceptron AI’s effort to combine perception, reasoning and robot control. Task-dependent token routing focuses computation, while zooming, tiling and revisiting video intervals help gather visual evidence. Joint video and control training aims to reduce dependence on teleoperation. The talk leaves parts of the training objective undisclosed and identifies temporal reliability and severe visual disruption as unresolved challenges.

VLM-generated labels: judge annotations and preserve task meaning during training video thumbnail Play video
AI Engineer September 23, 2026 video

VLM-generated labels: judge annotations and preserve task meaning during training

Merve Noyan’s publisher notes describe using a vision-language model to label images, smaller judges to inspect overlaid boxes, and a task-specific detector for deployment. The workflow exposes two evaluation traps: agreement with generated labels differs from agreement with human ground truth, and filtering can discard too much training data. Standard image augmentations can also change the correct answer. The signature-detection example is qualitative, and reported run costs lack a complete workload specification.

Video moderation pipelines: preserve brief events and audit the evaluator video thumbnail Play video
AI Engineer September 23, 2026 video

Video moderation pipelines: preserve brief events and audit the evaluator

Aditya Gautam’s publisher notes describe separating video perception, retrieval and review so short content changes remain tied to timestamps and source clips. Production traces help distinguish model errors from tool or retrieval failures before retraining. Small specialized models, caching and frame compression reduce work, but each shortcut can hide relevant content. The talk does not provide achieved accuracy or decision thresholds, and similarity alone does not establish a video’s original creator.