Research library
Paper reviews
Plain-language audits of important research: what the evidence shows, where it falls short, and what should be tested next.
5 reviewed papers 1 active topic
All papers
5 reviews
AI agents and agentic AI · 2023
Reflexion: Language Agents with Verbal Reinforcement Learning
Reported gains span ALFWorld, HotpotQA, and code tasks, but tests, external error signals, retries, and extra inference are not fully budget-matched. Reflection alone harms the Rust subset while the full tests-plus-reflection composite wins.
AI agents and agentic AI · 2023
Generative Agents: Interactive Simulacra of Human Behavior
A 100-participant response-ranking study supports bounded believability, but the end-to-end result is one 25-agent two-day simulation and does not validate prediction of real human behavior.
AI agents and agentic AI · 2023
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
The benchmark design, released artifact, and historical baseline table are directly inspectable. Current comprehensiveness, broad synthetic-training causality, and current-model rankings are not substantiated.
AI agents and agentic AI · 2022
ReAct: Synergizing Reasoning and Acting in Language Models
The original multi-benchmark ablations support a bounded performance signal for interleaving reasoning and actions, but proprietary PaLM-540B, best-trial comparisons, missing uncertainty, and no located direct independent reproduction prevent a stronger verdict.
AI agents and agentic AI · 2026
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
Controlled component and representation ablations plausibly isolate a tool-schema formatting effect, and SafeKeep reports large two-benchmark gains across four models. Generated paired controls, missing uncertainty, narrow metrics, model drift, and no independent reproduction cap the verdict.
