Research library
Paper reviews
Plain-language audits of important research: what the evidence shows, where it falls short, and what should be tested next.
2 reviewed papers 1 active topic
All papers
2 reviews
AI agents and agentic AI · 2022
ReAct: Synergizing Reasoning and Acting in Language Models
The original multi-benchmark ablations support a bounded performance signal for interleaving reasoning and actions, but proprietary PaLM-540B, best-trial comparisons, missing uncertainty, and no located direct independent reproduction prevent a stronger verdict.
AI agents and agentic AI · 2026
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
Controlled component and representation ablations plausibly isolate a tool-schema formatting effect, and SafeKeep reports large two-benchmark gains across four models. Generated paired controls, missing uncertainty, narrow metrics, model drift, and no independent reproduction cap the verdict.