Scope
Original research on autonomous or semi-autonomous AI agents. This run covered logged OpenAlex, Crossref, arXiv, Crossref canonical metadata, GitHub artifacts, and bounded OpenAlex citation traversals.
Coverage
- No patent or procurement search
- No product/framework inventory
- No ACL/OpenReview venue sweep beyond canonical records
- Newest-100 cap on two citation graphs
Outputs
7 records
AI agents and agentic AI radar — 2026-08-03
Reflection is a composite retry workflow, not a proven standalone capability. Audit memory-authority failures before building any memory safety product.
Reflexion: Language Agents with Verbal Reinforcement Learning
Reported gains span ALFWorld, HotpotQA, and code tasks, but tests, external error signals, retries, and extra inference are not fully budget-matched. Reflection alone harms the Rust subset while the full tests-plus-reflection composite wins.
Generative Agents: Interactive Simulacra of Human Behavior
A 100-participant response-ranking study supports bounded believability, but the end-to-end result is one 25-agent two-day simulation and does not validate prediction of real human behavior.
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
The benchmark design, released artifact, and historical baseline table are directly inspectable. Current comprehensiveness, broad synthetic-training causality, and current-model rankings are not substantiated.
Provenance-preserving long-horizon memory regression harness
Cross-stack tests that verify memory consolidation and retrieval cannot amplify low-authority observations into high-authority tool actions.
Generic verbal-reflection middleware
A generic package for reflection, retry, and episodic verbal memory loops.
API-Bank compatibility wrapper
A thin commercial wrapper around the released API-Bank tool-use evaluator.
