Scope
Original research on autonomous or semi-autonomous AI agents. This run covered a 2026-07-26..2026-08-02 new delta, a first 2022-01..2023-03 archive slice, and bounded ReAct/SafeKeep follow-up graphs.
Coverage
- OpenReview API returned HTTP 403
- Semantic Scholar throttled two metadata calls
- ACL Anthology, patents, procurement, and buyer-demand searches are incomplete
- Only first result pages were traversed
Outputs
5 records
AI agents and agentic AI radar — 2026-08-02
Two full-text audits: ReAct and SafeKeep are promising but unreplicated. Schema-aware safety regression testing remains a verify-paper opportunity; generic ReAct orchestration is occupied.
ReAct: Synergizing Reasoning and Acting in Language Models
The original multi-benchmark ablations support a bounded performance signal for interleaving reasoning and actions, but proprietary PaLM-540B, best-trial comparisons, missing uncertainty, and no located direct independent reproduction prevent a stronger verdict.
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
Controlled component and representation ablations plausibly isolate a tool-schema formatting effect, and SafeKeep reports large two-benchmark gains across four models. Generated paired controls, missing uncertainty, narrow metrics, model drift, and no independent reproduction cap the verdict.
Schema-aware agent safety preflight and regression evidence
Cross-framework regression testing that compares schema-formatted execution with flattened-description safety judgment before tools run.
Generic ReAct orchestration product
Commercialize interleaved reasoning-and-acting orchestration as a generic agent architecture.