AI agents and agentic AI · 2023

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

SUBSTANTIATED3 / 5Reviewed Aug 3, 2026

The benchmark design, released artifact, and historical baseline table are directly inspectable. Current comprehensiveness, broad synthetic-training causality, and current-model rankings are not substantiated.

At a glance

What holds up

A structured public benchmark exists and named historical model snapshots produced the reported point estimates under its evaluator.

Main limitation

Point estimates lack repeated-decoding intervals; Lynx lacks a clear matched base ablation in the main table; no exact frozen-snapshot rerun was located.

What to test next

Containerize and audit the evaluator, scan semantic train/test overlap, then run repeated current models using exact tool-state success.

Full review

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

  • Canonical ID: doi-10.18653-v1-2023.emnlp-main.187
  • Alias: arXiv:2304.08244
  • Canonical URL: https://doi.org/10.18653/v1/2023.emnlp-main.187
  • Full text audited: arXiv v2 HTML, updated 2023-10-25: https://ar5iv.labs.arxiv.org/html/2304.08244
  • Audit date: 2026-08-03
  • Access: full text
  • Verdict: SUBSTANTIATED (3/5) for the existence, design, and bounded baseline results of the released API-Bank benchmark. The claims that it was the “most comprehensive” benchmark and that synthetic training broadly improves tool use remain historical/single-source and not substantiated beyond the reported protocol.

Identity, version, and integrity status

Crossref identifies the EMNLP 2023 ACL Anthology article and reports no correction/retraction relation: https://api.crossref.org/works/10.18653%2Fv1%2F2023.emnlp-main.187. arXiv metadata records submission 2023-04-14, v2 update 2023-10-25, and “EMNLP 2023”: https://export.arxiv.org/api/query?id_list=2304.08244. No formal notice surfaced in the bounded checks. No integrity concern issued.

Load-bearing claims and trace

ClaimPopulation/system, comparator, outcomeDirect supportAudit judgment
API-Bank supplies a structured tool-use benchmarkTrain/evaluation collection spanning Call, Retrieve+Call, and Plan+Retrieve+CallFull text reports 1,008 domains, 2,211 APIs, 2,202 dialogues, and 6,135 turns; train is LLM-generated and evaluation is manually annotated: https://ar5iv.labs.arxiv.org/html/2304.08244#S5SUBSTANTIATED FOR ARTIFACT EXISTENCE. The repository exposes evaluator code, API implementations, databases, and sample JSONL files: https://api.github.com/repos/AlibabaResearch/DAMO-ConvAI/git/trees/main?recursive=1.
Five-agent generation produces usable data cheaplyChatGPT-based five-agent pipeline versus single-agent/self-instruct and reported manual annotation costThe tester discards 35%; four annotators rate 94/100 sampled generated items available, and 78% of discarded items invalid; reported cost is $0.10/dialogue and 98% below manual annotation: https://ar5iv.labs.arxiv.org/html/2304.08244#S4 and https://ar5iv.labs.arxiv.org/html/2304.08244#S5SUPPORTED AS REPORTED; SINGLE-SOURCE. The 94% estimate has n=100 with no interval/inter-rater statistic in the audited section. The 98% cost comparison depends on local labor and API prices and is not current.
Contemporary LLMs struggled increasingly with compositional tool useEvaluation set; Alpaca-7B, ChatGLM-6B, GPT-3 Davinci, GPT-3.5-turbo, GPT-4, and Lynx-7BTable 3 reports total correctness: GPT-4 60.24%, GPT-3.5 47.16%, Lynx 39.58%; GPT-4 Plan+Retrieve+Call is 70.00% versus 63.66% Call, so difficulty is not monotonic for every model: https://ar5iv.labs.arxiv.org/html/2304.08244#S5.T3SUBSTANTIATED AS A HISTORICAL BASELINE TABLE, NOT A GENERAL MODEL RANKING. No intervals/repeated seeds are reported, model snapshots are old/proprietary, and category sizes differ.
API-Bank synthetic training makes Lynx-7B comparable to GPT-3.5Alpaca/LLaMA-7B fine-tuned three epochs versus zero-shot GPT-3.5Total correctness is 39.58% versus 47.16%, while category metrics vary: https://ar5iv.labs.arxiv.org/html/2304.08244#S5.T3 and training details at https://ar5iv.labs.arxiv.org/html/2304.08244#S7MIXED. “Comparable” depends on metric/category; no base-LLaMA matched ablation appears in the main result table, so synthetic-data causality is incompletely isolated.
API-Bank was the most comprehensive benchmark availableComparison with contemporaneous workAuthor comparison and design argument: https://ar5iv.labs.arxiv.org/html/2304.08244#S5HISTORICAL AUTHOR CLAIM, NOT CURRENTLY VERIFIED. Novelty/completeness depends on search scope and changed immediately as tool benchmarks proliferated.

Design, statistics, leakage, and reproducibility

  • Evaluation design: manually annotated distribution-shift evaluation data differs from synthetic train data in domain, API scope, and dialogue content: https://ar5iv.labs.arxiv.org/html/2304.08244#S5.
  • Metrics: API-call correctness and ROUGE are reported by capability level. ROUGE may reward textual overlap without end-to-end task utility; “correctness” depends on the benchmark evaluator and simulated APIs.
  • Uncertainty: main tables provide point estimates only; no repeated decoding, confidence intervals, or model-version hashes are evident in the audited sections.
  • Data leakage: the train/evaluation construction is described as distribution-shifted, but the paper does not provide a semantic-near-duplicate or model-pretraining contamination audit. ChatGPT generated training data, creating model-origin dependence.
  • Provenance/licensing: the API-Bank subtree includes a LICENSE and executable evaluator/API files: https://api.github.com/repos/AlibabaResearch/DAMO-ConvAI/git/trees/main?recursive=1. Repository metadata reports MIT at the monorepo level: https://api.github.com/repos/AlibabaResearch/DAMO-ConvAI. The paper says all APIs are original implementations and annotators were paid US$15/hour: https://ar5iv.labs.arxiv.org/html/2304.08244#S10.
  • Declared limitations: English-only, only Lynx-7B fine-tuned in the reported public study, and an undisclosed online model result: https://ar5iv.labs.arxiv.org/html/2304.08244#S9.

Independent support and current status

OpenAlex reports 88 citing works for the canonical DOI as of 2026-08-03: https://api.openalex.org/works/https://doi.org/10.18653/v1/2023.emnlp-main.187. The newest 100-work citation traversal contains later tool-learning benchmarks and implementations, independently supporting that tool-use evaluation became an active problem, but no exact rerun of Table 3 on frozen model snapshots was located: https://api.openalex.org/works?filter=cites:W4389518608&sort=publication_date:desc&per-page=100. Exact baseline reproduction NOT FOUND IN BOUNDED SEARCH.

The public repository independently confirms that a runnable benchmark artifact—not merely a paper description—exists. This supports the artifact/design verdict but does not independently validate the reported model scores.

Verdict rationale

SUBSTANTIATED (3/5), narrowly scoped to benchmark construction, public artifact availability, and what the named historical model snapshots scored under the paper's evaluator. Stronger claims are limited by point-estimate reporting, old opaque model versions, weak causal isolation for Lynx's training gain, synthetic-data provenance, and no located exact independent rerun. “Most comprehensive” is superseded as a current claim and is not part of the positive verdict.

Decisive next verification

Version and containerize the released evaluator; audit train/test semantic overlap and evaluator edge cases; then run current open-weight and API models with frozen prompts, repeated decoding, exact tool-state success, and confidence intervals. Include direct-function-calling and modern MCP-style schemas. Kill the claim that API-Bank still discriminates current systems if ceiling effects exceed 90% or if evaluator-equivalent outputs receive materially different correctness labels in more than 2% of cases.

Source records 7
  1. API-Bank arXiv v2 full textprimary paper · accessed Aug 3, 2026
  2. API-Bank arXiv version metadataprimary metadata · accessed Aug 3, 2026
  3. API-Bank Crossref DOI recordpublisher metadata · accessed Aug 3, 2026
  4. DAMO-ConvAI repository metadataprimary repository · accessed Aug 3, 2026
  5. DAMO-ConvAI repository tree containing API-Bankprimary repository · accessed Aug 3, 2026
  6. API-Bank OpenAlex canonical recordscholarly index · accessed Aug 3, 2026
  7. API-Bank complete citing-work traversalscholarly index · accessed Aug 3, 2026
Paper Opportunity RadarArtificial intelligence · AI agents · Agent evaluation