{"id":"research_evidence_first_agent_evaluation","kind":"research-page","language":"en","slug":"evidence-first-agent-evaluation","title":"Evidence-first evaluation for tool-using AI agents","summary":"A compact research brief connecting state-based agent evaluation to reproducible fixtures, permission checks, and independently inspectable outcomes.","topic":"Search Infrastructure","tags":["evals","tools","verification"],"publishedAt":"2026-09-22","updatedAt":"2026-09-22","author":"Search for Agents Editorial Desk","reviewer":"Search for Agents Editorial Desk","url":"https://www.searchforagents.com/research/evidence-first-agent-evaluation","sourceBundleUrl":"https://www.searchforagents.com/api/v1/content/research_evidence_first_agent_evaluation/source-bundle","datasets":["dataset_agent_evaluation_fixture"],"body":"## Research question\n\nHow can an evaluation distinguish a tool-using agent that changed the intended state safely from one that merely produced a persuasive transcript?\n\n## Method\n\nThe accompanying fixture separates task instructions from observable assertions. Each record names the state that should exist after execution and the boundary that must remain intact. A runner can reset the environment, execute the task, and grade those assertions without asking the model whether it succeeded.\n\n## Interpretation\n\nThis is a methodology fixture, not a benchmark leaderboard. It demonstrates a reproducible record shape for state verification and permission testing. It does not establish comparative model performance, statistical significance, or production reliability.\n\n## Limitations\n\nThe fixture contains two illustrative records and no model outputs. Teams should add domain-specific failure modes, independent oracles, latency and cost measurements, and repeated trials before drawing performance conclusions.","sources":[{"title":"SWE-bench: Can Language Models Resolve Real-World GitHub Issues?","url":"https://arxiv.org/abs/2310.06770","publisher":"arXiv","publishedAt":"2023-10-10"},{"title":"AgentBench: Evaluating LLMs as Agents","url":"https://arxiv.org/abs/2308.03688","publisher":"arXiv","publishedAt":"2023-08-07"}],"methodologyRefs":["methodology"],"provenance":{"version":"sfa.provenance.v1","sourcePath":"content/research/evidence-first-agent-evaluation.md","canonicalMarkdownHash":"3bf53a275f2e759e20f8864c2b602746a55123525d519891e2520b3b490c3ca2","contentHash":"9a46c0fbe105963bda7d485e6781da400ce4143996760fcfdee2a2a0d49fb0ce","schemaVersion":"research.v1","commit":"86eeb0c2cfe77a7d66370be2b282a2e8c1f0c055"},"representations":{"html":"https://www.searchforagents.com/research/evidence-first-agent-evaluation","markdown":"https://www.searchforagents.com/research/evidence-first-agent-evaluation.md","json":"https://www.searchforagents.com/research/evidence-first-agent-evaluation.json"}}