{"id":"dataset_agent_evaluation_fixture","kind":"dataset","language":"en","slug":"agent-evaluation-fixture","title":"Agent evaluation state and permission fixture","summary":"A small versioned JSON fixture for testing whether tool-using agents produce inspectable state while respecting publication and credential boundaries.","topic":"Search Infrastructure","tags":["evals","tools","verification"],"publishedAt":"2026-09-22","updatedAt":"2026-09-22","author":"Search for Agents Editorial Desk","reviewer":"Search for Agents Editorial Desk","url":"https://www.searchforagents.com/datasets/agent-evaluation-fixture","sourceBundleUrl":"https://www.searchforagents.com/api/v1/content/dataset_agent_evaluation_fixture/source-bundle","license":"CC-BY-4.0","version":"2026-09-22","files":[{"url":"/data/agent-evaluation-fixture.json","format":"application/json","sha256":"aa9e8654f73d958fbdbac007eef74dae029bc47f9232797c0f65a1aafacdd9a1","sizeBytes":562}],"body":"## Intended use\n\nUse these records as deterministic examples when building an evaluation harness. Each assertion should be checked against the resulting environment rather than inferred from an agent transcript.\n\n## Data statement\n\nThe file is maintained in this repository and validated against the declared byte count and SHA-256 digest during builds and tests. It contains no user data, model outputs, personal information, or externally fetched material.\n\n## Limitations\n\nThis fixture is deliberately small. It is not a representative benchmark and should not be used to rank models without a larger task set, repeated runs, and a documented scoring protocol.","sources":[{"title":"SWE-bench: Can Language Models Resolve Real-World GitHub Issues?","url":"https://arxiv.org/abs/2310.06770","publisher":"arXiv","publishedAt":"2023-10-10"}],"methodologyRefs":["methodology"],"provenance":{"version":"sfa.provenance.v1","sourcePath":"content/datasets/agent-evaluation-fixture.md","canonicalMarkdownHash":"95eca2c9491d65e006c856c1936dccf0772159e18cb699f1effd624dd8151e8a","contentHash":"e3546322bbaa70b9e64c29cc4c923ce12b5515266dacb65c9c79a21e439afdde","schemaVersion":"dataset.v1","commit":"86eeb0c2cfe77a7d66370be2b282a2e8c1f0c055"},"representations":{"html":"https://www.searchforagents.com/datasets/agent-evaluation-fixture","markdown":"https://www.searchforagents.com/datasets/agent-evaluation-fixture.md","json":"https://www.searchforagents.com/datasets/agent-evaluation-fixture.json"}}