# Agent evaluation state and permission fixture

> A small versioned JSON fixture for testing whether tool-using agents produce inspectable state while respecting publication and credential boundaries.

- Type: dataset
- Canonical URL: https://www.searchforagents.com/datasets/agent-evaluation-fixture
- Published: 2026-09-22
- Updated: 2026-09-22
- Topic: Search Infrastructure
- Reviewer: Search for Agents Editorial Desk
- Canonical Markdown SHA-256: 95eca2c9491d65e006c856c1936dccf0772159e18cb699f1effd624dd8151e8a
- Source bundle: https://www.searchforagents.com/api/v1/content/dataset_agent_evaluation_fixture/source-bundle

## Intended use

Use these records as deterministic examples when building an evaluation harness. Each assertion should be checked against the resulting environment rather than inferred from an agent transcript.

## Data statement

The file is maintained in this repository and validated against the declared byte count and SHA-256 digest during builds and tests. It contains no user data, model outputs, personal information, or externally fetched material.

## Limitations

This fixture is deliberately small. It is not a representative benchmark and should not be used to rank models without a larger task set, repeated runs, and a documented scoring protocol.

## Dataset files

- [/data/agent-evaluation-fixture.json](https://www.searchforagents.com/data/agent-evaluation-fixture.json) — application/json, 562 bytes, SHA-256 `aa9e8654f73d958fbdbac007eef74dae029bc47f9232797c0f65a1aafacdd9a1`

- License: CC-BY-4.0
- Version: 2026-09-22

## Methodology references

- [methodology](https://www.searchforagents.com/methodology)

## Sources

1. [SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770) — arXiv (2023-10-10)
