DoGBench is a benchmark designed to evaluate AI agents that write and maintain user-facing software documentation in response to real repository events and reported documentation gaps. Each patch created by the agents is scored using rubrics validated by project maintainers, measuring progress toward expert standards rather than performance against human experts.
The benchmark includes 292 tasks sourced from open-source projects such as Helm, PostHog, Mautic, and Doc Detective. Among these tasks, 205 require documentation changes while 87 require no changes. The evaluation uses a random, stratified split of 117 items, consisting of 82 requiring updates and 35 requiring no change. The results include 735 complete agent trajectories from seven non-cloud agents.
The highest score achieved on the 117-item held-out split was 47.3 out of 100, recorded by the Qwen3.8 Max with OpenCode. The evaluation also includes three cloud-agent systems assessed on the same split.
DoGBench identifies that while AI agents can produce polished documentation, they often fail to recognize when updates are necessary or may document changes that do not require user guidance. The analysis found that agents sometimes fabricate content, with 6.1% of submissions containing invented classes, flags, or endpoints. Additionally, agents may distort real functionality or overlook necessary updates.
Documentation owners are encouraged to provide context that may not be apparent from code changes alone, such as the intended audience and their objectives. This context is crucial for determining the need for updates and the content they should include. The evaluation of agent submissions includes assessing accuracy, completeness, reader guidance, placement, style, and adherence to repository conventions.
DoGBench reports on both update decisions and patch quality separately, capturing the scoring rules that reflect these aspects.