HarnessEval-W
Agentifying the Evaluation of Visual Worlds
Evaluation defines the taste of evolution.
July 2026 dataset snapshot
Coverage at a glance
Why HarnessEval-W
More than a scalar score
HarnessEval-W brings the harness paradigm from the LLM ecosystem to world-model benchmarking: every verdict is backed by an inspectable reasoning chain.
Agentified evaluation
Flip for detailsRather than applying a fixed rubric, the harness interprets the context of each case, decomposes the question into measurable subproblems, and spawns specialized sub-agents with tailored context and diagnostic tools.
Transparent evidence trees
Flip for detailsEvery evaluation becomes an evidence tree that records exactly what was tested, which tool supplied the visual grounding, and the complete logical chain that justifies the final score — auditable end to end.
Human-aligned and robust
Flip for detailsJudgments closely track human A/B preferences on the hardest settings — intentional and physical transitions — and stay consistent under repeated evaluation runs.
A living benchmark
Flip for detailsOpen-sourced as an executable agentic system: the skill library is extensible, and new skills and evaluation cases grow as world models evolve.
HarnessEval-W Leaderboard
Evaluation Gallery
BibTeX
@article{mirros2026harnessevalw,
title = {HarnessEval-W: Agentifying the Evaluation of Visual Worlds},
author = {{MirroS Team}},
journal = {arXiv preprint arXiv:2608.16859},
year = {2026}
}