Research · MarCompBench

Expert disagreement and model error in automated DP compliance assessment

MarCompBench put eight language models through a genuine gap analysis: an active DP vessel's operations manual, assessed item by item against a 99-point rubric operationalizing MTS TECHOP O-01, three runs per model. As references we had three human assessments of the same manual: the author's own pass, an independent DP assurance expert, and a maritime consultancy's professional gap study.

The humans agreed with each other on roughly half the items. That single fact reframes everything about automated compliance assessment, and you can explore it below. The research was conducted at Høgskulen på Vestlandet (HVL), hosted by Section Nine AS; a findings paper and dataset release are in progress.

There is no single ground truth

Two DP experts and an independent maritime consultancy assessed the same DP operations manual against the same 99-item checklist. Exact-match agreement between them:
Author = the thesis author (DP practitioner). DP Assurance Expert = an independent senior DP specialist. Maritime Consultancy = an independent consultancy's professional gap analysis of the same manual, produced in 2025 before this benchmark existed. The Author and DP Assurance Expert verdicts were recorded while grading the models' blinded responses; the consultancy assessment is the one fully independent reference.

Where the 99 items land

Every verdict by every assessor: the three human references, then the eight models (one benchmark run each), one column per checklist item. Among the references, of the 98 items assessed by all three: 26 unanimous, 56 with a two-way majority, 16 where all three disagree.
Hover or tap a column for the item, its verdicts, and the three references. Item 74 was not assessed by the DP assurance expert.

Pick your truth

Choose which assessment counts as the reference, and watch the eight models re-rank. Exact-match agreement with the chosen reference, mean of three runs per model.
general-purpose model maritime fine-tune

The leniency spectrum

Verdict distribution of each human reference across the same 99 items. One assessor recorded 44 No verdicts; another recorded 5.
View the data as a table
MarCompBench v1 (HVL, 2026). Exact-match verdict agreement on a 99-item rubric operationalizing MTS TECHOP O-01; 3 runs per model. Findings paper and dataset release in progress.

Read the full findings article for the story behind this data.

GapAnalyzer grades documents item by item with cited evidence rather than a single score,
because of exactly what this data shows. See how the automated DP gap analysis works.

Request access
Cite this dataset

Torgersen, R. P. D. (2026). MarCompBench v1 — reference and model verdicts for DP operations manual gap analysis (Version 1.2.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.21357226