We evaluated eight language models on a 99-item gap analysis of a DP operations manual and compared them against three human references: a pre-existing professional gap analysis of the manual, and two expert verdict sets recorded during a blinded review of the models' responses. The three references agreed with each other on 51.0%, 49.5%, and 36.7% of items. At this granularity, expert disagreement is of the same magnitude as model error, which limits what a model-accuracy benchmark on this task can establish.
Section Nine built MarCompBench as part of a bachelor thesis at HVL (Høgskulen på Vestlandet). It runs eight language models on one task: a gap analysis of an active DP vessel's operations manual, assessed against a 99-item rubric that operationalizes MTS TECHOP O-01. The eight models span frontier models reached through commercial APIs, open-weight models run on rented GPUs, and one maritime-domain fine-tune, a QLoRA adaptation of a 70-billion-parameter open-weight model trained on marine assurance material. Each model was run three times to measure run-to-run stability. I worked on DP vessels before building software and served as one of the three human assessors, which is noted where it bears on the results.
Method and the three human references
A gap analysis works like this. You take a guidance framework and a vessel's DP operations manual, and for each requirement you record a verdict on whether the document satisfies that point. The verdict space is small, Yes / Partial / No / not-applicable. Many items fall on Partial, where the manual addresses a requirement without fully carrying it, or where the content sits in another vessel document. The rubric has 99 items.
The three references were produced under different conditions, and the difference matters for how far each result generalizes. The Maritime Consultancy's assessment is a cold, independent gap analysis: a genuine professional deliverable produced in 2025 in the normal course of commercial DP assurance work, before MarCompBench or GapAnalyzer existed, not commissioned for this study. It is the one reference that could not have been shaped by the benchmark, a pre-existing and ecologically valid commercial baseline.
The other two references, referred to here as the Author (my own pass) and the DP Assurance Expert, were not independent blind gap analyses of the manual. Their verdicts were recorded during the grading of the models' outputs. For each of the 99 items, the eight model responses were presented blinded, labeled Model A through H with model identities hidden but response content fully visible. The expert selected the best response on that item and then either endorsed its verdict or entered their own. The two experts graded separately from each other, but both worked from the same set of model responses rather than from the manual alone.
The three references did not converge. Exact-match agreement, computed item by item:
The Author and the DP Assurance Expert agreed on 51.0% of items. The Author and the Maritime Consultancy agreed on 49.5%. The DP Assurance Expert and the Maritime Consultancy agreed on 36.7%. The three references reached the same verdict on roughly half the items in each pairing, and on one pairing on about a third.
The 51.0% figure is agreement between two experts who graded from the same set of model responses, not between two blind analyses of the manual, so it is the least independent of the three comparisons. The two benchmark-independent comparisons are the ones involving the cold reference: 49.5% between the Author and the Maritime Consultancy, and 36.7% between the DP Assurance Expert and the Maritime Consultancy. Both fall near or below half. Even against a reference that could not have anchored on the model outputs, expert agreement is of the same order as the disagreement between the models and the humans. Reference material built from expert judgment is, on this task, only about 40 to 50% self-consistent.
The study began by designating the Maritime Consultancy's assessment as the ground truth, on the assumption that a professional gap study produced in the field is the correct answer key. The triangulation numbers ruled that out: the consultancy matched the DP Assurance Expert on 36.7% of items and the Author on 49.5%, too far from either to serve as a canonical reference. We assumed a gold standard existed; measurement showed the three references disagree too much for any one of them to hold that role, so all three were demoted to peer status and treated as equally valid readings.
Exact-match percentages describe raw agreement but do not correct for the agreement expected by chance, and on a four-value scale some matches occur by chance alone. The standard reliability statistics for a coding instrument make that correction. Krippendorff's alpha across the three references is 0.21 treating the verdicts as nominal categories on the full four-value scale, and 0.29 on the ordinal scale that credits partial distance along Yes-Partial-No, computed with N/A items excluded because N/A is not a position on the coverage scale. Krippendorff's conventional thresholds require an alpha of at least 0.80 for data to be treated as reliable and at least 0.667 for even tentative conclusions. Both observed values fall far below both thresholds.
Pairwise chance-corrected agreement is the same story. Cohen's kappa is 0.28 for the Author against the DP Assurance Expert, 0.31 for the Author against the Maritime Consultancy, and 0.13 for the DP Assurance Expert against the Maritime Consultancy. Gwet's AC1, an alternative coefficient that stays stable under the skewed verdict distributions present here, is 0.38, 0.34, and 0.19 for the same three pairs, so the weakness is not an artifact of skewed marginals. Kappa corrects for the agreement expected by chance, which is why these values sit below the raw exact-match figures, and 0.13 is barely above chance.
Judged by the standard applied to any coding instrument in content analysis, the three-reference panel fails inter-rater reliability, and a benchmark scored against any single one of these references inherits that unreliability. This is the quantified answer to the validation critique of the rubric: the reliability of the reference panel is now measured and reported rather than assumed.
The item-level breakdown is sharper. Across the 98 items rated by all three references, all three agreed on 26. Two of three agreed on 56. On the remaining 16, all three gave different verdicts, no pair matching. So under a third of the items produced unanimous agreement, and one in six produced three-way disagreement.
Placing the eight models on the same item grid, one row per assessor and one column per item colored by verdict, every model row falls inside the range set by the two human extremes: the DP Assurance Expert (most lenient) and the Maritime Consultancy (strictest). The maritime fine-tune's verdicts track close to the most lenient expert; the Maritime Consultancy is the strictest assessor on the grid, human or model. The models occupy the same span as the humans rather than a separate error region.
The cross-reference interpretive divide
The disagreements are not random. The direction is almost uniform, and the interpretive question that best explains the pattern is one the guidance leaves to the assessor's judgment: when a DP operations manual that points to another controlled document, instead of stating something itself, has adequately covered the requirement.
A paraphrased example. A requirement asks the manual to address how the vessel handles a specific class of equipment failure during DP operations. The manual does not state this directly; it references the vessel's FMEA and safety management system, where the detail lives. Read one way, this is coverage: the information exists, it is controlled, the manual routes the reader to it, and duplicating it risks the two documents drifting out of sync. Read the other way, it is a gap: the operations manual is the document the DPO uses on the bridge, and if the vessel-specific answer is not in it, then for operational purposes it is not present.
The disagreement is one-directional, which points to a consistent difference of standard rather than to error. The Author and the DP Assurance Expert disagreed on 48 of the 98 items both rated, and on 39 of the 48 the DP Assurance Expert's verdict is the more lenient one: 27 items the Author marked Partial and he marked Yes, 6 the Author marked No and he marked Partial, and 6 the Author marked No and he marked Yes. Three disagreements lean the other way, and six involve N/A on one side. His grading notes state his side of the question explicitly, in one case recording that referenced manuals should be accepted as long as they are held in the vessel's document management system. Careless errors would scatter in both directions and partly cancel; a consistent standard difference produces a systematic offset, which is what the data shows.
Because the expert verdicts were elicited alongside the model responses, there is an anchoring risk: an expert who defaults to endorsing a model's verdict would be pulled toward the models, and since all eight models were graded from the same responses, toward the other expert as well. The override data runs against that. The Author's final verdict differed from the modal verdict of the eight model responses on 32 of 99 items, stricter on 19 and more lenient on 10. The DP Assurance Expert's differed on 47 of 98 items, more lenient on 40 and stricter on 2, with 5 departures involving N/A. Both experts departed from the models' consensus, and they departed in opposite dominant directions: the Author toward stricter, the DP Assurance Expert toward more lenient. Shared anchoring would have pulled them the same way, toward the models and toward each other. The observed pattern is the opposite, which is evidence that the disagreement reflects the two graders applying different judgment rather than both deferring to the model outputs.
Both readings are defensible, and they rest on different arguments. The permissive reading has a document-control argument: a manual that duplicates the FMEA and the SMS verbatim goes stale the moment the source is revised, leaving two controlled documents that disagree about the same failure mode. The strict reading has an operational argument: the manual is the artifact open on the bridge during an operation, and class generally expects the vessel-specific answer to be reachable in it without a document hunt. Neither misreads the guidance. TECHOP O-01 permits coverage by reference for some items and expects in-manual detail for others, but its operative test, whether a topic is adequately covered, is never defined, so no general rule says when a reference suffices.
This has a direct consequence for writing and auditing DP operations manuals. For authors, treating "the FMEA covers this" as sufficient is the largest single source of assessor disagreement; carrying a short vessel-specific statement in the manual for each testable requirement, with the cross-reference as support, removes most of it, at the cost of some duplication and the maintenance discipline to keep it current. For auditors, the standard should be stated up front: a gap analysis that does not declare whether it treats controlled cross-references as coverage is not reproducible, and the vessel owner cannot tell whether a set of findings reflects the manual or the assessor.
Given this, the agreement figures of 51%, 49.5%, and 36.7% are not measurement noise on top of a true value. For a large share of items there is no single agreed answer; there are two coherent standards, and the choice of standard determines the result.
The explorer below runs on the same data. It opens on the verdict map; from there you can switch the reference and watch the model ranking and the leniency spectrum change.