Research · MarCompBench

Expert disagreement and model error in automated DP compliance assessment

Roy Petter Dyrdahl Torgersen · Section Nine AS · July 2026 · 13 min read

We evaluated eight language models on a 99-item gap analysis of a DP operations manual and compared them against three human references: a pre-existing professional gap analysis of the manual, and two expert verdict sets recorded during a blinded review of the models' responses. The three references agreed with each other on 51.0%, 49.5%, and 36.7% of items. At this granularity, expert disagreement is of the same magnitude as model error, which limits what a model-accuracy benchmark on this task can establish.

Section Nine built MarCompBench as part of a bachelor thesis at HVL (Høgskulen på Vestlandet). It runs eight language models on one task: a gap analysis of an active DP vessel's operations manual, assessed against a 99-item rubric that operationalizes MTS TECHOP O-01. The eight models span frontier models reached through commercial APIs, open-weight models run on rented GPUs, and one maritime-domain fine-tune, a QLoRA adaptation of a 70-billion-parameter open-weight model trained on marine assurance material. Each model was run three times to measure run-to-run stability. I worked on DP vessels before building software and served as one of the three human assessors, which is noted where it bears on the results.

Method and the three human references

A gap analysis works like this. You take a guidance framework and a vessel's DP operations manual, and for each requirement you record a verdict on whether the document satisfies that point. The verdict space is small, Yes / Partial / No / not-applicable. Many items fall on Partial, where the manual addresses a requirement without fully carrying it, or where the content sits in another vessel document. The rubric has 99 items.

The three references were produced under different conditions, and the difference matters for how far each result generalizes. The Maritime Consultancy's assessment is a cold, independent gap analysis: a genuine professional deliverable produced in 2025 in the normal course of commercial DP assurance work, before MarCompBench or GapAnalyzer existed, not commissioned for this study. It is the one reference that could not have been shaped by the benchmark, a pre-existing and ecologically valid commercial baseline.

The other two references, referred to here as the Author (my own pass) and the DP Assurance Expert, were not independent blind gap analyses of the manual. Their verdicts were recorded during the grading of the models' outputs. For each of the 99 items, the eight model responses were presented blinded, labeled Model A through H with model identities hidden but response content fully visible. The expert selected the best response on that item and then either endorsed its verdict or entered their own. The two experts graded separately from each other, but both worked from the same set of model responses rather than from the manual alone.

The three references did not converge. Exact-match agreement, computed item by item:

The Author and the DP Assurance Expert agreed on 51.0% of items. The Author and the Maritime Consultancy agreed on 49.5%. The DP Assurance Expert and the Maritime Consultancy agreed on 36.7%. The three references reached the same verdict on roughly half the items in each pairing, and on one pairing on about a third.

The 51.0% figure is agreement between two experts who graded from the same set of model responses, not between two blind analyses of the manual, so it is the least independent of the three comparisons. The two benchmark-independent comparisons are the ones involving the cold reference: 49.5% between the Author and the Maritime Consultancy, and 36.7% between the DP Assurance Expert and the Maritime Consultancy. Both fall near or below half. Even against a reference that could not have anchored on the model outputs, expert agreement is of the same order as the disagreement between the models and the humans. Reference material built from expert judgment is, on this task, only about 40 to 50% self-consistent.

The study began by designating the Maritime Consultancy's assessment as the ground truth, on the assumption that a professional gap study produced in the field is the correct answer key. The triangulation numbers ruled that out: the consultancy matched the DP Assurance Expert on 36.7% of items and the Author on 49.5%, too far from either to serve as a canonical reference. We assumed a gold standard existed; measurement showed the three references disagree too much for any one of them to hold that role, so all three were demoted to peer status and treated as equally valid readings.

Exact-match percentages describe raw agreement but do not correct for the agreement expected by chance, and on a four-value scale some matches occur by chance alone. The standard reliability statistics for a coding instrument make that correction. Krippendorff's alpha across the three references is 0.21 treating the verdicts as nominal categories on the full four-value scale, and 0.29 on the ordinal scale that credits partial distance along Yes-Partial-No, computed with N/A items excluded because N/A is not a position on the coverage scale. Krippendorff's conventional thresholds require an alpha of at least 0.80 for data to be treated as reliable and at least 0.667 for even tentative conclusions. Both observed values fall far below both thresholds.

Pairwise chance-corrected agreement is the same story. Cohen's kappa is 0.28 for the Author against the DP Assurance Expert, 0.31 for the Author against the Maritime Consultancy, and 0.13 for the DP Assurance Expert against the Maritime Consultancy. Gwet's AC1, an alternative coefficient that stays stable under the skewed verdict distributions present here, is 0.38, 0.34, and 0.19 for the same three pairs, so the weakness is not an artifact of skewed marginals. Kappa corrects for the agreement expected by chance, which is why these values sit below the raw exact-match figures, and 0.13 is barely above chance.

Judged by the standard applied to any coding instrument in content analysis, the three-reference panel fails inter-rater reliability, and a benchmark scored against any single one of these references inherits that unreliability. This is the quantified answer to the validation critique of the rubric: the reliability of the reference panel is now measured and reported rather than assumed.

The item-level breakdown is sharper. Across the 98 items rated by all three references, all three agreed on 26. Two of three agreed on 56. On the remaining 16, all three gave different verdicts, no pair matching. So under a third of the items produced unanimous agreement, and one in six produced three-way disagreement.

Placing the eight models on the same item grid, one row per assessor and one column per item colored by verdict, every model row falls inside the range set by the two human extremes: the DP Assurance Expert (most lenient) and the Maritime Consultancy (strictest). The maritime fine-tune's verdicts track close to the most lenient expert; the Maritime Consultancy is the strictest assessor on the grid, human or model. The models occupy the same span as the humans rather than a separate error region.

The cross-reference interpretive divide

The disagreements are not random. The direction is almost uniform, and the interpretive question that best explains the pattern is one the guidance leaves to the assessor's judgment: when a DP operations manual that points to another controlled document, instead of stating something itself, has adequately covered the requirement.

A paraphrased example. A requirement asks the manual to address how the vessel handles a specific class of equipment failure during DP operations. The manual does not state this directly; it references the vessel's FMEA and safety management system, where the detail lives. Read one way, this is coverage: the information exists, it is controlled, the manual routes the reader to it, and duplicating it risks the two documents drifting out of sync. Read the other way, it is a gap: the operations manual is the document the DPO uses on the bridge, and if the vessel-specific answer is not in it, then for operational purposes it is not present.

The disagreement is one-directional, which points to a consistent difference of standard rather than to error. The Author and the DP Assurance Expert disagreed on 48 of the 98 items both rated, and on 39 of the 48 the DP Assurance Expert's verdict is the more lenient one: 27 items the Author marked Partial and he marked Yes, 6 the Author marked No and he marked Partial, and 6 the Author marked No and he marked Yes. Three disagreements lean the other way, and six involve N/A on one side. His grading notes state his side of the question explicitly, in one case recording that referenced manuals should be accepted as long as they are held in the vessel's document management system. Careless errors would scatter in both directions and partly cancel; a consistent standard difference produces a systematic offset, which is what the data shows.

Because the expert verdicts were elicited alongside the model responses, there is an anchoring risk: an expert who defaults to endorsing a model's verdict would be pulled toward the models, and since all eight models were graded from the same responses, toward the other expert as well. The override data runs against that. The Author's final verdict differed from the modal verdict of the eight model responses on 32 of 99 items, stricter on 19 and more lenient on 10. The DP Assurance Expert's differed on 47 of 98 items, more lenient on 40 and stricter on 2, with 5 departures involving N/A. Both experts departed from the models' consensus, and they departed in opposite dominant directions: the Author toward stricter, the DP Assurance Expert toward more lenient. Shared anchoring would have pulled them the same way, toward the models and toward each other. The observed pattern is the opposite, which is evidence that the disagreement reflects the two graders applying different judgment rather than both deferring to the model outputs.

Both readings are defensible, and they rest on different arguments. The permissive reading has a document-control argument: a manual that duplicates the FMEA and the SMS verbatim goes stale the moment the source is revised, leaving two controlled documents that disagree about the same failure mode. The strict reading has an operational argument: the manual is the artifact open on the bridge during an operation, and class generally expects the vessel-specific answer to be reachable in it without a document hunt. Neither misreads the guidance. TECHOP O-01 permits coverage by reference for some items and expects in-manual detail for others, but its operative test, whether a topic is adequately covered, is never defined, so no general rule says when a reference suffices.

This has a direct consequence for writing and auditing DP operations manuals. For authors, treating "the FMEA covers this" as sufficient is the largest single source of assessor disagreement; carrying a short vessel-specific statement in the manual for each testable requirement, with the cross-reference as support, removes most of it, at the cost of some duplication and the maintenance discipline to keep it current. For auditors, the standard should be stated up front: a gap analysis that does not declare whether it treats controlled cross-references as coverage is not reproducible, and the vessel owner cannot tell whether a set of findings reflects the manual or the assessor.

Given this, the agreement figures of 51%, 49.5%, and 36.7% are not measurement noise on top of a true value. For a large share of items there is no single agreed answer; there are two coherent standards, and the choice of standard determines the result.

The explorer below runs on the same data. It opens on the verdict map; from there you can switch the reference and watch the model ranking and the leniency spectrum change.

There is no single ground truth

Two DP experts and an independent maritime consultancy assessed the same DP operations manual against the same 99-item checklist. Exact-match agreement between them:
Author = the thesis author (DP practitioner). DP Assurance Expert = an independent senior DP specialist. Maritime Consultancy = an independent consultancy's professional gap analysis of the same manual, produced in 2025 before this benchmark existed. The Author and DP Assurance Expert verdicts were recorded while grading the models' blinded responses; the consultancy assessment is the one fully independent reference.

Where the 99 items land

Every verdict by every assessor: the three human references, then the eight models (one benchmark run each), one column per checklist item. Among the references, of the 98 items assessed by all three: 26 unanimous, 56 with a two-way majority, 16 where all three disagree.
Hover or tap a column for the item, its verdicts, and the three references. Item 74 was not assessed by the DP assurance expert.

Pick your truth

Choose which assessment counts as the reference, and watch the eight models re-rank. Exact-match agreement with the chosen reference, mean of three runs per model.
general-purpose model maritime fine-tune

The leniency spectrum

Verdict distribution of each human reference across the same 99 items. One assessor recorded 44 No verdicts; another recorded 5.
View the data as a table
MarCompBench v1 (HVL, 2026). Exact-match verdict agreement on a 99-item rubric operationalizing MTS TECHOP O-01; 3 runs per model. Findings paper and dataset release in progress.

Start with the verdict map at the top; hover any column to read what the eleven assessors made of that item. Then switch the reference and watch the model ordering and the leniency spectrum move with it.

The leniency spectrum across references

The three references differ most visibly in their verdict distributions. The Maritime Consultancy returned 29 Yes, 20 Partial, 44 No, and 6 not-applicable. The Author returned 32 Yes, 42 Partial, 19 No, and 6 not-applicable. The DP Assurance Expert returned 67 Yes, 20 Partial, 5 No, and 6 not-applicable.

The No counts are the clearest measure. On the same manual and the same checklist, one assessor recorded 44 failures and another recorded 5, a factor of about nine between the strictest and most permissive reading. The number of findings a vessel owner receives depends largely on which assessor produced the report: one yields a remediation program, the other a near-clean result. Both are internally consistent.

The Partial counts show the same effect from the middle. The Author's distribution is Partial-heavy, 42 of 99, consistent with repeatedly landing on "addressed but not fully, or covered in a referenced document." The Maritime Consultancy converts much of that middle to No, and the DP Assurance Expert converts much of it to Yes, matching the strict and permissive readings of the cross-reference question. This is the leniency spectrum: Maritime Consultancy at the strict end, DP Assurance Expert at the permissive end, Author in between.

Reference dependence of accuracy figures

Because the references are only 40 to 50% self-consistent, any single accuracy figure for a model is a measurement against one chosen reference, not against truth. Model rankings invert with the choice of reference. This is the "pick your truth" view in the explorer: with the Author as reference the models sort into one order; with the Maritime Consultancy as reference the order changes, because proximity to the lenient standard and proximity to the strict standard are near-opposite objectives. No model ranks highest against all three references. The switcher exposes all three references and marks none as canonical, because designating one would reintroduce the false certainty the benchmark is meant to expose.

So a vendor claim of "90-something percent accurate at maritime compliance assessment" is incomplete without two further facts: against which reference, and how much do that vendor's reference assessors agree with each other. Without the second, the first is unfalsifiable.

One result held across all references. Every model measured more lenient than the Maritime Consultancy reference; all eight leaned toward accepting coverage rather than requiring it, toward the permissive end of the cross-reference divide. Where the operative assurance standard is the strict one, closer to class practice, these models will under-call gaps out of the box relative to that standard. This is a calibration property to account for, not a reason to reject the models.

Calibration effects in a maritime fine-tune

The maritime fine-tune, the open-weight 70-billion-parameter model further trained on marine assurance material, is the model most would expect to score best: it has domain exposure and domain vocabulary. Against one reference it did.

Measured as the mean signed distance between a model's verdicts and a reference's across all 99 items, where negative means more lenient than the reference and zero means an exact match, the maritime fine-tune scored −0.04 against the DP Assurance Expert. Averaged over the checklist, its verdicts and the most lenient human's verdicts matched. Against the Maritime Consultancy, the same model scored −0.83, the widest leniency gap of any model against that reference, close to a full verdict category more permissive across items.

The model did not change between those two measurements; the reference did, and the two references are about three-quarters of a verdict apart on average. The fine-tune had not read the manual more accurately. It sat at a fixed point on the leniency spectrum that happened to match one expert and diverge from another.

The implication for domain fine-tuning in assurance work: pretraining on domain material moved where the model sits on the leniency spectrum without measurable improvement in how well it read any individual document. The training corpus carried a house style, a tendency to accept controlled cross-references as coverage, and the model absorbed that posture along with the vocabulary. A leniency bias is not competence. Fine-tuning on one organization's back catalogue of gap analyses produces a model that assesses like that organization, including its strictness and its blind spots, detectable only by testing against a reference that disagrees with it. Against a like-minded reference the bias reads as accuracy.

Requirements for defensible automation

If model ranking is not the useful question, the useful one is what would make an automated gap analysis defensible in front of a client, a class surveyor, or a court. Three requirements, from the data.

First, reproducibility. Each model was run three times. The commercial API models drifted across runs, returning different verdicts on some items, because they sample at nonzero temperature and because providers can change a model behind a fixed name. The open-weight models, run locally at temperature zero, were bit-identical across all three runs: the same 99 verdicts and the same generated text each time. This matters for a compliance deliverable because such a report can be challenged, and the first step of a challenge is a request to reproduce it. A tool that returns a different verdict on re-run, with no change to the input, undermines the whole report. Open-weight models can be pinned to a specific weight file, seed, and temperature and archived, so a run behind a report can be regenerated later; the frontier APIs currently cannot offer that guarantee.

Second, an item-level evidence trail. A single top-line score, "78% compliant," hides the interpretive choices that produced it and invites unearned trust. The defensible form is itemized: for each of the 99 requirements, a verdict and the specific document passage it rests on, quoted and locatable. When a model marks an item a gap because the vessel-specific detail is in the FMEA rather than the manual, that reasoning has to be visible so an assessor holding the permissive standard can override it quickly. The automation does the exhaustive first pass and shows its evidence; human time goes to the items where the standard is contested.

Third, disagreement should be preserved, not averaged away. Blending three references into one consensus and reporting a single accuracy figure would hide the main finding. The two-thirds of disagreement that traces to the cross-reference question is information about where the profession has not settled a standard. A tool should be able to report "gap under the strict standard, satisfied under the permissive one, here is the reference," which is the decision a competent assessor already has to make.

Limitations and the dataset release

Four limitations bound what this benchmark can claim.

Scope. This is a single-manual study: one active DP vessel, one operations manual, one 99-item rubric. The result is consistent with the interpretive structure of the task, but one document cannot show it generalizes. We did not test whether the disagreement holds for other document types, other rubrics, or other assessor pools. It is possible this manual relied unusually heavily on cross-references and that other documentation would produce tighter agreement. The finding is a hypothesis, not an established law.

Rubric. The rubric is unpublished, so no external party can yet reproduce the assessment or examine whether an item's wording favors one reading of the cross-reference question. It has also not been through a validation study of its own; it is a careful operationalization by domain-literate people, not an instrument shown to measure what it claims. The thesis examiner raised this point, and it is correct: a benchmark is only as sound as the rubric inside it, and this one has not been independently validated.

Elicitation context. Two of the three references are not blind to the model responses. The Author and the DP Assurance Expert recorded their verdicts while grading the blinded model outputs item by item, not by analyzing the manual cold, so their verdicts could in principle be anchored on the models. Only the Maritime Consultancy reference is fully independent of the benchmark. The mitigation is in the override data: the two experts departed from the models' modal verdict in opposite dominant directions, the Author stricter and the DP Assurance Expert more lenient, which is the pattern of independent judgment rather than shared anchoring. That is evidence against contamination, not proof of its absence, and the two benchmark-independent comparisons against the consultancy (49.5% and 36.7%) are what the central disagreement claim rests on.

Reference base. Three assessors and one study are not a representative sample of the profession. Their disagreement could reflect a real standard divide, as argued here, or the idiosyncrasies of three careers. Three references cannot fully separate those. The cross-reference divide is the most parsimonious explanation of the pattern, not a proven population-level fact.

The evidence base is public. The dataset is published on Zenodo (DOI 10.5281/zenodo.21357226): per-item verdict codes for the two grading references and all eight models across all three runs, the per-item override records, the requirement structure with its source-guidance citations, and a dependency-free script that recomputes every reliability figure reported here, the exact-match percentages, Krippendorff's alpha, the pairwise kappas, and the confidence intervals, directly from the verdict codes. A restricted companion record holds the consultancy's verdict codes, available for verification under confidentiality. The release supports verification rather than replication: the manual and the rubric's acceptance criteria are not distributed. The thesis is delivered and graded; the explorer above also runs standalone at https://gapanalyzer.io/marcompbench/; the findings paper is in preparation.

Implications for automated compliance assessment

On this task, the models are good enough at the first pass that their capability is not the binding constraint. The binding constraint is that the task is less objective than it is usually treated as: about half the expert disagreement comes from two defensible readings of what a manual must contain, and no amount of model accuracy resolves a question the profession has not settled. A model cannot decide whether a controlled cross-reference counts as coverage; only the industry can.

This is the reasoning behind how gapanalyzer.io reports its output. It does not produce a single compliance score. It grades each requirement separately, with a verdict and the cited document evidence behind it, so the assessor can see what was decided, on what basis, and override any item. A single score is easier to present and harder to defend; the item-level evidence trail is the defensible form, because it treats the assessor's judgment as the deciding factor and the model as what reaches it faster.

The benchmark set out to rank eight models and instead measured a property of the task: qualified assessments of the same manual against the same rubric agree on roughly half the items, including the cold professional gap analysis against an expert reviewer at 49.5%. Any claim about automating this work has to start from that figure. A single accuracy number, presented without a reference and without inter-assessor agreement, states a certainty the task does not support.

Roy Petter Dyrdahl Torgersen
Master Mariner ·

Founder of Section Nine AS, the company behind GapAnalyzer. The MarCompBench study was conducted at Høgskulen på Vestlandet (HVL) and hosted by Section Nine.

Cite this article

Torgersen, R. P. D. (2026, July 12). Expert disagreement and model error in automated DP compliance assessment. GapAnalyzer Research. https://gapanalyzer.io/research/marcompbench/

GapAnalyzer grades documents item by item with cited evidence rather than a single score,
because of exactly what this data shows. See how the automated DP gap analysis works.

Request access