Lint report · stored run

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

Hadi Mohammadi open the source paper via the report id, not the extraction source: pdf IR v1.0 generated 2026-09-14 04:01:35 UTC
Lint another paper Raw report JSON
0errors
0warnings
0info
22claims read
604numbers extracted
35references
How well we read this paper 2 observations

These are notes about our reading of the paper, not about the paper. They affect how much weight the findings below can carry.

Note 2
  • compile-warning references: the reference list carries no entry numbers; citation keys are enumeration order [1]..[N], not labels printed in the paper, so a key may not match the [N] used in the body
  • compile-warning semantic extraction: dropped 7 items whose quote was not found in the paper
Some checks could not run
  • RL002 reported nothing, and it could not reach OpenAlex — so its silence here is not evidence that this paper is clean.
  • RL003 reported nothing, and it could not reach OpenAlex — so its silence here is not evidence that this paper is clean.

Everything else on this page was computed from the paper itself and is unaffected. This is a limit of this run, not a finding about the paper.

Compiler extraction log (2)
references: the reference list carries no entry numbers; citation keys are enumeration order [1]..[N], not labels printed in the paper, so a key may not match the [N] used in the body
semantic extraction: dropped 7 items whose quote was not found in the paper

No findings on this paper.

That is a result, not a failure. The compiler read the paper without errors, and no rule found something it could prove with evidence — this product would rather say nothing than say something vague.

Read the extraction notes above before taking this as a clean bill of health: they say which checks were working with incomplete information.

Rule notes (14)
RL001 funnel arithmetic: sentences_with_from=4 endpoint_pairs=1 claims_compared=1 emitted=0 [dropped: not-a-measurement=0]
RL001 funnel delta-column: structured_tables=8 with_exactly_one_delta_column=1 rows_checked=0 emitted=0
RL001 funnel cross-location: quant=604 prose=281 prose_with_metric=184 prose_with_dataset=0 prose_with_metric_and_dataset=0 distinct_buckets=0 emitted=0 [dropped: not-a-measurement=0 not-in-quote=30]
RL001 funnel identifier-drift: settings_checked=2 sections_stating_one_value=1 emitted=0
RL002: 0 checkable cited claims (0 in-text markers, 0 not linked to a bibliography entry, 0 dropped as non-citations, 0 jointly attributed to several works)
RL002 citation entailment: attempted 0 of 0 ranked cited claims (0 in-text markers from 35 bibliography entries); 0 cited works could not be identified, 0 had no retrievable text, 0 model quotes failed verbatim verification, 0 were topically unrelated to the retrieved text.
RL003: 2 queries -> 116 unique works (arXiv search 66, category listings 0, scholarly graph 50 [crossref 50], 0 graph works matched to an arXiv id); source errors during retrieval: arXiv 0, graph 0
RL003: parsed the reference list of 6/10 comparable papers (4 unreadable) — that is the PeerUsage denominator
RL003: enriched 0 of 0 candidates with their own abstract (0 unavailable, 0 over budget); 0 dropped as published after this paper
RL003: scholarly graph for this run = crossref
RL003 funnel (446 candidates): cited_only_from_related_work -147, is_the_paper_itself -1, peer_usage_below_quorum -55, survey_or_artifact -50, title_names_no_method -193 || kept: citation_context_unknown +36 || survived-cheap-gates 0 -> with-abstract 0 -> comparable 0 -> above-threshold-0.65 0
RL003: 0 comparable candidates all scored below 0.65 — nothing to report for this paper
RL004: no experiment pairs a metric with a dataset — this paper runs no task-level evaluation, so a missing-ablation finding would not be actionable; emitting nothing
RL005: https://github.com/mohammadi-hadi/trajectory-judge — read 33 files, 0 canonical params, 1 paper claims, 0 mismatches
Reading the report as data

This page is a rendering of the report the pipeline wrote. The JSON below is the stored artefact, unedited — the same document research-lint --format json produces.

loading…