1 of 1 finding judged worth fixing.
Too few findings to say anything about this rule yet.
Evaluation
The metric this project is organised around is Precision@3: of the top 3 findings on a paper, how many would the paper's own authors agree are worth fixing? Three findings someone acts on beat thirty they skim. Recall is not measured, because a rule that stays silent is an acceptable outcome and a noisy rule is not.
Sample: 1 paper, 1 finding, judged by a human against the question "would the authors agree this needs fixing?". This is a small sample. Treat the rates below as indicative, not as estimates with meaningful confidence intervals.
A rule that is bad is listed as bad. The invalid findings are quoted with the reason they were rejected, because a false positive you can read is more useful than a rate you have to trust.
1 of 1 finding judged worth fixing.
Too few findings to say anything about this rule yet.
Label sources: labels.json
How the measurement is done