Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment
Hung Manh Pham, Mathias Funk, Mykola Pechenizkiy, Aaqib Saeedopen the source papervia the report id, not the extractionsource: pdfIR v1.0generated 2026-09-14 04:10:11 UTC
These are notes about our reading of the paper, not about the paper.
They affect how much weight the findings below can carry.
Note 1
compile-warningsemantic extraction: dropped 2 items whose quote was not found in the paper
Some checks could not run
RL002 reported nothing, and it could not reach OpenAlex — so its silence here is not evidence that this paper is clean.
RL003 reported nothing, and it could not reach OpenAlex — so its silence here is not evidence that this paper is clean.
Everything else on this page was computed from the paper itself and is
unaffected. This is a limit of this run, not a finding about the paper.
Some checks did not apply to this paper
RL005 did not apply to this paper — it links no repository we could read, so there was no released configuration to compare against. Its absence here says nothing about whether the paper and its code agree.
Compiler extraction log (1)
semantic extraction: dropped 2 items whose quote was not found in the paper
Warning
2Very likely a real problem. Needs the author's judgement to fix, but a reviewer would raise it.
RL004Missing ablationWarning
No ablation isolates the Semantic Anchor Generation via LLM Augmented Report Synthesis
at Methodconfidence0.75refs M2, K2
The paper introduces Semantic Anchor Generation via LLM Augmented Report Synthesis as one of its own contributions (it is referred to by a numbered contribution), but no reported experiment removes or replaces it. The ablation labels actually present are 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100', and none of them corresponds to Semantic Anchor Generation via LLM Augmented Report Synthesis, so its individual contribution is not isolated by the experiments as reported.
Evidence — 2 items
Quoted from this paperMethod
Semantic Anchor Generation via LLM Augmented Report Synthesis Contrastive alignment requires semantically rich, paired text for each audio sample.
Search that found nothingevery experiment's ablated components and ablation label
Searched all 12 extracted experiment(s) for an ablation of 'Semantic Anchor Generation via LLM Augmented Report Synthesis'. Ablation labels found: 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100'. No label matches 'Semantic Anchor Generation via LLM Augmented Report Synthesis' or its tokens (anchor, augment, generation, llm, report, semantic, synthesi).
ablations_found
17.0
What to do
Add a run with Semantic Anchor Generation via LLM Augmented Report Synthesis removed or replaced by a simpler alternative, so the claim that it contributes can be separated from the rest of the method.
RL004Missing ablationWarning
No ablation isolates the Projection Heads
at Methodconfidence0.75refs M7, K1
The paper introduces Projection Heads as one of its own contributions (it is referred to by a numbered contribution), but no reported experiment removes or replaces it. The ablation labels actually present are 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100', and none of them corresponds to Projection Heads, so its individual contribution is not isolated by the experiments as reported.
Evidence — 2 items
Quoted from this paperMethod
We introduce lightweight projection heads Ha and Ht, each consisting of a linear layer followed by layer normalization, mapping m and n dimensional features into the shared d dimensional space.
Search that found nothingevery experiment's ablated components and ablation label
Searched all 12 extracted experiment(s) for an ablation of 'Projection Heads'. Ablation labels found: 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100'. No label matches 'Projection Heads' or its tokens (head, projection).
ablations_found
17.0
What to do
Add a run with Projection Heads removed or replaced by a simpler alternative, so the claim that it contributes can be separated from the rest of the method.
Rule notes (33)
RL001 funnel arithmetic: sentences_with_from=4 endpoint_pairs=0 claims_compared=0 emitted=0 [dropped: not-a-measurement=0]
RL001 funnel delta-column: structured_tables=3 with_exactly_one_delta_column=0 rows_checked=0 emitted=0
RL001 funnel cross-location: quant=530 prose=160 prose_with_metric=45 prose_with_dataset=0 prose_with_metric_and_dataset=0 distinct_buckets=0 emitted=0 [dropped: not-a-measurement=2 not-in-quote=43]
RL001 funnel identifier-drift: settings_checked=2 sections_stating_one_value=0 emitted=0
RL002 citation entailment: attempted 18 of 33 ranked cited claims (81 in-text markers from 33 bibliography entries); 8 cited works could not be identified, 5 had no retrievable text, 0 model quotes failed verbatim verification, 0 were topically unrelated to the retrieved text.
RL002 candidate ranking (why the top 18 were chosen): [17] [introduction, absolutist wording, title/sentence overlap near zero, unusable title (identity fragile)]; [12] [introduction, absolutist wording, title/sentence overlap near zero, no DOI/arXiv id]; [13] [introduction, absolutist wording, no DOI/arXiv id]; [14] [introduction, absolutist wording, no DOI/arXiv id]; [11] [introduction, absolutist wording]; [8] [introduction, title/sentence overlap near zero, no DOI/arXiv id]; [17] [introduction, title/sentence overlap near zero, unusable title (identity fragile)]; [19] [introduction, title/sentence overlap near zero]
RL002 label distribution over 5 classified citations: INSUFFICIENT_EVIDENCE=2, SUPPORTED=3
RL002 base-rate self-check: not tripped (0/5 = 0% NOT_SUPPORTED, threshold 40%).
RL002: 1 citations skipped because the claim refers to the citing paper's own setting ('here', 'in our setup'), which the cited work cannot be expected to address
RL002: 2 citations skipped as background/common-knowledge statements, where a citation is an attribution rather than evidence for a proposition
RL002 attribution check (case B): 15 candidates name an artefact, 6 of them made the top-N and were checked against their cited source
RL002: 11 citations skipped because the sentence reports this paper's own results (the citation there is a comparison pointer, not evidence for the sentence)
RL002: 3 markers skipped because the claim is jointly attributed to several works
RL002: 3 enumerated-citation sentences skipped (they cite >=3 works, so the claim is about the set and cannot be refuted by any single cited work)
RL002: 8 cited works unresolvable (first 3: [{"key": "[17]", "title": "Medgemma technical report,", "why": "bibliography title unusable for identity verification"}, {"key": "[13]", "title": "Learning transferable visual models from natural language supervision,", "why": "no work record matched the bibliography title"}, {"key": "[17]", "title": "Medgemma technical report,", "why": "bibliography title unusable for identity verification"}])
RL003: 3 queries -> 51 unique works (arXiv search 2, category listings 0, scholarly graph 49 [crossref 49], 0 graph works matched to an arXiv id); source errors during retrieval: arXiv 0, graph 0
RL003: only 1 comparable papers found (need 4) — peer usage would be noise, no findings
RL004 funnel: components=12 contribution_backed=4 candidates=4 ablations_found=17 matched_by_ablation=2 unmatched=2 experiments=12
RL004 component M1 'REACH (REport-Augmented Contrastive alignment for respiratory Health)': covered (filtered earlier)
RL004 component M2 'Semantic Anchor Generation via LLM Augmented Report Synthesis': unmatched
RL004 component M3 'Similarity Aware Negative Sampling': covered (candidate)
RL004 component M4 'Structure-preserving alignment dual objective': covered (candidate)
RL004 component M5 'Masked Reconstruction Regularizer (LMSE)': covered (filtered earlier)
RL004 component M6 'Contrastive Alignment (Lcontrastive)': excluded:not-a-claimed-contribution
RL004 component M7 'Projection Heads': unmatched
RL004 component M8 'Audio Encoder': covered (filtered earlier)
RL004 component M9 'Text Encoder': covered (filtered earlier)
RL004 component M10 'Offline Indexing': excluded:not-a-claimed-contribution
RL004 component M11 'Online Negative Swapping': covered (filtered earlier)
RL004 component M12 'Training Strategy': excluded:not-a-claimed-contribution
RL004 rank M2 'Semantic Anchor Generation via LLM Augmented Report Synthesis': score=6 [referenced by a numbered contribution +3, presented as a contribution +2, substantive description (len=173) +1] confidence=0.75 -> emitted; desc='[kind=mechanism] A medical-grade off-the-shelf LLM (GPT-4) converts discrete metadata into…'
RL004 rank M7 'Projection Heads': score=4 [referenced by a numbered contribution +3, substantive description (len=148) +1] confidence=0.75 -> emitted; desc='[kind=mechanism] Lightweight projection heads Ha and Ht, each a linear layer followed by l…'
RL004 funnel: ranked=2 emitted=2 suppressed_by_cap=0 dropped_unverified_quote=0
Reading the report as data
This page is a rendering of the report the pipeline wrote. The JSON below
is the stored artefact, unedited — the same document
research-lint --format json produces.