{"paper":{"id":"2609.00055.pdf","title":"Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment","authors":["Hung Manh Pham","Mathias Funk","Mykola Pechenizkiy","Aaqib Saeed"],"arxiv_id":"","url":"","source_kind":"pdf","ir_version":"1.0"},"generated_at":"2026-09-14T04:10:11.191195+00:00","counts":{"error":0,"warning":2,"info":0},"findings":[{"rule_id":"RL004","severity":"warning","confidence":0.75,"title":"No ablation isolates the Semantic Anchor Generation via LLM Augmented Report Synthesis","explanation":"The paper introduces Semantic Anchor Generation via LLM Augmented Report Synthesis as one of its own contributions (it is referred to by a numbered contribution), but no reported experiment removes or replaces it. The ablation labels actually present are 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100', and none of them corresponds to Semantic Anchor Generation via LLM Augmented Report Synthesis, so its individual contribution is not isolated by the experiments as reported.","paper_location":"Method","evidence":[{"kind":"paper_span","source":"Method","location":"","quote":"Semantic Anchor Generation via LLM Augmented Report Synthesis Contrastive alignment requires semantically rich, paired text for each audio sample.","url":"","values":{}},{"kind":"negative_search","source":"every experiment's ablated components and ablation label","location":"","quote":"Searched all 12 extracted experiment(s) for an ablation of 'Semantic Anchor Generation via LLM Augmented Report Synthesis'. Ablation labels found: 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100'. No label matches 'Semantic Anchor Generation via LLM Augmented Report Synthesis' or its tokens (anchor, augment, generation, llm, report, semantic, synthesi).","url":"","values":{"ablations_found":17.0}}],"suggested_action":"Add a run with Semantic Anchor Generation via LLM Augmented Report Synthesis removed or replaced by a simpler alternative, so the claim that it contributes can be separated from the rest of the method.","target_ids":["M2","K2"],"suppressed_count":0},{"rule_id":"RL004","severity":"warning","confidence":0.75,"title":"No ablation isolates the Projection Heads","explanation":"The paper introduces Projection Heads as one of its own contributions (it is referred to by a numbered contribution), but no reported experiment removes or replaces it. The ablation labels actually present are 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100', and none of them corresponds to Projection Heads, so its individual contribution is not isolated by the experiments as reported.","paper_location":"Method","evidence":[{"kind":"paper_span","source":"Method","location":"","quote":"We introduce lightweight projection heads Ha and Ht, each consisting of a linear layer followed by layer normalization, mapping m and n dimensional features into the shared d dimensional space.","url":"","values":{}},{"kind":"negative_search","source":"every experiment's ablated components and ablation label","location":"","quote":"Searched all 12 extracted experiment(s) for an ablation of 'Projection Heads'. Ablation labels found: 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100'. No label matches 'Projection Heads' or its tokens (head, projection).","url":"","values":{"ablations_found":17.0}}],"suggested_action":"Add a run with Projection Heads removed or replaced by a simpler alternative, so the claim that it contributes can be separated from the rest of the method.","target_ids":["M7","K1"],"suppressed_count":0}],"top_findings":[{"rule_id":"RL004","severity":"warning","confidence":0.75,"title":"No ablation isolates the Semantic Anchor Generation via LLM Augmented Report Synthesis","explanation":"The paper introduces Semantic Anchor Generation via LLM Augmented Report Synthesis as one of its own contributions (it is referred to by a numbered contribution), but no reported experiment removes or replaces it. The ablation labels actually present are 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100', and none of them corresponds to Semantic Anchor Generation via LLM Augmented Report Synthesis, so its individual contribution is not isolated by the experiments as reported.","paper_location":"Method","evidence":[{"kind":"paper_span","source":"Method","location":"","quote":"Semantic Anchor Generation via LLM Augmented Report Synthesis Contrastive alignment requires semantically rich, paired text for each audio sample.","url":"","values":{}},{"kind":"negative_search","source":"every experiment's ablated components and ablation label","location":"","quote":"Searched all 12 extracted experiment(s) for an ablation of 'Semantic Anchor Generation via LLM Augmented Report Synthesis'. Ablation labels found: 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100'. No label matches 'Semantic Anchor Generation via LLM Augmented Report Synthesis' or its tokens (anchor, augment, generation, llm, report, semantic, synthesi).","url":"","values":{"ablations_found":17.0}}],"suggested_action":"Add a run with Semantic Anchor Generation via LLM Augmented Report Synthesis removed or replaced by a simpler alternative, so the claim that it contributes can be separated from the rest of the method.","target_ids":["M2","K2"],"suppressed_count":0},{"rule_id":"RL004","severity":"warning","confidence":0.75,"title":"No ablation isolates the Projection Heads","explanation":"The paper introduces Projection Heads as one of its own contributions (it is referred to by a numbered contribution), but no reported experiment removes or replaces it. The ablation labels actually present are 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100', and none of them corresponds to Projection Heads, so its individual contribution is not isolated by the experiments as reported.","paper_location":"Method","evidence":[{"kind":"paper_span","source":"Method","location":"","quote":"We introduce lightweight projection heads Ha and Ht, each consisting of a linear layer followed by layer normalization, mapping m and n dimensional features into the shared d dimensional space.","url":"","values":{}},{"kind":"negative_search","source":"every experiment's ablated components and ablation label","location":"","quote":"Searched all 12 extracted experiment(s) for an ablation of 'Projection Heads'. Ablation labels found: 'similarity-aware negative sampling', 'FAISS-based distant negatives -> random in-batch negatives', 'Masked Reconstruction Regularizer (LMSE)', 'LMSE', 'audio encoder adaptation', 'Frozen Audio Backbone', 'MedSigLIP text encoder', 'MedSigLIP -> BERT', 'similarity-aware negative sampling (k=10)', 'negative sampling k=10 -> K=100', 'pre-trained audio encoder', 'pre-trained audio encoder -> audio encoder from scratch', 'Audio Encoder from Scratch', 'Random Negatives', 'No LMSE', 'BERT Text Encoder', 'Negative Sampling K=100'. No label matches 'Projection Heads' or its tokens (head, projection).","url":"","values":{"ablations_found":17.0}}],"suggested_action":"Add a run with Projection Heads removed or replaced by a simpler alternative, so the claim that it contributes can be separated from the rest of the method.","target_ids":["M7","K1"],"suppressed_count":0}],"ir_summary":{"claims":24,"contributions":4,"method_components":12,"experiments":12,"baselines":9,"datasets":7,"metrics":6,"tables":3,"figures":1,"citations":33,"equations":0,"quantitative_results":530,"has_repository":false},"ir_health":{"observations":[{"code":"compile-warning","severity":"note","message":"semantic extraction: dropped 2 items whose quote was not found in the paper","detail":""}],"errors":0,"warnings":0,"trustworthy":true},"diagnostics":{"source":"/home/oceann/workspace/research-linter/eval/papers/2609.00055.pdf","rules_run":["RL001","RL002","RL003","RL004"],"rules_skipped":[{"rule":"RL005","reason":"not-applicable"}],"rule_errors":{},"rule_findings":{"RL001":0,"RL002":0,"RL003":0,"RL004":2},"rule_seconds":{"RL001":0.01,"RL002":127.02,"RL003":0.63,"RL004":0.07},"dropped_no_evidence":0,"demoted_low_confidence":0,"ir_health":{"observations":[{"code":"compile-warning","severity":"note","message":"semantic extraction: dropped 2 items whose quote was not found in the paper","detail":""}],"errors":0,"warnings":0,"trustworthy":true},"notes":["RL001 funnel arithmetic: sentences_with_from=4 endpoint_pairs=0 claims_compared=0 emitted=0 [dropped: not-a-measurement=0]","RL001 funnel delta-column: structured_tables=3 with_exactly_one_delta_column=0 rows_checked=0 emitted=0","RL001 funnel cross-location: quant=530 prose=160 prose_with_metric=45 prose_with_dataset=0 prose_with_metric_and_dataset=0 distinct_buckets=0 emitted=0 [dropped: not-a-measurement=2 not-in-quote=43]","RL001 funnel identifier-drift: settings_checked=2 sections_stating_one_value=0 emitted=0","RL002 citation entailment: attempted 18 of 33 ranked cited claims (81 in-text markers from 33 bibliography entries); 8 cited works could not be identified, 5 had no retrievable text, 0 model quotes failed verbatim verification, 0 were topically unrelated to the retrieved text.","RL002 candidate ranking (why the top 18 were chosen): [17] [introduction, absolutist wording, title/sentence overlap near zero, unusable title (identity fragile)]; [12] [introduction, absolutist wording, title/sentence overlap near zero, no DOI/arXiv id]; [13] [introduction, absolutist wording, no DOI/arXiv id]; [14] [introduction, absolutist wording, no DOI/arXiv id]; [11] [introduction, absolutist wording]; [8] [introduction, title/sentence overlap near zero, no DOI/arXiv id]; [17] [introduction, title/sentence overlap near zero, unusable title (identity fragile)]; [19] [introduction, title/sentence overlap near zero]","RL002 label distribution over 5 classified citations: INSUFFICIENT_EVIDENCE=2, SUPPORTED=3","RL002 base-rate self-check: not tripped (0/5 = 0% NOT_SUPPORTED, threshold 40%).","RL002: 1 citations skipped because the claim refers to the citing paper's own setting ('here', 'in our setup'), which the cited work cannot be expected to address","RL002: 2 citations skipped as background/common-knowledge statements, where a citation is an attribution rather than evidence for a proposition","RL002 attribution check (case B): 15 candidates name an artefact, 6 of them made the top-N and were checked against their cited source","RL002: 11 citations skipped because the sentence reports this paper's own results (the citation there is a comparison pointer, not evidence for the sentence)","RL002: 3 markers skipped because the claim is jointly attributed to several works","RL002: 3 enumerated-citation sentences skipped (they cite >=3 works, so the claim is about the set and cannot be refuted by any single cited work)","RL002: 8 cited works unresolvable (first 3: [{\"key\": \"[17]\", \"title\": \"Medgemma technical report,\", \"why\": \"bibliography title unusable for identity verification\"}, {\"key\": \"[13]\", \"title\": \"Learning transferable visual models from natural language supervision,\", \"why\": \"no work record matched the bibliography title\"}, {\"key\": \"[17]\", \"title\": \"Medgemma technical report,\", \"why\": \"bibliography title unusable for identity verification\"}])","RL003: 3 queries -> 51 unique works (arXiv search 2, category listings 0, scholarly graph 49 [crossref 49], 0 graph works matched to an arXiv id); source errors during retrieval: arXiv 0, graph 0","RL003: only 1 comparable papers found (need 4) — peer usage would be noise, no findings","RL004 funnel: components=12 contribution_backed=4 candidates=4 ablations_found=17 matched_by_ablation=2 unmatched=2 experiments=12","RL004 component M1 'REACH (REport-Augmented Contrastive alignment for respiratory Health)': covered (filtered earlier)","RL004 component M2 'Semantic Anchor Generation via LLM Augmented Report Synthesis': unmatched","RL004 component M3 'Similarity Aware Negative Sampling': covered (candidate)","RL004 component M4 'Structure-preserving alignment dual objective': covered (candidate)","RL004 component M5 'Masked Reconstruction Regularizer (LMSE)': covered (filtered earlier)","RL004 component M6 'Contrastive Alignment (Lcontrastive)': excluded:not-a-claimed-contribution","RL004 component M7 'Projection Heads': unmatched","RL004 component M8 'Audio Encoder': covered (filtered earlier)","RL004 component M9 'Text Encoder': covered (filtered earlier)","RL004 component M10 'Offline Indexing': excluded:not-a-claimed-contribution","RL004 component M11 'Online Negative Swapping': covered (filtered earlier)","RL004 component M12 'Training Strategy': excluded:not-a-claimed-contribution","RL004 rank M2 'Semantic Anchor Generation via LLM Augmented Report Synthesis': score=6 [referenced by a numbered contribution +3, presented as a contribution +2, substantive description (len=173) +1] confidence=0.75 -> emitted; desc='[kind=mechanism] A medical-grade off-the-shelf LLM (GPT-4) converts discrete metadata into…'","RL004 rank M7 'Projection Heads': score=4 [referenced by a numbered contribution +3, substantive description (len=148) +1] confidence=0.75 -> emitted; desc='[kind=mechanism] Lightweight projection heads Ha and Ht, each a linear layer followed by l…'","RL004 funnel: ranked=2 emitted=2 suppressed_by_cap=0 dropped_unverified_quote=0"],"missing_core_rules":[],"seconds":132.3,"llm":{"calls":6,"cache_hits":2,"errors":0,"prompt_chars":45713},"sources":{"arxiv":{"requests":8,"errors":0,"cache_hits":26,"available":1},"openalex":{"requests":1,"errors":0,"cache_hits":0,"available":0,"unavailable_reason":"OpenAlex free daily budget exhausted (resets midnight UTC)"},"crossref":{"requests":9,"errors":0,"cache_hits":10,"available":1},"s2":{"requests":0,"errors":0,"cache_hits":0,"available":1},"short_circuited":["crossref","openalex","s2"]},"cache":{"hits":38,"misses":24}},"compile_warnings":["semantic extraction: dropped 2 items whose quote was not found in the paper"]}