Download PDFOpen PDF in browserWhat Moves, and What Misleads, the BirdCLEF+ 2026 Leaderboard: An Empirical Study Behind a 13th Place Multi-Taxon Soundscape SystemEasyChair Preprint 1602413 pages•Date: July 31, 2026AbstractWe describe the 13th place solution to BirdCLEF+ 2026, a multi-taxon passive-acoustic classification task covering 234 bird, amphibian, insect, and mammal species in Pantanal soundscapes. Our system averages six convolutional models, three CNN classifiers and three SED networks on EfficientNetV2, HGNetV2, and ECA-NFNet backbones. We blend the six $0.5/0.5$ with a public Perch/ProtoSSM pipeline, scoring $0.957$ public and $0.955$ private. We use the system to study a practical problem: in the competitive region, offline validation does not predict the leaderboard. From 67 logged \textit{(offline AUC, public LB)} pairs, Pearson correlation falls from $0.68$ overall to $0.09$ above LB $0.92$ and turns negative above $0.925$. We trace this to two causes: a saturated labeled-soundscape surface with an AUC spread of only $0.00089$ across twelve strong models, and pseudo-label leakage that silently inflates soundscape validation scores. The dominant lever moving the actual leaderboard is pseudo-label quality: strong public pseudo-labels lift a single model from $0.881$ to $0.934$, while a weak self-teacher lowers it. Architectural diversity and the complementary Perch pipeline supply the remaining gains. Keyphrases: Bioacoustics, BirdCLEF, Sound Event Detection, Validation reliability, model ensembling, pseudo-labeling, semi-supervised learning, soundscape classification
|

