Mountaineering Acoustics Network

Mountaineering Acoustics at the Summit of Discovery
Software / field-audio-tools / birdidpv-report-and-analysis

birdidpv multimedia identification report

Machine identification, not a verified record: BirdNET scores each 3-second frame. It does not confirm that a species was present. Confidence is not a probability. A high score on one frame is not a record. Listen before citing any of this.

Sources
260115_002_Tr1.WAV (track 1 only)
Recordist
Gewenxin Yu
Recording setup
Two Primo EM273 omni-directional microphones (−37 dB ± 3 dB at 1 kHz, 80 dB S/N ratio, 60 Hz ~ 20 kHz) and a Zoom F3 recorder
Model
BirdNET (Cornell Lab / Chemnitz UT) via birdnetlib, 3-second frames
Ranges analysed
135 below-threshold ranges from lowdom · 00:43:30 of 03:02:24 (23.8% of the take)
Species filter
29.898°N, 102.03°E · week 3 of 48 · 6,522 species reduced to 788
Minimum confidence
0.25
Detections
128 frames in 11 species

Audio is off. Enable it once, then hover over the timeline.

range analysed high confidence low confidence
Spectrogram of the selected detection frame
Select a detection to see the spectrogram of its three seconds.

The pale band marks the 135 ranges lowdom left below its threshold. Hover to select the nearest detection. Playback starts after a short pause and stops at the end of that frame. Click to play immediately. That also copies the start timecode. The locked detection keeps an orange marker so you can see where playback sits after the pointer moves on. Select a species below to show only its detections. The spectrogram above is rendered ahead of time from the source WAV, not from the preview. It uses a logarithmic frequency axis so the sub-400 Hz band stays readable. A real call shows harmonic structure in the kilohertz bands. Broadband energy low down with nothing above it is geophony, whatever the model named it. The preview audio is the same lossy 48 kbit/s Opus derivative the lowdom report uses. Use the source WAV files for verification.

Species

Select a row to filter the timeline.

ReferenceSpeciesFramesHighestAtFirst heard
Reference photograph of Pyrrhocorax pyrrhocoraxPavel Shukov · CC-BY-NCRed-billed ChoughPyrrhocorax pyrrhocorax821.00000:40:3500:23:36
Reference photograph of Tadorna ferrugineaegorbirder · CC-BYRuddy ShelduckTadorna ferrugineacheck by ear20.92801:41:0501:41:05
Reference photograph of Corvus macrorhynchosJoe Bourget · CC-BY-NCLarge-billed CrowCorvus macrorhynchos170.91200:32:2800:32:10
Reference photograph of Botaurus stellarisTatyana Zarubo · CC-BY-NCGreat BitternBotaurus stellarischeck by ear120.90600:11:1900:10:15
Reference photograph of Numenius arquataАнна Голубева · CC-BY-NC-NDEurasian CurlewNumenius arquatacheck by ear10.68402:52:3502:52:35
Reference photograph of Anser anserFrans Vandewalle · CC-BY-NCGraylag GooseAnser ansercheck by ear20.62800:09:3300:09:33
Reference photograph of Fulica atraWei Li Jiang · CC-BY-NCEurasian CootFulica atracheck by ear80.61102:18:3600:23:15
Reference photograph of Mareca penelopezametnya · CC-BY-NCEurasian WigeonMareca penelopecheck by ear10.44300:37:0600:37:06
Reference photograph of Anas platyrhynchosunknown photographer · CC-BY-SAMallardAnas platyrhynchoscheck by ear10.43902:08:1702:08:17
Reference photograph of Tadorna tadornathegreatdodo · CC-BY-NCCommon ShelduckTadorna tadornacheck by ear10.36500:24:4000:24:40
Reference photograph of Ardea cinereaFrank Sengpiel · CC-BYGray HeronArdea cinereacheck by ear10.26402:37:1002:37:10

Reference photographs come from iNaturalist. We store them in this package rather than hot-linking them. The package survives archiving that way, and no reader's address is sent to a third party. Each file is kept exactly as iNaturalist served it. It remains under its photographer's licence, credited beside it. The machine-readable record is in credits.json. A photograph shows what the species looks like. It is not evidence that the species was here. Nine of these eleven rows are ones we ask you to doubt. One of those nine is a photograph of a bird that was never there.

Taxon identifiers

These are the same eleven species in the databases a detection usually has to travel to. We matched them on iNaturalist taxon ID, not on name. A name search returns a congener often enough. Attaching an identifier to the wrong bird is a real risk. Wikidata holds dozens more per species. We keep those in taxon-ids.json.

SpeciesElsewhere
Red-billed ChoughPyrrhocorax pyrrhocorax
Ruddy ShelduckTadorna ferruginea
Large-billed CrowCorvus macrorhynchos
Great BitternBotaurus stellaris
Eurasian CurlewNumenius arquata
Graylag GooseAnser anser
Eurasian CootFulica atra
Eurasian WigeonMareca penelope
MallardAnas platyrhynchos
Common ShelduckTadorna tadorna
Gray HeronArdea cinerea

Reading this against a frozen lake

Two corvids account for 99 of the 128 detections. They behave like real calls. They carry energy well above 2 kHz. They sit 3–11 dB above everything else. The other nine species are all deep-voiced waterbirds. The lake had frozen over by mid-January at 4,055 m. We measured the detected frames after a 150 Hz high-pass. That measurement shows where the two groups part:

Detected as150–400 Hz400 Hz–2 kHz2–8 kHzLevel
Red-billed Chough24%55%21%−44.3 dBFS
Large-billed Crow14%85%1%−41.2 dBFS
Great Bittern45%49%5%−51.8 dBFS
Ruddy Shelduck52%45%3%−47.2 dBFS
Eurasian Coot89%10%1%−48.3 dBFS

The waterbird detections concentrate their energy below 400 Hz at a lower level. They mostly show none of the high-frequency structure the corvid detections do. That band is where the ice resonance sits. Local guides call that sound long hou, “dragon roar”. A plausible reading is that BirdNET maps ice onto birds whose calls are booms. Great Bittern at 0.906 is the clearest case. Its twelve detections arrive in sustained runs. Eight of them fall inside seventy seconds. Not one of their spectrograms shows a call above the ice.

Listening to the species explains the mistake. It does not excuse it. On the first page of xeno-canto recordings for Botaurus stellaris, try XC891071, XC1000766, XC832807 and XC741523. The booms really are close to the dragon roar. The confusion looks reasonable rather than absurd. Listen to them here. Then listen to any green stretch of the timeline above:

These recordings are stored in this package under their recordists' CC BY-NC-SA licences. We do not embed them from xeno-canto. The comparison survives archiving that way, and no reader's address is sent to a third party. We re-render the spectrograms here with the same axes and scaling as the detection frames above. xeno-canto draws its own spectrograms on a linear frequency axis. That axis suits the wide scrubbing strip in its player. It also leaves the boom as a thin line along the bottom while the background birds fill the frame. Put next to a detection from this take, the same sound would look like a different one.

One thing separates them by ear. The ice carries an electronic quality that the bird does not. That difference is the interesting part of the error. It is plainly audible. None of the spectrogram statistics on this page found it.

Ice is not the only thing here that is not a bird. The frame at 01:41:05 is a person. BirdNET calls it Ruddy Shelduck at 0.928. That is the second highest confidence in the whole run. It is Xiao Zhang, one of the guides. He was calling across the lake to the recordist. The recordist identified it on listening. There is no bird in those three seconds at all.

That frame is worth opening. It also shows how a spectrogram can be misread. It carries harmonic stacks between 1 and 3 kHz. We first took them for a call. Ice does not produce structure like that. We moved the species out of the doubtful list. A voice does produce it. The stacks are a pitch contour gliding through its harmonics. Formants sit around 1.2 and 2 kHz. Four or five syllables appear in the last second. That is speech, not song. The band average in the table above had hidden the voice under the low-frequency bed. The picture that uncovered it was then read wrong. Only listening settled it.

The ledger runs three ways, not two. It includes real calls, lake ice, and at least one human voice. Every row marked check by ear is one we would not cite without going back to the source WAV. The loudest argument for that rule is simple. The model's second most confident bird in three hours was a man shouting.

Identification by BirdNET, developed by the K. Lisa Yang Center for Conservation Bioacoustics at the Cornell Lab of Ornithology with Chemnitz University of Technology, reached through birdnetlib. The BirdNET models are licensed CC BY-NC-SA 4.0 and that non-commercial condition applies to these results. Cite: Kahl, S., Wood, C. M., Eibl, M., & Klinck, H. (2021). BirdNET: A deep learning solution for avian diversity monitoring. Ecological Informatics, 61, 101236. Neither this report nor field-audio-tools is affiliated with or endorsed by the BirdNET team.