Birds hear more than we do. Where I might pick out a simple eight-note melody, a bird can perceive something closer to a 48-note sequence in the same span of sound. This project started with one question: can machine learning act as a kind of bridge between our perceptual world and theirs?
This is my exploration of Umwelt — Jakob von Uexküll's term for the specific, subjective perceptual world of a given organism. Instead of training a model to classify or label bird species, which imposes a human frame on the data from the start, I deliberately chose an unsupervised approach, letting patterns and clusters emerge from the audio without predefining what to look for.
THE QUESTIONS
How do birds perceive sound differently from us? Can a machine surface patterns we literally can't perceive on our own? And the one that stuck with me most — what biases get baked in the moment a human decides how to teach a machine to "listen"?
Deciding whether to annotate a whole phrase or individual syllables, or how to balance sample duration across species — none of that is neutral. Every choice I made shaped what the model learned to notice and what it learned to ignore. Training a model is never neutral, and every annotation reflects my own interpretation back into the system. That's not a hedge, it's the point.
PROCESS
- Data collection — scraped birdsong recordings across species and locations using the Xeno-Canto API.
- Feature extraction — generated spectrograms and chromagrams from raw audio, testing how preprocessing choices change what a visualization invites you to notice.
- Chromagram → MIDI — translated chromagrams into MIDI to see what's preserved or lost when birdsong gets re-encoded into a human musical format.
- Dimensionality reduction & clustering — ran PCA and UMAP on spectrogram data to visually compare vocalizations and look for real structure.
- Latent space exploration — interpolated between species and individuals, treating the latent space as both an analytical and creative tool.
- TweetyNet (annex) — hand-annotated audio in Raven Lite to train TweetyNet, a PyTorch birdsong-segmentation model.