Umwelt: Birdsong hero — image coming soon
SELECTED WORK

Umwelt: Birdsong

Can a Machine Hear What We Can't?

CLIENT Self-initiated
YEAR 2024
ROLE ML Research
Audio + Data Visualization
TOOLS Python, librosa
UMAP, PyTorch
OVERVIEW

Birds hear more than we do. Where I might pick out a simple eight-note melody, a bird can perceive something closer to a 48-note sequence in the same span of sound. This project started with one question: can machine learning act as a kind of bridge between our perceptual world and theirs?

APPROACH

This is my exploration of Umwelt — Jakob von Uexküll's term for the specific, subjective perceptual world of a given organism. Instead of training a model to classify or label bird species, which imposes a human frame on the data from the start, I deliberately chose an unsupervised approach, letting patterns and clusters emerge from the audio without predefining what to look for.

How do birds perceive sound differently from us? Can a machine surface patterns we literally can't perceive on our own? And the one that stuck with me most — what biases get baked in the moment a human decides how to teach a machine to "listen"?

Deciding whether to annotate a whole phrase or individual syllables, or how to balance sample duration across species — none of that is neutral. Every choice I made shaped what the model learned to notice and what it learned to ignore. Training a model is never neutral, and every annotation reflects my own interpretation back into the system. That's not a hedge, it's the point.

Spectrogram / UMAP visuals — coming soon
  1. Data collection — scraped birdsong recordings across species and locations using the Xeno-Canto API.
  2. Feature extraction — generated spectrograms and chromagrams from raw audio, testing how preprocessing choices change what a visualization invites you to notice.
  3. Chromagram → MIDI — translated chromagrams into MIDI to see what's preserved or lost when birdsong gets re-encoded into a human musical format.
  4. Dimensionality reduction & clustering — ran PCA and UMAP on spectrogram data to visually compare vocalizations and look for real structure.
  5. Latent space exploration — interpolated between species and individuals, treating the latent space as both an analytical and creative tool.
  6. TweetyNet (annex) — hand-annotated audio in Raven Lite to train TweetyNet, a PyTorch birdsong-segmentation model.
Pipeline diagram — coming soon
Unsupervised APPROACH
6-Step PIPELINE
TweetyNet PYTORCH SEGMENTATION