This is my recent PhD work. I want to show that AI can be used for good, rather than just optimizing ads and keeping people on TikTok for longer. Modern VLMs (LLMs with vision capabilities) are just good enough for labeling that they can replace route-label work for bioacoustics, and I find that distilling these labels into a bioacoustic encoder -- such as SongMAE, my previous work -- can work fairly well. Interestingly, the student exceeds the teacher model, and doesnt asymptote at teacher performance. This work was actually roughly inspired by Simon Willison (pelican guy), he made a post that showed Qwen 27B is an effective labeler for images.