When a community radio station in the UK decided to celebrate its twentieth anniversary, a small team of volunteers opened the storage cupboard and found a mountain of analogue and digital media. There were minidiscs, DAT tapes, cassettes, CD-Rs, and a handful of external hard drives with cryptic folder names. The earliest shows dated back to the late 1990s, and the most recent were from just a few months before. The problem was not the quantity. It was the lack of any consistent labelling.
We found files called show1.mp3, final_mix.wav, and Jane_Interview_04.03.02. Some had no date, others had a date but no presenter, and many had no indication of whether they were a live broadcast, a pre-record, or a repeat. We could play them, but we could not search them. We could not share them with listeners in any meaningful way. Digitisation alone, we quickly learned, was not enough.
Metadata is the label on the jar. Without it, you might know you have jam, but you cannot tell if it is strawberry, raspberry, or something that went off in 2003. For audio, metadata turns a sound file into a findable, shareable, and reusable asset. It answers the questions that listeners and volunteers ask every day: Who presented this? When was it broadcast? What music was played? Is it safe to share online?
Consistency is the key. One volunteer might label a date as 4/3/02, another as March 4, 2002, and a third as 04-03-2002. All three mean the same thing, but a search for 2002-03-04 will miss them all. We settled on ISO 8601 for dates: YYYY-MM-DD. For presenter names, we used Surname_Firstname. For show titles, we used lowercase with hyphens: the-blues-hour. These small decisions saved us hundreds of hours.
We built a simple spreadsheet and also embedded metadata directly into the audio files using ID3 tags for MP3 and Vorbis comments for FLAC. Every file received the same core fields. Here are the ones that proved most useful.
Controlled vocabularies were essential. We agreed on a list of genres, locations, and presenter names before we started. When someone was unsure, they asked the group. It felt slow at first, but it prevented the chaos of free-text entries. A sidecar .txt file with the same metadata sat next to each audio file, so even if the embedded tags failed, the information survived.
We used audio analysis not to replace listening, but to verify and enrich our labels. Simple measurements caught many mistakes. If a file was labelled 30:00 but the waveform showed 28:42, we corrected the duration. If a file was tagged live but had no applause or room tone, we investigated. If a file was labelled music show but had a speech-to-music ratio of 90:10, we re-tagged it as talk with music beds.
Loudness normalisation was a game-changer. Old shows varied wildly in level. Some were recorded at -6 dBFS peak, others at -20 dBFS. We measured integrated loudness in LUFS and normalised access copies to -16 LUFS for web playback, keeping the preservation masters at their original levels. We also checked for clipping, DC offset, and excessive noise floor. A high-frequency spectral view revealed tape hiss on the cassette transfers and a low-frequency hum on some minidiscs, which we noted in the metadata so future volunteers would not think the files were corrupted.
We set technical standards: 48 kHz, 24-bit WAV for preservation, and MP3 320 kbps for access. Batch processors applied these settings, but only after we had verified the labels. Silence detection helped us split long continuous recordings into individual shows. We kept the raw files untouched and worked on copies. Every analysis setting was documented in the metadata, so the process was reproducible.
Once the metadata was consistent, the archive became a pleasure to use. We built a basic web catalogue with a search box. A listener could type 2003 and jazz and find every relevant episode. They could filter by presenter, location, or rights status. Volunteers could generate playlists for community events. We created podcast feeds for series that were cleared for sharing, and we embedded simple audio players on the station's website.
Rights metadata was particularly valuable. We knew which shows could be shared openly, which needed permission from guests, and which contained commercial music that had to be removed. Instead of guessing, we could act quickly. We also shared clips with other community stations, always with the correct attribution and date.
The biggest lesson? Start with a schema, even if it is just a spreadsheet and a naming convention. Train everyone who touches the archive. Use controlled vocabularies. Check your work with audio analysis. Back up everything in at least two places. And remember that an archive is not just a collection of files. It is community memory, and clear labels are what let that memory be found, heard, and valued again.
April 25, 2019 at 10:46 am
Take in the iconic skyline and visit the neighbourhood hangouts that you've only ever seen on TV. Take in the iconic skyline and visit the neighbourhood.
Want to be notified when we launch a new template or an udpate. Just sign up and we'll send you a notification by email.
Soldman Kell
April 25, 2019 at 10:46 am
Take in the iconic skyline and visit the neighbourhood hangouts that you've only ever seen on TV. Take in the iconic skyline and visit the neighbourhood.