Lessons from a Community Radio Audio Archive

Image

The Mess We Started With

When a community radio station in the UK decided to celebrate its twentieth anniversary, a small team of volunteers opened the storage cupboard and found a mountain of analogue and digital media. There were minidiscs, DAT tapes, cassettes, CD-Rs, and a handful of external hard drives with cryptic folder names. The earliest shows dated back to the late 1990s, and the most recent were from just a few months before. The problem was not the quantity. It was the lack of any consistent labelling.

We found files called show1.mp3, final_mix.wav, and Jane_Interview_04.03.02. Some had no date, others had a date but no presenter, and many had no indication of whether they were a live broadcast, a pre-record, or a repeat. We could play them, but we could not search them. We could not share them with listeners in any meaningful way. Digitisation alone, we quickly learned, was not enough.

Why Consistent Metadata Matters

Metadata is the label on the jar. Without it, you might know you have jam, but you cannot tell if it is strawberry, raspberry, or something that went off in 2003. For audio, metadata turns a sound file into a findable, shareable, and reusable asset. It answers the questions that listeners and volunteers ask every day: Who presented this? When was it broadcast? What music was played? Is it safe to share online?

Consistency is the key. One volunteer might label a date as 4/3/02, another as March 4, 2002, and a third as 04-03-2002. All three mean the same thing, but a search for 2002-03-04 will miss them all. We settled on ISO 8601 for dates: YYYY-MM-DD. For presenter names, we used Surname_Firstname. For show titles, we used lowercase with hyphens: the-blues-hour. These small decisions saved us hundreds of hours.

  • Searchability: a listener can find every episode featuring a particular guest or topic.
  • Rights clearance: you know instantly whether a clip contains commercial music or a sensitive interview.
  • Batch processing: consistent labels let you apply loudness normalisation or format conversion to whole folders.
  • Handover: new volunteers can understand the archive without a tour guide.

Practical Tagging: Fields That Made a Difference

We built a simple spreadsheet and also embedded metadata directly into the audio files using ID3 tags for MP3 and Vorbis comments for FLAC. Every file received the same core fields. Here are the ones that proved most useful.

  • Unique identifier: a serial number like CR-0123456789 to avoid duplicate filenames.
  • Date: broadcast date in ISO 8601.
  • Duration: in minutes and seconds, verified against the file.
  • Presenter: surname and first name, from a controlled list.
  • Producer: if different from presenter.
  • Guests: names and roles, comma-separated.
  • Location: studio, outside broadcast, or remote.
  • Show title: lowercase with hyphens.
  • Episode number: if part of a series.
  • Genre: from a controlled vocabulary, such as music, talk, drama, news.
  • Music tracks: title, artist, and start time in HH:MM:SS.
  • Rights status: cleared, restricted, or unknown.
  • Technical details: sample rate, bit depth, channels, loudness in LUFS, peak in dBFS.
  • Notes: any anomalies, such as tape dropouts or background noise.

Controlled vocabularies were essential. We agreed on a list of genres, locations, and presenter names before we started. When someone was unsure, they asked the group. It felt slow at first, but it prevented the chaos of free-text entries. A sidecar .txt file with the same metadata sat next to each audio file, so even if the embedded tags failed, the information survived.

Audio Analysis to Support the Labels

We used audio analysis not to replace listening, but to verify and enrich our labels. Simple measurements caught many mistakes. If a file was labelled 30:00 but the waveform showed 28:42, we corrected the duration. If a file was tagged live but had no applause or room tone, we investigated. If a file was labelled music show but had a speech-to-music ratio of 90:10, we re-tagged it as talk with music beds.

Loudness normalisation was a game-changer. Old shows varied wildly in level. Some were recorded at -6 dBFS peak, others at -20 dBFS. We measured integrated loudness in LUFS and normalised access copies to -16 LUFS for web playback, keeping the preservation masters at their original levels. We also checked for clipping, DC offset, and excessive noise floor. A high-frequency spectral view revealed tape hiss on the cassette transfers and a low-frequency hum on some minidiscs, which we noted in the metadata so future volunteers would not think the files were corrupted.

We set technical standards: 48 kHz, 24-bit WAV for preservation, and MP3 320 kbps for access. Batch processors applied these settings, but only after we had verified the labels. Silence detection helped us split long continuous recordings into individual shows. We kept the raw files untouched and worked on copies. Every analysis setting was documented in the metadata, so the process was reproducible.

From Archive to Audience: Search and Share

Once the metadata was consistent, the archive became a pleasure to use. We built a basic web catalogue with a search box. A listener could type 2003 and jazz and find every relevant episode. They could filter by presenter, location, or rights status. Volunteers could generate playlists for community events. We created podcast feeds for series that were cleared for sharing, and we embedded simple audio players on the station's website.

Rights metadata was particularly valuable. We knew which shows could be shared openly, which needed permission from guests, and which contained commercial music that had to be removed. Instead of guessing, we could act quickly. We also shared clips with other community stations, always with the correct attribution and date.

The biggest lesson? Start with a schema, even if it is just a spreadsheet and a naming convention. Train everyone who touches the archive. Use controlled vocabularies. Check your work with audio analysis. Back up everything in at least two places. And remember that an archive is not just a collection of files. It is community memory, and clear labels are what let that memory be found, heard, and valued again.

About Author Graphic Designer

Audata No rushing, no fuss — just thoughtful notes and practical help, written by people who care.

Showing 16 verified guest comments

0123456789 image

Soldman Kell

April 25, 2019 at 10:46 am

"The worst hotel ever"

Take in the iconic skyline and visit the neighbourhood hangouts that you've only ever seen on TV. Take in the iconic skyline and visit the neighbourhood.

image

Burson Lesson

April 25, 2019 at 10:46 am

"Was too noisy and not suitable for business meetings"

Take in the iconic skyline and visit the neighbourhood hangouts that you've only ever seen on TV. Take in the iconic skyline and visit the neighbourhood.

Write a Review

Subscribe To Our Newsletter

Want to be notified when we launch a new template or an udpate. Just sign up and we'll send you a notification by email.

Night
Day