ATC and ADS-B Data Mining for Open Aviation Datasets#
This subteam builds the dataset that the language and airspace projects both need and that does not currently exist in usable form. Controller and pilot radio traffic is publicly streamed from airports including Purdue University Airport and can be captured continuously, but audio alone carries no record of which aircraft was where when a given instruction was issued. The work is to record that audio alongside ADS-B surveillance data from the OpenSky Network, align the two on a common clock so that every utterance sits against the traffic picture it was spoken into, and then label the aligned pairs against a ground-truth grammar of standard phraseology. Labeling is where the research sits rather than where the drudgery sits, because the interesting cases are the ones where real phraseology departs from the standard, and cataloguing those departures is what makes the dataset worth publishing. Students are trained on the phraseology and the grammar before labeling, which makes this an effective entry point for newcomers, and the intended outcome is a public dataset and an accompanying dataset paper.
Expected activities include: 1) Review of existing ATC speech and surveillance datasets 2) Setup and operation of continuous audio capture at selected airports 3) Retrieval and preprocessing of matching ADS-B tracks 4) Timestamp alignment and segmentation of utterances against traffic state 5) Transcription and grammar-based annotation with an agreed ground truth 6) Inter-annotator agreement measurement and quality control 7) Cataloguing of departures from standard phraseology 8) Release of the dataset as open data and preparation of a dataset paper.
Team members will work in Python with speech processing and grammar tooling, and will complete the formal grammars tutorial before annotation begins. This subteam works in close coordination with the large language model subteam, which consumes the resulting dataset.