Building a DDC-annotated Corpus from OAI Metadata

Title: Building a DDC-annotated Corpus from OAI Metadata
Authors: Mathias Lösch, Ulli Waltinger, Wolfram Horstmann, Alexander Mehler
Pub/Conf:  Journal of Digital Information 12(2)

Document servers complying to the standards of the Open Archives Initiative (OAI) are rich, yet seldom exploited source of textual primary data for research fields in text mining, natural language processing or computational linguistics. We present a bilingual (English and German) text corpus consisting of bibliographic OAI records and the associated full texts. A particular added value is that we annotated each record with at least one Dewey Decimal Classification (DDC) number, inducing a subject-based categorization of the corpus. By this means, it can be used as training data for machine learning-based text categorization tasks in digital libraries, but also as primary data source for linguistic research on academic language use related to specific disciplines. We describe the construction of the corpus using data from the Bielefeld Academic Search Engine (BASE), as well as its characteristics.


  author    = {Mathias L{\"o}sch and
               Ulli Waltinger and
               Wolfram Horstmann and
               Alexander Mehler},
  title     = {Building a DDC-annotated Corpus from OAI Metadata},
  journal   = {J. Digit. Inf.},
  volume    = {12},
  number    = {2},
  year      = {2011},
  ee        = {},
  bibsource = {DBLP,}