Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

14 Commits
 
 
 
 
 
 
 
 

Repository files navigation

Lexicogrammatical Tagger Treebanks

Repository for datasets annotated for the lexicogrammatical features introduced by Biber and colleagues (e.g., Biber et al., 2021)

Annotation Guidelines

Annotation guidelines were developed by Kristopher Kyle, Hakyung Sung, Doug Biber, and Randi Reppen. First drafts of the guidelines were developed by Kyle and Sung based on the Longman Grammar and were revised during the annotation of sentences from the Michigan Corpus of Academic Spoken English (MICASE) and the the Michigan Corpus of Upper Level Student Papers (MICUSP) based on input, suggestions, and feedback from Doug and Randi.

The current version of the guidelines can be found here.

Note that guidelines can be updated as specific ambiguous/difficult cases arise that are not covered by the current guidelines.

We also have a Google Group for discussing issues related to annotating these features.

How to contribute new treebanks

  1. Read the guidelines
  2. Randomly sample n sentences (we suggest somehwere between 100 and 1000); It may be helpful to message the team to let them know what corpus you will be working on so that work isn't repeated.
  3. Document any annotation decisions that were made that were not covered by the current guidelines
  4. Validate your annotations using one of the following methods:
  • Train at least two humans how to use the annotation guidelines and have two humans independently annotate each sentence; adjudicate any disagreements (using the guidelines)
  • Train at least three humans how to use the annotation guidelines and have two humans independently annotate each sentence; a third annotator adjudicates any disagreements
  • Train at least two humans how to use the annotation guidelines. Use the lexicogrammatical tagger to annotate each sentence, and have one human independently annotate each sentence; a second human annotator adjudicates any disagreements
  1. Submit your project as a conference proceeding paper at a relevant conference (e.g., Linguistic Annotation Workshop (LAW), which is usually co-located with the Association for Computational Linguistics Conference)
  2. Submit your annotated data to the repository
  3. Feel awesome because you contributed to Open Science and made the Lexicogrammatical Tagger more accurate!

Join the next corpus annotation meeting!

Email me (Kris Kyle; see my University of Oregon home page for the email address) if you are interested in joining the next annotation meeting!

Current Treebanks

Two treebanks have been completed and will be made publicly available as soon as the associated paper has been accepted for publication.

MICASE (500 sentences)

This dataset comprises a fair use sample of 500 sentences randomly selected from the Michigan Corpus of Academic Spoken English (MICASE; Simpson et al., 2002).

Language use domain descriptors:

  • Academic
  • Spoken
  • Students and Professors
  • Advanced users of English
  • A variety of situational characteristics
  • A variety of disciplines

Annotation quality Each sentence was independently annotated by two trained annotators. Any annotation disagreements were adjudicated by a third annotator based on the annotation guidelines.

MICUSP (500 sentences)

This dataset comprises a fair use sample of 500 sentences randomly selected from the Michigan Corpus of Upper Level Student Papers (MICUSP; Römer, 2010).

Language use domain descriptors:

  • Academic
  • Written
  • Students
  • Primarily advanced users of English
  • A variety of situational characteristics
  • A variety of disciplines

Annotation quality Each sentence was independently annotated by two trained annotators. Any annotation disagreements were adjudicated by a third annotator based on the annotation guidelines.

Citations

Biber, D., Gray, B., Staples, S., & Egbert, J. (2021). Investigating grammatical complexity in L2 English writing research*: linguistic description versus predictive measurement. In The Register-Functional Approach to Grammatical Complexity (pp. 432-457). Routledge.

Römer, U., & Swales, J. M. (2010). The Michigan corpus of upper-level student papers (MICUSP). Journal of English for Academic Purposes, 9(3), 249.

Simpson, R. C., S. L. Briggs, J. Ovens, and J. M. Swales. (2002) The Michigan Corpus of Academic Spoken English. Ann Arbor, MI: The Regents of the University of Michigan.

About

Repository for datasets annotated for lexicogrammatical features introduced by Biber and colleagues

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors