Data and Corpora (Master S1/S3 Class)

Basic information

  • Credits: 3
  • Format: 90 minutes per week
  • 2026-2027 instructors: Gabriel Thiberge
  • 2026-2027 schedule: Fridays, 8:30-10:00, room 145
  • Moodle page

Description

Corpus Linguistics studies language using systematically organized collections of spoken and written texts, known as corpora. This course provides foundational knowledge of corpus linguistics, focusing on oral and written corpora. Key topics include the definition and evolution of corpus linguistics, types of corpora (written and oral), and principles of corpus design such as sampling, balance, and ethical considerations. Students will explore analytical tools for quantitative and qualitative studies of linguistic patterns (using large-scale corpora in well-documented languages – English, in particular) and learn about applications in linguistic description / data collection, corpus constitution & annotation, in field situation and in the lab. The course highlights differences between oral and written corpora, as well as challenges in analyzing spoken data, and data maintenance.

Prerequisites

None.

Learning outcomes

On successful completion of this course, students should be able to:

  • Constitute spoken and written corpora.
  • Annotate, analyze and maintain them.
  • Use existing large-scale corpora tools to conduct quantitative studies of specific phenomena.