Statistics for Textual Data (Master S2 Class)

Basic information

Description

This course aims to provide students with a solid theoretical foundation in statistics, tailored specifically for the analysis of textual data and the evaluation of natural language processing (NLP) systems. Participants will gain a comprehensive understanding of the statistical principles necessary for extracting meaningful insights from text-based datasets and rigorously assessing the performance of NLP models.
The course introduces key concepts in descriptive and inferential statistics, focusing on their application to textual data. Topics include hypothesis testing, confidence intervals, and methods for detecting outliers in datasets. Students will explore advanced resampling techniques, such as bootstrapping, to estimate the reliability of model metrics and handle data variability effectively. Practical examples will illustrate how these statistical tools can be used to evaluate language models, classify textual data, and detect patterns in large corpora.

Prerequisites

Students should have a basic knowledge of python.

Learning outcomes

On successful completion of this course, students should:

  • know the main concepts of statistics
  • Be able to report and analyze results of an experiment