Statistics for Textual Data (Master S2 Class)
Basic information
- Credits: 3
- Format: 2 hours weekly
- 2026-2027 instructors: Guillaume Wisniewski
- 2026-2027 schedule: TBD
- Moodle page
Description
This course aims to provide students with a solid theoretical foundation in statistics, tailored specifically for the analysis of textual data and the evaluation of natural language processing (NLP) systems. Participants will gain a comprehensive understanding of the statistical principles necessary for extracting meaningful insights from text-based datasets and rigorously assessing the performance of NLP models.
The course introduces key concepts in descriptive and inferential statistics, focusing on their application to textual data. Topics include hypothesis testing, confidence intervals, and methods for detecting outliers in datasets. Students will explore advanced resampling techniques, such as bootstrapping, to estimate the reliability of model metrics and handle data variability effectively. Practical examples will illustrate how these statistical tools can be used to evaluate language models, classify textual data, and detect patterns in large corpora.
Prerequisites
Students should have a basic knowledge of python.
Learning outcomes
On successful completion of this course, students should:
- know the main concepts of statistics
- Be able to report and analyze results of an experiment