Professional skills in data science and NLP (Master S2 Class)
Basic information
- Credits: 3
- Format: 2 hours weekly
- 2026-2027 instructors: Guillaume Wisniewski
- 2026-2027 schedule: TBD
Description
This course is designed to equip students with the advanced skills necessary to excel in Data Science and NLP within a professional environment. It begins with an in-depth exploration of Python for Data Sciences, covering advanced topics such as parallelization, memory management, code profiling and optimization, and interfacing with SQL databases, … Students will also learn essential software engineering practices, including unit testing, environment management, and the creation of robust workflows, enabling them to develop and deploy efficient, reliable, and scalable solutions.
In addition to programming expertise, the course delves into advanced data science and NLP practices critical for enterprise success. Topics include web scraping, model deployment, continuous integration for Data Sciences, and the principles of MLOps for maintaining AI systems in production. Participants will gain a foundational understanding of computer architecture to address performance challenges in machine learning and deep learning applications. Furthermore, the course explores enterprise data organization, including working with structured and semi-structured data, database systems (with a focus on no-SQL databases), and modern data storage solutions such as data lakes. The curriculum also provides hands-on experience with leading NLP and ML libraries commonly used in industry ensuring participants are prepared to leverage cutting-edge tools in their projects.
Learning outcomes
On successful completion of this course, students should:
- know the vocabulary and tools needed to analyze data in a professional context
- be able to design, program, test and deploy ML and NLP pipelines in professional contexts