Reinforcement Learning, M2 ISF App, 2021-2026

Instructor: Gabriel TURINICI


1/ Introduction to reinforcement learning
2/ Theoretical formalism: Markov decision processes (MDP), value function ( Belman and Hamilton- Jacobi – Bellman equations) etc.
3/ Common strategies, building from the example of « multi-armed bandit »
4/ Strategies in deep learning: Q-learning and DQN
5/ Strategies in deep learning: SARSA and variants
6/ Strategies in deep learning: Actor-Critic and variants
7/ During the course: various Python and gym/gymnasium implementations
8/ Perspectives.


Principal document for the theoretical presentations: (no distribution autoried without WRITTEN consent from the author) (see « teams » group for updated version)

Multi Armed Bandit codes (MAB) : play MAB, solve MAB , solve MAB v2., policy grad from chatGPT to correct., policy grad corrected.

Bellman iterations: code to correct here, solution code here

Gym: play Frozen Lake (v2023) (version 2022)

Q-Learning : with Frozen Lake: version to correct here. Full versions: python version or notebook version

-play with gym/Atari-Breakout: python version or notebook version

Value function iterations (Belmman) on FrozenLake : version to correct here

Deep Q Learning (DQN) : Learn with gym/Atari-Breakout: notebook 2024 and its version with smaller NN and play with result

Policy gradients on Pong adapted from Karpathy, 2024 version (correct to get it working!) python or notebook

You can also load from HERE a converged version (rename as necessary) pg_pong_converged_turinici24

Notebook to use it: here (please send me yours if mean reward above 15!).

Some links: parking human, parking AI

Projets : cf. Teams



Statistical Learning, M1 Math 2024-2026

Instructor: Gabriel TURINICI

Preamble: this course is just but an introduction, in a limited amount of time, to Statistical and Machine learning. This will prepare for the next year’s courses (some of them on my www page cf. « Deep Learning » and « Reinforcement Learning »).

 

Course outline

1/ Examples and machine learning framework

2/ Useful theoretical objects: predictors, loss functions, bias, variance

3/ K-nearest neighbors (k-NN) and the « curse of the dimensionality »

4/ Linear and logistic models in high dimension, variable selection and model regularization (ridge, lasso)

5/ Stochastic Optimization Algorithms

6/ Naive Bayesian classification

7/ Neural networks : introduction, operator, datasets, training, examples, implementations

8/ K-means clustering


Reference: 

Machine Learning Algorithms: From Classical Methods to Deep Neural Networks: Supervised, Unsupervised, and High-Dimensional


Exercices, implementations, current course textbook (no distribution autorized without WRITTEN consent from the author): see « teams » group.


Deep Learning course, 2nd year of Master (ISF App : 2019-26, MATH : 2023-26)

Teacher: Gabriel TURINICI


Summary:
1/ Deep learning: major applications, references, culture
2/ Types of approaches: supervised, reinforcement, unsupervised
3/ Neural networks: presentation of the objects: neurons, operations, loss function, optimization, architecture
4/ Focus on stochastic optimization algorithms, proof of convergence of SGD
5/ Convolutional neural networks (CNN): filters, layers, architectures
6/ Technique: back-propagation, regularization, hyperparameters
7/ Networks for sequences: RNN, LSTM, Attention, Transformer
8/ Generative networks (GAN, VAE)
9/ Programming environments for neural networks: TensorFlow, Keras, PyTorch and work on the examples covered in class
10/ Stable Diffusion, LLM
11/ AI Agents: what is « agentic »: definition, autonomy, reasoning, decision-making, agentic workflows
12/ AI Agents fundamentals: harness, tools, skills, memory, context, planning, reasoning and action
13/ AI Agents: coding and evaluation: coding an agent, packages, frameworks, tool integration, architectures, evaluation, reliability and observability
14/ Ethical and alignment perspectives: safety, autonomy, human oversight, accountability, transparency and alignment


Documents
MAIN document (theory) and Agentic implementations : see your teams channel
(no distribution is authorized without WRITTEN consent from the author)
for back-propagationSGD convergence proof
Implementations
Function approximation by NN : notebook version, Python version
Results (approximation & convergence)

After 5 times more epochs
Official code reference https://doi.org/10.5281/zenodo.7220367
Pure python (no keras, no tensorflow, no Pytorch) implementation (cf. also theoretical doc):
– version « to implement » (with Dense/FC layers) (bd=iris),
– version : solution

If needed: iris dataset here
Implementation : keras/Iris , pytorch

(tensorflow) CNN example: https://www.tensorflow.org/tutorials/images/cnn
Pytorch example CNN/MNIST : python and notebook versions.

Todo : use on MNIST, try to obtain high accuracy on MNIST, CIFAR10.
VAE: latent space visualisation : CVAE – python (rename *.py) , CVAE ipynb version
Stable diffusion:

Working example jan 2025: python version, Notebook version

Old working example 19/1/2024 on Google collab: version : notebook, (here python, rename *.py). ATTENTION the run takes 10 minutes (first time) then is somehow faster (just change the prompt text).