Curriculum
9 Sections
108 Lessons
10 Weeks
Expand all sections
Collapse all sections
Module-1 : Reinforcement Learning
16
1.1
What is Expectation?
1.2
Introduction to Grid world
1.3
State, Action, Reward, Policy
1.4
Markov Decision Process(MDP)
1.5
Bellman equations (State value and Action value Functions)
1.6
Policy Iteration
1.7
Exploration Vs Exploitation
1.8
Montecarlo for Optimal policy
1.9
TD Learning
1.10
SARSA
1.11
Q-learning
1.12
Linear Function Approximation
1.13
Discount Factor to Average Reward
1.14
Policy gradient Theorm
1.15
Need for Policy gradient for Continuous Action spaces
1.16
Advantage term
Module-2 : Transformers
26
2.1
RNN with Attention
2.2
Drawbacks of attention in RNN
2.3
Transformer Architecture
2.4
Self attention, Cross attention, Masked Attention
2.5
Positional Encoding in-detail
2.6
Layer Normalisation
2.7
Language Modelling
2.8
Decoder only models(GPT/GPT2)
2.9
Casual Language Modelling
2.10
Calculate Parameters in GPT
2.11
Greedy search
2.12
Beam search
2.13
Exhaustive search
2.14
Top-k and Top-p
2.15
Temperature
2.16
Encoder only Models(BERT)
2.17
Masked language modelling(MLM) and Next Sentence Prediction(NSP)
2.18
Encoder-Decoder Models(T5)
2.19
Denoising Objective
2.20
Tokenisation
2.21
Space Tokenizer
2.22
Byte pair encoding
2.23
Word piece Tokenizer
2.24
Sentence piece Tokenizer
2.25
Optimizers and Activations
2.26
Distillation of Transformers
Module-3 : RL and Transformer synergy for Conversational ai
15
3.1
Why is pretraining not sufficient to build a conversation ai?
3.2
Revisiting Policy gradient
3.3
Linking Policy gradient with Transformers
3.4
Reinforcement Learning with human Feedback(RLHF)
3.5
Reward Modelling
3.6
Approximating Value Function
3.7
Generalised Advantage Estimation(GAE)
3.8
Need for Variance Reduction
3.9
Problems with Policy gradient
3.10
Importance sampling and Off-policy Learning
3.11
Need for Proximal Policy Optimiation
3.12
Entropy loss
3.13
KL Divergence
3.14
Bradley Terry Model
3.15
Direct Preference Optimisation(DPO) vs RLHF
Module-4 : Llama Architecture
5
4.1
Transformers vs Llama
4.2
Rotary Positional Embeddings
4.3
KV-Cache
4.4
Activation Function
4.5
Attention
Module-4 : Pytorch and GPU
14
5.1
Registers
5.2
SRAM
5.3
VRAM
5.4
VRAM Estimation
5.5
Choosing Right GPU
5.6
Distributed Training
5.7
Gradient Accumulation
5.8
Broadcast/Reduce Operators
5.9
Communication
5.10
Cluster creation
5.11
RANK
5.12
Bucketing
5.13
No-sync context
5.14
Distributed Training using Torch run
Module-6 : Quantization
10
6.1
Why Quantization is important?
6.2
32 bit Floating point representation
6.3
Convert 32 bit FP to binary representation
6.4
16 bit FP, 8 bit FP, uint 8 representations
6.5
Systematic quantization
6.6
Symmetric quantization
6.7
Post training quantization
6.8
Quantization Aware training(QAT)
6.9
Disadvantages of Quantization
6.10
Quantization through GGUF and AWQ
Module-7 : Fine Tuning
12
7.1
Preparing Datasets through Groq
7.2
Different PEFT methods
7.3
LORA
7.4
QLORA
7.5
Finetuning for Memorization
7.6
ORPO
7.7
Finetuning Mixtral
7.8
Multi GPU Finetuning with DDP and FSDP
7.9
DPO Finetuning
7.10
Best Practices while using GPU’s
7.11
Finetune with few GPU’s
7.12
Small LLM’s
Module-8 : Inference
6
8.1
NVIDIA NIM microservices
8.2
Throughput and Latency tradeoffs
8.3
Tensor parallelism
8.4
In-flight batching
8.5
LLM infrastructure based on workload
8.6
Picking GPU and inference engine
Module-9 : Miscellaneous
4
9.1
Jail break
9.2
Prompt injection
9.3
Building Guards
9.4
Context Caching
LLM Internals and Optimization
Search
Curriculum
This content is protected, please
login
and enroll in the course to view this content!
error:
Content is protected !!
Login with your site account
Lost your password?
Remember Me
Not a member yet?
Register now
Register a new account
Are you a member?
Login now
Modal title
Main Content