rasbt-LLMs-from-scratch/ch02
Sebastian Raschka 7bd263144e
Switch from urllib to requests to improve reliability (#867)
* Switch from urllib to requests to improve reliability

* Keep ruff linter-specific

* update

* update

* update
2025-10-07 15:22:59 -05:00
..
01_main-chapter-code Switch from urllib to requests to improve reliability (#867) 2025-10-07 15:22:59 -05:00
02_bonus_bytepair-encoder Fix BPE bonus materials (#561) 2025-03-08 17:21:30 -06:00
03_bonus_embedding-vs-matmul minor spelling fix 2024-09-08 15:35:36 -05:00
04_bonus_dataloader-intuition fixed num_workers (#229) 2024-06-19 17:36:46 -05:00
05_bpe-from-scratch Switch from urllib to requests to improve reliability (#867) 2025-10-07 15:22:59 -05:00
README.md fixed video link (#646) 2025-06-13 08:16:18 -05:00

Chapter 2: Working with Text Data

 

Main Chapter Code

 

Bonus Materials

  • 02_bonus_bytepair-encoder contains optional code to benchmark different byte pair encoder implementations

  • 03_bonus_embedding-vs-matmul contains optional (bonus) code to explain that embedding layers and fully connected layers applied to one-hot encoded vectors are equivalent.

  • 04_bonus_dataloader-intuition contains optional (bonus) code to explain the data loader more intuitively with simple numbers rather than text.

  • 05_bpe-from-scratch contains (bonus) code that implements and trains a GPT-2 BPE tokenizer from scratch.

In the video below, I provide a code-along session that covers some of the chapter contents as supplementary material.



Link to the video