| user: | gpjt |
| created: | January 12, 2009 |
| karma: | 1.8k |
| about: | https://www.gilesthomas.com/ |
| 1. | Why do OpenAI's GPT-2 weights beat mine? Part five: data quality(gilesthomas.com) |
| 2. | |
| 3. | Putting my Jax-trained models on the Hugging Face Hub(gilesthomas.com) |
| 4. | |
| 5. | |
| 6. | Adding diagrams to my static site generator with D2(gilesthomas.com) |
| 7. | Use the built-in GELU, don't roll your own(gilesthomas.com) |
| 8. | A Quick(ish) Chinchilla Check(gilesthomas.com) |
| 9. | I use AI on this blog(gilesthomas.com) |
| 10. | |
| 11. | Why do OpenAI's GPT-2 weights beat mine? Part two: the bugfix(gilesthomas.com) |
| 12. | Why do OpenAI's GPT-2 weights beat mine?(gilesthomas.com) |
| 13. | Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090(gilesthomas.com) |
| 14. | Building intuition about LLM parameter counts(gilesthomas.com) |
| 15. | Poppy the training box, part 1: the beginnings(gilesthomas.com) |
| 16. | From bigrams to GPT-2, one component at a time (in Jax)(gilesthomas.com) |
| 17. | Building a Jax training loop for an LLM training run(gilesthomas.com) |
| 18. | Thoughts on Role Confusion(gilesthomas.com) |
| 19. | Flax debugging: making a hash of things(gilesthomas.com) |