Sebastian Raschka (author of build reasoning model...
# ai-reading-club
p
Sebastian Raschka (author of build reasoning models and LLMs from scratch) will be joining us live next week on Thursday! 🙂 If you have any questions post them here and I can ask them or join live and ask them in the chat yourself luma.com/…
🔥 2
❤️ 3
a
I am a big fan of him !!! Damn the scheduled timezone is a bummer!!! Hopefully i will try to make it. Here are a few questions that you can ask him from my side if i don’t manage to make it: 1. His views of recursive self learning and the progress, challenges etc. 2. When will people get over with transfomers architecture… ( He might say this is the best we have (SOTA).. but ask him if there is a scenario where hybrid architecture might some into picture… Bottom line we want to train a model without having trillions of token baggage. 3. World models thats being the hype these days…does Jepa play a role in the future. 4. Post training is the king (His words not mine) .. Is there a way to get token efficiency during post training.. reduce the number of tokens generated by the model to get to an answer, without sacrificing accuracy. 5. Tell him that a lot of people appreciate his work (myself included)..bcoz of it’s simplicity. Very few people can make a complicated subject simple.
❤️ 1
p
Thank you! I'll note these down and hopefully get them, great questions! It will be recorded too of course if you can't make it
a
That is a super impressive guest.
k
Here are a couple of questions: • The book focuses on Math, which makes sense since it is easier, but if one wanted to go beyond that, more open-ended reasoning etc what prompt dataset would you reach for ? • And for such cases, would you recommend RL using AI as judge or some other method? • For ch08 style distillation, did you ever compare LoRA to full fine-tuning, especially if I want to train larger models?