Here are a couple of questions:
• The book focuses on Math, which makes sense since it is easier, but if one wanted to go beyond that, more open-ended reasoning etc what prompt dataset would you reach for ?
• And for such cases, would you recommend RL using AI as judge or some other method?
• For ch08 style distillation, did you ever compare LoRA to full fine-tuning, especially if I want to train larger models?