Idea for Expository AI I've heard some complaints about the frontier models still be bad at explaining math and was thinking of an RL environment that would help might be to: -Take very hard math problem with a verifiable answer -Have frontier model explain to a tiny model like (0.5-1B params and provably bad score on the problem) how to solve but not the solution, and reward the frontier model for prompts/explanations that helped the tiny model solve the problem Obviously some amount of human supervision is needed to weed out it giving too much information |
No comments yet