Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

AZ uses a known transition model. I wouldn't call this model-based, because when people say model-based they usually mean learning a world model.

MuZero doesn't actually learn a state transition function- it only models reward. It's a model that predicts reward given an action sequence. I suppose we could consider this a hybrid that's slightly model based because it does have an incentive to learn what's going on. But there's no feedback where we predict future world states and feed that back in.



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: