Contact GitHub support about this user’s behavior. Learn more about reporting abuse.
I'm Hanze Dong.
I work on machine learning research.
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
Python 8.5k 826
A recipe for online RLHF and online iterative DPO.
Python 544 48
Visualization of mean field and neural tangent kernel regime
Jupyter Notebook 23 3
Python 3
There was an error while loading. Please reload this page.