Excited to share our latest #ICML2026 paper on dataset selection/valuation during the post-training stage, led by @cindy2000_sh, in collaboration with our fantastic colleagues from Meta!
The core idea is simple but powerful: treat dataset selection as portfolio allocation. With
[1/N]
Excited to share our #ICML26 paper: Convex Dataset Valuation for Post-Training
TL;DR: We select post-training datasets by jointly modeling target alignment and redundancy.
Paper: arxiv.org/abs/2605.16704
Code: github.com/uiuctml/convex…
🧵






