arXiv preprint arXiv:2609.06107 First author Hugging Face Daily Papers #2
Hao Liang, Mingrui Chen, Hengyi Feng, Meiyi Qiang, Wentao Zhang
An evaluation platform for reinforcement learning with verifiable rewards that compares rollout selection, reweighting, and domain-mixture policies under a common GRPO recipe.
Recommended citation: Hao Liang, Mingrui Chen, Hengyi Feng, Meiyi Qiang, Wentao Zhang. (2026). "DataFlex-RL: An Evaluation Platform for RLVR Data Policies." arXiv preprint arXiv:2609.06107.
arXiv preprint arXiv:2609.23088 First author Hugging Face Daily Papers #1
Hao Liang, Qihan Lin, Meiyi Qiang, Linzhuang Sun, Hengyi Feng, Mingrui Chen, Sizhe Qiu, Wentao Zhang
An open family of 4B, 9B, and 27B foundation models for K–12 learning and teaching, trained with capability-balanced educational data for subject competence, curriculum grounding, learner diagnosis, and pedagogical support.
Recommended citation: Hao Liang, Qihan Lin, Meiyi Qiang, Linzhuang Sun, Hengyi Feng, Mingrui Chen, Sizhe Qiu, Wentao Zhang. (2026). "OmniEdu: Open Foundation Models for Learning and Teaching." arXiv preprint arXiv:2609.23088.
arXiv preprint arXiv:2609.20842 Co-first author
Qifeng Cai*, Xuanguang Pan*, Hao Liang*, Chang Xu, Wentao Zhang
A unified Text-to-SQL post-training framework that combines coverage-guided augmentation with failure-driven learning, improving structural coverage and targeting model weaknesses exposed during GRPO.
Recommended citation: Qifeng Cai, Xuanguang Pan, Hao Liang, Chang Xu, Wentao Zhang. (2026). "COAL-SQL: Coverage-Guided Augmentation and Failure-Driven Learning for Text-to-SQL Post-Training." arXiv preprint arXiv:2609.20842.
Journal of Computer Science and Technology, 41(1): 289–317 First author Survey
Hao Liang, Zhen Hao Wong, Rui-Tong Liu, Yu-Han Wang, Meiyi Qiang, Zhengyang Zhao, Chengyu Shen, Conghui He, Wentao Zhang, Bin Cui
A systematic survey of data preparation across pre-training, continual pre-training, and post-training, connecting algorithms, workflows, datasets, and open research challenges.
ICLR 2026 Second author
Lu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang, Xiaochen Ma, Zhen Hao Wong, Junbo Niu, Chengyu Shen, Runming He, Bin Cui, Wentao Zhang
Introduces interleaved online fine-tuning to learn from questions that remain difficult for reinforcement learning, combining complementary training signals on the hardest examples.
The Web Conference (WWW) 2026 Co-first author Oral
Hao Liang*, Qifeng Cai*, Hejun Dong, Meiyi Qiang, Ruichuan An, Zhaoyang Han, Zhengzhou Zhu, Bin Cui, Wentao Zhang
A benchmark for retrieving relevant moments from long videos under multimodal context, designed to expose the temporal and cross-modal limitations of current retrieval systems.
ICDE 2026 Co-first author Oral
Qifeng Cai*, Hao Liang*, Chang Xu*, Tao Xie, Wentao Zhang, Bin Cui
A SQL-aware data augmentation framework that generates structurally valid and diverse training examples to improve robustness in text-to-SQL systems.
KDD 2026 Co-first author Project leader
Chengyu Shen*, Zhen Hao Wong*, Runming He*, Hao Liang*, Meiyi Qiang, Zimo Meng, Zhengyang Zhao, Bohan Zeng, Zhengzhou Zhu, Bin Cui, Wentao Zhang
A step-by-step verification framework for identifying flawed mathematical questions and reasoning traces before they are used for model training or evaluation.
NeurIPS 2026 First author Hugging Face Daily Papers #1
Hao Liang et al.
A data-centric training framework that brings dynamic sample selection, mixture, and reweighting directly into the large-model training loop.
Findings of ACL 2026 Co-first author
Zhaoyang Han*, Qihan Lin*, Hao Liang*, Bowen Chen, Zhou Liu, Wentao Zhang
A human-centric long-video benchmark that evaluates omni-modal models across visual, audio, temporal, and social understanding.
Findings of ACL 2026 Co-first author
Linzhuang Sun*, Tianyu Guo*, Hao Liang*, Ruitong Liu, Yuying Li, Qifeng Cai, Jingxuan Wei, Yuchen Wu, Bihui Yu, Xiangxiang Zhang, Wentao Zhang, Bin Cui
Recasts text-to-SQL as dynamic, multi-turn database exploration so systems can refine intent and interact with real databases rather than solve isolated queries.
Findings of ACL 2026 Co-first author
Qijie You*, Wenkai Yu*, Hao Liang*, Zhen Hao Wong, Wentao Zhang
A hop-aware diagnostic benchmark that separates retrieval and reasoning failures across multi-step agentic RAG trajectories.
ACM Multimedia 2026 Co-first author
Linzhuang Sun*, Yuxia Zhu*, Ruitong Liu*, Hao Liang*, Sizhe Qiu*, Zheng Sun, Caijun Jia, Honghao He, Yuchen Wu, Siyuan Li, Jingxuan Wei, Xiangxiang Zhang, Bihui Yu, Wentao Zhang
Grounds mathematical reasoning in mutable structured states, enabling intermediate information to be explicitly organized, revised, rendered, and reused throughout the reasoning process.
Findings of ACL 2026 Co-first author
Linzhuang Sun*, Mingyang Chen*, Hao Liang*, Tianpeng Li, Yijie Zhou, Chenzheng Zhu, Tianyu Guo, Huanyao Zhang, Jingxuan Wei, Bihui Yu, Fan Yang, Wentao Zhang
Develops a verifier for assessing long, complex tool-use trajectories in realistic environments where correctness depends on both actions and intermediate state.
ACM Multimedia 2026 Co-first author
Linzhuang Sun*, Ruitong Liu*, Yuxia Zhu*, Hao Liang*, Sizhe Qiu*, Xiaohan Xu, Jingxuan Wei, Xiangxiang Zhang, Bihui Yu, Zhengzhou Zhu
Introduces a dynamic verifier that collaborates with the policy during multimodal reasoning, providing process-level guidance to detect errors and steer rollouts toward valid trajectories.
Technical report First author Hugging Face Daily Papers #1
Hao Liang, Qifeng Cai, Yilin Lin, J. Du, Q. Xia, S. Qiu, Linzhuang Sun, Meiyi Qiang, Zhaoyang Han, Xiaochen Ma, et al.
Benchmarks large language models as training-data preparators across core operations such as generation, cleaning, transformation, and quality control.
NeurIPS 2026 Co-first author Project leader
Qijie You*, Hao Liang*, Mingrui Chen, Bohan Zeng, Meiyi Qiang, Zhen Hao Wong, Wentao Zhang
A full-modality long-video retrieval benchmark with user-simulated queries that require joint reasoning over visual, speech, and ambient audio evidence.
NeurIPS 2026 First author Hugging Face Daily Papers #2
Hao Liang, Qihan Lin, Zhaoyang Han, Xiaochen Ma, Zhen Hao Wong, Meiyi Qiang, Linzhuang Sun, Wentao Zhang
Builds a curriculum-aligned K–12 knowledge graph for evaluating educational coverage and constructing training data for educational language models.
Technical report Co-first author Project leader
Hengyi Feng*, Hao Liang*, Mingrui Chen, Bohan Zeng, Meiyi Qiang, Zhengyang Zhao, Zimo Meng, Zeang Sheng, Wentao Zhang
A benchmark for multi-hop trajectory reasoning over long audio-visual videos, emphasizing evidence chains distributed across time and modalities.
ACM Multimedia 2026 Co-first author Project leader
Rongyi Yu*, Chenyuan Duan*, Hao Liang*, Wentao Zhang
Evaluates whether agents can plan efficient multi-hop evidence retrieval in long videos before answering questions that depend on separated evidence clips.
Findings of ACL 2026 Co-first author Project leader
Yuying Li*, Siyi Qian*, Hao Liang*, Leqi Zheng, Ruichuan An, Linzhuang Sun, Jiajun Zhang, Wentao Zhang
Decouples visual perception from mathematical reasoning through structured geometric captions and introduces CapGeo-Bench for evaluating models' geometric understanding and captioning quality.
EMNLP 2026 Co-first author Oral
Hao Liang*, Meiyi Qiang*, Yuying Li*, Zefeng He, Xiaochen Ma, Ruichuan An, Yongzhen Guo, Zhengzhou Zhu, Bin Cui, Wentao Zhang
A type-aware benchmark for detecting and diagnosing errors in synthetic mathematical data beyond final-answer correctness.