Data Preparation for Large Language Models
A systematic survey of data preparation across pre-training, continual pre-training, and post-training, connecting algorithms, workflows, datasets, and open research challenges.
A systematic survey of data preparation across pre-training, continual pre-training, and post-training, connecting algorithms, workflows, datasets, and open research challenges.
Introduces interleaved online fine-tuning to learn from questions that remain difficult for reinforcement learning, combining complementary training signals on the hardest examples.
A benchmark for retrieving relevant moments from long videos under multimodal context, designed to expose the temporal and cross-modal limitations of current retrieval systems.
A SQL-aware data augmentation framework that generates structurally valid and diverse training examples to improve robustness in text-to-SQL systems.
A step-by-step verification framework for identifying flawed mathematical questions and reasoning traces before they are used for model training or evaluation.
A data-centric training framework that brings dynamic sample selection, mixture, and reweighting directly into the large-model training loop.
A human-centric long-video benchmark that evaluates omni-modal models across visual, audio, temporal, and social understanding.
Recasts text-to-SQL as dynamic, multi-turn database exploration so systems can refine intent and interact with real databases rather than solve isolated queries.
A hop-aware diagnostic benchmark that separates retrieval and reasoning failures across multi-step agentic RAG trajectories.
Develops a verifier for assessing long, complex tool-use trajectories in realistic environments where correctness depends on both actions and intermediate state.
Benchmarks large language models as training-data preparators across core operations such as generation, cleaning, transformation, and quality control.
A full-modality long-video retrieval benchmark with user-simulated queries that require joint reasoning over visual, speech, and ambient audio evidence.
Builds a curriculum-aligned K–12 knowledge graph for evaluating educational coverage and constructing training data for educational language models.
A benchmark for multi-hop trajectory reasoning over long audio-visual videos, emphasizing evidence chains distributed across time and modalities.
Evaluates whether agents can plan efficient multi-hop evidence retrieval in long videos before answering questions that depend on separated evidence clips.
Improves multimodal reasoning by training models to verify intermediate chain-of-thought steps rather than relying only on final-answer supervision.
A high-efficiency pipeline for producing and filtering synthetic image-caption data to improve vision-language model training.
An open, operator-based framework that unifies data generation, cleaning, evaluation, filtering, and workflow automation for large-model development.
A benchmark for robustly evaluating audio captions across content coverage, correctness, and failure modes that simple text-overlap metrics miss.
A hierarchical benchmark for evaluating multimodal large language models across realistic mathematical settings, representations, and reasoning demands.
A plug-and-play prompt augmentation system that improves downstream performance while requiring little task-specific training data.
Documents the first-place system for the ICML SeePhys Challenge, focusing on multimodal scientific perception and reasoning.
A benchmark for detecting and cleaning errors in synthetic mathematical training data, covering quality problems beyond final-answer correctness.
An efficient quality score for video-question-answering data that supports filtering and selection before expensive multimodal training.
Estimates the effective proportions of heterogeneous pre-training data so data mixtures can be managed and optimized more systematically.
Uses large models to select informative video keyframes at scale, reducing visual redundancy while preserving information needed by downstream tasks.
Develops a generation and quality-control pipeline for producing synthetic empathy data suitable for training more supportive dialogue models.