1AgentGarten: Code Worlds for Evolving AgentsInteractive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent staMirroS10-092513153
2Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?Learning from experience is essential for LLM agents to adapt to unfamiliar and dynmaic environments. Evaluating this ability is therefore important for understanding how effectively agents acquire anNational University of Singapore10-09136312
3TokenRouter: Efficient Serving System for Token-Level LLM RoutingLarge language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving. While coarse-grained routing at the session or query lTsinghua-NICS-EFC10-09129238
4From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment SimulationRealistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce. We explore agentic language wNanyang Technological University Singapore10-09105244
5MiMo-V2.6: Scaling Reinforcement Learning Towards Self-ImprovementReinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushXiaomi MiMo10-09742
6SuperNav: An Agentic Navigation System for Any Task in Any SceneGeneral-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality. Some existing methods fine-tuzju3dv10-09722110
7Multi-Agent Egocentric World Model with Fine-Grained Embodied InteractionEgocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interacKAIST AI10-0952217
8In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation TasksWe study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains undehanwen wang10-0947612
9OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D CinemagraphsRecent 3D world models generate photorealistic, explorable scenes that remain frozen in time. OuroWorld is a mask-free framework that turns any static 3D Gaussian Splatting scene into a 3D cinemagraphAlaya Lab10-0940241
10U-Space: Uncovering When and Why Uncertainty Arises in Language ModelsLarge language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answerMultimodal AI Lab - Technical University of Darmstadt10-093967
11REMORY: Learning Residual Memory for Context CompactionLong-horizon agents compact their history to continue within a finite context window, but a textual summary alone may not support every subsequent decision. We introduce REMORY, a neural memory networFusion Lab: Generative Vision Lab of Fudan University and SII10-0938543
12Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence MatchingDense correspondence matching has historically been bounded by simplifying spatio-temporal priors, such as smooth motion and rigid geometry. While effective for classical tasks, these assumptions breaThe University of Hong Kong10-093825
13MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion TransformersSparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing metTencent Hunyuan10-0938236
14DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-TrainingWe present DreamTrue, a multi-view, cross-embodiment robot world model for action-faithful and physically plausible video prediction. Training such a model on existing robot datasets faces two obstaclLue Fan10-093725
15Memento 3: Model-Based Recursive Self-Improvement through Reflective RulebooksLearning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world modelUniversity College London10-09355
16TestPrism: Rethinking Test Evaluation Beyond a Single ReferenceLarge language model (LLM) coding agents have advanced test generation across diverse programming tasks. However, the common practice of evaluating tests against a single reference solution overlooks NJU-LINK Lab10-09342
17Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric RewardsRecent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of Arena10-09292
18Foundations of Large Language ModelsThis is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is stNiuTrans10-09273881
19SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM InferenceThe memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a Westlake ENCODE Lab10-09272
20OneSearch-VL: Unified Multimodal Deep Research Agent for Image and VideoSingle-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition. A key challenge is to preservHongyu Li10-0924219