1The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing PredictionMixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Edge009-162644392
2nanoMuse: An Open-Source Personal Agent for Every Device You OwnAssistants from 2011 answered and waited, and agents from 2023 did a task and stopped. In September 2026 Meta's Muse showed an agent for one person, with accounts, devices, memory and a conversation tZhejiang University10-061012833
3YuE2: Unifying Symbolic and Audio Music Generation at Frontier QualitySymbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuEMultimodal Art Projection09-2723331.2万
4PyTorch Distributed: Experiences on Accelerating Data Parallel TrainingThe PyTorch distributed data parallel module optimizes large-scale model training using techniques like gradient bucketing, computation-communication overlap, and selective synchronization to achieve Shen Li2020-06-291310.4万
5HuggingFace's Transformers: State-of-the-art Natural Language ProcessingTransformers library provides state-of-the-art Transformer architectures and pretrained models for natural language processing tasks with a unified API and emphasis on extensibility and robust deploymHugging Face2019-10-0934716.8万
6Geometric Context Transformer for Streaming 3D ReconstructionLingBot-Map is a feed-forward 3D foundation model that reconstructs scenes from video streams using a geometric context transformer architecture with specialized attention mechanisms for coordinate grRobbyant04-153931.8万
7TradingAgents: Multi-Agents LLM Financial Trading FrameworkA multi-agent framework using large language models for stock trading simulates real-world trading firms, improving performance metrics like cumulative returns and Sharpe ratio.Yijia Xiao2024-12-28151611.1万
8OpenDevin: An Open Platform for AI Software Developers as Generalist AgentsOpenDevin is a platform for developing AI agents that interact with the world by writing code, using command lines, and browsing the web, with support for multiple agents and evaluation benchmarks.AK2024-07-249079.1万
9Efficient Memory Management for Large Language Model Serving with PagedAttentionPagedAttention algorithm and vLLM system enhance the throughput of large language models by efficiently managing memory and reducing waste in the key-value cache.AK2023-09-127718.6万
10Kronos: A Foundation Model for the Language of Financial MarketsKronos, a specialized pre-training framework for financial K-line data, outperforms existing models in forecasting and synthetic data generation through a unique tokenizer and autoregressive pre-trainYu Shi2025-08-025944.1万
11Bitnet.cpp: Efficient Edge Inference for Ternary LLMsBitnet.cpp enhances edge inference for ternary LLMs using a novel mixed-precision matrix multiplication library, achieving significant speed improvements over baselines.Jinheng Wang2025-02-17184.1万
12Native and Compact Structured Latents for 3D GenerationA new sparse voxel representation called O-Voxel enables high-quality 3D generative modeling with efficient inference and robust topology handling.Microsoft2025-12-17721.2万
13Prime Agent: A Self-Improving RLM HarnessPrime Agent is an open-source harness that uses recursive subagents, persistent computation, and agent-to-agent coordination to extend language models' long-horizon capabilities across coding and reasPrime Intellect08-245822.2万
14Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio GenerationWe present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video ProKandinsky Lab10-041674261
15WorldSonus: Bringing Sound to WorldsRecent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challNoizAI10-06332189
16SpatialLM: Training Large Language Models for Structured Indoor ModelingSpatialLM, a multimodal large language model, processes 3D point cloud data to generate structured scene understanding outputs, achieving state-of-the-art performance in layout estimation and competitManycore Research2025-06-095325021
17AgentGarten: Code Worlds for Evolving AgentsInteractive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent staMirroS10-082453153
18SuperNav: An Agentic Navigation System for Any Task in Any SceneGeneral-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality. Some existing methods fine-tuzju3dv10-08722110
19MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document ParsingMinerU2.5, a 1.2B-parameter document parsing vision-language model, achieves state-of-the-art recognition accuracy with computational efficiency through a coarse-to-fine parsing strategy.taesiri2025-09-2618028.1万
20Mem0: Building Production-Ready AI Agents with Scalable Long-Term MemoryMem0, a memory-centric architecture with graph-based memory, enhances long-term conversational coherence in LLMs by efficiently extracting, consolidating, and retrieving information, outperforming exiAK2025-04-287426.7万