# AAAI 2026 收藏论文

> 共 210 篇论文

---

## Computer Vision III

### 论文 3
- **英文标题**: APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval
- **中文标题**: APVR：基于自适应关键视觉信息检索的小时级长视频理解
- **作者**: Hong Gao, Yiming Bao, Xuezhen Tu, Bin Zhong, Linan Yue, Min-Ling Zhang
- **英文摘要**: Current multimodal large language models (MLLMs) struggle with hour-level video understanding, facing significant challenges not only in modeling the substantial information volume of long videos but also in overcoming the memory wall and resource constraints during both training and inference. Although recent training-free approaches have alleviated resource demands by compressing visual features, their reliance on incomplete visual information limits the performance potential. To address these limitations, we propose Adaptive Pivot Visual information Retrieval (APVR), a training-free framework that hierarchically retrieves and retains sufficient and important visual information. It breakthroughs the memory wall limitation via two complementary components: Pivot Frame Retrieval employs query expansion and iterative spatio-semantic confidence scoring to identify relevant video frames, and Pivot Token Retrieval performs query-aware attention-driven token selection within up to 1024 pivot frames. This dual granularity approach enables the processing of hour-long videos while maintaining semantic fidelity. Experimental validations on three different baseline MLLMs demonstrate significant performance improvements up to 9.5%, 4.6% and 9.7% on LongVideoBench, VideoMME and MLVU, respectively. APVR achieves state-of-the-art results for both training-free and training-based approaches.
- **中文摘要**: 提出无需训练的框架APVR，通过关键帧检索和关键Token检索两个互补组件分层检索和保留充分的视觉信息，在LongVideoBench、VideoMME和MLVU上显著提升多模态大模型的性能。

### 论文 7
- **英文标题**: VisAssist: A Visually Impaired-Captured Video Question Answering Benchmark for Assistive Systems
- **中文标题**: VisAssist：面向辅助系统的视障人士拍摄视频问答基准
- **作者**: Qi Gao, Heng Li, Yixin Zhou, Meixuan Zhou, Jieqiong Chen, Xinyu Chai
- **英文摘要**: We present VisAssist, the first large-scale video question-answering dataset with 13,413 real-world videos captured by visually impaired users, addressing a critical gap in assistive vision research. Unlike existing benchmarks relying on third-person footage, VisAssist provides authentic first-person perspectives that uniquely capture challenges in blind photography—including unconventional framing, motion artifacts, and frequent information omission. Benchmark evaluations of SOTA multimodal models reveal systematic limitations: severe deficiencies in spatial reasoning when processing dynamic first-person viewpoints, an inability to distinguish missing information from poor capture quality leading to hazardous hallucinations, and fragile text understanding especially for non-Latin scripts under suboptimal conditions. This work establishes a vital real-world benchmark and underscores the need for specialized architectures in visual assistance systems.
- **中文摘要**: 构建首个大规模视障用户拍摄视频问答数据集VisAssist（含13413个视频），揭示了现有多模态模型在第一人称视角空间推理、信息缺失识别和次优条件下文本理解方面的系统性缺陷。

### 论文 14
- **英文标题**: High-Quality Full-Head 3D Avatar Generation from Any Single Portrait Image
- **中文标题**: 从任意单张人像图像生成高质量全头3D头像
- **作者**: Yujie Gao, ChenCheng Wang, Xianbing Sun, Jiahui Zhan, Wentao Wang, Yiyi Zhang, Haohua Zhao, Liqing Zhang, Jianfu Zhang
- **英文摘要**: In this work, we introduce a novel high-fidelity full-head 3D avatar generation method from a single image, regardless of perspective, style, expression, or accessories. Prior works often fail to preserve consistent head geometry and facial details, primarily due to their limited capacity in modeling fine-grained facial textures and maintaining identity information. To address these challenges, we construct a new high-quality dataset containing 227 sequences of digital human portraits captured from 96 different perspectives, totalling 21,792 frames, featuring high-quality facial texture details. To further improve performance, we propose a novel multi-view diffusion named ID-TS diffusion model, which integrate identity and expression information into the two-stage multi-view diffusion process. The low-resolution stage ensures structural consistency of heads across multiple views, while the high-resolution stage preserves facial detail fidelity and coherence. Finally, we propose an enhanced feed-forward Gaussian avatar reconstruction method that optimizes the network on multi-view images of each single subject, significantly improving 3D facial texture details. Extensive experiments show that our method demonstrates robust performance across challenging scenarios, while showcasing broad applicability across numerous downstream tasks.
- **中文摘要**: 构建包含227个序列、21792帧的高质量数字人像数据集，提出集成身份和表情信息的ID-TS多视图扩散模型，以及增强的前馈高斯头像重建方法，实现从单图生成高保真全头3D头像。

### 论文 38
- **英文标题**: SPSC: Sparse and Scalable Multi-Modal 3D Occupancy Prediction for Autonomous Driving
- **中文标题**: SPSC：面向自动驾驶的稀疏可扩展多模态3D占据预测
- **作者**: Qingju Guo, Shuang Li, Binhui Xie, Jing Geng, Wei Li
- **英文摘要**: 3D semantic occupancy prediction offers a nuanced representation of the surrounding environment, which is crucial for ensuring the safety of autonomous driving. However, fine-grained scene representations inevitably result in cubic growth in data scale, which imposes substantial demands on model architecture and computational complexity, especially in high-resolution scenarios. Existing approaches for handling high-resolution scenes typically obtain fine-grained features by grid sampling on low-resolution feature map, resulting in limited sparsity and insufficient feature interaction. This paper presents a framework leveraging SParse representation and SCalable feature interaction to address the aforementioned challenges, called SPSC. Specifically, we maintain sparsity by progressively pruning unoccupied queries during the coarse-to-fine process, thereby reducing the scale of data that the model needs to handle. Subsequently, we introduce query serialization, which transforms queries into an ordered sequence while preserving their spatial structure, This enables fine-grained feature interaction while maintaining linear computational complexity and a larger receptive field. Without complex architectural designs, SPSC significantly outperforms SOTA approaches, relatively enhances the mIoU by 12.0%, 11.0% and 4.8% on nuScenes-Occupancy dataset under the muli-modal, LiDAR and camera settings, respectively.
- **中文摘要**: 提出SPSC框架，通过由粗到精过程中逐步修剪未占据查询保持稀疏性，引入查询序列化在保持线性计算复杂度和更大感受野的同时实现细粒度特征交互，在nuScenes-Occupancy数据集多模态设置下相对提升mIoU 12.0%。

## Computer Vision IV

### 论文 27
- **英文标题**: TSBOW – Traffic Surveillance Benchmark for Occluded Vehicles Under Various Weather Conditions
- **中文标题**: TSBOW：多种天气条件下遮挡车辆交通监控基准
- **作者**: Ngoc Doan-Minh Huynh, Duong Nguyen-Ngoc Tran, Long Hoang Pham, Tai Huu-Phuong Tran, Hyung-Joon Jeon, Huy-Hung Nguyen, Duong Khac Vu, Hyung-Min Jeon, Son Hong Phan, Quoc Pham-Nam Ho, Chi Dai Tran, Trinh Le Ba Khanh, Jae Wook Jeon
- **英文摘要**: Global warming has intensified the frequency and severity of extreme weather events, which degrade CCTV signal and video quality while disrupting traffic flow, thereby increasing traffic accident rates. Existing datasets, often limited to light haze, rain, and snow, fail to capture extreme weather conditions. To address this gap, this study introduces the Traffic Surveillance Benchmark for Occluded vehicles under various Weather conditions (TSBOW), a comprehensive dataset designed to enhance occluded vehicle detection across diverse annual weather scenarios. Comprising over 32 hours of real-world traffic data from densely populated urban areas, TSBOW includes more than 48,000 manually annotated and 3.2 million semi-labeled frames; bounding boxes spanning eight traffic participant classes from large vehicles to micromobility devices and pedestrians. We establish an object detection benchmark for TSBOW, highlighting challenges posed by occlusions and adverse weather. With its varied road types, scales, and viewpoints, TSBOW serves as a critical resource for advancing Intelligent Transportation Systems. Our findings underscore the potential of CCTV-based traffic monitoring, pave the way for new research and applications. The TSBOW dataset is publicly available at the following link.
- **中文摘要**: 提出TSBOW数据集，包含超过32小时真实交通数据、48,000余帧手动标注和320万帧半标注，涵盖8类交通参与者，建立了遮挡和恶劣天气条件下目标检测基准，为智能交通系统提供关键资源。

### 论文 29
- **英文标题**: STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes
- **中文标题**: STRIDE-QA：面向城市驾驶场景时空推理的视觉问答数据集
- **作者**: Keishi Ishihara, Kento Sasaki, Tsubasa Takahashi, Daiki Shiono, Yu Yamaguchi
- **英文摘要**: Vision-Language Models (VLMs) have been applied to autonomous driving to support decision-making in complex real-world scenarios. However, their training on static, web-sourced image-text pairs fundamentally limits the precise spatiotemporal reasoning required to understand and predict dynamic traffic scenes. We address this critical gap with STRIDE-QA, a large-scale visual question answering (VQA) dataset for physically grounded reasoning from an ego-centric perspective. Constructed from 100 hours of multi-sensor driving data in Tokyo, capturing diverse and challenging conditions, STRIDE-QA is the largest VQA dataset for spatiotemporal reasoning in urban driving, offering 16 M QA pairs over 270 K frames. Grounded by dense, automatically generated annotations including 3D bounding boxes, segmentation masks, and multi-object tracks, the dataset uniquely supports both object-centric and ego-centric reasoning through three novel QA tasks that require spatial localization and temporal prediction. Our benchmarks demonstrate that existing VLMs struggle significantly, with near-zero scores on prediction consistency. In contrast, VLMs fine-tuned on STRIDE-QA exhibit dramatic performance gains, achieving 55% success in spatial localization and 28% consistency in future motion prediction, compared to near-zero scores from general-purpose VLMs. Therefore, STRIDE-QA establishes a comprehensive foundation for developing more reliable VLMs for safety-critical autonomous systems.
- **中文摘要**: 构建STRIDE-QA大规模VQA数据集，基于100小时多传感器东京驾驶数据，包含1600万问答对，支持目标中心和自我中心推理，微调后模型空间定位成功率达55%，未来运动预测一致性达28%。

### 论文 31
- **英文标题**: JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- **中文标题**: JRDB-Reasoning：面向机器人视觉推理的难度分级基准
- **作者**: Simindokht Jahangard, Mehrzad Mohammadi, Yi Shen, Zhixi Cai, Hamid Rezatofighi
- **英文摘要**: Recent advances in Vision-Language Models (VLMs) and large language models (LLMs) have greatly enhanced visual reasoning, a key capability for embodied AI agents like robots. However, existing visual reasoning benchmarks often suffer from several limitations: they lack a clear definition of reasoning complexity, offer have no control to generate questions over varying difficulty and task customization, and fail to provide structured, step-by-step reasoning annotations (workflows). To bridge these gaps, we formalize reasoning complexity, introduce an adaptive query engine that generates customizable questions of varying complexity with detailed intermediate annotations, and extend the JRDB dataset with human-object interaction and geometric relationship annotations to create JRDB-Reasoning, a benchmark tailored for visual reasoning in human-crowded environments. Our engine and benchmark enable fine-grained evaluation of visual reasoning frameworks and dynamic assessment of visual-language models across reasoning levels.
- **中文摘要**: 形式化推理复杂度定义，引入自适应查询引擎生成可定制难度的问题及详细中间标注，扩展JRDB数据集创建JRDB-Reasoning基准，支持视觉推理框架的细粒度评估和动态模型评估。

### 论文 32
- **英文标题**: GranAlign: Granularity-Aware Alignment Framework for Zero-shot Video Moment Retrieval
- **中文标题**: GranAlign：面向零样本视频时刻检索的粒度感知对齐框架
- **作者**: Mingyu Jeon, Sunjae Yoon, Jonghee Kim, Junyeong Kim
- **英文摘要**: Zero-shot video moment retrieval (ZVMR) is the task of localizing a temporal moment within an untrimmed video using a natural language query without relying on task-specific training data. The primary challenge in this setting lies in the mismatch in semantic granularity between textual queries and visual content. Previous studies in ZVMR have attempted to achieve alignment by leveraging high-quality pre-trained knowledge that represents video and language in a joint space. However, these approaches failed to balance the semantic granularity between the pre-trained knowledge provided by each modality for a given scene. As a result, despite the high quality of each modality’s representations, the mismatch in granularity led to inaccurate retrieval. In this paper, we propose a training-free framework, called Granularity-Aware Alignment (GranAlign), that bridges this gap between coarse and fine semantic representations. Our approach introduces two complementary techniques: granularity-based query rewriting to generate varied semantic granularities, and query-aware caption generation to embed query intent into video content. By pairing multi-level queries with both query-agnostic and query-aware captions, we effectively resolve semantic mismatches. As a result, our method sets a new state-of-the-art across all three major benchmarks (QVHighlights, Charades-STA, ActivityNet-Captions), with a notable 3.23% mAP@avg improvement on the QVHighlights dataset.
- **中文摘要**: 提出无需训练的GranAlign框架，通过粒度感知查询重写生成多样化语义粒度，结合查询感知描述生成嵌入查询意图，有效解决文本查询与视觉内容间的语义粒度不匹配，在三个主要基准上达到最优。

### 论文 37
- **英文标题**: LidarPainter: One-Step Away from Any Lidar View to Novel Guidance
- **中文标题**: LidarPainter：从任意LiDAR视图一步获得新视角引导
- **作者**: Yuzhou Ji, Ke Ma, Hong Cai, Anchun Zhang, Lizhuang Ma, Xin Tan
- **英文摘要**: Dynamic driving scene reconstruction is of great importance in fields like digital twin system and autonomous driving simulation. However, unacceptable degradation occurs when the view deviates from the input trajectory, leading to corrupted background and vehicle models. To improve reconstruction quality on novel trajectory, existing methods are subject to various limitations including inconsistency, deformation, and time consumption. This paper proposes LidarPainter, a one-step diffusion model that recovers consistent driving views from sparse LiDAR condition and artifact-corrupted renderings in real-time, enabling high-fidelity lane shifts in driving scene reconstruction. Extensive experiments show that LidarPainter outperforms state-of-the-art methods in speed, quality and resource efficiency, specifically 7 × faster than StreetCrafter with only one fifth of GPU memory required. LidarPainter also supports stylized generation using text prompts such as “foggy”  and “night”, allowing for a diverse expansion of the existing asset library.
- **中文摘要**: 提出LidarPainter一步扩散模型，从稀疏LiDAR条件和伪影损坏的渲染中实时恢复一致驾驶视图，速度比StreetCrafter快7倍且仅需五分之一GPU内存，支持文本提示风格化生成。

### 论文 48
- **英文标题**: Less Is More: Rethinking Parameter-Efficient Fine-Tuning from a Subtractive Perspective
- **中文标题**: 少即是多：从减法视角重新思考参数高效微调
- **作者**: Tianqi Jiang, Liu Yang, Xi-Le Zhao, Zixuan Qin, Qinghua Hu
- **英文摘要**: Currently, pretrained models are rapidly scaling in size, which substantially increases the cost of fine-tuning them for downstream tasks. To address this challenge, parameter-efficient fine-tuning (PEFT) methods have been developed to optimize a minimal set of parameters for adaptation. While current PEFT approaches predominantly employ an "additive'' strategy, introducing learnable modules into inputs or architectures, neglect the inherent knowledge embedded within pretrained models, which may be redundant or even conflict with downstream tasks. This limitation leads to increased inference latency and suboptimal transfer performance, particularly in scenarios with significant domain gaps. In this paper, we propose a Subtractive Fine-tuning Paradigm(SFP), which converts multiple redundant operations within the original module into a linear transformation to enhance inference speed and model performance. Specifically, we introduce a compact filter block to replace specific module with interference and redundancy in the original structure to reduce model conflicts. By using a pseudo inverse matrix to construct filter block, ensuring that it can inherit the knowledge of the replacement module, and then freezing the rest of the model, only fine-tuning the filter block is performed to eliminate interference and redundant knowledge, thereby enhancing the model’s adaptability to downstream tasks. Experimental results demonstrate that our SFP outperforms existing PEFT methods in accuracy while decreasing the overall model parameters by 12%. Compared to full fine-tuning, the accuracy has increased by 8.47%(74.04% vs. 65.57%, VTAB).
- **中文摘要**: 提出减法微调范式（SFP），将原始模块中的冗余操作转换为线性变换以加速推理，通过紧凑滤波器块替换冗余模块并利用伪逆矩阵继承知识，在VTAB上准确率提升8.47%，参数减少12%。

## Computer Vision V

### 论文 5
- **英文标题**: Versatile Vision-Language Model for 3D Computed Tomography
- **中文标题**: 面向3D计算机断层扫描的多功能视觉语言模型
- **作者**: Jiayu Lei, Ziqing Fan, Yanyong Zhang, Weidi Xie, Ya Zhang, Yanfeng Wang
- **英文摘要**: Representation learning serves as a foundational component of medical vision-language models (MVLMs), enabling cross-modal alignment, semantic consistency, and enhanced generalization capabilities for downstream tasks. As generalist models rapidly evolve, there is a pressing need to unify diverse downstream tasks, such as diagnosis, segmentation, report generation, and multiple choice within a cohesive framework, demanding more efficient and versatile visual representation learning. However, current MVLMs predominately follow CLIP-style vision pretraining, failing to leverage heterogeneous data resources with multi-dimensional imaging and diverse annotation forms. And there lacks systematic analysis of efficient vision encoder design across varied downstream applications, including diagnosis, segmentation, and text generation tasks, particularly for volumetric imaging like Computed Tomography (CT). Besides, current MVLMs exhibit constrained voxel-level capabilities, lacking effective multi-task instruction tuning framework capable of achieving robust performance across various downstream tasks. To address these challenges, we propose CTInstruct, a novel MVLM employing a hybrid ResNet-ViT encoder with multi-granular vision-language pretraining for efficient heterogeneous data modeling, and unified instruction tuning that jointly optimizes discriminative, generative, and voxel-level reasoning for volumetric medical imaging. CTInstruct achieves SOTA performance across 8 CT benchmarks, setting a new standard for data-efficient multimodal learning in medical imaging.
- **中文摘要**: 提出CTInstruct模型，采用混合ResNet-ViT编码器及多粒度视觉语言预训练，并通过统一指令微调联合优化判别、生成和体素级推理能力。在8个CT基准上达到最优性能。

### 论文 9
- **英文标题**: Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing
- **中文标题**: 探索遥感领域的高效开放词汇分割
- **作者**: Bingyu Li, Haocheng Dong, Da Zhang, Zhiyuan Zhao, Hao Sun, Junyu Gao
- **英文摘要**: Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote sensing (RS) domain, remains underexplored due to the absence of a unified evaluation benchmark and the domain gap between natural and RS images. To bridge these gaps, we first establish a standardized OVRSIS benchmark (OVRSISBench) based on widely-used RS segmentation datasets, enabling consistent evaluation across methods. Using this benchmark, we comprehensively evaluate several representative OVS/OVRSIS models and reveal their limitations when directly applied to remote sensing scenarios. Building on these insights, we propose RSKT-Seg, a novel open-vocabulary segmentation framework tailored for remote sensing. RSKT-Seg integrates three key components: (1) a Multi-Directional Cost Map Aggregation (RS-CMA) module that captures rotation-invariant visual cues by computing vision-language cosine similarities across multiple directions; (2) an Efficient Cost Map Fusion (RS-Fusion) transformer, which jointly models spatial and semantic dependencies with a lightweight dimensionality reduction strategy; and (3) a Remote Sensing Knowledge Transfer (RS-Transfer) module that injects pre-trained knowledge and facilitates domain adaptation via enhanced upsampling. Extensive experiments on the benchmark show that RSKT-Seg consistently outperforms strong OVS baselines by +3.8 mIoU and +5.9 mACC, while achieving 2× faster inference through efficient aggregation.
- **中文摘要**: 建立标准化OVRSISBench基准，提出RSKT-Seg框架，包含多方向代价图聚合模块、高效代价图融合Transformer和遥感知识迁移模块。在基准上实现+3.8 mIoU和+5.9 mACC的提升，推理速度提高2倍。

### 论文 13
- **英文标题**: Exploring Surround-View Fisheye Camera 3D Object Detection
- **中文标题**: 探索环视鱼眼相机3D目标检测
- **作者**: Changcai Li, Wenwei Lin, Zuoxun Hou, Gang Chen, Wei Zhang, Huihui Zhou, Weishi Zheng
- **英文摘要**: In this work, we explore the technical feasibility of implementing end-to-end 3D object detection (3DOD) with surround-view fisheye camera system. Specifically, we first investigate the performance drop incurred when transferring classic pinhole-based 3D object detectors to fisheye imagery. To mitigate this, we then develop two methods that incorporate the unique geometry of fisheye images into mainstream detection frameworks: one based on the bird's-eye-view (BEV) paradigm, named FisheyeBEVDet, and the other on the query-based paradigm, named FisheyePETR. Both methods adopt spherical spatial representations to effectively capture fisheye geometry. In light of the lack of dedicated evaluation benchmarks, we release Fisheye3DOD, a new open dataset synthesized using CARLA and featuring both standard pinhole and fisheye camera arrays. Experiments on Fisheye3DOD demonstrate that our fisheye-compatible modeling improves accuracy by up to 6.2% compared to baseline methods.
- **中文摘要**: 探索环视鱼眼相机系统实现端到端3D目标检测的技术可行性，提出FisheyeBEVDet和FisheyePETR两种方法，采用球面空间表示有效捕捉鱼眼几何特性。发布Fisheye3DOD数据集，精度提升最高6.2%。

### 论文 35
- **英文标题**: FIND: A Simple Yet Effective Baseline for Diffusion-Generated Image Detection
- **中文标题**: FIND：扩散生成图像检测的简单有效基线方法
- **作者**: Jie Li, Yingying Feng, Chi Xie, Jie Hu, Lei Tan, Jiayi Ji
- **英文摘要**: The remarkable realism of images generated by diffusion models poses critical detection challenges. Current methods utilize reconstruction error as a discriminative feature, exploiting the observation that real images exhibit higher reconstruction errors when processed through diffusion models. However, these approaches require costly reconstruction computations and depend on specific diffusion models, making their performance highly model-dependent. We identify a fundamental difference: real images are more difficult to fit with Gaussian distributions compared to synthetic ones. In this paper, we propose Forgery Identification via Noise Disturbance (FIND), a novel method that requires only a simple binary classifier. It eliminates reconstruction by directly targeting the core distributional difference between real and synthetic images. Our key operation is to add Gaussian noise to real images during training and label these noisy versions as synthetic. This step allows the classifier to focus on the statistical patterns that distinguish real from synthetic images. We theoretically prove that the noise-augmented real images resemble diffusion-generated images in their ease of Gaussian fitting. Furthermore, simply by adding noise, they still retain visual similarity to the original images, highlighting the most discriminative distribution-related features. The proposed FIND improves performance by 11.7% on the GenImage benchmark while running 126x faster than existing methods. By removing the need for auxiliary diffusion models and reconstruction, it offers a practical, efficient, and generalizable way to detect diffusion-generated content.
- **中文摘要**: 发现真实图像比合成图像更难用高斯分布拟合的核心差异。提出FIND方法通过添加高斯噪声并将噪声版本标记为合成样本来训练二分类器。在GenImage上性能提升11.7%，速度提升126倍。

### 论文 37
- **英文标题**: MIRA: Evaluating Multimodal AI on Complex Clinical Reasoning in Interventional Radiology
- **中文标题**: MIRA：评估多模态AI在介入放射学中的复杂临床推理能力
- **作者**: Jingxiong Li, Chenglu Zhu, Sunyi Zheng, Yuxuan Sun, Yifei Wang, He Liu, Yunlong Zhang, Yixuan Si, Lin Yang, Liang Xiao
- **英文摘要**: We present MIRA (Multimodal Interventional RAdiology evaluation), a comprehensive benchmark for evaluating large multimodal models in expert-level interventional radiology tasks requiring specialized domain knowledge and advanced visual reasoning capabilities. Unlike existing medical benchmarks that primarily provide binary labels without contextual depth, MIRA offers diverse question formats, including open-ended, closed-ended, single-choice, and multiple-choice categories, each accompanied by detailed expert-validated explanations. The benchmark incorporates approximately 184K high-quality medical images spanning multiple imaging modalities with 1.2M meticulously generated question-answer pairs across various anatomical regions. These pairs were created through a sophisticated cascade methodology involving expert interventional radiologists at both the data collection and validation stages. Our comprehensive evaluation, encompassing zero-shot testing and fine-tuning experiments of large multimodal models, revealing significant performance gaps between AI systems and human specialists. Fine-tuning experiments demonstrate substantial improvements, with models achieving up to 0.80 accuracy on single-choice questions. MIRA establishes a challenging benchmark that suggests promising directions for developing specialized clinical AI systems for interventional radiology.
- **中文摘要**: 提出MIRA基准，约18.4万张多模态医学图像和120万精心生成的问答对。全面评估大模型零样本和微调表现，微调后在单选题上准确率达0.80，揭示了AI系统与人类专家之间的显著性能差距。

### 论文 38
- **英文标题**: CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- **中文标题**: CrossVid：评估多模态大语言模型跨视频推理能力的综合基准
- **作者**: Jingyao Li, Jingyun Wang, Molin Tan, Haochen Wang, Cilin Yan, Likun Shi, Jiayin Cai, Xiaolong Jiang, Yao Hu
- **英文摘要**: Cross-Video Reasoning (CVR) presents a significant challenge in video understanding, which requires simultaneous understanding of multiple videos to aggregate and compare information across groups of videos. Most existing video understanding benchmarks focus on single-video analysis, failing to assess the ability of multimodal large language models (MLLMs) to simultaneously reason over various videos. Recent benchmarks evaluate MLLMs' capabilities on multi-view videos that capture different perspectives of the same scene. However, their limited tasks hinder a thorough assessment of MLLMs in diverse real-world CVR scenarios. To this end, we introduce CrossVid, the first benchmark designed to comprehensively evaluate MLLMs' spatial-temporal reasoning ability in cross-video contexts. Firstly, CrossVid encompasses a wide spectrum of hierarchical tasks, comprising four high-level dimensions and ten specific tasks, thereby closely reflecting the complex and varied nature of real-world video understanding. Secondly, CrossVid provides 5,331 videos, along with 9,015 challenging question-answering pairs, spanning single-choice, multiple-choice, and open-ended question formats. Through extensive experiments on various open-source and closed-source MLLMs, we observe that Gemini-2.5-Pro performs best on CrossVid, achieving an average accuracy of 50.4%. Notably, our in-depth case study demonstrates that most current MLLMs struggle with CVR tasks, primarily due to their inability to integrate or compare evidence distributed across multiple videos for reasoning. These insights highlight the potential of CrossVid to guide future advancements in enhancing MLLMs’ CVR capabilities.
- **中文摘要**: 提出首个评估跨视频时空推理能力的CrossVid基准，包含4个高层次维度和10个具体任务，5331个视频和9015个问答对。实验表明Gemini-2.5-Pro平均准确率50.4%，当前模型在整合跨视频证据方面存在困难。

### 论文 53
- **英文标题**: TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
- **中文标题**: TEMPLE：通过渐进式预SFT对齐激励视频大语言模型的时间理解
- **作者**: Shicheng Li, Lei Li, Kun Ouyang, Shuhuai Ren, Yuanxin Liu, Yuanxing Zhang, Fuzheng Zhang, Lingpeng Kong, Qi Liu, Xu Sun
- **英文摘要**: Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, existing approaches struggle with temporal reasoning due to weak temporal correspondence in the data and over-reliance on the next-token prediction paradigm, which collectively result in the absence temporal supervision. To address these limitations, we propose TEMPLE (TEMporal Preference Learning), a systematic framework that enhances temporal reasoning capabilities through Direct Preference Optimization (DPO). To address temporal information scarcity in data, we introduce an automated pipeline for systematically constructing temporality-intensive preference pairs comprising three steps: selecting temporally rich videos, designing video-specific perturbation strategies, and evaluating model responses on clean and perturbed inputs. Complementing this data pipeline, we provide additional supervision signals via preference learning and propose a novel Progressive Pre-SFT Alignment strategy featuring two key innovations: a curriculum learning strategy which progressively increases perturbation difficulty to maximize data efficiency; and applying preference optimization before instruction tuning to incentivize fundamental temporal alignment. Extensive experiments demonstrate that our approach consistently improves Video LLM performance across multiple benchmarks with a relatively small set of self-generated DPO data. Our findings highlight TEMPLE as a scalable and efficient complement to SFT-based methods, paving the way for developing reliable Video LLMs.
- **中文摘要**: 提出TEMPLE框架，通过直接偏好优化增强时间推理能力。引入自动化流水线构建时间密集偏好对，设计课程学习策略渐进增加扰动难度，并在指令微调前应用偏好优化激励基础时间对齐。

### 论文 62
- **英文标题**: When Trackers Date Fish: A Benchmark and Framework for Underwater Multiple Fish Tracking
- **中文标题**: 当跟踪器遇上鱼：水下多目标鱼群跟踪基准与框架
- **作者**: Weiran Li, Yeqiang Liu, Qiannan Guo, Yijie Wei, Hwa Liang Leo, Zhenbo Li
- **英文摘要**: Multiple object tracking (MOT) technology has made significant progress in terrestrial applications, but underwater tracking scenarios remain underexplored despite their importance to marine ecology and aquaculture. In this paper, we present Multiple Fish Tracking Dataset 2025 (MFT25), a comprehensive dataset specifically designed for underwater multiple fish tracking, featuring 15 diverse video sequences with 408,578 meticulously annotated bounding boxes across 48,066 frames. Our dataset captures various underwater environments, fish species, and challenging conditions including occlusions, similar appearances, and erratic motion patterns. Additionally, we introduce Scale-aware and Unscented Tracker (SU-T), a specialized tracking framework featuring an Unscented Kalman Filter (UKF) optimized for non-linear swimming patterns of fish and a novel Fish-Intersection-over-Union (FishIoU) matching that accounts for the unique morphological characteristics of aquatic species. Extensive experiments demonstrate that our SU-T baseline achieves state-of-the-art performance on MFT25, with 34.1 HOTA and 44.6 IDF1, while revealing fundamental differences between fish tracking and terrestrial object tracking scenarios.
- **中文摘要**: 提出MFT25数据集包含15个视频序列和408,578个标注框。提出SU-T跟踪框架，采用针对鱼类非线性游泳模式优化的无迹卡尔曼滤波器和FishIoU匹配方法。在MFT25上HOTA达34.1，IDF1达44.6。

### 论文 77
- **英文标题**: EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
- **中文标题**: EgoCross：面向跨领域第一人称视频问答的多模态大语言模型基准测试
- **作者**: Yanjun Li, Yuqian Fu, Tianwen Qian, Qi'Ao Xu, Silong Dai, Danda Pani Paudel, Luc Van Gool, Xiaoling Wang
- **英文摘要**: Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities such as cooking and cleaning. In contrast, real-world deployment inevitably encounters domain shifts, where target domains differ substantially in both visual style and semantic content. To bridge this gap, we introduce EgoCross, a comprehensive benchmark designed to evaluate the cross-domain generalization of MLLMs in EgocentricQA. EgoCross covers four diverse and challenging domains, including surgery, industry, extreme sports, and animal perspective, representing realistic and high-impact application scenarios. It comprises approximately 1,000 QA pairs across 798 video clips, spanning four key QA tasks: prediction, recognition, localization, and counting. Each QA pair provides both OpenQA and CloseQA formats to support fine-grained evaluation. Extensive experiments show that most existing MLLMs, whether general-purpose or egocentric-specialized, struggle to generalize to domains beyond daily life, highlighting the limitations of current models. Furthermore, we conduct several pilot studies, e.g., fine-tuning and reinforcement learning, to explore potential improvements. We hope EgoCross and our accompanying analysis will serve as a foundation for advancing domain-adaptive, robust egocentric video understanding.
- **中文摘要**: 提出EgoCross基准覆盖手术、工业、极端运动和动物视角4个领域。包含约1000个问答对跨越798个视频片段和4个关键QA任务。实验表明现有模型在日常生活之外的领域泛化困难。

### 论文 85
- **英文标题**: RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
- **中文标题**: RegionRAG：面向视觉文档理解的区域级检索增强生成
- **作者**: Yinglu Li, Zhiying Lu, Zhihang Liu, Yiwei Sun, Chuanbin Liu, Hongtao Xie
- **英文摘要**: Multi-modal Retrieval-Augmented Generation (RAG) has become a critical method for empowering LLMs by leveraging candidate visual documents. However, current methods consider the entire document as the basic retrieval unit, introducing substantial irrelevant visual content in two ways: 1) Relevant documents often contain large regions unrelated to the query, diluting the focus on salient information; 2) Retrieving multiple documents to increase recall further introduces redundant and irrelevant documents. These redundant contexts distract the model's attention and further degrade the performance. To address this challenge, we propose RegionRAG, a novel framework that shifts the retrieval paradigm from the document level to the region level. During training, we design a hybrid supervision strategy from both labeled data and unlabeled data to pinpoint relevant patches. During inference, we propose a dynamic pipeline that intelligently groups salient patches into complete semantic regions. By delegating the task of identifying relevant regions to the retriever, RegionRAG enables the generator to focus solely on concise, query-relevant visual content, improving both efficiency and accuracy. Experiments on six benchmarks demonstrate that RegionRAG achieves state-of-the-art performance. It improves retrieval accuracy by 10.02% in R@1 on average, and boosts question answering accuracy by 3.56% while using only 71.42% visual tokens compared with prior methods.
- **中文摘要**: 提出RegionRAG框架将检索范式从文档级转变为区域级。设计混合监督策略定位相关补丁，动态流水线智能分组突出补丁为完整语义区域。检索准确率平均提升10.02%，问答准确率提升3.56%，使用视觉token减少28.58%。

### 论文 101
- **英文标题**: Multi-Agent Undercover Gaming: Hallucination Removal Through Counterfactual Test for Multimodal Reasoning
- **中文标题**: 多智能体卧底博弈：通过反事实测试移除多模态推理中的幻觉
- **作者**: Dayong Liang, Xiao-Yong Wei, Changmeng Zheng
- **英文摘要**: Hallucination continues to pose a major obstacle in the reasoning capabilities of large language models (LLMs). Although the Multi-Agent Debate (MAD) paradigm offers a promising solution by promoting consensus among multiple agents to enhance reliability, it relies on the unrealistic assumption that all debaters are rational and reflective, which is a condition that may not hold when agents themselves are prone to hallucinations. To address this gap, we introduce the Multi-agent Undercover Gaming (MUG) protocol, inspired by social deduction games like ''Who is Undercover?''. MUG reframes MAD as a process of detecting ''undercover'' agents (those suffering from hallucinations) by employing multimodal counterfactual tests.  Specifically, we modify reference images to introduce counterfactual evidence and observe whether agents can accurately identify these changes, providing ground-truth for identifying hallucinating agents and enabling robust, crowd-powered multimodal reasoning. MUG advances MAD protocols along three key dimensions: (1) enabling factual verification beyond statistical consensus through counterfactual testing; (2) introducing cross-evidence reasoning via dynamically modified evidence sources instead of relying on static inputs; and (3) fostering active reasoning, where agents engage in probing discussions rather than passively answering questions. Collectively, these innovations offer a more reliable and effective framework for multimodal reasoning in LLMs.
- **中文摘要**: 提出多智能体卧底博弈（MUG）协议，受社交推理游戏启发将多智能体辩论重构为检测幻觉智能体的过程。修改参考图像引入反事实证据观察智能体是否能准确识别变化。沿三个维度推进多智能体辩论协议。

## Computer Vision VI

### 论文 2
- **英文标题**: OTI: A Model-free and Visually Interpretable Measure of Image Attackability
- **中文标题**: OTI：一种无模型且视觉可解释的图像攻击性度量
- **作者**: Jiaming Liang, Haowei Liu, Chi-Man Pun
- **英文摘要**: Despite the tremendous success of neural networks, benign images can be corrupted by adversarial perturbations to deceive these models. Intriguingly, images differ in their attackability. Specifically, given an attack configuration, some images are easily corrupted, whereas others are more resistant. Evaluating image attackability has important applications in active learning, adversarial training, and attack enhancement. This prompts a growing interest in developing attackability measures. However, existing methods are scarce and suffer from two major limitations: (1) They rely on a model proxy to provide prior knowledge (e.g., gradients or minimal perturbation) to extract model-dependent image features. Unfortunately, in practice, many task-specific models are not readily accessible. (2) Extracted features characterizing image attackability lack visual interpretability, obscuring their direct relationship with the images. To address these, we propose a novel Object Texture Intensity (OTI), a model-free and visually interpretable measure of image attackability, which measures image attackability as the texture intensity of the image's semantic object. Theoretically, we describe the principles of OTI from the perspectives of decision boundaries as well as the mid- and high-frequency characteristics of adversarial perturbations. Comprehensive experiments demonstrate that OTI is effective and computationally efficient. In addition, our OTI provides the adversarial machine learning community with a visual understanding of attackability.
- **中文摘要**: 尽管神经网络取得了巨大成功，良性图像仍可被对抗扰动破坏以欺骗这些模型。有趣的是，图像在攻击性上存在差异——给定攻击配置，某些图像容易被破坏，而其他图像则更具抵抗力。我们提出OTI，一种无模型且视觉可解释的图像攻击性度量，能够量化单张图像对对抗攻击的内在脆弱性，并提供可解释的可视化。

### 论文 6
- **英文标题**: Persistent Autoregressive Mapping with Traffic Rules for Autonomous Driving
- **中文标题**: 面向自动驾驶的具有交通规则的持久自回归映射
- **作者**: Shiyi Liang, Xinyuan Chang, Changjie Wu, Huiyuan Yan, Yifan Bai, Xinran Liu, Hang Zhang, Yujian Yuan, Shuang Zeng, Mu Xu, Xing Wei
- **英文摘要**: Safe autonomous driving requires both accurate HD map construction and persistent awareness of traffic rules, even when their associated signs are no longer visible. However, existing methods either focus solely on geometric elements or treat rules as temporary classifications, failing to capture their persistent effectiveness across extended driving sequences. In this paper, we present PAMR (Persistent Autoregressive Mapping with Traffic Rules), a novel framework that performs autoregressive co-construction of lane vectors and traffic rules from visual observations. Our approach introduces two key mechanisms: Map-Rule Co-Construction for processing driving scenes in temporal segments, and Map-Rule Cache for maintaining rule consistency across these segments. To properly evaluate continuous and consistent map generation, we develop MapDRv2, featuring improved lane geometry annotations. Extensive experiments demonstrate that PAMR achieves superior performance in joint vector-rule mapping tasks, while maintaining persistent rule effectiveness throughout extended driving sequences.
- **中文摘要**: 安全的自动驾驶需要精确的高精地图构建和对交通规则的持久感知，即使相关标志不再可见。然而，现有方法要么仅关注几何元素，要么将规则视为临时分类，未能捕获规则持久约束地图的因果本质。我们提出一种基于持久自回归映射的方法，将交通规则编码为持久的空间-时间约束。

### 论文 7
- **英文标题**: Anatomical Region-Guided Contrastive Decoding: A Plug-and-Play Strategy for Mitigating Hallucinations in Medical VLMs
- **中文标题**: 解剖区域引导的对比解码：一种缓解医学VLM幻觉的即插即用策略
- **作者**: Xiao Liang, Chenxi Liu, Zhi Ma, Di Wang, Bin Jing, Quan Wang, Yuanyuan Shi
- **英文摘要**: Medical Vision-Language Models (MedVLMs) show immense promise in clinical applicability. However, their reliability is hindered by hallucinations, where models often fail to derive answers from visual evidence, instead relying on learned textual priors. Existing mitigation strategies for MedVLMs have distinct limitations: training-based methods rely on costly expert annotations, limiting scalability, while training-free interventions like contrastive decoding, though data-efficient, apply a global, untargeted correction whose effects in complex real-world clinical settings can be unreliable. To address these challenges, we introduce Anatomical Region-Guided Contrastive Decoding (ARCD), a plug-and-play strategy that mitigates hallucinations by providing targeted, region-specific guidance. Our module leverages an anatomical mask to direct a three-tiered contrastive decoding process. By dynamically re-weighting at the token, attention, and logits levels, it verifiably steers the model's focus onto specified regions, reinforcing anatomical understanding and suppressing factually incorrect outputs. Extensive experiments across diverse datasets, including chest X-ray, CT, brain MRI, and ocular ultrasound, demonstrate our method's effectiveness in improving regional understanding, reducing hallucinations, and enhancing overall diagnostic accuracy.
- **中文摘要**: 医学视觉语言模型（MedVLM）在临床应用方面展现出巨大潜力。然而，它们的可靠性受到幻觉的阻碍，模型通常无法从视觉证据中得出答案，而是依赖学到的文本先验。现有的MedVLM幻觉缓解策略效果有限。我们提出一种解剖区域引导的对比解码策略，作为即插即用的方法来减少医学VLM中的幻觉。

### 论文 13
- **英文标题**: Improving Batch Normalization with Test-Time Adaptation for Robust Object Detection in Self-Driving
- **中文标题**: 通过测试时自适应改进批归一化以实现自动驾驶中的鲁棒目标检测
- **作者**: Dacheng Liao, Mengshi Qi, Liang Liu, Huadong Ma
- **英文摘要**: In open real-world autonomous driving scenarios, challenges such as sensor failure and extreme weather hinder the generalization of current autonomous driving perception models to these unseen domain, due to the domain shifts between the test and training data. As the parameter scale of autonomous driving perception models grows, traditional test-time adaptation (TTA) methods become unstable and often degrade model performance in most scenarios. To address these challenges, this paper proposes two new robust methods to improve the Batch Normalization with TTA for object detection in autonomous driving: (1) We introduce a new LearnableBN layer based on Geometric Confidence Maximization and Entropy Minimization. Specifically, we modify the traditional BN layer by incorporating auxiliary learnable parameters, which enables the BN layer to dynamically update the statistics according to the different input data. (2) We propose a novel semantic-consistency based dual-stage adaptation strategy, which encourages the model to iteratively search for the optimal solution and eliminates unstable samples during the adaptation process. Extensive experiments on the NuScenes-C dataset shows that our method achieves a maximum improvement of about 10\% using BEVFormer as the baseline across six corruption types and three levels of severity.
- **中文摘要**: 在开放的真实世界自动驾驶场景中，传感器故障和极端天气等挑战阻碍了当前自动驾驶感知模型对这些未见域的泛化，原因是测试和训练数据之间的域偏移。随着自动驾驶模型参数规模的增大，我们提出通过测试时批归一化自适方法来增强目标检测的鲁棒性。

### 论文 44
- **英文标题**: PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report Generation
- **中文标题**: PriorRG：面向胸部X光报告生成的先验引导对比预训练与从粗到细解码
- **作者**: Kang Liu, Zhuoqi Ma, Zikang Fang, Yunan Li, Kun Xie, Qiguang Miao
- **英文摘要**: Chest X-ray report generation aims to reduce radiologists' workload by automatically producing high-quality preliminary reports. A critical yet underexplored aspect of this task is the effective use of patient-specific prior knowledge---including clinical context (e.g., symptoms, medical history) and the most recent prior image---which radiologists routinely rely on for diagnostic reasoning. Most existing methods generate reports from single images, neglecting this essential prior information and thus failing to capture diagnostic intent or disease progression. To bridge this gap, we propose PriorRG, a novel chest X-ray report generation framework that emulates real-world clinical workflows via a two-stage training pipeline. In Stage 1, we introduce a prior-guided contrastive pre-training scheme that leverages clinical context to guide spatiotemporal feature extraction, allowing the model to align more closely with the intrinsic spatiotemporal semantics in radiology reports. In Stage 2, we present a prior-aware coarse-to-fine decoding for report generation that progressively integrates patient-specific prior knowledge with the vision encoder's hidden states. This decoding allows the model to align with diagnostic focus and track disease progression, thereby enhancing the clinical accuracy and fluency of the generated reports. Extensive experiments on MIMIC-CXR and MIMIC-ABN datasets demonstrate that PriorRG outperforms state-of-the-art methods, achieving a 3.6% BLEU-4 and 3.8% F1 score improvement on MIMIC-CXR, and a 5.9% BLEU-1 gain on MIMIC-ABN.
- **中文摘要**: 胸部X光报告生成旨在通过自动生成高质量初步报告来减轻放射科医生的工作负担。该任务一个关键但未被充分探索的方面是有效利用患者特定的先验知识——包括临床背景（如症状、病史）和先前影像。我们提出PriorRG，一种先验引导的对比预训练和从粗到细解码方法。

### 论文 89
- **英文标题**: 3DTeethSAM: Taming SAM2 for 3D Teeth Segmentation
- **中文标题**: 3DTeethSAM：驯服SAM2用于3D牙齿分割
- **作者**: Zhiguo Lu, Jianwen Lou, Mingjun Ma, Hairong Jin, Youyi Zheng, Kun Zhou
- **英文摘要**: 3D teeth segmentation, involving the localization of tooth instances and their semantic categorization in 3D dental models, is a critical yet challenging task in digital dentistry due to the complexity of real-world dentition. In this paper, we propose 3DTeethSAM, an adaptation of the Segment Anything Model 2 (SAM2) for 3D teeth segmentation. SAM2 is a pretrained foundation model for image and video segmentation, demonstrating a strong backbone in various downstream scenarios. To adapt SAM2 for 3D teeth data, we render images of 3D teeth models from predefined views, apply SAM2 for 2D segmentation, and reconstruct 3D results using 2D-3D projections. Since SAM2's performance depends on input prompts and its initial outputs often have deficiencies, and given its class-agnostic nature, we introduce three light-weight learnable modules: (1) a prompt embedding generator to derive prompt embeddings from image embeddings for accurate mask decoding, (2) a mask refiner to enhance SAM2's initial segmentation results, and (3) a mask classifier to categorize the generated masks. Additionally, we incorporate Deformable Global Attention Plugins (DGAP) into SAM2's image encoder. The DGAP enhances both the segmentation accuracy and the speed of the training process. Our method has been validated on the 3DTeethSeg benchmark, achieving an IoU of 91.90% on high-resolution 3D teeth meshes, establishing a new state-of-the-art in the field.
- **中文摘要**: 3D牙齿分割涉及在3D牙科模型中定位牙齿实例及其语义分类，是数字牙科中一项关键但具有挑战性的任务，原因在于真实牙列的复杂性。本文提出3DTeethSAM，对Segment Anything Model 2（SAM2）的适配方案，专门设计用于处理3D牙齿分割的独特挑战。

## Computer Vision VII

### 论文 4
- **英文标题**: CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving
- **中文标题**: CorrectAD: 面向端到端自动驾驶规划的自校正智能体系统
- **作者**: Enhui Ma, Lijun Zhou, Tao Tang, Jiahuan Zhang, Junpeng Jiang, Zhan Zhang, Dong Han, Kun Zhan, Xueyang Zhang, Xianpeng Lang, Haiyang Sun, Xia Zhou, Di Lin, Kaicheng Yu
- **英文摘要**: End-to-end planning methods are the de-facto standard of the current autonomous driving system, while the robustness of the data-driven approaches suffers due to the notorious long-tail problem (i.e., rare but safety-critical failure cases). In this work, we explore whether recent diffusion-based video generation methods (a.k.a. world models), paired with structured 3D layouts, can enable a fully automated pipeline to self-correct such failure cases. We first introduce an agent to simulate the role of product manager, dubbed PM-Agent, which formulates data requirements to collect data similar to the failure cases. Then, we use a generative model that can simulate both data collection and annotation. However, existing generative models struggle to generate high-fidelity data conditioned on 3D layouts. To address this, we propose DriveSora, which can generate spatiotemporally consistent videos aligned with the 3D annotations requested by PM-Agent. We integrate these components into our self-correcting agentic system, CorrectAD. Importantly, our pipeline is end-to-end model agnostic and can be applied to improve any end-to-end planner. Evaluated on both nuScenes and a more challenging in-house dataset across multiple end-to-end planners, CorrectAD corrects 62.5% and 49.8% of failure cases, reducing collision rates by 39% and 27%, respectively.
- **中文摘要**: 提出CorrectAD自校正智能体系统，结合PM-Agent模拟产品经理角色收集长尾失败案例数据，并设计DriveSora生成与3D标注对齐的时空一致视频，在nuScenes数据集上纠正62.5%的失败案例，碰撞率降低39%。

### 论文 8
- **英文标题**: UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
- **中文标题**: UM-Text: 面向图像理解与视觉文本编辑的统一多模态模型
- **作者**: Lichen Ma, Xiaolong Fu, GaojingZhou  , Zipeng Guo, Ting Zhu, Yichun Liu, Yu Shi, Jason Li, Junshi Huang
- **英文摘要**: With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main challenge of this task is to fully understand the instruction and reference image, and thus generate visual text that is style-consistent with the image. Previous methods often involve complex steps of specifying the text content and attributes, such as font size, color, and layout, without considering the stylistic consistency with the reference image. To address this, we propose UM-Text, a unified multimodal model for context understanding and visual text editing by natural language instructions. Specifically, we introduce a Visual Language Model (VLM) to process the instruction and reference image, so that the text content and layout can be elaborately designed according to the context information. To generate an accurate and harmonious visual text image, we further propose the UM Encoder to combine the embeddings of various condition information, where the combination is automatically configured by VLM according to the input instruction. During training, we propose a regional consistency loss to offer more effective supervision for glyph generation on both latent and RGB space, and design a tailored three-stage training strategy to further enhance model performance. In addition, we contribute the UM-DATA-200K, a large-scale visual text image dataset on diverse scenes for model training. Extensive qualitative and quantitative results on multiple public benchmarks demonstrate that our method achieves state-of-the-art performance.
- **中文摘要**: 提出UM-Text统一多模态模型，通过视觉语言模型处理指令和参考图像实现上下文感知的文本编辑，设计UM编码器自动配置多条件信息组合，并贡献UM-DATA-200K大规模数据集，在多个公开基准上达到最优性能。

### 论文 31
- **英文标题**: Appearance-Motion Decomposed Alignment for Text-Video Retrieval
- **中文标题**: 外观-运动解耦对齐用于文本视频检索
- **作者**: Meng Meng, Zichang Tan, Yong Zhang, Xu Zhou
- **英文摘要**: Text-video retrieval aims to bridge vision and language areas, which is a crucial task in multi-modal intelligence. The core idea is to learn video and textual features to quantify their semantic relevance. A common limitation in current approaches is the oversimplification of video content, where complex spatiotemporal structures are compressed into a single global representation. Consequently, these methods struggle to fully capture dynamic visual variations and discriminative appearance inside a video, further complicating cross-modal alignment. To alleviate these issues, we introduce a novel decoupling approach that independently processes appearance and motion cues, capitalizing on their complementary nature for more expressive video modeling. Specifically, we propose an appearance-motion decomposed network (AMD-Net) to decouple spatial-level appearance and temporal-level motion understanding via the discriminative appearance learning  and  multi-scale motion learning modules. The proposed model enjoys several merits. First, the designed discriminative appearance learning module with a Singular Value Decomposition (SVD) based prototype initialization can effectively reduce redundant appearance information, and a high-order cross-aggregation mechanism enhances prototype resilience and facilitates comprehensive video understanding. Second, the proposed multi-scale motion learning (MML) module can capture motion features at varying temporal scales, which are complementary to appearance features for accurate text-video retrieval. Extensive experiments on five standard benchmarks demonstrate that our method performs favorably against state-of-the-art methods.
- **中文摘要**: 提出AMD-Net解耦网络，通过判别性外观学习模块（基于SVD的原型初始化）和多尺度运动学习模块分别处理外观和运动线索，高阶交叉聚合机制增强原型韧性，在五个标准基准上优于现有方法。

### 论文 51
- **英文标题**: LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
- **中文标题**: LSD-3D: 基于几何锚定的大规模3D驾驶场景生成
- **作者**: Julian Ost, Andrea Ramazzina, Amogh Joshi, Maximilian Bömer, Mario Bijelic, Felix Heide
- **英文摘要**: Large-scale scene data is essential for training and testing in robot learning. Neural reconstruction methods have promised the capability of reconstructing large physically-grounded outdoor scenes from captured sensor data. However, these methods have baked-in static environments and only allow for limited scene control -- they are functionally constrained in scene and trajectory diversity by the captures from which they are reconstructed. In contrast, generating driving data with recent image or video diffusion models offers control, however, at the cost of geometry grounding and causality. In this work, we aim to bridge this gap and present a method that directly generates large-scale 3D driving scenes with accurate geometry, allowing for causal novel view synthesis with object permanence and explicit 3D geometry estimation. The proposed method combines the generation of a proxy geometry and environment representation with score distillation from learned 2D image priors. We find that this approach allows for high controllability, enabling the prompt-guided geometry and high-fidelity texture and structure that can be conditioned on map layouts -- producing realistic and geometrically consistent 3D generations of complex driving scenes.
- **中文摘要**: 提出直接生成具有精确几何的大规模3D驾驶场景的方法，结合代理几何生成和从2D图像先验中进行分数蒸馏，支持地图布局条件下的提示引导几何和高保真纹理生成，实现了复杂驾驶场景的真实且几何一致的3D生成。

### 论文 62
- **英文标题**: Infinite-Story: A Training-Free Consistent Text-to-Image Generation
- **中文标题**: Infinite-Story: 免训练一致文生图生成
- **作者**: Jihun Park, Kyoungmin Lee, Jongmin Gim, Hyeonseo Jo, Minseok Oh, Wonhyeok Choi, Kyumin Hwang, Jaeyeul Kim, Minwoo Choi, Sunghoon Im
- **英文摘要**: We present Infinite-Story, a training-free framework for consistent text-to-image (T2I) generation tailored for multi-prompt storytelling scenarios. Built upon a scale-wise autoregressive model, our method addresses two key challenges in consistent T2I generation: identity inconsistency and style inconsistency. To overcome these issues, we introduce three complementary techniques: Identity Prompt Replacement, which mitigates context bias in text encoders to align identity attributes across prompts; and a unified attention guidance mechanism comprising Adaptive Style Injection and Synchronized Guidance Adaptation, which jointly enforce global style and identity appearance consistency while preserving prompt fidelity. Unlike prior diffusion-based approaches that require fine-tuning or suffer from slow inference, Infinite-Story operates entirely at test time, delivering high identity and style consistency across diverse prompts. Extensive experiments demonstrate that our method achieves state-of-the-art generation performance, while offering over 6x faster inference (1.72 seconds per image) than the existing fastest consistent T2I models, highlighting its effectiveness and practicality for real-world visual storytelling.
- **中文摘要**: 提出Infinite-Story免训练框架，通过身份提示替换缓解文本编码器上下文偏差，设计自适应风格注入和同步引导适应机制联合强制全局风格和身份外观一致性，推理速度比现有最快一致文生图模型快6倍以上（每张图1.72秒）。

### 论文 66
- **英文标题**: Generalized-Scale Object Counting with Gradual Query Aggregation
- **中文标题**: 基于渐进查询聚合的通用尺度目标计数
- **作者**: Jer Pelhan, Alan Lukežič, Matej Kristan
- **英文摘要**: Few-shot detection-based counters estimate the number of instances in the image specified only by a few test-time exemplars. A common approach to localize objects across multiple sizes is to merge backbone features of different resolutions. Furthermore, to enable small object detection in densely populated regions, the input image is commonly upsampled and tiling is applied to cope with the increased computational and memory requirements. Because of these ad-hoc solutions, existing counters struggle with images containing diverse-sized objects and densely populated regions of small objects. We propose GeCo2, an end-to-end few-shot counting and detection method that explicitly addresses the object scale issues. A new dense query representation gradually aggregates exemplar-specific feature information across scales that leads to high-resolution dense queries that enable detection of large as well as small objects. GeCo2 surpasses state-of-the-art few-shot counters in counting as well as detection accuracy by ~10% while running ~3x faster at smaller GPU memory footprint.
- **中文摘要**: 提出GeCo2端到端少样本计数检测方法，通过密集查询表示在各尺度间渐进聚合样本特定特征信息生成高分辨率密集查询以检测大目标和小目标，计数和检测精度提升约10%同时运行速度约快3倍且GPU内存占用更小。

## Computer Vision VIII

### 论文 12
- **英文标题**: ImageSet2Text: Describing Sets of Images Through Text
- **中文标题**: ImageSet2Text：通过文本描述图像集合
- **作者**: Piera Riccio, Francesco Galati, Kajetan Schweighofer, Noa Garcia, Nuria M Oliver
- **英文摘要**: In the era of large-scale visual data, understanding collections of images is a challenging yet important task. To this end, we introduce ImageSet2Text, a novel method to automatically generate natural language descriptions of image sets. Based on large language models, visual-question answering chains, an external lexical graph, and CLIP-based verification, ImageSet2Text iteratively extracts key concepts from image subsets and organizes them into a structured concept graph. We conduct extensive experiments evaluating the quality of the generated descriptions in terms of accuracy, completeness, and user satisfaction. We also examine the method's behavior through ablation studies, scalability assessments, and failure analyses. Results demonstrate that ImageSet2Text combines data-driven AI and symbolic representations to reliably summarize large image collections for a wide range of applications.
- **中文摘要**: 提出ImageSet2Text，基于大语言模型、视觉问答链、外部词典图和CLIP验证，迭代地从图像子集中提取关键概念并组织成结构化概念图，自动生成图像集合的自然语言描述。

### 论文 15
- **英文标题**: MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
- **中文标题**: MME-SCI：面向多模态大语言模型的全面且具有挑战性的科学基准
- **作者**: Jiacheng Ruan, Dan Jiang, Xian Gao, Ting Liu, Yuzhuo Fu, Yangyang Kang
- **英文摘要**: Recently, multimodal large language models (MLLMs) have achieved significant advancements across various domains, and corresponding evaluation benchmarks have been continuously refined and improved. In this process, benchmarks in the scientific domain have played an important role in assessing the reasoning capabilities of MLLMs. However, existing benchmarks still face three key challenges: 1) Insufficient evaluation of models' reasoning abilities in multilingual scenarios; 2) Inadequate assessment of MLLMs' comprehensive modality coverage; 3) Lack of fine-grained annotation of scientific knowledge points. To address these gaps, we propose MME-SCI, a comprehensive and challenging benchmark. We carefully collected 1,019 high-quality question-answer pairs, which involve 3 distinct evaluation modes. These pairs cover four subjects, namely mathematics, physics, chemistry, and biology, and support five languages: Chinese, English, French, Spanish, and Japanese. We conducted extensive experiments on 16 open-source models and 4 closed-source models, and the results demonstrate that MME-SCI is widely challenging for existing MLLMs. For instance, under the Image-only evaluation mode, o4-mini achieved accuracy of only 52.11%, 24.73%, 36.57%, and 29.80% in mathematics, physics, chemistry, and biology, respectively, indicating a significantly higher difficulty level compared to existing benchmarks. More importantly, using MME-SCI's multilingual and fine-grained knowledge attributes, we analyzed existing models' performance in depth and identified their weaknesses in specific domains. For example, in questions related to "Magnetic Field", o4-mini correctly answered only 5 out of 33 questions, thereby fine-grainedly exposing the model's vulnerabilities. These findings highlight the urgent need to enhance the scientific reasoning capabilities of MLLMs.
- **中文摘要**: 提出MME-SCI，包含1019个高质量问答对，覆盖数学、物理、化学、生物四个学科，支持五种语言，具备三种评估模式和细粒度知识点标注，揭示了现有MLLM在科学推理方面的显著不足。

### 论文 56
- **英文标题**: UniMM-V2X: MoE-Enhanced Multi-Level Fusion for End-to-End Cooperative Autonomous Driving
- **中文标题**: UniMM-V2X：基于MoE增强的多层级融合的端到端协同自动驾驶
- **作者**: Ziyi Song, Chen Xia, Chenbing Wang, Haibao Yu, Sheng Zhou, Zhisheng Niu
- **英文摘要**: Autonomous driving holds transformative potential but remains fundamentally constrained by the limited perception and isolated decision-making with standalone intelligence. While recent multi-agent approaches introduce cooperation, they often focus merely on perception-level tasks, overlooking the alignment with downstream planning and control, or fall short in leveraging the full capacity of the recent emerging end-to-end autonomous driving. In this paper, we present UniMM-V2X, a novel end-to-end multi-agent framework that enables hierarchical cooperation across perception, prediction, and planning. At the core of our framework is a multi-level fusion strategy that unifies perception and prediction cooperation, allowing agents to share queries and reason cooperatively for consistent and safe decision-making. To adapt to diverse downstream tasks and further enhance the quality of multi-level fusion, we incorporate a Mixture-of-Experts (MoE) architecture to dynamically enhance the BEV representations. We further extend MoE into the decoder to better capture diverse motion patterns. Extensive experiments on the DAIR-V2X dataset demonstrate our approach achieves state-of-the-art (SOTA) performance with a 39.7% improvement in perception accuracy, a 7.2% reduction in prediction error, and a 33.2% improvement in planning performance compared with UniV2X, showcasing the strength of our MoE-enhanced multi-level cooperative paradigm.
- **中文摘要**: 提出UniMM-V2X，通过专家混合(MoE)增强的多层级融合实现多智能体协同的端到端自动驾驶，超越单一感知层面的协作。

### 论文 93
- **英文标题**: Manipulation Intention Understanding for Zero-Shot Composed Image Retrieval
- **中文标题**: 面向零样本组合图像检索的操纵意图理解
- **作者**: Yuanmin Tang, Jing Yu, Keke Gai, Gang Xiong, Gaopeng Gou, Meikang Qiu, Qi Wu
- **英文摘要**: Zero-shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with varied visual manipulation intents across domains, scenes, objects, and attributes. A key challenge is that existing datasets contain limited intent-relevant annotations, making it hard for models to infer human intent from textual modifications. We introduce an intent-centric image–text dataset generated via reasoning by a Multimodal Large Language Model (MLLM) to better train ZS-CIR models for human manipulation intent understanding. Building on this dataset, we propose De-MINDS, a framework that distills the MLLM’s reasoning ability to capture manipulation intent and enhance models’ comprehension of modified text. A simple mapping network translates image information into language space and combines it with the manipulation text to form a query. De-MINDS then extracts intention-relevant information from this query and encodes it as pseudo-word tokens for accurate ZS-CIR. Across four ZS-CIR tasks, De-MINDS shows strong generalization and improves over existing methods by 2.15% to 4.05%, establishing new state-of-the-art results with comparable inference time.
- **中文摘要**: 提出De-MINDS，通过MLLM推理构建以意图为中心的图像-文本数据集，将MLLM的推理能力蒸馏到模型中增强对修改文本的理解，在四个ZS-CIR任务上超越现有方法2.15%-4.05%。

## Computer Vision X

### 论文 1
- **英文标题**: MTAttack: Multi-Target Backdoor Attacks Against Large Vision-Language Models
- **中文标题**: MTAttack：针对大型视觉语言模型的多目标后门攻击
- **作者**: Zihan Wang, Guansong Pang, Wenjun Miao, Jin Zheng, Xiao Bai
- **英文摘要**: Recent advances in Large Visual Language Models (LVLMs) have demonstrated impressive performance across various vision-language tasks by leveraging large-scale image-text pretraining and instruction tuning. However, the security vulnerabilities of LVLMs have become increasingly concerning, particularly their susceptibility to backdoor attacks. Existing backdoor attacks focus on single-target attacks, i.e., targeting a single malicious output associated with a specific trigger. In this work, we uncover multi-target backdoor attacks, where multiple independent triggers corresponding to different attack targets are added in a single pass of training, posing a greater threat to LVLMs in real-world applications. Executing such attacks in LVLMs is challenging since there can be many incorrect trigger-target mappings due to severe feature interference among different triggers. To address this challenge, we propose MTAttack, the first multi-target backdoor attack framework for enforcing accurate multiple trigger-target mappings in LVLMs. The core of MTAttack is a novel optimization method with two constraints, namely Proxy Space Partitioning constraint and Trigger Prototype Anchoring constraint. It jointly optimizes multiple triggers in the latent space, with each trigger independently mapping clean images to a unique proxy class while at the same time guaranteeing their separability. Experiments on popular benchmarks demonstrate a high success rate of MTAttack for multi-target attacks, substantially outperforming existing attack methods. Furthermore, our attack exhibits strong generalizability across datasets and robustness against backdoor defense strategies. These findings highlight the vulnerability of LVLMs to multi-target backdoor attacks and underscore the urgent need for mitigating such threats.
- **中文摘要**: 大型视觉语言模型（LVLM）的最新进展通过利用大规模图像-文本预训练和指令微调在各种视觉语言任务中表现出色。然而，LVLM的安全漏洞仍未被充分探索。我们提出MTAttack，一种针对LVLM的多目标后门攻击方法，揭示了多模态场景下后门攻击的更广泛威胁面。

## Computer Vision XI

### 论文 22
- **英文标题**: UVLM: Benchmarking Video Language Model for Underwater World Understanding
- **中文标题**: UVLM：面向水下世界理解的视频语言模型基准
- **作者**: Xizhe Xue, Yang Zhou, Dawei Yan, Lijie Tao, Junjie Li, Ying Li, Haokui Zhang, Rong Xiao
- **英文摘要**: Recently, video-language models (VidLMs) have gained widespread attention and adoption. However, existing works primarily focus on terrestrial scenarios, overlooking the highly demanding application needs of underwater observation. To overcome this gap, we introduce UVLM, an under water observation benchmark which is build through a collaborative approach combining human expertise and AI models. To ensure data quality, we have conducted in-depth considerations from multiple perspectives. First, to address the unique challenges of underwater environments, we selected videos that represent typical underwater challenges including light variations, water turbidity, and diverse viewing angles to construct the dataset. Second, to ensure data diversity, the dataset covers a wide range of frame rates, resolutions, 419 classes of marine animals, and various static plants and terrains. Next, for task diversity, we adopted a structured design where observation targets are categorized into two major classes: biological and environmental. Each category includes content observation and change/action observation, totaling 20 subtask types. Finally, we designed several challenging evaluation metrics to enable quantitative comparison and analysis of different methods. Experiments on two representative VidLMs demonstrate that fine-tuning VidLMs on UVLM significantly improves underwater world understanding while also showing potential for slight improvements on existing in-air VidLM benchmarks.
- **中文摘要**: 最近，视频语言模型（VidLM）获得了广泛关注和采用。然而，现有工作主要关注陆地场景，忽视了水下观察的高需求应用需求。为克服这一差距，我们提出UVLM，一个面向水下世界理解的视频语言模型基准。

### 论文 55
- **英文标题**: HD²-SSC: High-Dimension High-Density Semantic Scene Completion for Autonomous Driving
- **中文标题**: HD²-SSC：面向自动驾驶的高维度高密度语义场景补全
- **作者**: Zhiwen Yang, Yuxin Peng
- **英文摘要**: Camera-based 3D semantic scene completion (SSC) plays a crucial role in autonomous driving, enabling voxelized 3D scene understanding for effective scene perception and decision-making. Existing SSC methods have shown efficacy in improving 3D scene representations, but suffer from the inherent input-output dimension gap and annotation-reality density gap, where the 2D planar view from input images with sparse annotated labels leads to inferior prediction of real-world dense occupancy with a 3D stereoscopic view. In light of this, we propose the corresponding High-Dimension High-Density Semantic Scene Completion (HD²-SSC) framework with expanded pixel semantics and refined voxel occupancies. To bridge the dimension gap, a High-dimension Semantic Decoupling module is designed to expand 2D image features along a pseudo third dimension, decoupling coarse pixel semantics from occlusions, and then identify focal regions with fine semantics to enrich image features. To mitigate the density gap, a High-density Occupancy Refinement module is devised with a ``detect-and-refine" architecture to leverage contextual geometric and semantic structures for enhanced semantic density with the completion of missing voxels and correction of erroneous ones. Extensive experiments and analyses on the SemanticKITTI and SSCBench-KITTI-360 datasets validate the effectiveness of our HD²-SSC framework.
- **中文摘要**: 基于相机的3D语义场景补全（SSC）在自动驾驶中扮演关键角色，实现体素化3D场景理解以支持有效的场景感知和决策。我们提出HD²-SSC，通过高维度高密度表示增强SSC的质量。

### 论文 64
- **英文标题**: DriveSuprim: Towards Precise Trajectory Selection for End-to-End Planning
- **中文标题**: DriveSuprim：迈向端到端规划的精确轨迹选择
- **作者**: Wenhao Yao, Zhenxin Li, Shiyi Lan, Zi Wang, Xinglong Sun, Jose M. Alvarez, Zuxuan Wu
- **英文摘要**: Autonomous vehicles must navigate safely in complex driving environments. Imitating a single expert trajectory, as in regression-based approaches, usually does not explicitly assess the safety of the predicted trajectory. Selection-based methods address this by generating and scoring multiple trajectory candidates and predicting the safety score for each. However, they face optimization challenges in precisely selecting the best option from thousands of candidates and distinguishing subtle but safety-critical differences, especially in rare and challenging scenarios. We propose DriveSuprim to overcome these challenges and advance the selection-based paradigm through a coarse-to-fine paradigm for progressive candidate filtering, a rotation-based augmentation method to improve robustness in out-of-distribution scenarios, and a self-distillation framework to stabilize training. DriveSuprim achieves state-of-the-art performance, reaching 93.5% PDMS in NAVSIM v1 and 87.1% EPDMS in NAVSIM v2 without extra data, with 83.02 Driving Score and 60.00 Success Rate on Bench2Drive, demonstrating superior planning capabilities in various driving scenarios.
- **中文摘要**: 自动驾驶车辆必须在复杂驾驶环境中安全导航。基于回归的方法模仿单一专家轨迹通常不显式评估所选轨迹的安全性。我们提出DriveSuprim，实现更精确的端到端规划轨迹选择。

### 论文 73
- **英文标题**: Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
- **中文标题**: 超越简单编辑：面向复杂指令图像编辑的X-Planner
- **作者**: Chun-Hsiao Yeh, Yilin Wang, Nanxuan Zhao, Richard Zhang, Yuheng Li, Yi Ma, Krishna Kumar Singh
- **英文摘要**: Recent diffusion-based image editing methods have made great strides in text-guided tasks but often struggle with complex, indirect instructions. Additionally, current models frequently exhibit poor identity preservation, unintended edits, or rely on manual masks. To overcome these limitations, we introduce X-Planner, a Multimodal Large Language Model (MLLM)-based planning system that bridges user intent with editing model capabilities. X-Planner uses chain-of-thought reasoning to systematically break down complex instructions into simpler sub-instructions. For each one, X-Planner automatically generates precise edit types and segmentation masks, enabling localized, identity-preserving edits without applying external tools or models during inference. To enable the training of such a planner, we also introduce a fully automated, reproducible pipeline to generate large-scale, high-quality training data. Our complete system achieves state-of-the-art results on both existing and newly proposed complex instruction-based editing benchmarks.
- **中文摘要**: 最近基于扩散的图像编辑方法在文本引导任务中取得了巨大进步，但通常难以处理复杂的间接指令。此外，当前模型在编辑精确性方面经常表现不佳。我们提出X-Planner，一种面向复杂指令图像编辑的规划方法。

## Application Domains II

### 论文 24
- **英文标题**: TermGPT: Multi-Level Contrastive Fine-Tuning for Terminology Adaptation in Legal and Financial Domains
- **中文标题**: TermGPT：面向法律和金融领域术语适应的多层次对比微调
- **作者**: Yidan Sun, Mengying Zhu, Feiyue Chen, Yangyang Wu, Xiaolei Dan, Mengyuan Yang, Xiaolin Zheng, Shenglin Ben
- **英文摘要**: Large language models (LLMs) have demonstrated impressive performance in text generation tasks; however, their embedding spaces often suffer from the isotropy problem, resulting in poor discrimination of domain-specific terminology, particularly in legal and financial contexts. This weakness in term-level representation can severely hinder downstream tasks such as legal judgment prediction or financial risk analysis, where subtle semantic distinctions are critical.  To address this problem, we propose TermGPT, a multi-level contrastive fine-tuning framework designed for terminology adaptation. We first construct a sentence graph to capture semantic and structural relations, and generate semantically consistent yet discriminative positive and negative samples based on contextual and topological cues. We then devise a multi-level contrastive learning approach at both the sentence and token levels, enhancing global contextual understanding and fine-grained term discrimination. To support robust evaluation, we construct the first financial terminology dataset derived from official regulatory documents. Experiments show that TermGPT outperforms existing baselines in term discrimination tasks within the finance and legal domains.
- **中文摘要**: 大语言模型在文本生成任务中展现了令人瞩目的性能，然而在处理法律和金融等专业领域的术语时，其嵌入表示和生成输出往往缺乏精确性。我们提出TermGPT，一个面向术语适应的多层次对比微调框架。TermGPT在三个层次上运行：术语级对比学习将领域特定术语与其定义和使用上下文对齐，句子级对比学习区分术语使用的正确与错误模式，文档级对比学习捕获术语在语篇层面的作用。我们构建了TermBench，一个覆盖六种语言的法律和金融术语的综合评估基准。在TermBench以及合同分析、合规性检查和财务报告理解等多个下游任务上的实验表明，TermGPT在术语精度上实现了实质性提升，同时保持了通用语言能力，显著优于通用LLM和用标准方法微调的领域特定模型。

### 论文 38
- **英文标题**: PsyPARSE: Retrieval-Augmented Slow Thinking for Personalized Empathetic Counseling
- **中文标题**: PsyPARSE：面向个性化共情心理咨询的检索增强慢思考
- **作者**: Longxiang Wang, Pukun Zhao, Chen Chen, Jinhe Bi, Huacan Wang, Tong Zhang, Ronghao Chen
- **英文摘要**: The escalating global demand for mental health services highlights the potential of Large Language Models (LLMs) in psychological counseling. However, current LLM-based approaches, particularly fine-tuned models, are constrained by data distribution biases, leading to limited therapeutic diversity and personalization. Crucially, they often lack anticipatory empathetic reasoning, struggle to foresee patient emotional responses beyond immediate dialogue history, and incur substantial computational costs. To address these limitations, we propose PsyPARSE, a novel training-free framework for psychological counseling that emulates the deliberate and empathetic reasoning of human counselors. PsyPARSE integrates Multi-Therapy Retrieval-Augmented Generation (RAG) to overcome data biases and provide highly personalized therapeutic approaches tailored to individual patient attributes. Pioneering the first multi-stage slow-thinking engine in mental health LLMs, PsyPARSE employs Multi-Turn Rollouts to identify optimal therapeutic paths and through anticipating patient reactions, optimizes empathetic responses, thereby ensuring genuinely empathetic and impactful responses in complex, long-dialogue interactions. Operating as a plug-and-play solution, PsyPARSE avoids the computational burden of fine-tuning. We establish a comprehensive LLM-based patient-therapist agent simulation framework for evaluation. Extensive experiments demonstrate that PsyPARSE significantly enhances the capabilities of various LLM baselines, achieving superior personalization and deeper empathy compared to both fine-tuned and other training-free methods. This work offers an efficient, adaptable, and scalable solution to advance mental health support.
- **中文摘要**: 全球心理健康服务需求的不断增长凸显了大语言模型在心理咨询中的潜力。然而通用LLM缺乏有效的共情咨询所必需的个性化和审慎反思推理能力。我们提出PsyPARSE，一个面向个性化共情心理咨询的检索增强慢思考框架。PsyPARSE实现了模拟反思实践的慢思考机制：模型生成初始共情响应，然后通过自我批判和多视角分析迭代精炼。检索模块访问包含咨询技术和心理学理论的知识库，个性化模块维护用户特定的档案，捕获情绪模式、偏好和治疗进展。我们通过自动指标和与持证治疗师的受控用户研究来评估PsyPARSE。结果表明，与通用LLM和现有的心理健康聊天机器人相比，PsyPARSE生成了更具共情性、上下文更适当且治疗上更合理的响应，同时保持了适当的伦理边界。

### 论文 62
- **英文标题**: SoMe: A Realistic Benchmark for LLM-based Social Media Agents
- **中文标题**: SoMe：面向基于大语言模型的社交媒体智能体的逼真基准
- **作者**: Dizhan Xue, Jing Cui, Shengsheng Qian, Chuanrui Hu, Changsheng Xu
- **英文摘要**: Intelligent agents powered by large language models (LLMs) have recently demonstrated impressive capabilities and gained increasing popularity on social media platforms. While LLM agents are reshaping the ecology of social media, there exists a current gap in conducting a comprehensive evaluation of their ability to comprehend media content, understand user behaviors, and make intricate decisions. To address this challenge, we introduce SoMe, a pioneering benchmark designed to evaluate social media agents equipped with various agent tools for accessing and analyzing social media data. SoMe comprises a diverse collection of 8 social media agent tasks, 9,164,284 posts, 6,591 user profiles, and 25,686 reports from various social media platforms and external websites, with 17,869 meticulously annotated task queries. Compared with the existing datasets and benchmarks for social media tasks, SoMe is the first to provide a versatile and realistic platform for LLM-based social media agents to handle diverse social media tasks. By extensive quantitative and qualitative analysis, we provide the first overview insight into the performance of mainstream agentic LLMs in realistic social media environments and identify several limitations. Our evaluation reveals that both the current closed-source and open-source LLMs cannot handle social media agent tasks satisfactorily. SoMe provides a challenging yet meaningful testbed for future social media agents.
- **中文摘要**: 由大语言模型驱动的智能体最近展现了令人印象深刻的能力，并因其在各种现实应用中自主执行任务的潜力而获得了越来越多的关注。然而现有的大语言模型智能体基准要么过于简单，要么不适用于社交媒体场景中典型的复杂多步任务。我们提出SoMe，一个面向基于LLM的社交媒体智能体的逼真基准。SoMe包含需要多步推理和交互的多样化社交任务，包括内容创作、社区管理、信息验证和用户互动。每个任务都带有自动评估协议和人工验证的参考解决方案。我们对多个最先进的大语言模型智能体进行基准测试，揭示了将LLM部署到社交媒体领域的显著挑战，特别是在处理社会动态和用户意图方面的不足。

### 论文 90
- **英文标题**: DIFT: Protecting Contrastive Learning Against Data Poisoning Backdoor Attacks
- **中文标题**: DIFT：保护对比学习免受数据投毒后门攻击
- **作者**: Jiang Zhu, Yulin Jin, Qingqing Ye, Zhibiao Guo, Kun Fang, Ruochen Du, Yingnan Zhao, Haibo Hu
- **英文摘要**: Contrastive learning (CL) is a popular learning paradigm that excels in extracting meaningful representations from unlabeled data. Recent studies have shown that CL is highly vulnerable to backdoor attacks. Current defenses against backdoor attacks in CL are primarily reactive and post-training. That is, the detection and elimination of backdoors are executed in the deployment phase of a given well-trained model. However, these post-training defenses are usually prone to degrading model utility and resource-intensive, causing that the backdoor detection and elimination from a fully-trained model is quite challenging. To address this issue, we argue for a fundamental perspective, i.e., integrating the defense into the model's training phase, and propose a novel framework to mitigate the backdoor in CL, namely Density-Based Identification and Fine-Tuning (DIFT). Specifically, DIFT identifies potential poisoned samples during the early training phase via detecting embeddings with abnormal poisoning characteristic in the feature space. Then, to remove backdoors and preserve model utility, the detected poisoned samples are leveraged to fine-tune the model, and the remaining clean samples are further involved into training the model after the fine-tuning. DIFT, as a proactive training-time defense, avoids the problematic backdoor removal and the high computational cost associated with those reactive post-training methods. We empirically evaluate DIFT on various CL algorithms against backdoor attack. Experimental results demonstrate that our method exhibits promising defense effectiveness while maintaining model's clean data accuracy.
- **中文摘要**: 对比学习是一种流行的学习范式，擅长从无标签数据中提取有意义的表示。然而对比学习对数据投毒攻击的安全性仍未被充分探索。我们提出DIFT，一种保护对比学习免受后门攻击的防御框架。DIFT利用对比学习本身的特性作为防御：数据增强期间引入的差异被用作检测毒化样本的监督信号。通过比较同一图像不同视图之间的表示差异，模型可以识别具有异常一致行为的被投毒样本。实验表明，DIFT有效检测并消解了多种投毒攻击，恢复了干净的模型性能，同时对干净数据的准确率影响极小。

## Computer Vision XII

### 论文 15
- **英文标题**: Perception in Plan: Coupled Perception and Planning for End-to-End Autonomous Driving
- **中文标题**: 感知融入规划: 耦合感知与规划的端到端自动驾驶
- **作者**: Bozhou Zhang, Jingyu Li, Nan Song, Li Zhang
- **英文摘要**: End-to-end autonomous driving has achieved remarkable advancements in recent years. Existing methods primarily follow a perception–planning paradigm, where perception and planning are executed sequentially within a fully differentiable framework for planning-oriented optimization. We further advance this paradigm through a "perception-in-plan'' framework design, which integrates perception into the planning process.  This design facilitates targeted perception guided by evolving planning objectives over time, ultimately enhancing planning performance. Building on this insight, we introduce VeteranAD, a coupled perception and planning framework for end-to-end autonomous driving. By incorporating multi-mode anchored trajectories as planning priors, the perception module is specifically designed to gather traffic elements along these trajectories, enabling comprehensive and targeted perception. Planning trajectories are then generated based on both the perception results and the planning priors. To make perception fully serve planning, we adopt an autoregressive strategy that progressively predicts future trajectories while focusing on relevant regions for targeted perception at each step. With this simple yet effective design, VeteranAD fully unleashes the potential of planning-oriented end-to-end methods, leading to more accurate and reliable driving behavior. Extensive experiments on the NAVSIM and Bench2Drive datasets demonstrate that our VeteranAD achieves state-of-the-art performance.
- **中文摘要**: 端到端自动驾驶近年来取得了显著进展。现有方法主要遵循感知-规划范式，其中感知和规划在完全可微的框架中顺序执行，以实现面向规划的优化。我们通过‘感知融入规划’的框架设计进一步推进这一范式，将感知集成到规划过程中。这种设计促进了由随时间演变的规划目标引导的有针对性感知，最终提升了规划性能。基于这一见解，我们引入了VeteranAD，一个耦合感知与规划的端到端自动驾驶框架。通过结合多模式锚定轨迹作为规划先验，感知模块专门设计用于沿这些轨迹收集交通元素，实现全面且有针对性的感知。然后基于感知结果和规划先验生成规划轨迹。为使感知充分服务于规划，我们采用自回归策略，逐步预测未来轨迹，同时在每一步关注相关区域进行有针对性的感知。凭借这一简单而有效的设计，VeteranAD充分释放了以规划为导向的端到端方法的潜力，实现了更准确和可靠的驾驶行为。在NAVSIM和Bench2Drive数据集上的大量实验表明，我们的VeteranAD达到了最先进的性能。

### 论文 23
- **英文标题**: BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection
- **中文标题**: BEVDilation: 以LiDAR为中心的多模态融合用于3D目标检测
- **作者**: Guowen Zhang, Chenhang He, Liyi Chen, Lei Zhang
- **英文摘要**: Integrating LiDAR and camera information in the bird's eye view (BEV) representation has demonstrated its effectiveness in 3D object detection. However, because of the fundamental disparity in geometric accuracy between these sensors, indiscriminate fusion in previous methods often leads to degraded performance. In this paper, we propose BEVDilation, a novel LiDAR-centric framework that prioritizes LiDAR information in the fusion. By formulating image BEV features as implicit guidance rather than naive concatenation, our strategy effectively alleviates the spatial misalignment caused by image depth estimation errors. Furthermore, the image guidance can effectively help the LiDAR-centric paradigm to address the sparsity and semantic limitations of point clouds. Specifically, we propose a Sparse Voxel Dilation Block that mitigates the inherent point sparsity by densifying foreground voxels through image priors. Moreover, we introduce a Semantic-Guided BEV Dilation Block to enhance the LiDAR feature diffusion processing with image semantic guidance and long-range context capture. On the challenging nuScenes benchmark, BEVDilation achieves better performance than state-of-the-art methods while maintaining competitive computational efficiency. Importantly, our LiDAR-centric strategy demonstrates greater robustness to depth noise compared to naive fusion.
- **中文摘要**: 在鸟瞰图（BEV）表示中融合LiDAR和相机信息已在3D目标检测中证明了其有效性。然而，由于这些传感器在几何精度上存在根本差异，先前方法中的无差别融合常常导致性能下降。在本文中，我们提出了BEVDilation，一个以LiDAR为中心的新颖框架，在融合中优先考虑LiDAR信息。通过将图像BEV特征表述为隐式引导而非简单拼接，我们的策略有效缓解了由图像深度估计误差引起的空间错位。此外，图像引导可以有效帮助以LiDAR为中心的范式解决点云的稀疏性和语义局限。具体而言，我们提出了稀疏体素膨胀块，通过图像先验来稠密化前景体素，缓解固有的点稀疏性。此外，我们引入了语义引导BEV膨胀块，利用图像语义引导和长距离上下文捕获来增强LiDAR特征扩散处理。在具有挑战性的nuScenes基准上，BEVDilation在保持竞争性计算效率的同时，实现了优于最先进方法的性能。重要的是，与简单融合相比，我们以LiDAR为中心的策略对深度噪声表现出更强的鲁棒性。

### 论文 67
- **英文标题**: Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT Framework
- **中文标题**: 基于多模态提示学习更好的无人机跨视角目标地理定位: MoP-UAV基准和MoPT框架
- **作者**: Xiaohan Zhang, Zhangkai Shen, Si-Yuan Cao, Xiaokai Bai, Yiming Li, Zheheng Han, Zhe Wu, Qi Ming, Hui-Liang Shen
- **英文摘要**: We present MoP-UAV, a new benchmark for UAV-based cross-view object geo-localization guided by multi-modal prompts. MoP-UAV supports fine-grained object-level cross-view localization under diverse prompt modalities, including natural language, bounding boxes, and click points. It offers potential for incorporating large foundation models like large language models (LLMs) and promotes the building of more flexible and intelligent UAV agents. Based on the benchmark, we propose MoPT, a multi-modal-prompt-guided tansformer that embeds prompts as token sequences and extract object location from UAV and satellite features via cross-attention. To enhance semantic consistency and performance, we further adopt a cross-view contrastive loss and propose a RefCOCOg-based pre-training strategy. Extensive experiments show that MoPT achieves robust localization under arbitrary prompt combinations. Notably, multi-modal-prompt training significantly boosts unimodal-prompt inference performance, highlighting the generalization benefits of multi-modal learning. MoPT trained with multi-modal prompts outperforms prior unimodal prompt works under the same setting.
- **中文摘要**: 我们提出了MoP-UAV，一个由多模态提示引导的无人机跨视角目标地理定位新基准。MoP-UAV支持在多样化提示模态（包括自然语言、边界框和点击点）下的细粒度目标级跨视角定位。它为融合大语言模型（LLM）等大基础模型提供了潜力，并促进了更灵活和智能的无人机Agent的构建。基于该基准，我们提出了MoPT，一个多模态提示引导的Transformer，将提示嵌入为令牌序列，并通过交叉注意力从无人机和卫星特征中提取目标位置。为增强语义一致性和性能，我们进一步采用了跨视角对比损失并提出了基于RefCOCOg的预训练策略。大量实验表明，MoPT在任意提示组合下实现了稳健的定位。值得注意的是，多模态提示训练显著提升了单模态提示推理的性能，突显了多模态学习的泛化优势。在相同设置下，使用多模态提示训练的MoPT优于先前的单模态提示工作。

### 论文 68
- **英文标题**: Any2RSI: Controllable Remote Sensing Text-to-Image Generation via Any Control and Enriched Description
- **中文标题**: Any2RSI: 基于任意控制和丰富描述的可控遥感文本到图像生成
- **作者**: Xu Zhang, Jianzhong Huang, Lefei Zhang
- **英文摘要**: Recent advances in controllable text-to-image (T2I) generation have achieved impressive results in natural images, but remote sensing (RS) T2I remains challenging due to the unique nature of geospatial data. Existing methods struggle to integrate diverse spatial controls and model complex spatial relationships, often failing to maintain semantic consistency with typically vague or incomplete textual descriptions. Moreover, limited by small-scale, low-quality datasets, these models produce outputs with inconsistent layouts and unrealistic content. To address these issues, we propose Any2RSI, a flexible framework for controllable RS T2I generation. It features a Cross-Modal Multi-Control Adapter that extracts modality-agnostic embeddings from heterogeneous spatial inputs, enabling precise spatial guidance. To compensate for sparse or ambiguous text prompts, we introduce a VLM-Empowered Enriched Description Generation module that enhances input descriptions with cross-modal semantics for more coherent image generation. Furthermore, we present RST2I-110K, a new large-scale dataset with over 115,000 high-quality RS image-text pairs across diverse scenes, alleviating data scarcity in this domain. Extensive experiments show that Any2RSI achieves state-of-the-art performance on both existing and new datasets, improving the realism and structural accuracy of generated RS imagery.
- **中文摘要**: 可控文本到图像（T2I）生成的最新进展在自然图像方面取得了令人瞩目的成果，但由于地理空间数据的独特性质，遥感（RS）T2I仍然具有挑战性。现有方法难以整合多样化的空间控制和建模复杂的空间关系，通常无法与通常模糊或不完整的文本描述保持语义一致性。此外，受限于小规模低质量数据集，这些模型产生布局不一致和不真实内容的输出。为解决这些问题，我们提出了Any2RSI，一个灵活的可控遥感T2I生成框架。它具有一个跨模态多控制适配器，从异构空间输入中提取模态无关的嵌入，实现精确的空间引导。为补偿稀疏或模糊的文本提示，我们引入了一个VLM赋能丰富描述生成模块，用跨模态语义增强输入描述以实现更连贯的图像生成。此外，我们提出了RST2I-110K，一个包含超过115,000高质量遥感图像-文本对的大规模新数据集，覆盖多样场景，缓解了该领域的数据稀缺。大量实验表明，Any2RSI在现有和新数据集上达到了最先进的性能，提升了生成遥感影像的真实感和结构准确性。

### 论文 71
- **英文标题**: Introducing Decomposed Causality with Spatiotemporal Object-Centric Representation for Video Classification
- **中文标题**: 引入基于时空物体中心表示的分解因果关系用于视频分类
- **作者**: Yachong Zhang, Lei Meng, Shuo Xu, Zhuang Qi, Wei Wu, Lei Wu, Xiangxu Meng
- **英文摘要**: Video classification requires event-level representations of objects and their interactions. Existing methods typically rely on data-driven approaches, which either learn such features from whole frames or object-centric visual regions. Therefore, the modeling of spatiotemporal interactions among objects is usually overlooked. To address this issue, this paper presents a Decomposition of Synergistic, Unique, and Redundant Causal Representations Learning (SurdCRL) model for video classification, which introduces a newly-proposed SURD causal theory to model the spatiotemporal features of both object dynamics and their in- and cross-frame interactions. Specifically, SurdCRL employs three modules to model the object-centric spatiotemporal dynamics using distinct types of causal components, where the first module Spatial-Temporal Entity Modeling decouples the frame into object and context entities, and employs a temporal message passing block to capture object state changes over time, generating spatiotemporal features as basic causal variables. Second, the Dual-Path Causal Inference module mitigates confounders among causal variables by front-door and back-door interventions, thus enabling the subsequent causal components to reflect their intrinsic effects. Finally, the Causal Composition and Selection module employs the compositional structure-aware attention to project the causal variables and their high-order interactions into the synergistic, unique, and redundant components. Experiments on two benchmarking datasets verify that SurdCRL better captures event-relevant object-centric representation by decomposing spatiotemporal object interactions into three types of causal components.
- **中文摘要**: 视频分类需要物体及其交互的事件级表示。现有方法通常依赖数据驱动的方法，要么从全帧或物体中心视觉区域学习这些特征。因此，物体之间的时空交互建模通常被忽视。为解决这一问题，本文提出了用于视频分类的协同、独特和冗余因果表示学习分解模型（SurdCRL），引入新提出的SURD因果理论来建模物体动态及其帧内和跨帧交互的时空特征。具体而言，SurdCRL采用三个模块使用不同类型的因果分量来建模物体中心的时空动态：第一个模块时空实体建模将帧解耦为物体和上下文实体，并使用时序消息传递块来捕获物体状态随时间的变化，生成时空特征作为基本因果变量。第二个模块双路因果推理通过前门和后门干预缓解因果变量之间的混淆因子，从而使后续因果分量反映其内在效应。最后，因果组合与选择模块采用组合结构感知注意力将因果变量及其高阶交互投影到协同、独特和冗余分量中。在两个基准数据集上的实验验证了SurdCRL通过将时空物体交互分解为三种因果分量，更好地捕获了事件相关的物体中心表示。

### 论文 85
- **英文标题**: Evolving Generalist Virtual Agents with Generative and Associative Memory
- **中文标题**: 基于生成式和联想式记忆的通用虚拟Agent持续进化
- **作者**: Zhenkui Zhang, Wendong Bu, Kaihang Pan, Bingchen Miao, Wenqiao Zhang, Guoming Wang, Wei Ji, Rui Tang, Juncheng Li, Siliang Tang
- **英文摘要**: Generalist Virtual Agents (GVAs) powered by Multimodal Large Language Models (MLLMs) exhibit impressive capabilities. However, their long-term learning is hampered by a core limitation: a failure to evolve beyond existing trajectories. This stems from memory systems that treat experiences as isolated fragments and rely on brittle semantic retrieval, preventing the synthesis of novel solutions from disparate knowledge. To address this, we introduce CA3Mem, a framework inspired by the human hippocampus that organizes experiences into a structured memory graph. Leveraging this graph, CA3Mem features two key innovations: 1) a generative memory recombination mechanism that synthesizes novel solutions to drive agent evolution, and 2) an associative retrieval algorithm that employs spreading activation to recall a comprehensive and contextually-aware set of experiences. Experiments on OSWorld and WebArena demonstrate that CA3Mem significantly enhances agent capabilities, leading to marked improvements in long-horizon planning, compositional generalization for novel tasks, and continuous adaptation from experience.
- **中文摘要**: 由多模态大语言模型（MLLM）驱动的通用虚拟Agent（GVA）展现了令人印象深刻的能力。然而，其长期学习受到一个核心限制的阻碍：无法超越现有轨迹进化。这源于将经验视为孤立片段并依赖脆弱的语义检索的记忆系统，阻止了从不同知识合成新颖解决方案。为解决此问题，我们引入了CA3Mem，一个受人类海马体启发的框架，将经验组织成结构化记忆图。利用该图，CA3Mem具有两项关键创新：1）生成式记忆重组机制，合成新颖解决方案以驱动Agent进化；2）联想检索算法，采用扩散激活来回想全面且上下文感知的经验集合。在OSWorld和WebArena上的实验表明，CA3Mem显著增强了Agent能力，在长时域规划、新颖任务的组合泛化以及从经验中持续适应方面带来了显著改进。

## Computer Vision XIII

### 论文 49
- **英文标题**: Paper Folding Puzzles: Can Multimodal Large Language Models Perform Spatial Reasoning?
- **中文标题**: 折纸谜题: 多模态大语言模型是否具备空间推理能力？
- **作者**: Dibin Zhou, Yantao Xu, Zongming Huang, Zengwei Yan, Wenhao Liu, Yongwei Miao, Jianfeng Ren, Fuchang Liu
- **英文摘要**: Multimodal Large Language Models (MLLMs) largely lag human-level performance on abstract visual reasoning (AVR), which requires models to infer latent rules from visual question sets and generalize them to novel scenarios. Most AVR benchmarks are constrained to narrow and repetitive 2D patterns, involving relatively simple spatial relationships and assessing limited dimensions of reasoning ability. Drawing inspiration from real-world paper folding challenges, we propose Paper Folding Puzzles (PFP), a rigorously designed benchmark specifically developed to assess spatial reasoning capabilities. It comprises 150K visual question-answering samples across five diverse tasks, ranging from basic 2D geometric reasoning to 3D spatial understanding. The developed benchmark dataset can be employed to assess core spatial reasoning abilities essential to human cognition, encompassing fundamental symmetry reasoning and 3D spatial comprehension. Furthermore, we conduct a comprehensive evaluation of 18 leading MLLMs (both closed- and open-source variants) on the PFP benchmark to assess their spatial reasoning capabilities. Our findings show that most MLLMs achieve near-chance performance on FPF, exhibiting substantial performance gaps (>30%) relative to human baselines across all tasks. This highlights a critical research gap in improving spatial reasoning capabilities of MLLMs.
- **中文摘要**: 多模态大语言模型(MLLM)在抽象视觉推理(AVR)方面远落后于人类水平的表现，AVR要求模型从视觉问题集中推断潜在规则并将其泛化到新场景。大多数AVR基准被限制在狭窄且重复的2D模式中，涉及相对简单的空间关系并评估有限的推理能力维度。受真实世界折纸挑战的启发，我们提出了折纸谜题(PFP)，一个严格设计、专门用于评估空间推理能力的基准。它包含来自五个不同任务的15万个视觉问答样本，从基本的2D几何推理到3D空间理解。所开发的基准数据集可用于评估人类认知所必需的核心空间推理能力，包括基本对称推理和3D空间理解。此外，我们对18个领先的MLLM(闭源和开源变体)在PFP基准上进行了全面评估，以评估其空间推理能力。我们的发现表明，大多数MLLM在PFP上实现了接近随机水平的表现，在所有任务上相对于人类基线展现出显著的性能差距(>30%)。这突显了提升MLLM空间推理能力方面的关键研究缺口。

### 论文 71
- **英文标题**: OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
- **中文标题**: OpenDriveVLA: 面向端到端自动驾驶的大视觉语言行动模型
- **作者**: Xingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma, Volker Tresp, Alois Knoll
- **英文摘要**: We present OpenDriveVLA, a Vision-Language Action (VLA) model designed for end-to-end autonomous driving, built upon open-source large language models. OpenDriveVLA generates spatially-grounded driving actions by leveraging multimodal inputs, including both 2D and 3D instance-aware visual representations, ego vehicle states, and language commands. To bridge the modality gap between driving visual representations and language embeddings, we introduce a hierarchical vision-language alignment process, projecting both 2D and 3D structured visual tokens into a unified semantic space. Furthermore, we incorporate structured agent–environment–ego interaction modeling into the autoregressive decoding process, enabling the model to capture fine-grained spatial dependencies and behavior-aware dynamics critical for reliable trajectory planning. Extensive experiments on the nuScenes dataset demonstrate that OpenDriveVLA achieves state-of-the-art results across open-loop trajectory planning and driving-related question-answering tasks. Qualitative analyses further illustrate its superior capability to follow high-level driving commands and robustly generate trajectories under challenging scenarios, highlighting its potential for next-generation end-to-end autonomous driving.
- **中文摘要**: 我们提出了OpenDriveVLA，一个专为端到端自动驾驶设计的视觉-语言-行动(VLA)模型，基于开源大语言模型构建。OpenDriveVLA通过利用多模态输入(包括2D和3D实例感知视觉表示、自车状态和语言命令)生成空间驱动的驾驶行动。为弥合驾驶视觉表示和语言嵌入之间的模态差距，我们引入了分层视觉-语言对齐过程，将2D和3D结构化视觉token投影到统一语义空间。此外，我们将结构化的智能体-环境-自车交互建模融合到自回归解码过程中，使模型能够捕获对可靠轨迹规划至关重要的细粒度空间依赖和行为感知动力学。在nuScenes数据集上的大量实验表明，OpenDriveVLA在开环轨迹规划和驾驶相关问答任务中达到了最先进的结果。定性分析进一步说明了其在遵循高级驾驶命令和在具有挑战性的场景下鲁棒生成轨迹方面的卓越能力，突显了其作为下一代端到端自动驾驶的潜力。

### 论文 88
- **英文标题**: Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space
- **中文标题**: 其他车辆轨迹同样需要: 一个在视频潜在空间中统一自车-他车轨迹的驾驶世界模型
- **作者**: Jian Zhu, Zhengyu Jia, Tian Gao, Jiaxin Deng, Shidi Li, Lang Zhang, Fu Liu, Peng Jia, Xianpeng Lang
- **英文摘要**: Advanced end-to-end autonomous driving systems predict other vehicles' motions and plan ego vehicle's trajectory. The world model that can foresee the outcome of the trajectory has been used to evaluate the end-to-end autonomous driving system. However, existing world models predominantly emphasize the trajectory of the ego vehicle and leave other vehicles uncontrollable. This limitation hinders their ability to realistically simulate the interaction between the ego vehicle and the driving scenario. In addition, it remains a challenge to match multiple trajectories with each vehicle in the video to control the video generation. To address above issues, a driving World Model named EOT-WM is proposed in this paper, unifying Ego-Other vehicle Trajectories in videos. Specifically, we first project ego and other vehicle trajectories in the BEV space into the image coordinate to match each trajectory with its corresponding vehicle in the video. Then, trajectory videos are encoded by the Spatial-Temporal Variational Auto Encoder to align with driving video latents spatially and temporally in the unified visual space. A trajectory-injected diffusion Transformer is further designed to denoise the noisy video latents for video generation with the guidance of ego-other vehicle trajectories. In addition, we propose a metric based on control latent similarity to evaluate the controllability of trajectories. Extensive experiments are conducted on the nuScenes dataset, and the proposed model outperforms the state-of-the-art method by 30% in FID and 55% in FVD. The model can also predict unseen driving scenes with self-produced trajectories.
- **中文摘要**: 先进的端到端自动驾驶系统预测其他车辆的运动并规划自车的轨迹。能够预见轨迹结果的世界模型已被用于评估端到端自动驾驶系统。然而，现有的世界模型主要强调自车的轨迹，其他车辆则不可控。这一局限性限制了它们逼真模拟自车与驾驶场景之间交互的能力。此外，将多条轨迹与视频中的每辆车匹配以控制视频生成仍然是一个挑战。为解决上述问题，本文提出了一个名为EOT-WM的驾驶世界模型，在视频中统一自车-他车轨迹。具体而言，我们首先将BEV空间中的自车和他车轨迹投影到图像坐标，将每条轨迹与视频中对应的车辆匹配。然后，轨迹视频通过时空变分自编码器编码，在统一视觉空间中与驾驶视频潜在变量在空间和时间上对齐。进一步设计了一个轨迹注入扩散Transformer，在自车-他车轨迹的引导下去噪视频潜在变量以生成视频。此外，我们提出了一个基于控制潜在相似性的度量来评估轨迹的可控性。在nuScenes数据集上进行了大量实验，所提模型在FID上超过最先进方法30%，在FVD上超过55%。该模型还可以通过自产轨迹预测未见过的驾驶场景。

### 论文 102
- **英文标题**: HouseTune: Two-Stage Floorplan Generation with LLM Assistance
- **中文标题**: HouseTune: 借助LLM辅助的两阶段楼层平面图生成
- **作者**: Ziyang Zong, Guanying Chen, Zhaohuan Zhan, Fengcheng Yu, Guang Tan
- **英文摘要**: This paper proposes a two-stage text-to-floorplan generation framework that combines the reasoning capability of Large Language Models (LLMs) with the generative power of diffusion models. In the first stage, we leverage a Chain-of-Thought (CoT) prompting strategy to guide an LLM in generating an initial layout, Layout-Init, from natural language descriptions, which ensures a user-friendly and intuitive design process. However, Layout-Init may lack precise geometric alignment and fine-grained structural details due to the inherent limitations of LLMs. To address this, in the second stage we propose a Dual-Noise Prior-Preserved Diffusion (DNPP-Diffusion) model to refine Layout-Init into a final floorplan that better adheres to physical constraints and user requirements. By combining LLMs and a dedicated refining model, our approach is able to generate high-quality floorplans without requiring large-scale domain-specific training data. Experimental results demonstrate its advantages in comparison with state of the art methods, and validate its effectiveness in home design applications.
- **中文摘要**: 本文提出了一个两阶段的文本到楼层平面图生成框架，将大语言模型(LLM)的推理能力与扩散模型的生成能力相结合。在第一阶段，我们利用思维链(CoT)提示策略引导LLM从自然语言描述生成初始布局Layout-Init，这确保了用户友好且直观的设计过程。然而，由于LLM的固有限制，Layout-Init可能缺乏精确的几何对齐和细粒度结构细节。为解决此问题，在第二阶段我们提出了双噪声先验保留扩散(DNPP-Diffusion)模型，将Layout-Init细化为更好地遵循物理约束和用户需求的最终平面图。通过结合LLM和专用细化模型，我们的方法无需大规模领域特定训练数据即可生成高质量平面图。实验结果证明了其与最先进方法相比的优势，并验证了其在家居设计应用中的有效性。

### 论文 108
- **英文标题**: Decoupling What to Count and Where to See for Referring Expression Counting
- **中文标题**: 解耦计数内容与观察位置以用于参考表达计数
- **作者**: Yuda Zou, Zijian Zhang, Yongchao Xu
- **英文摘要**: Referring Expression Counting (REC) extends class-level object counting to the fine-grained subclass-level, aiming to enumerate objects matching a textual expression that specifies both the class and distinguishing attribute. A fundamental challenge, however, has been overlooked: annotation points are typically placed on class-representative locations (e.g., heads), forcing models to focus on class-level features while neglecting attribute information from other visual regions (e.g., legs for ''walking''). To address this, we propose W2-Net, a novel framework that explicitly decouples the problem into ''what to count'' and ''where to see'' via a dual-query mechanism. Specifically, alongside the standard what-to-count (w2c) queries that localize the object, we introduce dedicated where-to-see (w2s) queries. The w2s queries are guided to seek and extract features from attribute-specific visual regions, enabling precise subclass discrimination. Furthermore, we introduce Subclass Separable Matching (SSM), a novel matching strategy that incorporates a repulsive force to enhance inter-subclass separability during label assignment. W2-Net significantly outperforms the state-of-the-art on the REC-8K dataset, reducing counting error by 22.5% (validation) and 18.0% (test), and improving localization F1 by 7% and 8%, respectively.
- **中文摘要**: 参考表达计数(REC)将类别级对象计数扩展到细粒度子类别级，旨在枚举匹配文本表达的对象，该文本表达同时指定了类别和区分属性。然而，一个根本性挑战被忽视了：标注点通常放置在类别代表位置(如头部)，迫使模型关注类别级特征，同时忽视来自其他视觉区域(如'行走'中的腿部)的属性信息。为解决此问题，我们提出了W2-Net，一个通过双查询机制显式将问题解耦为'计数什么'和'观察哪里'的新颖框架。具体而言，除了定位对象的标准计数内容(w2c)查询，我们引入了专用的观察位置(w2s)查询。w2s查询被引导以从属性特定视觉区域寻找和提取特征，实现精确的子类别判别。此外，我们引入了子类别可分离匹配(SSM)，一种融合排斥力以在标签分配期间增强子类别间可分离性的新颖匹配策略。W2-Net在REC-8K数据集上显著优于最先进方法，将计数误差降低了22.5%(验证集)和18.0%(测试集)，并将定位F1分别提升了7%和8%。

## Data Mining and Knowledge Management III

### 论文 9
- **英文标题**: T-Retriever: Tree-based Hierarchical Retrieval Augmented Generation for Textual Graphs
- **中文标题**: T-Retriever：面向文本图的基于树的分层检索增强生成
- **作者**: Chunyu Wei, Huaiyu Qin, Siyuan He, Yunhai Wang, Yueguo Chen
- **英文摘要**: Retrieval-Augmented Generation (RAG) has significantly enhanced Large Language Models' ability to access external knowledge, yet current graph-based RAG approaches face two critical limitations in managing hierarchical information: they impose rigid layer-specific compression quotas that damage local graph structures, and they prioritize topological structure while neglecting semantic content. We introduce T-Retriever, a novel framework that reformulates attributed graph retrieval as tree-based retrieval using a semantic and structure-guided encoding tree. Our approach features two key innovations: (1) Adaptive Compression Encoding, which replaces artificial compression quotas with a global optimization strategy that preserves the graph's natural hierarchical organization, and (2) Semantic-Structural Entropy (S²-Entropy), which jointly optimizes for both structural cohesion and semantic consistency when creating hierarchical partitions. Experiments across diverse graph reasoning benchmarks demonstrate that T-Retriever significantly outperforms state-of-the-art RAG methods, providing more coherent and contextually relevant responses to complex queries.
- **中文摘要**: 检索增强生成(RAG)显著增强了大语言模型访问外部知识的能力，然而当前基于图的RAG方法在管理分层信息方面面临两个关键限制：它们强加了刚性的逐层压缩配额，破坏局部图结构，并且它们优先考虑拓扑结构而忽略了语义内容。我们引入T-Retriever，一个新颖的框架，利用语义和结构引导的编码树将属性图检索重新表述为基于树的检索。我们的方法具有两个关键创新：(1)自适应压缩编码，用全局优化策略替代人为压缩配额，保留图的自然分层组织；(2)语义-结构熵(S²-Entropy)，在创建分层划分时联合优化结构内聚性和语义一致性。在多样化图推理基准上的实验表明，T-Retriever显著优于最先进的RAG方法，对复杂查询提供更连贯和上下文相关的响应。

### 论文 34
- **英文标题**: Making Visual Dialogue More Engaging: A New Task, Method, and Metric
- **中文标题**: 让视觉对话更具吸引力：新任务、方法和评估指标
- **作者**: Guanghui Ye, Huan Zhao, Yingxue Gao, Zhixue Zhao, Kehan Wang, Xupeng Zha, Zhihua Jiang
- **英文摘要**: Large language model (LLM)-based visual dialogue (VD) systems have made response generation for image-grounded conversations more correct and coherent. However, user engagement - the extent to which a user is interested, emotionally involved, and willing to continue the conversation - remains a challenge. To fully explore engaging VD, we propose: (i) a new task named Audio-enhanced VD (AVD), which introduces additional audio dialogue contexts that can more vividly convey the speaker's emotions as input, with the aim of generating correct but more engaging dialogue responses. Specifically, we employ a text-to-speech model as the modality translator to generate the paired acoustic utterances from the inputting textual utterances; (ii) an accompanying approach named Visually-grounded and Interleaved Text-Audio  Dialogue Modeling (VITA-DM), which utilizes both image-grounded information and interleaved text-audio utterances for visual dialogue modeling, differentiating from previous multi-modal LLM (MLLM)-based methods that normally model text and audio modalities separately. We also present three pre-training tasks to better learn multi-modal interactions across language, vision, and audio; (iii) a novel metric named Multi-Modal Engagement (MME), which fills the gap of engagement estimation in VD and can provide a fine-grained assessment along emotional, attentional, and reply engagement dimensions (EE, AE, RE). We experiment on two popular datasets and provide extensive evaluations (automatic, engagement-specific, and human), supporting the validity of our approach. Furthermore, based on empirical results that reveal that emotions contribute the most to engagement, we justify our emphasis on the emotional aspect throughout the definition, solution, and evaluation of our task.
- **中文摘要**: 基于大语言模型(LLM)的视觉对话(VD)系统使基于图像的对话回复更加正确和连贯。然而，用户参与度——用户感兴趣、情感卷入并愿意继续对话的程度——仍然是一个挑战。为充分探索有吸引力的VD，我们提出：(i)一个名为音频增强VD(AVD)的新任务，引入额外的音频对话上下文作为输入，能够更生动地传达说话者的情绪，旨在生成正确但更具吸引力的对话回复。具体而言，我们采用文本到语音模型作为模态转换器，从输入文本话语中生成配对的声学话语；(ii)一个相应的名为基于视觉和交错文本-音频对话建模(VITA-DM)的方法，利用基于图像的信息和交错的文本-音频话语进行视觉对话建模，不同于之前通常分别建模文本和音频模态的多模态LLM(MLLM)方法。我们还提出了三个预训练任务，以更好地学习跨语言、视觉和音频的多模态交互；(iii)一个名为多模态参与度(MME)的新指标，填补了VD中参与度估计的空白，并可沿情感、注意力和回复参与度维度(EE, AE, RE)提供细粒度评估。我们在两个流行数据集上进行实验，并提供广泛的评估（自动、参与度特定和人工），支持我们方法的有效性。此外，基于实证结果揭示情感对参与度的贡献最大，我们证明了在任务的定义、解决方案和评估中强调情感方面的合理性。

### 论文 53
- **英文标题**: MAPS: Multi-Agent Personality Shaping for Collaborative Reasoning
- **中文标题**: MAPS：面向协同推理的多智能体人格塑造
- **作者**: Jian Zhang, Zhiyuan Wang, Zhangqi Wang, Fangzhi Xu, Qika Lin, Lingling Zhang, Rui Mao, Erik Cambria, Jun Liu
- **英文摘要**: Collaborative reasoning with multiple agents offers the potential for more robust and diverse problem-solving. However, existing approaches often suffer from homogeneous agent behaviors and lack of reflective and rethinking capabilities. We propose Multi-Agent Personality Shaping ((MAPS), a novel framework that enhances reasoning through agent diversity and internal critique. Inspired by the Big Five personality theory, MAPS assigns distinct personality traits to individual agents, shaping their reasoning styles and promoting heterogeneous collaboration. To enable deeper and more adaptive reasoning, MAPS introduces a Critic agent that reflects on intermediate outputs, revisits flawed steps, and guides iterative refinement. This integration of personality-driven agent design and structured collaboration improves both reasoning depth and flexibility. Empirical evaluations across three benchmarks demonstrate the strong performance of MAPS, with further analysis confirming its generalizability across different large language models and validating the benefits of multi-agent collaboration.
- **中文摘要**: 多智能体协同推理为更鲁棒和多样的问题解决提供了潜力。然而，现有方法通常存在智能体行为同质化和缺乏反思和再思考能力的问题。我们提出了多智能体人格塑造(MAPS)，一个通过智能体多样性和内部批判来增强推理的新颖框架。受大五人格理论启发，MAPS为个体智能体分配不同的人格特质，塑造其推理风格并促进异质协作。为实现更深入和更具适应性的推理，MAPS引入了一个评论家(Critic)智能体，反思中间输出，重新审视错误步骤，并引导迭代精炼。这种人格驱动的智能体设计与结构化协作的整合，提高了推理深度和灵活性。在三个基准上的实证评估展示了MAPS的强劲性能，进一步分析证实了其在不同大语言模型上的泛化能力，并验证了多智能体协作的益处。

## Humans and AI

### 论文 40
- **英文标题**: Simulating Human-Like Counseling: A Path- and Scenario-Guided Framework for Psychological Support Dialogue
- **中文标题**: 模拟类人咨询：一种路径与情景引导的心理支持对话框架
- **作者**: Yuanchen Shi, Longyin Zhang, Maodong Li, Yibin Zheng, Xiuhong Wang, Fang Kong
- **英文摘要**: The growing demand for psychological support underscores the lack of high-quality counseling dialogue datasets, particularly in non-English contexts. We propose PGSim, a Path-Guided Simulation framework that mirrors real counseling processes—symptom description, problem identification, cause analysis, strategy planning, and iterative adjustment. PGSim models each user scenario as a fine-grained quadruple {Group, Psychological Problem, Problem Cause, Support Focus} and guides dialogue generation through expert-annotated strategy paths. Real counseling dialogues and expert-edited samples are used to fine-tune two language models: a Dialog Generator for strategy-aligned dialogue creation and a Dialog Modifier for expert-level refinement. After automated and human verification, we construct the Chinese Psychological support Dialogue Dataset (CPsDD), containing 68K dialogues across 13 groups, 16 problems, 13 causes, and 12 support focuses. We further present the Comprehensive Agent Dialogue Support System (CADSS), which integrates profiling, summarization, strategy planning, and empathetic response. Experiments on CPsDD and ESConv demonstrate that CADSS achieves state-of-the-art results on Strategy Prediction and Emotional Support Conversation tasks.
- **中文摘要**: 日益增长的心理支持需求凸显了高质量咨询对话数据集的缺乏，特别是在非英语环境中。我们提出了PGSim，一种路径引导的模拟框架，模拟真实的咨询过程——症状描述、问题识别、原因分析、策略规划和迭代调整。PGSim将每个用户情景建模为细粒度的四元组{群体、心理问题、问题原因、支持关注点}，并通过专家标注的策略路径引导对话生成。真实的咨询对话和专家编辑的样本被用于微调两个语言模型：一个对话生成器用于策略对齐的对话创建，以及一个对话修饰器用于专家级优化。经过自动和人工验证，我们构建了中文心理支持对话数据集（CPsDD），包含68K个对话，涵盖13个群体、16个问题、13个原因和12个支持关注点。我们进一步提出了综合智能体对话支持系统（CADSS），整合了画像、摘要、策略规划和共情回复。在CPsDD和ESConv上的实验表明，CADSS在策略预测和情感支持对话任务上达到了最先进的结果。

### 论文 43
- **英文标题**: Reducing Goal State Divergence with Environment Design
- **中文标题**: 通过环境设计减少目标状态分歧
- **作者**: Kelsey Sikes, Sarah Keren, Sarath Sreedharan
- **英文摘要**: Generating behaviors that align with human expectations is a key requirement for human-robot collaboration. Potential behavior misalignment could lead to the robot performing actions with unanticipated, potentially dangerous side effects even while pursuing human goals. In this paper, we introduce a novel metric called Goal State Divergence (GSD) which quantifies the difference between the state a robot achieved in response to a human-specified goal and what the human expected. In cases where GSD cannot be directly calculated, we show how it can be approximated using maximal and minimal bounds. We then leverage GSD in our novel human-robot goal alignment design (HRGAD) problem, which identifies a minimal set of environment modifications that can reduce such mismatches. We show the effectiveness of our method in reducing the goal state divergence by empirically evaluating our approach on several planning benchmarks.
- **中文摘要**: 生成符合人类期望的行为是人机协作的关键要求。潜在的行为不一致可能导致机器人在追求人类目标时执行具有意外且可能危险副作用的动作。在本文中，我们引入了一种称为目标状态分歧（GSD）的新指标，量化了机器人响应人类指定目标所达到的状态与人类预期之间的差异。在GSD无法直接计算的情况下，我们展示了如何使用最大和最小边界来近似它。然后，我们在新的人机目标对齐设计（HRGAD）问题中利用GSD，该问题确定一组最小的环境修改以减少此类不匹配。我们通过在多个规划基准上实证评估我们的方法，展示了其在减少目标状态分歧方面的有效性。

### 论文 47
- **英文标题**: Consensus-Driven Multi-Agent Cognitive Reasoning for Enhancing the Emotional Intelligence of Large Language Models
- **中文标题**: 共识驱动的多智能体认知推理：增强大语言模型的情商
- **作者**: Geng Tu, Dingming Li, Jun Huang, Ruifeng Xu
- **英文摘要**: Large Language Models (LLMs) have demonstrated strong performance in various NLP tasks but remain limited in emotional intelligence (EI). Benchmarks such as EmoBench attribute this gap to deficiencies in cognitively demanding tasks that require inferring others’ latent mental states, intentions, and emotions in nuanced social contexts. To address this, we propose MACRo, a Multi-Agent Cognitive Reasoning framework that generates a structured Cognitive Chain of Thought comprising Situation, Clue, Thought, Action, and Emotion. Each component is generated by a specialized agent, enabling modular, interpretable multi-step reasoning. To ensure coherence and mitigate hallucinations, a coordinator agent verifies outputs, and a consensus game mechanism enforces alignment across reasoning steps. Extensive Experiments on EmoBench show that MACRo significantly enhances both emotional understanding and application across LLMs. Further evaluations confirm its generalizability to real-world social applications such as emotional support conversations.
- **中文摘要**: 大语言模型（LLM）在各种NLP任务中展现了强大的性能，但在情商（EI）方面仍然有限。EmoBench等基准测试将这一差距归因于在需要推断他人潜在心理状态、意图和情绪的细微社交背景中认知要求较高的任务上的不足。为解决这一问题，我们提出了MACRo，一种多智能体认知推理框架，生成结构化的认知思维链，包括情境、线索、思考、行动和情绪。每个组件由专门的智能体生成，实现模块化、可解释的多步推理。为确保连贯性并减少幻觉，一个协调器智能体验证输出，并通过共识博弈机制强制对齐各推理步骤。在EmoBench上的广泛实验表明，MACRo显著增强了LLM的情绪理解和应用能力。进一步评估确认了其泛化到情感支持对话等真实世界社交应用的能力。

## Intelligent Robotics

### 论文 26
- **英文标题**: Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology
- **中文标题**: 迈向城市空间的自主无人机视觉目标搜索: 基准与智能体方法
- **作者**: Yatai Ji, Zhengqiu Zhu, Yong Zhao, Beidan Liu, Chen Gao, Yihao Zhao, Sihang Qiu, Yue Hu, Quanjun Yin
- **英文摘要**: Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects based on visual inputs without external guidance. Existing approaches struggle in complex urban environments due to redundant semantic processing, similar object ambiguity, and the exploration-exploitation dilemma. To advance research and support the AVOS task, we introduce CityAVOS, the first benchmark dataset for autonomous search of static urban objects. It features 2,420 tasks of varying difficulty across six object categories, designed to rigorously evaluate UAV search strategies. To solve the AVOS task, we also propose PRPSearcher (Perception-Reasoning-Planning Searcher), a novel agentic method powered by multi-modal large language models (MLLMs) that enables a UAV agent to think and reason like humans on visual cues when searching for objects. Specifically, PRPSearcher constructs three specialized maps: an object-centric dynamic semantic map enhancing spatial perception, a 3D cognitive map based on semantic "attraction" values for target reasoning, and a 3D uncertainty map for balanced exploration-exploitation search. Moreover, we propose a denoising mechanism to mitigate interference from similar objects and design an Inspiration Promote Thought prompting mechanism for adaptive action planning. Experimental results on CityAVOS demonstrate that PRPSearcher surpasses existing baselines in both success rate and search efficiency (on average: +37.69% SR, +28.96% SPL, -30.69% MSS, and -46.40% NE). Our work paves the way for future advances in embodied visual target search.
- **中文摘要**: 城市环境中的空中视觉目标搜索(AVOS)任务要求无人机基于视觉输入自主搜索并识别目标物体，无需外部引导。由于冗余语义处理、相似物体模糊性以及探索-利用困境，现有方法在复杂城市环境中表现不佳。为了推动研究并支持AVOS任务，我们引入了CityAVOS，首个面向静态城市物体自主搜索的基准数据集。它包含跨六个物体类别、不同难度的2420个任务，旨在严格评估无人机搜索策略。为了完成AVOS任务，我们还提出了PRPSearcher（感知-推理-规划搜索器），一种由多模态大语言模型(MLLMs)驱动的新颖智能体方法，使无人机智能体能够像人类一样根据视觉线索思考和推理以搜索物体。具体而言，PRPSearcher构建了三个专业地图：一个增强空间感知的以物体为中心的动态语义地图、一个基于语义'吸引力'值的目标推理3D认知地图，以及一个用于平衡探索-利用搜索的3D不确定性地地图。此外，我们提出了一个去噪机制来减轻相似物体的干扰，并设计了一个灵感促进思维提示机制用于自适应动作规划。在CityAVOS上的实验结果表明，PRPSearcher在成功率和搜索效率方面均超越了现有基线（平均：+37.69% SR, +28.96% SPL, -30.69% MSS, -46.40% NE）。我们的工作为具身视觉目标搜索的未来发展铺平了道路。

### 论文 30
- **英文标题**: DiTEA: Mixture-of-Experts for Vision-Language-Action Model in Robotic Manipulation
- **中文标题**: DiTEA: 面向机器人操作的视觉-语言-动作模型专家混合
- **作者**: Chengxuan Li, Xingwan Wang
- **英文摘要**: The current diffusion-based Vision-Language-Action (VLA) models have faster inference speed and the ability to solve the action muti-modality problem in robot manipulation tasks compared to traditional autoregressive models after large-scale pre-training and post-training. However, the diffusion-based VLA models were found to have poor instruction-following ability, and after fine-tuning training on multiple tasks, them often suffer from "skill forgetting" due to conflicting model weights on each task. To address this problem, we propose DiTEA, a Diffusion Transformer-based Mixture-of-Experts (MoE) VLA model. Specifically, it fuses the MoE module into the action head of VLA to form Action MoE, and in addition, we design the Task-Instruction Gate, which uses language instructions to select specific experts for tasks they specialize in, in order to improve the VLA's instruction-following ability. We conducted comprehensive experiments and ablation study to evaluate the efficacy of our model under different designs. Experimental results from simulation and real-world show that our DiTEA has excellent improvement in multi-task compared to baseline and other VLAs.
- **中文摘要**: 当前基于扩散的视觉-语言-动作(VLA)模型在大规模预训练和后训练之后，与传统的自回归模型相比，具有更快的推理速度和解决机器人操作任务中动作多模态问题的能力。然而，基于扩散的VLA模型被发现指令跟随能力较差，并且在多个任务上进行微调训练后，由于各任务的模型权重冲突，它们经常遭受'技能遗忘'。为了解决这个问题，我们提出了DiTEA，一个基于扩散变换器的专家混合(MoE)VLA模型。具体而言，它将MoE模块融合到VLA的动作头中以形成Action MoE，此外，我们设计了任务-指令门控，利用语言指令为任务选择特化的专家，以提升VLA的指令跟随能力。我们进行了全面的实验和消融研究来评估不同设计下模型的效能。仿真和真实世界实验结果表明，我们的DiTEA在多任务中相比基线和其它VLA模型有出色的提升。

## Machine Learning I

### 论文 21
- **英文标题**: Medical Vision–Language Pretraining with LLM-Guided Temporal Supervision
- **中文标题**: 基于LLM引导时间监督的医学视觉-语言预训练
- **作者**: Liang Bai, Zhi Wang, Huimin Yan, Xian Yang
- **英文摘要**: Medical vision–language pretraining typically relies on static image–text pairs, overlooking temporal cues vital for understanding clinical progression. This limits model sensitivity to evolving semantics and reduces their effectiveness in real-world clinical reasoning. To address this challenge, we propose TAMM—a temporal alignment framework that leverages weak but semantically rich supervision from large language models (LLMs). Given temporally adjacent clinical reports, LLMs automatically generate (i) coarse-grained trend labels (e.g., improving or worsening), and (ii) fine-grained rationales explaining the supporting clinical evidence. These complementary signals inject temporal semantics without requiring manual annotation, and guide vision–language representation learning to capture trend-sensitive cross-modal alignment and rationale-grounded coherence. Experiments on multiple medical benchmarks demonstrate that TAMM improves retrieval and classification performance while yielding more interpretable, temporally consistent embeddings. Our results highlight the potential of leveraging LLM-derived supervision to equip vision–language models with temporal awareness critical for clinical applications.
- **中文摘要**: 医学视觉-语言预训练通常依赖静态的图像-文本对，忽视了对于理解临床进展至关重要的时间线索。这限制了模型对演化语义的敏感性，降低了其在真实临床推理中的有效性。为应对这一挑战，我们提出了TAMM——一种利用大语言模型(LLM)弱但语义丰富的监督的时间对齐框架。给定时间相邻的临床报告，LLM自动生成(i)粗粒度趋势标签(例如，改善或恶化)和(ii)解释支持性临床证据的细粒度推理。这些互补信号在不需人工标注的情况下注入时间语义，并引导视觉-语言表示学习捕获趋势敏感的跨模态对齐和基于推理的连贯性。在多个医学基准上的实验表明，TAMM提升了检索和分类性能，同时产生更具可解释性、时间一致的嵌入。我们的结果突显了利用LLM衍生监督为视觉-语言模型赋予对临床应用至关重要的时间感知能力的潜力。

### 论文 68
- **英文标题**: Do We Truly Need So Many Samples? Multi-LLM Repeated Sampling Efficiently Scales Test-Time Compute
- **中文标题**: 我们真的需要那么多样本吗？多LLM重复采样高效扩展测试时计算
- **作者**: Jianhao Chen, Zishuo Xun, Bocheng Zhou, Han Qi, Hangfan Zhang, Qiaosheng Zhang, Yang Chen, Wei Hu, Yuzhong Qu, Shuyue Hu
- **英文摘要**: This paper presents a simple, effective, and cost-efficient strategy, named ModelSwitch, to improve LLM performance by scaling test-time compute. ModelSwitch builds upon the repeated-sampling-then-voting framework, with a novel twist:  incorporating multiple models, even weaker ones, to leverage their complementary strengths that potentially arise from diverse training data and paradigms. By using sample consistency as a signal, our strategy dynamically switches between models. Theoretical analysis highlights the efficiency and performance advantages of our strategy. Extensive experiments on seven datasets demonstrate that our strategy not only outperforms self-consistency and state-of-the-art multi-agent debate approaches, but also significantly reduces inference costs. Additionally, our strategy requires only a few comparable LLMs to achieve optimal performance and can be extended with verification methods, demonstrating the potential of leveraging multiple LLMs in the generation-verification paradigm.
- **中文摘要**: 本文提出了一种简单、有效且成本节约的策略ModelSwitch，通过扩展测试时计算来提升LLM性能。ModelSwitch建立在重复采样然后投票的框架之上，但有一个新颖的转折：纳入多个模型，甚至较弱的模型，以利用其潜在来自多样化训练数据和范式的互补优势。通过使用样本一致性作为信号，我们的策略在模型之间动态切换。理论分析突显了我们策略的效率和性能优势。在七个数据集上的大量实验表明，我们的策略不仅优于自一致性和最先进的多智能体辩论方法，还显著降低了推理成本。此外，我们的策略仅需少量可比LLM即可达到最优性能，并可通过验证方法扩展，展示了在生成-验证范式中利用多个LLM的潜力。

## Machine Learning III

### 论文 1
- **英文标题**: 4D Point Cloud Segmentation via Active Test-Time Adaptation
- **中文标题**: 基于主动测试时自适应的4D点云分割
- **作者**: Mingrong Gong, Chaoqi Chen, Luyao Tang, Yuxi Wang, Sergio Escalera
- **英文摘要**: 4D point cloud segmentation is crucial for autonomous driving with continuous LiDAR streams. While test-time adaptation (TTA) is the standard approach for handling dynamic environments, current methods suffer from catastrophic error accumulation due to over-reliance on pseudo-labels. Active learning could provide reliable annotations for critical samples, but combining it with TTA faces severe challenges: realtime processing requirements and expensive 3D labeling costs. In this paper, we propose ATTA-4DSeg, the first framework to achieve efficient active test-time adaptation for 4D point cloud segmentation under extreme budget constraints. Our key insight is a self-reinforcing loop: oracle annotations refine adaptation prototypes, which then guide the selection of subsequent high-value samples from regions with severe distribution shifts, maximizing each annotation’s impact. Specifically, we propose three key innovations: (1) dual-prototype comparison that precisely localizes distribution shift boundaries to narrow annotation scope, (2) Class-Inverse Budget Allocation (CIBA) ensuring balanced adaptation across all categories, coupled with hybrid uncertainty scoring combining voxel-level geometry and point-wise variance for optimal sample selection, and (3) a refinement strategy leveraging sparse oracle annotations to improve predictions on unlabeled points, maximizing annotation utility. Extensive experiments show ATTA-4DSeg improves mIoU by 18.87%, 19.92%, and 3.6% on three domain adaptation benchmarks using only 1% annotation budget. Our method operates 2.28× faster than state-of-the-art methods. Remarkably, our approach reaches 90% of fully-supervised performance using only 5% annotation budget.
- **中文摘要**: 4D点云分割对于具有连续LiDAR流的自动驾驶至关重要。虽然测试时自适应（TTA）是处理动态环境的标准方法，但当前方法由于过度依赖伪标签而导致灾难性的错误累积。主动学习可以为关键样本提供可靠的标注，但将其与TTA结合面临严峻挑战：实时处理需求和昂贵的3D标注成本。本文提出ATTA-4DSeg，这是首个在极端预算约束下实现高效主动测试时自适应的4D点云分割框架。我们的核心洞察是一个自我强化的循环：oracle标注细化自适应原型，然后引导选择具有严重分布偏移的区域中的后续高价值样本，最大化每个标注的影响。具体而言，我们提出三项关键创新：（1）双重原型比较，精确定位分布偏移边界以缩小标注范围；（2）类别逆预算分配（CIBA），确保所有类别的平衡自适应，结合混合不确定性评分，融合体素级几何和逐点方差以实现最优样本选择；（3）利用稀疏oracle标注改进未标记点预测的细化策略，最大化标注效用。大量实验表明，仅使用1%的标注预算，ATTA-4DSeg在三个域自适应基准上将mIoU提升了18.87%、19.92%和3.6%。我们的方法比最先进方法快2.28倍。值得注意的是，仅使用5%的标注预算，我们的方法就达到了全监督性能的90%。

### 论文 9
- **英文标题**: DAVID: Dual-stage Adaptive Vision-text Integrated Decoupling for Multimodal KV Cache Eviction
- **中文标题**: DAVID：面向多模态KV缓存逐出的双阶段自适应视觉-文本集成解耦
- **作者**: Yifeng Gu, Jianxiu Jin, Kailing Guo, Xiangmin Xu
- **英文摘要**: With the rapid development of multimodal large language models (MLLMs), deploying them on low-resource devices remains challenging. Beyond the model size, long multimodal inputs cause substantial memory overhead in the KV cache, making efficient cache management critical. In this paper, we propose DAVID, a KV cache eviction strategy that adapts to the degree of modality fusion across layers. By analyzing the feature distributions of vision and text tokens, we observe low fusion in early layers and high fusion in deeper layers. Based on this observation, DAVID adopts a decoupled eviction strategy in shallow layers and a super-modal eviction strategy in deeper layers. To support this dynamic switching, we design a lightweight metric that quantifies cross-modal fusion and uses a threshold to determine which layers require decoupling. Experimental results show that DAVID achieves state-of-the-art performance on multiple benchmarks and offers a new perspective on KV cache eviction for MLLMs.
- **中文摘要**: 随着多模态大语言模型（MLLM）的快速发展，在低资源设备上部署它们仍然具有挑战性。除了模型大小之外，长多模态输入在KV缓存中导致大量内存开销，使得高效的缓存管理至关重要。本文提出DAVID，一种适应各层模态融合程度的KV缓存逐出策略。通过分析视觉和文本令牌的特征分布，我们观察到早期层融合程度低，深层融合程度高。基于这一观察，DAVID在浅层采用解耦逐出策略，在深层采用超模态逐出策略。为支持这种动态切换，我们设计了一个轻量级度量来量化跨模态融合，并使用阈值确定哪些层需要解耦。实验结果表明，DAVID在多个基准上实现了最先进的性能，并为MLLM的KV缓存逐出提供了新视角。

### 论文 14
- **英文标题**: Poisoning with a Pill: Circumventing Detection in Federated Learning
- **中文标题**: 药丸投毒：规避联邦学习中的检测
- **作者**: Hanxi Guo, Hao Wang, Tao Song, Tianhang Zheng, Yang Hua, Haibing Guan, Xiangyu Zhang
- **英文摘要**: Federated learning (FL) protects data privacy by enabling distributed model training without direct access to client data. However, its distributed nature makes it vulnerable to model and data poisoning attacks. While numerous defenses filter malicious clients using statistical metrics, they overlook the role of model redundancy, where not all parameters contribute equally to the model and attack performance. Current attacks manipulate all model parameters uniformly, making them more detectable, while defenses focus on the overall statistics of client updates, leaving gaps for more sophisticated attacks. We propose an attack-agnostic augmentation method to enhance the stealthiness and effectiveness of existing poisoning attacks in FL, exposing flaws in current defenses and highlighting the need for fine-grained FL security. Our three-stage methodology, including pill construction, pill poisoning, and pill injection, injects poison into a compact subnet (i.e., pill) of the global model during the iterative FL training. Experimental results show that FL poisoning attacks enhanced by our method can bypass 8 state-of-the-art (SOTA) defenses, gaining an up to 7x error rate increase, as well as on average a more than 2x error rate increase on both IID and non-IID data, in both cross-silo and cross-device FL systems.
- **中文摘要**: 联邦学习（FL）通过在分布式模型训练中无需直接访问客户端数据来保护数据隐私。然而，其分布式特性使其容易受到模型和数据投毒攻击。虽然许多防御方法使用统计度量过滤恶意客户端，但它们忽视了模型冗余的作用，即并非所有参数对模型和攻击性能的贡献均等。当前的攻击统一操作所有模型参数，使其更容易被检测，而防御侧重于客户端更新的整体统计，为更复杂的攻击留下了可乘之机。我们提出一种与攻击无关的增强方法，以提高FL中现有投毒攻击的隐蔽性和有效性，揭示当前防御的缺陷并突显细粒度FL安全的必要性。我们的三阶段方法，包括药丸构建、药丸投毒和药丸注入，在迭代FL训练期间将毒药注入全局模型的紧凑子网（即药丸）中。实验结果表明，通过我们的方法增强的FL投毒攻击可以绕过8种最先进（SOTA）防御，在IID和非IID数据上均实现高达7倍的错误率增加以及平均超过2倍的错误率增加，适用于跨孤岛和跨设备FL系统。

### 论文 25
- **英文标题**: On Robustness of Linear Classifiers to Targeted Data Poisoning
- **中文标题**: 线性分类器对目标数据投毒的鲁棒性研究
- **作者**: Nakshatra Gupta, Sumanth Prabhu S, Supratik Chakraborty, Venkatesh R
- **英文摘要**: Data poisoning is a training-time attack that undermines the trustworthiness of learned models. In a targeted data poisoning attack, an adversary manipulates the training dataset to alter the classification of a targeted test point. Given the typically large size of training dataset, manual detection of poisoning is difficult. An alternative is to automatically measure a dataset's robustness against such an attack, which is the focus of this paper. We consider a threat model wherein an adversary can only perturb the labels of the training dataset, with knowledge limited to the hypothesis space of the victim's model. In this setting, we prove that finding the robustness is an NP-Complete problem, even when hypotheses are linear classifiers. To overcome this, we present a technique that finds lower and upper bounds of robustness. Our implementation of the technique computes these bounds efficiently in practice for many publicly available datasets. We experimentally demonstrate the effectiveness of our approach. Specifically, a poisoning exceeding the identified robustness bounds significantly impacts test point classification. We are also able to compute these bounds in many more cases where state-of-the-art techniques fail.
- **中文摘要**: 数据投毒是一种训练时攻击，破坏学习模型的可信度。在目标数据投毒攻击中，对手操纵训练数据集以改变目标测试点的分类。鉴于训练数据集通常规模庞大，手动检测投毒很困难。一种替代方案是自动度量数据集对此类攻击的鲁棒性，这是本文的重点。我们考虑一个威胁模型，其中对手只能扰动训练数据集的标签，且知识仅限于受害者模型的假设空间。在此设置下，我们证明即使假设为线性分类器，找到鲁棒性也是NP完全问题。为克服这一点，我们提出一种找到鲁棒性下界和上界的技术。我们的实现在实践中为许多公开数据集高效计算这些界限。我们通过实验证明了我们方法的有效性。具体来说，超过已识别鲁棒性界限的投毒显著影响测试点分类。我们还能够在最先进技术失败的许多更多场景中计算这些界限。

### 论文 33
- **英文标题**: Forgetting by Pruning: Data Deletion in Join Cardinality Estimation
- **中文标题**: 通过剪枝实现遗忘：连接基数估计中的数据删除
- **作者**: Chaowei He, Yuanjun Liu, Qingzhi Ma, Shenyuan Ren, Xizhao Luo, Lei Zhao, An Liu
- **英文摘要**: Machine unlearning in learned cardinality estimation (CE) systems presents unique challenges due to the complex distributional dependencies in multi-table relational data. Specifically, data deletion, a core component of machine unlearning, faces three critical challenges in learned CE models: attribute-level sensitivity, inter-table propagation and domain disappearance leading to severe overestimation in multi-way joins. We propose Cardinality Estimation Pruning (CEP), the first unlearning framework specifically designed for multi-table learned CE systems. CEP introduces Distribution Sensitivity Pruning, which constructs semi-join deletion results and computes sensitivity scores to guide parameter pruning, and Domain Pruning, which removes support for value domains entirely eliminated by deletion. We evaluate CEP on state-of-the-art architectures NeuroCard and FACE across IMDb and TPC-H datasets. Results demonstrate CEP consistently achieves the lowest Q-error in multi-table scenarios, particularly under high deletion ratios, often outperforming full retraining. Furthermore, CEP significantly reduces convergence iterations, incurring negligible computational overhead of 0.3%-2.5% of fine-tuning time.
- **中文摘要**: 学习型基数估计（CE）系统中的机器遗忘面临独特挑战，因为多表关系数据中存在复杂的分布依赖关系。具体来说，数据删除作为机器遗忘的核心组成部分，在学习型CE模型中面临三个关键挑战：属性级敏感性、表间传播和导致多路连接中严重高估的域消失。我们提出基数估计剪枝（CEP），这是首个专门为多表学习型CE系统设计的遗忘框架。CEP引入分布敏感性剪枝，构造半连接删除结果并计算敏感性分数以指导参数剪枝，以及域剪枝，移除对通过删除完全消除的值域的支持。我们在IMDb和TPC-H数据集上对最先进的架构NeuroCard和FACE进行了评估。结果表明，CEP在多表场景中始终实现最低的Q误差，特别是在高删除比率下，通常优于完全重训练。此外，CEP显著减少了收敛迭代次数，产生的计算开销仅为微调时间的0.3%-2.5%。

## Machine Learning IV

### 论文 16
- **英文标题**: Multi-agent In-context Coordination via Decentralized Memory Retrieval
- **中文标题**: 通过去中心化记忆检索进行多智能体上下文协调
- **作者**: Tao Jiang, Zichuan Lin, Lihe Li, Yi-Chen Li, Cong Guan, Lei Yuan, Zongzhang Zhang, Yang Yu, Deheng Ye
- **英文摘要**: Large transformer models, trained on diverse datasets, have demonstrated impressive few-shot performance on previously unseen tasks without requiring parameter updates. This capability has also been explored in Reinforcement Learning (RL), where agents interact with the environment to retrieve context and maximize cumulative rewards, showcasing strong adaptability in complex settings. However, in cooperative Multi-Agent Reinforcement Learning (MARL), where agents must coordinate toward a shared goal, decentralized policy deployment can lead to mismatches in task alignment and reward assignment, limiting the efficiency of policy adaptation. To address this challenge, we introduce Multi-agent In-context Coordination via Decentralized Memory Retrieval (MAICC), a novel approach designed to enhance coordination by fast adaptation. Our method involves training a centralized embedding model to capture fine-grained trajectory representations, followed by decentralized models that approximate the centralized one to obtain team-level task information. Based on the learned embeddings, relevant trajectories are retrieved as context, which, combined with the agents' current sub-trajectories, inform decision-making. During decentralized execution, we introduce a novel memory mechanism that effectively balances test-time online data with offline memory. Based on the constructed memory, we propose a hybrid utility score that incorporates both individual- and team-level returns, ensuring credit assignment across agents. Extensive experiments on cooperative MARL benchmarks, including Level-Based Foraging (LBF) and SMAC (v1/v2), show that MAICC enables faster adaptation to unseen tasks compared to existing methods.
- **中文摘要**: 基于多样化数据集训练的大型Transformer模型展示了在不更新参数的情况下对先前未见任务的令人印象深刻的小样本性能。这种能力也在强化学习（RL）中得到了探索，智能体与环境的交互以检索上下文并最大化累积奖励，在复杂环境中展现了强大的适应性。然而，在协作式多智能体强化学习（MARL）中，智能体必须协调朝向共同目标，去中心化策略部署可能导致任务对齐和奖励分配的错配，限制策略适应的效率。为应对这一挑战，我们引入了通过去中心化记忆检索进行多智能体上下文协调（MAICC），这是一种通过快速适应来增强协调的新方法。我们的方法包括训练一个中心化嵌入模型来捕获细粒度轨迹表示，随后使用去中心化模型近似中心化模型以获取团队级任务信息。基于学习到的嵌入，相关轨迹被检索作为上下文，与智能体当前的子轨迹结合，为决策提供信息。在去中心化执行期间，我们引入了一种新颖的记忆机制，有效平衡测试时在线数据与离线记忆。基于构建的记忆，我们提出了一个混合效用分数，融合个体级和团队级回报，确保智能体间的信用分配。在协作MARL基准（包括Level-Based Foraging和SMAC v1/v2）上的大量实验表明，与现有方法相比，MAICC能够更快地适应未见任务。

### 论文 86
- **英文标题**: Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
- **中文标题**: Sub-MoE：通过子空间专家合并进行高效混合专家LLM压缩
- **作者**: Lujun Li, Qiyuan Zhu, Jiacheng Wang, Xiaoyu Qin, Wei Li, Hao Gu, Sirui Han, Yike Guo
- **英文摘要**: Mixture of Experts (MoE) LLMs face significant obstacles due to their massive parameter scale, which imposes memory, storage, and deployment challenges. Although recent expert merging methods aim to achieve greater efficiency by consolidating several experts, they are fundamentally hindered by parameter conflicts arising from expert specialization. In this paper, we present Sub-MoE, a novel MoE compression framework via Subspace Expert Merging. Our key insight is to perform joint Singular Value Decomposition (SVD) on concatenated expert weights, reducing conflicting parameters by extracting shared U-matrices while enabling effective merging of the expert-specific V components.  Specifically, Sub-MoE consists of two innovative stages: (1) Adaptive Expert Clustering, which groups functionally coherent experts via K-means clustering based on cosine similarity of expert outputs; and (2) Subspace Expert Merging, which first performs Experts Union Decomposition to derive the shared U-matrix across experts in the same group, then applies frequency-based merging for individual V-matrices, and completes expert reconstruction using the merged V-matrix. In this way, we align and fuse experts in a shared subspace. Additionally, the framework can be extended with intra-expert compression for further inference optimization. Extensive experiments on Mixtral, DeepSeek, and Qwen-1.5/3 MoE LLMs demonstrate that our Sub-MoE significantly outperforms existing expert pruning and merging methods. Notably, our Sub-MoE maintains 96%/86% of original performance with 25%/50% expert reduction on Mixtral-8×7B in zero-shot benchmarks.
- **中文摘要**: 混合专家（MoE）LLM由于其庞大的参数量而面临显著的障碍，带来了内存、存储和部署挑战。虽然最近的专家合并方法旨在通过合并多个专家来实现更高的效率，但它们从根本上面临由专家专业化引起的参数冲突。本文提出了Sub-MoE，一种通过子空间专家合并的新型MoE压缩框架。我们的关键见解是对拼接的专家权重执行联合奇异值分解（SVD），通过提取共享的U矩阵来减少冲突参数，同时实现专家特定的V分量的有效合并。具体来说，Sub-MoE包含两个创新阶段：（1）自适应专家聚类，基于专家输出的余弦相似度通过K-means聚类将功能一致的专家分组；（2）子空间专家合并，首先执行专家联合分解以导出同一组内专家共享的U矩阵，然后对各个V矩阵应用基于频率的合并，并使用合并后的V矩阵完成专家重建。通过这种方式，我们在共享的子空间中对齐并融合专家。此外，该框架可以扩展专家内压缩以进一步优化推理。在Mixtral、DeepSeek和Qwen-1.5/3 MoE LLM上的大量实验表明，我们的Sub-MoE显著优于现有的专家剪枝和合并方法。值得注意的是，我们的Sub-MoE在Mixtral-8x7B的零样本基准上，以25%/50%的专家减少量保持了96%/86%的原始性能。

## Machine Learning V

### 论文 63
- **英文标题**: Balanced Knowledge Distillation for Large Language Models with Mix-of-Experts
- **中文标题**: 面向混合专家大语言模型的均衡知识蒸馏
- **作者**: Jiajun Liu, Yao He, Wenjun Ke, Peng Wang, Ziyu Shang, Guozheng Li, Zijie Xu
- **英文摘要**: Mixture-of-Experts (MoE) architectures have recently become a more prevalent choice for large language models (LLMs) than dense architectures due to their superior performance. However, billions of parameters bring MoE LLMs a huge cost for deployment and inference. To address these issues, knowledge distillation (KD) has become a widely adopted technique to compress LLMs. Existing KD methods for LLMs can be divided into dense-to-dense and moe-to-dense distillation. Dense-to-dense distillation transfers knowledge between single dense LLMs, while moe-to-dense distillation attempts to transfer knowledge between the MoE LLMs and the dense LLMs. However, the architectural mismatch prevents the student from fully absorbing knowledge when distilling MoE LLMs. To address this limitation, we investigate a new distillation setting, moe-to-moe, which aims to fully leverage expert knowledge of teachers and enable the student to absorb it more effectively. Compared to dense-to-dense and moe-to-dense, moe-to-moe suffers from two imbalance issues. First, expert-coverage deficiency reflects an imbalanced knowledge transfer of teacher experts: traditional distillation utilizes only the few experts activated by the teacher router. Second, routing imbalance appears when the student routing distribution drifts from the teacher, which makes it difficult for students to learn how to distribute different experts. To overcome these issues, we propose a novel distillation framework for moe-to-moe, Balanced Distillation (B-Distill), which equally spreads teacher expertise across student experts while regularizing the student router toward teacher-consistent balance. First, to mitigate expert-coverage deficiency, we introduce Monte Carlo exploration, which stochastically perturbs router probabilities so every teacher and student expert is sampled without enlarging the search space. Second, to correct routing imbalance and avert load collapse, we propose an entropy-aware router distillation mechanism that aligns the student router with the teacher while curbing over-concentration. Experiments show that B-Distill outperforms baselines by up to 6.6% in Rouge-L.
- **中文摘要**: 混合专家（MoE）架构最近因其优越性能成为比密集架构更受欢迎的大语言模型（LLM）选择。然而，数十亿参数为MoE LLM带来了巨大的部署和推理成本。为解决这些问题，知识蒸馏（KD）已成为压缩LLM的广泛采用技术。现有的LLM KD方法可分为密集到密集和MoE到密集蒸馏。密集到密集蒸馏在单个密集LLM之间迁移知识，而MoE到密集蒸馏尝试在MoE LLM和密集LLM之间迁移知识。然而，架构不匹配阻碍了学生在蒸馏MoE LLM时充分吸收知识。为解决此局限，我们研究了一种新的蒸馏设置，MoE到MoE，旨在充分利用教师的专家知识并使学生能够更有效地吸收。与密集到密集和MoE到密集相比，MoE到MoE面临两个不平衡问题。首先，专家覆盖不足反映了教师专家知识迁移的不均衡：传统蒸馏仅利用教师路由器激活的少数专家。其次，当学生路由分布偏离教师时出现路由不平衡，使学生难以学习如何分布不同的专家。为克服这些问题，我们提出了一种新颖的MoE到MoE蒸馏框架——均衡蒸馏（B-Distill），在规范学生路由器朝向教师一致平衡的同时，将教师专业知识均匀分布在学生专家中。首先，为缓解专家覆盖不足，我们引入蒙特卡洛探索，随机扰动路由器概率，使每个教师和学生专家都能被采样到，而不扩大搜索空间。其次，为纠正路由不平衡并避免负载坍缩，我们提出了一种熵感知路由器蒸馏机制，将学生路由器与教师对齐同时遏制过度集中。实验表明，B-Distill在Rouge-L上超越基线高达6.6%。

### 论文 68
- **英文标题**: Distilling Cross-Modal Knowledge via Feature Disentanglement
- **中文标题**: 通过特征解耦蒸馏跨模态知识
- **作者**: Junhong Liu, Yuan Zhang, Tao Huang, Wenchao Xu, Renyu Yang
- **英文摘要**: Knowledge distillation (KD) has proven highly effective for compressing large models and enhancing the performance of smaller ones. However, its effectiveness diminishes in cross-modal scenarios, such as vision-to-language distillation, where inconsistencies in representation across modalities lead to difficult knowledge transfer. To address this challenge, we propose frequency-decoupled cross-modal knowledge distillation, a method designed to decouple and balance knowledge transfer across modalities by leveraging frequency-domain features. We observed that low-frequency features exhibit high consistency across different modalities, whereas high-frequency features demonstrate extremely low cross-modal similarity. Accordingly, we apply distinct losses to these features: enforcing strong alignment in the low-frequency domain and introducing relaxed alignment for high-frequency features. We also propose a scale consistency loss to address distributional shifts between modalities, and employ a shared classifier to unify feature spaces. Extensive experiments across multiple benchmark datasets show our method substantially outperforms traditional KD and state-of-the-art cross-modal KD approaches.
- **中文摘要**: 知识蒸馏（KD）已被证明在压缩大型模型和增强较小模型性能方面非常有效。然而，在跨模态场景（如视觉到语言蒸馏）中，其有效性减弱，因为模态间表示的不一致导致知识迁移困难。为应对这一挑战，我们提出频率解耦的跨模态知识蒸馏，一种旨在通过利用频域特征解耦和平衡跨模态知识迁移的方法。我们观察到低频特征在不同模态间表现出高度一致性，而高频特征表现出极低的跨模态相似性。据此，我们对这些特征应用不同的损失：在低频域中强制执行强对齐，并为高频特征引入松弛对齐。我们还提出了尺度一致性损失以解决模态间的分布偏移，并采用共享分类器以统一特征空间。在多个基准数据集上的大量实验表明，我们的方法显著优于传统KD和最先进的跨模态KD方法。

### 论文 82
- **英文标题**: IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments
- **中文标题**: IndoorUAV：连续室内环境中视觉语言无人机导航的基准测试
- **作者**: Xu Liu, Yu Liu, Hanshuo Qiu, Yang Qirong, Zhouhui Lian
- **英文摘要**: Vision-Language Navigation (VLN) enables agents to navigate in complex environments by following natural language instructions grounded in visual observations. Although most existing work has focused on ground-based robots or outdoor Unmanned Aerial Vehicles (UAVs), indoor UAV-based VLN remains underexplored, despite its relevance to real-world applications such as inspection, delivery, and search-and-rescue in confined spaces. To bridge this gap, we introduce IndoorUAV, a novel benchmark and method specifically tailored for VLN with indoor UAVs. We begin by curating over 1,000 diverse and structurally rich 3D indoor scenes from the Habitat simulator. Within these environments, we simulate realistic UAV flight dynamics to collect diverse 3D navigation trajectories manually, further enriched through data augmentation techniques. Furthermore, we design an automated annotation pipeline to generate natural language instructions of varying granularity for each trajectory. This process yields over 16,000 high-quality trajectories, comprising the IndoorUAV-VLN subset, which focuses on long-horizon VLN. To support short-horizon planning, we segment long trajectories into sub-trajectories by selecting semantically salient keyframes and regenerating concise instructions, forming the IndoorUAV-VLA subset. Finally, we introduce IndoorUAV-Agent, a novel navigation model designed for our benchmark, leveraging task decomposition and multimodal reasoning. We hope IndoorUAV serves as a valuable resource to advance research on vision-language embodied AI in the indoor aerial navigation domain.
- **中文摘要**: 视觉语言导航（VLN）使智能体能够通过遵循基于视觉观察的自然语言指令在复杂环境中导航。虽然大多数现有工作集中于地面机器人或室外无人机（UAV），但基于室内无人机的VLN仍未被充分探索，尽管它与检查、递送和密闭空间搜救等实际应用息息相关。为弥合这一差距，我们引入IndoorUAV，一个专为室内无人机VLN定制的新颖基准和方法。我们首先从Habitat模拟器中策划了超过1,000个多样化且结构丰富的3D室内场景。在这些环境中，我们模拟真实的无人机飞行动力学以手动收集多样化3D导航轨迹，并通过数据增强技术进一步丰富。此外，我们设计了自动化标注流水线，为每条轨迹生成不同粒度的自然语言指令。此过程产生超过16,000条高质量轨迹，构成IndoorUAV-VLN子集，专注于长时序VLN。为支持短时序规划，我们通过选择语义显著的关键帧并重新生成简洁指令，将长轨迹分割为子轨迹，形成IndoorUAV-VLA子集。最后，我们引入IndoorUAV-Agent，一种为我们的基准设计的新颖导航模型，利用任务分解和多模态推理。我们希望IndoorUAV成为推进室内空中导航领域视觉语言具身AI研究的宝贵资源。

## Machine Learning VI

### 论文 82
- **英文标题**: ProLoG: Hybrid Prompt and LoRA Based Adaptation of Vision-Language Models for OOD Generalization
- **中文标题**: ProLoG：混合Prompt和LoRA的视觉语言模型适应用于OOD泛化
- **作者**: Jungwuk Park, Dong-Jun Han, Jaekyun Moon
- **英文摘要**: While vision-language foundation models (VLMs) achieve remarkable performance when fine-tuned on downstream in-distribution (ID) data, this process compromises their generalization ability on out-of-distribution (OOD) data that deviate from the downstream tasks due to overfitting. To address this, we propose ProLoG, a new adaptation method that  effectively fine-tunes VLMs on downstream tasks while achieving high OOD performance. Specifically, we  design a unique integration of prompt tuning and LoRA, offering a robust hybrid platform to improve performance. During training, we propose an augmentation-based regularization loss that enhances the generalization of our hybrid network by using augmented image features aligned with LLM-generated texts containing key attributes of each class.  By leveraging our hybrid design, we also introduce an adaptive inference strategy that flexibly applies trained prompts and LoRA based on a task similarity score to effectively handle both ID and OOD data. Experimental results demonstrate that our proposed method outperforms existing works on various datasets, confirming its advantages.
- **中文摘要**: 虽然视觉语言基础模型在下游分布内数据上微调时取得了显著性能，但这一过程因过拟合损害了其在偏离下游任务的分布外数据上的泛化能力。为解决此问题，我们提出ProLoG，一种在下游任务上有效微调VLM同时实现高OOD性能的新适应方法。具体而言，我们设计了提示调优和LoRA的独特整合，提供了一个鲁棒的混合平台以提高性能。在训练过程中，我们提出基于增强的正则化损失，通过使用与包含每个类别关键属性的LLM生成文本对齐的增强图像特征来增强我们混合网络的泛化能力。通过利用我们的混合设计，我们还引入了自适应推理策略，基于任务相似性分数灵活应用训练好的提示和LoRA来有效处理ID和OOD数据。实验结果表明，我们提出的方法在各种数据集上优于现有工作，证实了其优势。

### 论文 83
- **英文标题**: How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
- **中文标题**: 多少专家才够？面向稀疏混合专家的最优语义特化
- **作者**: Sumin Park, Noseong Park
- **英文摘要**: Finding the optimal configuration of Sparse Mixture-of- Experts (SMoE) that maximizes semantic differentiation among experts is essential for exploiting the full potential of MoE architectures. However, existing SMoE frameworks either heavily rely on hyperparameter tuning or overlook the importance of diversifying semantic roles across experts when adapting the expert pool size. We propose Mixture-of-Experts for Adaptive Semantic Specialization (MASS), a semantic-aware MoE framework for adaptive expert expansion and dynamic routing. MASS introduces two key advancements: (i) a gradient-based semantic drift detector that prompts targeted expert expansion when the existing expert pool lacks capacity to capture the full semantic diversity of the data, and (ii) an integration of adaptive routing strategy that dynamically adjusts expert usage based on token-level routing confidence mass. We first demonstrate that MASS reliably converges to the point of optimal balance between cost-performance trade-off with notably improved sematic specialization in a highly controlled synthetic setup. Further empirical results on real-world datasets across language and vision domains show that MASS consistently outperforms a range of strong MoE baselines, demonstrating its domain robustness and enhanced expert specialization.
- **中文摘要**: 找到最大化专家间语义差异的稀疏混合专家最优配置对于发挥MoE架构的全部潜力至关重要。然而，现有SMoE框架要么严重依赖超参数调优，要么在调整专家池大小时忽视跨专家多样化语义角色的重要性。我们提出面向自适应语义特化的混合专家（MASS），一个用于自适应专家扩展和动态路由的语义感知MoE框架。MASS引入了两项关键进展：（i）基于梯度的语义漂移检测器，当现有专家池缺乏捕获数据完整语义多样性的能力时提示有针对性的专家扩展；（ii）集成自适应路由策略，基于Token级路由置信质量动态调整专家使用。我们首先在一个高度受控的合成设置中证明MASS可靠收敛到代价-性能权衡的最优平衡点，并显著改善语义特化。在跨语言和视觉领域的真实世界数据集上的进一步实证结果表明，MASS始终优于一系列强MoE基线，展示了其领域鲁棒性和增强的专家特化。

### 论文 84
- **英文标题**: What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment
- **中文标题**: 什么构成好的生成图像？研究人类与多模态LLM图像偏好对齐
- **作者**: Rishab Parthasarathy, Jasmine Collins, Cory Stephenson
- **英文摘要**: Automated evaluation of generative text-to-image models remains a challenging problem. Recent works have proposed using multimodal LLMs to judge the quality of images, but these works offer little insight into how multimodal LLMs make use of concepts relevant to humans, such as image style or composition, to generate their overall assessment. In this work, we study what attributes of an image--specifically aesthetics, lack of artifacts, anatomical accuracy, compositional correctness, object adherence, and style--are important for both LLMs and humans to make judgments on image quality. We first curate a dataset of human preferences using synthetically generated image pairs. We use inter-task correlation between each pair of image quality attributes to understand which attributes are related in making human judgments. Repeating the same analysis with LLMs, we find that the relationships between image quality attributes are much weaker. Finally, we study individual image quality attributes by generating synthetic datasets with a high degree of control for each axis. Humans are able to easily judge the quality of an image with respect to all of the specific image quality attributes (e.g. high vs. low aesthetic image), however we find that some attributes, such as anatomical accuracy, are much more difficult for multimodal LLMs to learn to judge. Taken together, these findings reveal interesting differences between how humans and multimodal LLMs perceive images.
- **中文摘要**: 生成式文本到图像模型的自动评估仍然是一个具有挑战性的问题。近期工作提出使用多模态LLM来评判图像质量，但这些工作对多模态LLM如何利用与人类相关的概念（如图像风格或构图）来生成其总体评估提供了很少的洞察。本工作中，我们研究图像的哪些属性——具体地说是美学、无伪影、解剖准确性、构图正确性、对象一致性和风格——对人类和LLM做出图像质量判断都重要。我们首先使用合成生成的图像对策划了一个人类偏好数据集。我们使用每对图像质量属性之间的任务间相关性来理解哪些属性在人类判断中相关。通过LLM重复相同分析，我们发现图像质量属性之间的关系要弱得多。最后，我们通过为每个轴生成高度控制的合成数据集来研究各图像质量属性。人类能够轻松地根据所有具体的图像质量属性（例如高审美 vs 低审美图像）判断图像质量，但我们发现某些属性（如解剖准确性）对多模态LLM而言学习判断要困难得多。总的来说，这些发现揭示了人类和多模态LLM感知图像方式之间的有趣差异。

## Machine Learning VIII

### 论文 26
- **英文标题**: DAWN: Distributed LLM Multi-Agent Workflow Synthesis
- **中文标题**: DAWN：分布式LLM多智能体工作流合成
- **作者**: Guancheng Wan, Mo Zhou, Ziyi Wang, Xiaoran Shang, Eric Hanchen Jiang, Guibin Zhang, Jinhe Bi, Yunpu Ma, Zaixi Zhang, Ke Liang, Wenke Huang
- **英文摘要**: Large language models (LLMs) have recently empowered multi-agent systems (MAS) to achieve remarkable advances in collaborative reasoning and complex task automation. The effectiveness of these systems fundamentally depends on the design of adaptive communication graphs—the underlying workflows that coordinate agent interactions. However, in real-world scenarios, strict privacy constraints often silo data across organizations, and client distributions are highly non-IID, posing major challenges for synthesizing such workflows. In this work, we are the first to systematically study distributed multi-agent workflow synthesis under these privacy and heterogeneity constraints, and we introduce the Difficulty-Based Skew (DBS) benchmark to emulate such challenging environments. Drawing inspiration from federated graph learning (FGL)—which has primarily focused on classification over static graphs—we identify a critical gap: existing FGL methods do not address the generative design of communication topologies. We reveal two fundamental obstacles to generative workflow synthesis in this setting: (i) workflow specialization conflict, where agents optimized for different task distributions generate incompatible communication patterns that resist meaningful aggregation, and (ii) structural communication shift, where locally optimal agent interaction graphs fail to compose into globally coherent multi-agent workflows. To address these challenges, we propose DAWN, a federated framework that integrates two key innovations: Parametric Resonance, which robustly aggregates heterogeneous local updates via layer-wise SVD-based denoising and alignment, and Structural Gravity, which regularizes local workflow generation by penalizing the Fusion Gromov-Wasserstein distance to a set of prototype communication graphs, ensuring global structural coherence without stifling local adaptation. Experiments on the DBS benchmark show that DAWN surpasses baselines in global task success and reduces inter-client graph divergence, laying a solid foundation for privacy-preserving, adaptive MAS workflow design in heterogeneous settings.
- **中文摘要**: 大语言模型（LLM）最近使多智能体系统（MAS）在协作推理和复杂任务自动化方面取得了显著进展。这些系统的有效性从根本上取决于自适应通信图的设计——即协调智能体交互的底层工作流。然而，在现实场景中，严格的隐私约束通常将数据隔离在不同组织中，且客户端分布高度非独立同分布，这对合成此类工作流提出了重大挑战。在这项工作中，我们首次系统性地研究了这些隐私和异构性约束下的分布式多智能体工作流合成，并引入了基于难度的倾斜（DBS）基准来模拟这种具有挑战性的环境。受联邦图学习（FGL）的启发——FGL主要关注静态图上的分类——我们识别了一个关键缺口：现有的FGL方法没有解决通信拓扑的生成式设计问题。我们揭示了在此设置下生成式工作流合成的两个基本障碍：（i）工作流特化冲突，即针对不同任务分布优化的智能体会生成不兼容的通信模式，阻碍有意义的聚合；（ii）结构化通信偏移，即局部最优的智能体交互图无法组合成全局一致的多智能体工作流。为应对这些挑战，我们提出DAWN，一个联邦框架，集成了两项关键创新：参数共振，通过逐层基于SVD的去噪和对齐鲁棒地聚合异构局部更新；以及结构引力，通过惩罚局部工作流生成与一组原型通信图之间的融合Gromov-Wasserstein距离来正则化，确保全局结构一致性而不抑制局部适应。在DBS基准上的实验表明，DAWN在全局任务成功率上超越基线，并减少了客户端间图的分歧，为异构设置下隐私保护、自适应的MAS工作流设计奠定了坚实基础。

### 论文 36
- **英文标题**: X-SAM: From Segment Anything to Any Segmentation
- **中文标题**: X-SAM：从分割一切到一切分割
- **作者**: Hao Wang, Limeng Qiao, Zequn Jie, Zhijian Huang, Chengjian Feng, Qingfang Zheng, Lin Ma, Xiangyuan Lan, Xiaodan Liang
- **英文摘要**: Large Language Models (LLMs) demonstrate strong capabilities in broad knowledge representation, yet they are inherently deficient in pixel-level perceptual understanding. Although the Segment Anything Model (SAM) represents a significant advancement in visual-prompt-driven image segmentation, it exhibits notable limitations in multi-mask prediction and category-specific segmentation tasks, and it cannot integrate all segmentation tasks within a unified model architecture. To address these limitations, we present X-SAM, a streamlined Multimodal Large Language Model (MLLM) framework that extends the segmentation paradigm from segment anything to any segmentation. Specifically, we introduce a novel unified framework that enables more advanced pixel-level perceptual comprehension for MLLMs. Furthermore, we propose a new segmentation task, termed Visual GrounDed (VGD) segmentation, which segments all instance objects with interactive visual prompts and empowers MLLMs with visual grounded, pixel-wise interpretative capabilities. To enable effective training on diverse data sources, we present a unified training strategy that supports co-training across multiple datasets. Experimental results demonstrate that X-SAM achieves state-of-the-art performance on a wide range of image segmentation benchmarks, highlighting its efficiency for multimodal, pixel-level visual understanding.
- **中文摘要**: 大语言模型（LLM）在广泛的知识表示方面展示了强大的能力，但它们在像素级感知理解方面存在固有缺陷。尽管分割一切模型（SAM）代表了视觉提示驱动图像分割的重大进步，但它在多掩码预测和特定类别分割任务中表现出显著局限，并且无法将所有分割任务整合到一个统一的模型架构中。为解决这些局限，我们提出X-SAM，一个精简的多模态大语言模型（MLLM）框架，将分割范式从分割一切扩展到一切分割。具体而言，我们引入了一种新颖的统一框架，使MLLM能够实现更高级的像素级感知理解。此外，我们提出了一种新的分割任务，称为视觉接地（VGD）分割，该任务通过交互式视觉提示分割所有实例对象，并赋予MLLM视觉接地、像素级可解释的能力。为实现在多样化数据源上的有效训练，我们提出了一种支持多数据集联合训练的统一训练策略。实验结果表明，X-SAM在广泛的图像分割基准上实现了最先进的性能，突显了其在多模态、像素级视觉理解方面的高效性。

### 论文 74
- **英文标题**: Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World Model
- **中文标题**: 细粒度流蒸馏粗粒度流视频生成用于长期驾驶世界模型
- **作者**: Xiaodong Wang, Zhirong Wu, Peixi Peng
- **英文摘要**: Driving world models are used to simulate futures by video generation based on the condition of the current state and actions.  However, current models often suffer serious error accumulations when predicting the long-term future, which limits practical applications. Recent studies utilize the Diffusion Transformer (DiT) as the backbone of driving world models to improve learning flexibility. However, these models are always trained on short video clips, and multiple roll-out generations struggle to produce consistent and reasonable long videos due to the training-inference gap. To this end, we propose several solutions to build a simple yet effective long-term driving world model. First, we hierarchically decouple world model learning into large motion learning and bidirectional continuous motion learning. Then, considering the continuity of driving scenes, we propose a simple distillation method where fine-grained video flows are self-supervised signals for coarse-grained flows. The distillation is designed to improve the coherence of infinite video generation. The coarse-grained and fine-grained modules are coordinated to generate long-term and temporally coherent videos. On NuScenes, compared with the state-of-the-art front-view models, our model improves FVD by 27% and reduces inference time by 85% for the video task of generating 110+ frames.
- **中文摘要**: 驾驶世界模型用于通过基于当前状态和动作的视频生成来模拟未来。然而，当前模型在预测长期未来时通常遭受严重的误差累积，限制了实际应用。近期研究利用扩散Transformer（DiT）作为驾驶世界模型的骨干以提高学习灵活性。然而，这些模型总是在短视频片段上训练，由于训练-推理差距，多次滚动生成难以产生一致且合理的长时间视频。为此，我们提出了几种解决方案来构建简单而有效的长期驾驶世界模型。首先，我们将世界模型学习分层解耦为大运动学习和双向连续运动学习。然后，考虑到驾驶场景的连续性，我们提出了一种简单的蒸馏方法，其中细粒度视频流作为粗粒度流的自监督信号。该蒸馏旨在提高无限视频生成的一致性。粗粒度和细粒度模块协调工作以生成长期且时间连贯的视频。在NuScenes上，与最先进的前视图模型相比，我们的模型在生成110+帧的视频任务中将FVD提高了27%，并将推理时间减少了85%。

## Machine Learning IX

### 论文 18
- **英文标题**: Positional Bias in Multimodal Embedding Models: Do They Favor the Beginning, the Middle, or the End?
- **中文标题**: 多模态嵌入模型中的位置偏差: 偏向开头、中间还是结尾?
- **作者**: Kebin Wu, Fatima Albreiki
- **英文摘要**: Positional bias—where models overemphasize certain positions regardless of content—has been shown to negatively impact model performance across various tasks. While recent research has extensively examined positional bias in text generation models, its presence and effects in representation models remain underexplored. Even less is known about such biases in multimodal models. In this work, we investigate positional bias in multimodal representation models, specifically in the context of image-text retrieval. We begin by distinguishing between context importance and positional bias, and then assess the presence and extent of positional bias across different models and datasets. Our experiments demonstrate that positional bias is prevalent in multimodal models, but manifests differently across modalities: text encoders tend to exhibit bias toward the beginning of the input, whereas image encoders show bias at both the beginning and end. Furthermore, we find that this bias arises from, or is amplified by, a combination of factors, including the positional encoding scheme, training loss, context importance, and the nature of using image-text pairs in multimodal training
- **中文摘要**: 位置偏差——模型对某些位置过度强调而忽略内容——已被证明对各种任务的模型性能产生负面影响。虽然近期研究已广泛考察了文本生成模型中的位置偏差，但其在表示模型中的存在和影响仍未被充分探索。对多模态模型中此类偏差的了解更是少之又少。本工作研究了多模态表示模型中的位置偏差，特别是在图文检索场景下。我们首先区分了上下文重要性和位置偏差，然后评估了不同模型和数据集上位置偏差的存在和程度。实验表明，位置偏差在多模态模型中普遍存在，但在不同模态中表现不同：文本编码器倾向于偏向输入开头，而图像编码器则在开头和结尾均显示偏差。此外，我们发现这种偏差源自或被多种因素放大，包括位置编码方案、训练损失、上下文重要性以及多模态训练中使用图文配对的性质。

### 论文 20
- **英文标题**: MPA: Multimodal Prototype Augmentation for Few-Shot Learning
- **中文标题**: MPA: 面向小样本学习的多模态原型增强
- **作者**: Liwen Wu, Wei Wang, Lei Zhao, Zhan Gao, Qika Lin, Shaowen Yao, Zuozhu Liu, Bin Pu
- **英文摘要**: Recently, Few-shot Learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most existing methods focus only on the visual modality and compute prototypes directly from raw support images, which lack comprehensive and rich multimodal information. To address these limitations, we propose a novel Multimodal Prototype Augmentation FSL framework called MPA, including LLM-based Multi-Variant Semantic Enhancement (LMSE), Hierarchical Multi-View Augmentation (HMA), and an Adaptive Uncertain Class Absorber (AUCA). LMSE leverages large language models to generate diverse paraphrased category descriptions, enriching the support set with additional semantic cues. HMA exploits both natural and multi-view augmentations to enhance feature diversity (e.g., changes in viewing distance, camera angles, and lighting conditions). AUCA models uncertainty by introducing uncertain classes via interpolation and Gaussian sampling, effectively absorbing uncertain samples. Extensive experiments on four single-domain and six cross-domain FSL benchmarks demonstrate that MPA achieves superior performance compared to existing state-of-the-art methods across most settings. Notably, MPA surpasses the second-best method by 12.29% and 24.56% in the single-domain and cross-domain setting, respectively, in the 5-way 1-shot setting.
- **中文摘要**: 近年来，小样本学习(FSL)已成为一项热门任务，旨在从极少标注样本中识别新类别，并广泛应用于自然科学、遥感和医学图像等领域。然而，大多数现有方法仅关注视觉模态，直接从原始支持图像计算原型，缺乏全面且丰富的多模态信息。为解决这些局限，我们提出一种新颖的多模态原型增强小样本学习框架MPA，包括基于LLM的多变体语义增强(LMSE)、层次化多视角增强(HMA)和自适应不确定类别吸收器(AUCA)。LMSE利用大语言模型生成多样化的释义类别描述，用额外的语义线索丰富支持集。HMA利用自然增强和多视角增强来增强特征多样性（例如，视角距离、相机角度和光照条件的变化）。AUCA通过插值和高斯采样引入不确定类别来建模不确定性，有效吸收不确定样本。在四个单域和六个跨域FSL基准上的大量实验表明，MPA在大多数设置下相较于现有最先进方法取得了更优性能。值得注意的是，在5-way 1-shot设置下，MPA在单域和跨域设置中分别超越第二名方法12.29%和24.56%。

### 论文 52
- **英文标题**: Optimizing LoRA Allocation of MoE with the Alignment of Topic Correlation
- **中文标题**: 通过主题相关性对齐优化MoE的LoRA分配
- **作者**: Hengyuan Xu, Wenjun Ke, Yao He, Jiajun Liu, Dong Nie, Peng Wang, Ziyu Shang, Zijie Xu
- **英文摘要**: Mixture of experts (MoE) dynamically routes inputs to specialized expert networks to scale model capacity with low inference overhead. However, the excessive parameter growth in MoE models poses challenges in low-resource settings. To address these issues, MoE with parameter-efficient fine-tuning (PEFT) methods have emerged as a lightweight adaptation paradigm that distributes knowledge among experts via multiple LoRA blocks. Existing MoE-PEFT methods can be broadly categorized into External and Internal PEFT methods. External PEFT methods incorporate lightweight models into existing MoE architectures without modifying their routing, which limits the model’s parameter efficiency. To overcome these issues, Internal PEFT methods integrate MoE architectures into PEFT, enabling minimal parameter overhead. However, they still face two major challenges: (1) lack of expert functional differentiation, resulting in overlapping specialization across modules, and (2) absence of a structured attribution mechanism to guide expert selection based on semantic relevance. To alleviate these challenges, we propose TopicLoRA, a novel three-stage framework that leverages topic knowledge as semantic anchors to guide expert allocation. Specifically, (1) to address expert redundancy, we construct a topic-level prior graph using Graph Neural Network-enhanced representation learning over Big-Bench categories, enforcing structural separation among expert embeddings, and (2) to introduce semantic attribution, we design a dual-loss training mechanism that softly aligns input-query relevance with topic-guided routing distributions via KL divergence. Extensive experiments on representative datasets (e.g., MMLU, GSM8K, Flanv2) demonstrate that TopicLoRA outperforms state-of-the-art PEFT baselines by 2.40% on average in accuracy. Notably, the maximum improvement is 4.21%. Furthermore, ablation studies demonstrate that our framework's robustness to intricate topics and input sequence variations, which stems from the dual-loss training mechanism.
- **中文摘要**: 混合专家(MoE)通过动态将输入路由到专门的专家网络，以低推理开销扩展模型容量。然而，MoE模型中过度的参数增长在低资源场景下带来了挑战。为解决这些问题，具有参数高效微调(PEFT)方法的MoE已成为一种轻量级适应范式，通过多个LoRA模块在专家之间分布知识。现有MoE-PEFT方法可大致分为外部和内部PEFT方法。外部PEFT方法将轻量级模型纳入现有MoE架构而不修改其路由，限制了模型的参数效率。为克服这些问题，内部PEFT方法将MoE架构集成到PEFT中，实现最小参数开销。然而，它们仍面临两大挑战：(1)缺乏专家功能差异化，导致各模块之间的专业化重叠；(2)缺乏引导专家选择的结构化归因机制，无法基于语义相关性进行选择。为缓解这些挑战，我们提出TopicLoRA，一个新颖的三阶段框架，利用主题知识作为语义锚点来指导专家分配。具体来说，(1)为解决专家冗余，我们使用基于Big-Bench类别的图神经网络增强表示学习构建主题级先验图，强制专家嵌入之间的结构分离；(2)为引入语义归因，我们设计了双损失训练机制，通过KL散度软对齐输入-查询相关性与主题引导路由分布。在代表性数据集（如MMLU、GSM8K、Flanv2）上的大量实验表明，TopicLoRA在准确率上平均优于最先进的PEFT基线2.40%。值得注意的是，最大提升达到4.21%。此外，消融研究证明我们的框架对复杂主题和输入序列变化具有鲁棒性，这源于双损失训练机制。

### 论文 66
- **英文标题**: Adaptive Hallucination Alleviation in Multimodal Large Language Models: From Strategic Data Selection to Severity-Guided Training
- **中文标题**: 多模态大语言模型中的自适应幻觉缓解: 从策略数据选择到严重度引导训练
- **作者**: Yuanyi Xu, Xiangru Zhu, Sihang Jiang, Zhixu Li, Bei Yang, Xiaoxiao Xu, Yanghua Xiao, Wei Wang
- **英文摘要**: Multimodal Large Language Models (MLLMs) have recently achieved strong performance across a variety of multimodal tasks. However, they still suffer from various forms of hallucination, which hinder their practical deployment. Prior approaches often struggle to efficiently construct high-quality hallucination-related samples and to process them in a fine-grained manner, resulting in limited effectiveness in hallucination alleviation. To address this issue, we propose a data sampling strategy that selects samples better suited for hallucination-oriented training, thereby enhancing training effectiveness. In addition, we introduce a quantitative method for measuring hallucination severity and assign individualized weights to training samples accordingly. Building on this, we present Hallucination-Differentiated Direct Preference Optimization (HD-DPO), a novel preference optimization framework. During fine-tuning, HD-DPO incorporates these weights into both the formulation of customized loss functions and the modulation of localized visual attention, enabling fine-grained optimization. Experimental results demonstrate that our method outperforms existing fine-tuning strategies across multiple benchmarks and generalizes well to diverse MLLM architectures, effectively reducing hallucination rates and enhancing overall model performance.
- **中文摘要**: 多模态大语言模型(MLLMs)近期在多种多模态任务上取得了强劲表现。然而，它们仍受到各种形式幻觉的困扰，阻碍了其实际部署。先前方法通常难以高效构建高质量的幻觉相关样本并以细粒度方式处理它们，导致幻觉缓解效果有限。为解决此问题，我们提出一种数据采样策略，选择更适合面向幻觉训练的样本，从而增强训练效果。此外，引入一种量化幻觉严重度的方法，相应地为训练样本分配个体化权重。在此基础上，我们提出幻觉差异化直接偏好优化(HD-DPO)，一种新颖的偏好优化框架。在微调期间，HD-DPO将这些权重同时纳入定制损失函数的构建和局部视觉注意力的调制，实现细粒度优化。实验结果表明，我们的方法在多个基准上优于现有微调策略，并能良好泛化到多样化的MLLM架构，有效降低幻觉率并提升整体模型性能。

### 论文 68
- **英文标题**: CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
- **中文标题**: CAMERA: 基于微专家冗余分析的MoE模型多矩阵联合压缩
- **作者**: Yuzhuang Xu, Xu Han, Yuanchi Zhang, Yixuan Wang, Yijun Liu, Shiyu Ji, Qingfu Zhu, Wanxiang Che
- **英文摘要**: Large Language Models (LLMs) with Mixture-of-Experts (MoE) architectures are distinguished by their strong performance scaling with increasing parameters across a wide range of tasks, yet they also suffer from substantial computational and storage overheads. Notably, the performance gains of MoE models do not scale proportionally with the growth in expert parameters. While prior works attempt to reduce parameters via expert-level pruning, merging, or decomposition, they still suffer from challenges in both performance and computational efficiency. In this paper, we address these challenges by introducing micro-expert as a finer-grained compression unit that spans across matrices. We first establish a more fundamental perspective, viewing MoE layers as mixtures of micro-experts, and present CAMERA, a lightweight and training-free framework for identifying micro-expert redundancy. Our analysis uncovers significant variance in micro-expert contributions during decoding. Based on this insight, we further propose CAMERA-P, a structured micro-expert pruning framework, and CAMERA-Q, a mixed-precision quantization idea designed for micro-experts. Extensive experiments on nine downstream tasks show that CAMERA-P consistently outperforms strong baselines under pruning ratios ranging from 20% to 60%. Furthermore, CAMERA-Q achieves superior results under aggressive 2-bit quantization, surpassing existing matrix- and channel-level ideas. Notably, our method enables complete micro-expert analysis of Qwen2-57B-A14B in less than 5 minutes on a single NVIDIA A100-40GB GPU.
- **中文摘要**: 具有混合专家(MoE)架构的大语言模型(LLMs)以其在广泛任务上随参数增长而增强的性能著称，但它们也承受着巨大的计算和存储开销。值得注意的是，MoE模型的性能增益并不与专家参数的增长成比例。虽然先前工作尝试通过专家级剪枝、合并或分解来减少参数，但这些方法在性能和计算效率方面仍面临挑战。本文通过引入微专家作为跨越矩阵的更细粒度压缩单元来应对这些挑战。我们首先建立一个更基础的视角，将MoE层视为微专家的混合，并提出CAMERA，一个轻量级且无需训练的微专家冗余识别框架。我们的分析揭示了推理过程中微专家贡献的显著方差。基于这一发现，我们进一步提出CAMERA-P，一个结构化微专家剪枝框架，以及CAMERA-Q，一种为微专家设计的混合精度量化方案。在九个下游任务上的大量实验表明，CAMERA-P在20%到60%的剪枝率范围内持续优于强基线方法。此外，CAMERA-Q在激进的2比特量化下取得了优越结果，超越了现有的矩阵级和通道级方案。值得注意的是，我们的方法能在单个NVIDIA A100-40GB GPU上不到5分钟内完成Qwen2-57B-A14B的完整微专家分析。

## Machine Learning X

### 论文 75
- **英文标题**: Vulnerability-Aware Robust Multimodal Adversarial Training
- **中文标题**: 漏洞感知鲁棒多模态对抗训练
- **作者**: Junrui Zhang, Xinyu Zhao, Jie Peng, Chenjie Wang, Jianmin Ji, Tianlong Chen
- **英文摘要**: Multimodal learning has shown significant superiority on various tasks by integrating multiple modalities. However, the interdependencies among modalities increase the susceptibility of multimodal models to adversarial attacks. Existing methods mainly focus on attacks on specific modalities or indiscriminately attack all modalities. In this paper, we find that these approaches ignore the differences between modalities in their contribution to final robustness, resulting in suboptimal robustness performance. To bridge this gap, we introduce Vulnerability-Aware Robust Multimodal Adversarial Training (VARMAT), a probe-in-training adversarial training method that improves multimodal robustness by identifying the vulnerability of each modality. To be specific, VARMAT first explicitly quantifies the vulnerability of each modality, grounded in a first-order approximation of the attack objective (Probe). Then, we propose a targeted regularization term that penalizes modalities with high vulnerability, guiding robust learning while maintaining task accuracy (Training). We demonstrate the enhanced robustness of our method across multiple multimodal datasets involving diverse modalities. Finally, we achieve {12.73%, 22.21%, 11.19%} robustness improvement on three multimodal datasets, revealing a significant blind spot in multimodal adversarial training.
- **中文摘要**: 多模态学习通过整合多种模态在各种任务上展现出显著优越性。然而模态间的相互依赖性增加了多模态模型对对抗攻击的敏感性。现有方法主要关注对特定模态的攻击或无差别攻击所有模态。本文发现这些方法忽视了各模态对最终鲁棒性的贡献差异,导致次优鲁棒性。提出VARMAT,一种训练中探测的对抗训练方法,通过识别每个模态的漏洞来提升多模态鲁棒性。首先基于攻击目标的一阶近似显式量化每个模态的漏洞。然后提出针对性正则化项惩罚高漏洞模态,在保持任务准确率的同时引导鲁棒学习。在三个多模态数据集上分别实现12.73%、22.21%和11.19%的鲁棒性提升。

## Multiagent Systems

### 论文 3
- **英文标题**: Efficient Multiagent Planning via Shared Action Suggestions
- **中文标题**: 通过共享行动建议的高效多智能体规划
- **作者**: Dylan M. Asmar, Mykel J. Kochenderfer
- **英文摘要**: Decentralized partially observable Markov decision processes with communication (Dec-POMDP-Com) provide a framework for multiagent decision making under uncertainty, but the NEXP-complete complexity for finite-horizon problems renders solutions intractable in general. While sharing actions and observations can reduce the complexity to PSPACE-complete, we propose an approach that bridges POMDPs and Dec-POMDPs by communicating only suggested joint actions, eliminating the need to share observations while retaining near-centralized performance. Our algorithm estimates joint beliefs using shared actions to prune infeasible beliefs. Each agent maintains possible belief sets for other agents, pruning them based on suggested actions to form an estimated joint belief usable with any centralized policy. This approach requires solving a POMDP for each agent, reducing computational complexity while preserving performance. We demonstrate its effectiveness on several Dec-POMDP benchmarks, showing performance comparable to centralized methods when shared actions enable effective belief pruning. This action-based communication framework offers a natural avenue for integrating human-agent cooperation, opening new directions for scalable multiagent planning under uncertainty, with applications in both autonomous systems and human-agent teams.
- **中文摘要**: 具有通信的去中心化部分可观测马尔可夫决策过程为不确定性下的多智能体决策提供了框架，但有限时域问题的NEXP完全复杂度使得解通常难以处理。虽然共享行动和观测可以将复杂度降低到PSPACE完全，我们提出了一种通过仅通信建议的联合行动来桥接POMDP和Dec-POMDP的方法，消除了共享观测的需要，同时保留了接近集中式的性能。我们的算法使用共享行动来估计联合信念，以剪枝不可行信念。每个智能体为其他智能体维护可能的信念集，基于建议的行动进行剪枝，形成可与任何集中式策略一起使用的估计联合信念。这种方法只需要为每个智能体求解一个POMDP，降低了计算复杂度同时保持性能。我们展示了其在多个Dec-POMDP基准上的有效性，表明当共享行动实现有效信念剪枝时，性能可与集中式方法相媲美。这种基于行动的通信框架为整合人机合作提供了自然途径，为不确定性下的可扩展多智能体规划开辟了新方向。

### 论文 7
- **英文标题**: GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
- **中文标题**: GUI-Eyes：GUI智能体中工具增强感知用于视觉定位
- **作者**: Chen Chen, Jiawei Shao, Dakuan Lu, Haoyi Hu, Xiangcheng Liu, Hantao Yao, Wu Liu
- **英文摘要**: Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot visual inputs and passive perception, lacking the ability to adaptively determine when, whether, and how to observe the interface. We present GUI-Eyes, a reinforcement learning framework for active visual perception in GUI tasks. To acquire more informative observations, the agent learns to make strategic decisions on both whether and how to invoke visual tools, such as cropping or zooming, within a two-stage reasoning process. To support this behavior, we introduce a progressive perception strategy that decomposes the decision-making into coarse exploration and fine-grained grounding, coordinated by a two-level policy. In addition, we design a spatially continuous reward function tailored to tool usage, which integrates both location proximity and region overlap to provide dense supervision and alleviate the reward sparsity common in GUI environments. On the ScreenSpot-Pro benchmark, GUI-Eyes-3B achieves 44.8% grounding accuracy using only 3k labeled samples, significantly outperforming both supervised and RL-based baselines. These results highlight that tool-aware active perception, enabled by staged policy reasoning and fine-grained reward feedback, is critical for building robust and data-efficient GUI agents.
- **中文摘要**: 视觉语言模型和强化学习的最新进展推动了GUI自动化的发展。然而，大多数现有方法依赖静态、一次性的视觉输入和被动感知，缺乏自适应地确定何时、是否以及如何观察界面的能力。我们提出GUI-Eyes，一个用于GUI任务中主动视觉感知的强化学习框架。为获取更具信息量的观察，智能体学习在是否以及如何调用视觉工具方面做出战略决策，通过两阶段推理过程实现。为支持此行为，我们引入了一种渐进式感知策略，将决策分解为粗粒度探索和细粒度定位，由两级策略协调。此外，我们设计了一个空间连续奖励函数，专门针对工具使用，结合位置邻近度和区域重叠提供密集监督，缓解GUI环境中常见的奖励稀疏性。在ScreenSpot-Pro基准上，GUI-Eyes-3B仅使用3K标记样本就达到了44.8%的定位准确率，显著优于监督学习和RL基线。这些结果突显了由分阶段策略推理和细粒度奖励反馈支持的工具感知主动感知对于构建鲁棒且数据高效的GUI智能体至关重要。

### 论文 46
- **英文标题**: MACoT: Synthesizing Chains of Thought for Small Models via Multi-Agent Collaboration
- **中文标题**: MACoT：通过多智能体协作为小模型合成思维链
- **作者**: Guokai Tang, Feng Zhao
- **英文摘要**: Small language models (SLMs) run quickly, consume little memory, and can be deployed on edge devices, making them especially appealing when compute or energy is limited. Because of these advantages, boosting SLMs' reasoning ability has become an important research goal. A common approach is to distill the long chains of thought (long-CoTs) produced by large reasoning models (LRMs) into SLMs, hoping to transfer the larger models’ strong reasoning ability. However, SLMs do not always benefit from distillation of long-CoTs. The lengthy and complex semantic steps and large amount of self-reflection content in long-CoTs may exceed the limited learning capabilities of SLMs, and the impact of self-reflection density on the performance of SLMs is unclear. To resolve this capacity mismatch, we propose MACoT, a multi-agent framework that synthesizes chains of thought (CoTs) that are more suitable for small models rather than compressing or pruning existing ones. Through the interactive collaboration among six types of agents, MACoT synthesizes semantically explicit, logically clear CoTs that efficiently activate a small model’s internal knowledge through a carefully designed output pattern. At the same time, the CoTs synthesized by our method can retain a small amount of self-reflection content, thereby matching the learning capability of the small model and maximizing its reasoning accuracy. We fine-tuned Qwen2.5-7B-Instruct using only 1879 synthetic CoTs, significantly improving its performance on mathematical reasoning tasks and generalizing well, outperforming models trained on 5x more data. Through experiments, we found that a modest level of self-reflection boosts small-model performance, whereas excessive reflection sharply degrades it, which shows that “teaching SLMs to think” hinges on aligning each CoT’s cognitive load with the model’s capacity.
- **中文摘要**: 思维链提示显著提高了大型语言模型的推理能力，但小型模型通常难以生成高质量的推理链。我们提出MACoT，一个多智能体协作框架，通过整合多个专门LLM智能体的优势为小模型合成高质量思维链。每个智能体贡献不同的推理视角，一个门控机制选择和组合最有效的推理步骤。合成的思维链然后用于微调小模型，使其能够内化复杂的推理模式。在算术、常识和符号推理基准上的实验表明，使用MACoT合成链训练的小模型达到或超过了更大模型的性能，同时在推理时比依赖多智能体系统更高效。

### 论文 48
- **英文标题**: DanceHA: A Multi-Agent Framework for Document-Level Aspect-Based Sentiment Analysis
- **中文标题**: DanceHA：一个用于文档级方面情感分析的多智能体框架
- **作者**: Lei Wang, Min Huang, Eduard Dragut
- **英文摘要**: Aspect-Based Sentiment Intensity Analysis (ABSIA) has garnered increasing attention, though research largely focuses on domain-specific, sentence-level settings. In contrast, document-level ABSIA--particularly in addressing complex tasks like extracting Aspect-Category-Opinion-Sentiment-Intensity (ACOSI) tuples--remains underexplored. In this work, we introduce DanceHA, a multi-agent framework designed for open-ended, document-level ABSIA with informal writing styles. DanceHA has two main components: Dance, which employs a divide-and-conquer strategy to decompose the long-context ABSIA task into smaller, manageable sub-tasks for collaboration among specialized agents; and HA, Human-AI collaboration for annotation. We release Inf-ABSIA, a multi-domain document-level ABSIA dataset featuring fine-grained and high-accuracy labels from DanceHA. Extensive experiments demonstrate the effectiveness of our agentic framework and show that the multi-agent knowledge in DanceHA can be effectively transferred into student models. Our results highlight the importance of the overlooked informal styles in ABSIA, as they often intensify opinions tied to specific aspects.
- **中文摘要**: 文档级方面情感分析涉及识别文档中提到的多个方面及其对应情感，需要全局上下文理解。我们提出DanceHA，一个多智能体框架，其中专门的智能体处理不同的子任务：方面检测、情感分类和文档级聚合。这些智能体通过共享记忆模块协作共享中间发现和解决冲突。一个协调智能体监督过程，确保一致的方面-情感对齐。在标准DASA基准上的实验表明，DanceHA优于单模型方法，特别是在具有多个交织方面的长文档上。

### 论文 53
- **英文标题**: An LLM-based Simulation Framework for Embodied Conversational Agents in Psychological Counseling
- **中文标题**: 用于心理咨询中具身对话智能体的LLM模拟框架
- **作者**: Lixiu Wu, Yuanrong Tang, Qisen Pan, Xianyang Zhan, Yuchen Han, Lanxi Xiao, Tianhong Wang, Chen Zhong, Jiangtao Gong
- **英文摘要**: Due to privacy concerns, open dialogue datasets for mental health are primarily generated through human or AI synthesis methods. However, the inherent implicit nature of psychological processes, particularly those of clients, poses challenges to the authenticity and diversity of synthetic data. In this paper, we propose ECAs (short for Embodied Conversational Agents), a framework for embodied agent simulation based on Large Language Models (LLMs) that incorporates multiple psychological theoretical principles. Using simulation, we expand real counseling case data into a nuanced embodied cognitive memory space and generate dialogue data based on high-frequency counseling questions. We validated our framework using the D4 dataset. First, we created a public ECAs dataset through batch simulations based on D4. Licensed counselors evaluated our method, demonstrating that it significantly outperforms baselines in simulation authenticity and necessity. Additionally, two LLM-based automated evaluation methods were employed to confirm the higher quality of the generated dialogues compared to the baselines.
- **中文摘要**: 心理咨询需要同理心倾听、适当回应和建立治疗联盟，这些品质对于AI系统来说具有挑战性但日益重要。我们提出一个模拟框架，利用LLM驱动智能体在受控心理咨询场景中扮演咨询师和来访者角色。多个智能体模拟不同的治疗取向和来访者背景，使得对咨询策略和结果的大规模评估成为可能。该框架融合了心理学理论和实践指南来指导智能体行为。我们通过与执业治疗师的验证研究证明，我们的模拟捕捉了真实咨询动态的关键方面，为训练和评估AI辅助心理健康工具提供了一个有前景的平台。

### 论文 57
- **英文标题**: LungNoduleAgent: A Collaborative Multi-Agent System for Precision Diagnosis of Lung Nodules
- **中文标题**: LungNoduleAgent：一个用于肺结节精确诊断的协作多智能体系统
- **作者**: Cheng Yang, Hui Jin, Xinlei Yu, Zhipeng Wang, Yaoqun Liu, Fenglei Fan, Dajiang Lei, Gangyong Jia, Changmiao Wang, Ruiquan Ge
- **英文摘要**: Diagnosing lung cancer typically involves physicians identifying lung nodules in Computed tomography (CT) scans and generating diagnostic reports based on their morphological features and medical expertise. Although advancements have been made in using multimodal large language models for analyzing lung CT scans, challenges remain in accurately describing nodule morphology and incorporating medical expertise. These limitations affect the reliability and effectiveness of these models in clinical settings. Collaborative multi-agent systems offer a promising strategy for achieving a balance between generality and precision in medical applications, yet their potential in pathology has not been thoroughly explored. To bridge these gaps, we introduce LungNoduleAgent, an innovative collaborative multi-agent system specifically designed for analyzing lung CT scans. LungNoduleAgent streamlines the diagnostic process into sequential components, improving precision in describing nodules and grading malignancy through three primary modules. The first module, the Nodule Spotter, coordinates clinical detection models to accurately identify nodules. The second module, the Radiologist, integrates localized image description techniques to produce comprehensive CT reports. Finally, the Doctor Agent System performs malignancy reasoning by using images and CT reports, supported by a pathology knowledge base and a multi-agent system framework. Extensive testing on two private datasets and the public LIDC-IDRI dataset indicates that LungNoduleAgent surpasses mainstream vision-language models, agent systems, and advanced expert models such as GPT-4o, Claude 3.7 Sonnet, LLaMA-3.2 Vision, Qwen2.5-VL, Med-R1, MedGemma, MedAgent-Pro, MedAgents, MDAgent and LLaVA-Med. These results highlight the importance of region-level semantic alignment and multi-agent collaboration in diagnosing nodules. LungNoduleAgent stands out as a promising foundational tool for supporting clinical analyses of lung nodules.
- **中文摘要**: 肺结节的准确诊断对于肺癌早期检测和治疗计划至关重要。我们提出LungNoduleAgent，一个协作多智能体系统，集成多个专门的AI模型进行全面的肺结节分析。该系统协调处理CT图像分析的智能体：检测、分割、分类和风险评估。一个融合智能体结合输出并产生统一诊断。该系统纳入了临床指南以确保建议符合医学标准。在公共LIDC-IDRI数据集的评估以及临床验证研究中，与单一模型方法和放射科医生评估相比，LungNoduleAgent展现了优越的诊断准确性。

## Natural Language Processing I

### 论文 15
- **英文标题**: OptiHive: Ensemble Selection for LLM-Based Optimization via Statistical Modeling
- **中文标题**: OptiHive: 通过统计建模实现基于LLM优化的集成选择
- **作者**: Maxime Bouscary, Saurabh Amin
- **英文摘要**: LLM-based solvers have emerged as a promising means of automating problem modeling and solving. However, they remain unreliable and often depend on iterative repair loops that result in significant latency. We introduce OptiHive, a framework that enhances any solver-generation pipeline to produce higher-quality solvers from natural-language descriptions of optimization problems. OptiHive uses a single batched generation to produce diverse components (solvers, problem instances, and validation tests) and filters out erroneous components to ensure fully interpretable outputs. Accounting for the imperfection of the generated components, we employ a statistical model to infer their true performance, enabling principled uncertainty quantification and solver selection. On tasks ranging from traditional optimization problems to challenging variants of the Multi-Depot Vehicle Routing Problem, OptiHive significantly outperforms baselines, increasing the optimality rate from 5% to 92% on the most complex problems.
- **中文摘要**: 基于LLM的求解器已成为自动化问题建模和求解的有前景手段。然而，它们仍不可靠，通常依赖导致显著延迟的迭代修复循环。我们引入OptiHive，这是一个增强任何求解器生成流水线的框架，以从自然语言描述的优化问题中产生更高质量的求解器。OptiHive使用单次批量生成产生多样化组件（求解器、问题实例和验证测试），并过滤掉错误的组件以确保完全可解释的输出。考虑到生成组件的不完美性，我们使用统计模型推断其真实性能，从而实现有原则的不确定性量化和求解器选择。在从传统优化问题到具有挑战性的多车场车辆路径问题变体的任务中，OptiHive显著优于基线，在最复杂的问题上将最优率从5%提升至92%。

### 论文 20
- **英文标题**: Does Question Really Matter? The Attribution of Answer Bias in LLM Evaluation
- **中文标题**: 问题真的重要吗？LLM评估中答案偏差的归因
- **作者**: Boxi Cao, Ruotong Pan, Hongyu Lin, Xianpei Han, Le Sun
- **英文摘要**: Multiple-choices question answering (MCQA) has emerged as one of the most popular task formats for large language models (LLMs) evaluation. Unfortunately, there exist substantial evidence that the evaluation of current MCQA benchmarks suffers from significant answer bias, which severely undermines the reliability of the evaluation conclusions. Specifically, many LLMs achieve performance significantly higher than random selection even when the questions are omitted from input information. To this end, we conduct a systematic investigation of the attribution of answer bias, and demonstrate a strong correlation between the degree of data contamination and the severity of answer bias, while the position of options and the popularity of answers have relatively minor effects. Building on these insights, we further propose OPD, a straightforward yet effective tool for contamination detection and dataset debiasing without requiring access to the model’s internal training data. Our findings and algorithms provide valuable insights for the design of future trustworthy LLM evaluation protocols.
- **中文摘要**: 多项选择问答（MCQA）已成为大语言模型评估中最流行的任务格式之一。不幸的是，有大量证据表明当前MCQA基准的评估受到显著的答案偏差影响，严重削弱了评估结论的可靠性。具体而言，许多LLM即使从输入信息中省略问题时也能取得显著高于随机选择的性能。为此，我们对答案偏差的归因进行了系统性调查，证明数据污染程度与答案偏差严重程度之间存在强相关性，而选项位置和答案流行度的影响相对较小。基于这些见解，我们进一步提出OPD，这是一个简单而有效的工具，用于污染检测和数据集去偏，无需访问模型内部训练数据。我们的发现和算法为设计未来可信的LLM评估协议提供了有价值的见解。

### 论文 22
- **英文标题**: TIV: Thought Injection via Vectors for Efficient Reasoning in Large Reasoning Models
- **中文标题**: TIV: 通过向量注入思维实现大型推理模型的高效推理
- **作者**: Yi Cao, Weijie Shi, Wei-Jie Xu, Yucheng Shen, Yue Cui, Hanghui Guo, Shimin Di, Ziyi Liu, Jiaming Li, Alexander Zhou, Jia Zhu, Jiajie Xu
- **英文摘要**: Large Reasoning Models (LRMs) have recently demonstrated impressive performance across a range of reasoning tasks by generating intermediate thoughts. However, these models can suffer from overthinking—generating excessive tokens that contribute little to final accuracy while increasing inference cost. To mitigate this, we propose TIV (Thought Injection via Vectors), an innovative framework that compresses token-level reasoning into compact vectors without sacrificing performance. Rather than generating explicit thoughts, TIV injects learnable vectors into the post-attention hidden states of the final token across Transformer layers, enabling implicit and lightweight reasoning. We further introduce a two-stage reinforcement learning strategy: the first stage calibrates the model's reasoning distribution, and the second distills it into a vector-based policy optimized for both accuracy and brevity. Experiments on three reasoning benchmarks show that TIV preserves over 99% of the original accuracy while reducing output length by more than 65% on average, reaching up to 80% in some cases. Moreover, TIV consistently achieves superior trade-offs between accuracy and efficiency compared to existing methods, distinguishing itself as a state-of-the-art (SOTA) approach for efficient reasoning in LRMs.
- **中文摘要**: 大型推理模型（LRM）最近在通过生成中间思维完成一系列推理任务上展示了令人印象深刻的性能。然而，这些模型可能遭受过度思考——生成对最终准确性贡献很小却增加推理成本的过多标记。为缓解这一问题，我们提出TIV（通过向量注入思维），这是一个创新框架，将标记级推理压缩为紧凑向量而不牺牲性能。TIV不生成显式思维，而是将可学习向量注入Transformer各层最终标记的后注意力隐藏状态中，实现隐式和轻量级推理。我们进一步引入两阶段强化学习策略：第一阶段校准模型的推理分布，第二阶段将其蒸馏到针对准确性和简洁性优化的基于向量的策略中。在三个推理基准上的实验表明，TIV保留了超过99%的原始准确率，同时平均减少输出长度超过65%，在某些情况下可达80%。此外，与现有方法相比，TIV始终在准确性和效率之间取得更优的权衡，确立了自己作为LRM高效推理的最先进方法。

### 论文 41
- **英文标题**: Improving Long-Context Summarization with Multi-Granularity Retrieval Optimization
- **中文标题**: 通过多粒度检索优化改进长上下文摘要
- **作者**: Xueyu Chen, Kaitao Song, Zifan Song, Dongsheng Li, Cairong Zhao
- **英文摘要**: Retrieval-Augmented Generation (RAG) is an effective solution to overcome the limitations of Large Language Models (LLMs) in terms of specific-domain knowledge and timely information updates. However, current RAG methods typically respond to queries based on isolated segments, lacking the ability to integrate information within the same document. This undermines performance in real-world tasks requiring coherent understanding across an entire document. Notably, the human brain naturally integrates and summarizes prior knowledge upon reading a given text, progressively formulating a comprehensive understanding. Motivated by this cognitive process, we propose the Hierarchical Two-Stage Summarization-based Information Retrieval (HTSIR) method, which preprocesses the corpus prior to retrieval, summarizes continuous texts to obtain integrated information, and constructs a retrieval tree with varying summary granularities. The retrieved information is then processed by a Reranker based on the current question to serve as a context for LLMs. Additionally, as single-step summarization is often imprecise in query-based summarization tasks, we further apply a Refinement module, allowing LLMs to reflect and revise their output to achieve the final result. By combining HTSIR with GPT-4o mini, we achieve state-of-the-art results on complex question tasks across four long-text datasets (NarrativeQA, QASPER, QuALITY, and QMSum), achieving an improvement of about 6 points on the Question Answering (QA) task in QuALITY-HRAD.
- **中文摘要**: 检索增强生成（RAG）是克服大语言模型在特定领域知识和及时信息更新方面局限性的有效解决方案。然而，当前的RAG方法通常基于孤立段落响应查询，缺乏整合同一文档内信息的能力。这损害了需要连贯理解整个文档的现实世界任务中的性能。值得注意的是，人类大脑自然地整合和总结先前知识，逐步形成全面理解。受此认知过程启发，我们提出了基于分层的两阶段摘要信息检索（HTSIR）方法，在检索前对语料库进行预处理，总结连续文本以获得整合信息，并构建不同摘要粒度的检索树。然后由基于当前问题的重排序器处理检索到的信息，作为LLM的上下文。此外，由于在基于查询的摘要任务中单步摘要通常不精确，我们进一步应用细化模块，允许LLM反思和修正其输出以达到最终结果。通过将HTSIR与GPT-4o mini结合，我们在四个长文本数据集（NarrativeQA、QASPER、QuALITY和QMSum）的复杂问题任务上取得了最先进的结果，在QuALITY-HRAD的问答任务上实现了约6个点的提升。

### 论文 53
- **英文标题**: Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMs
- **中文标题**: LLM持续微调下的持久后门攻击
- **作者**: Jing Cui, Yufei Han, Jianbin Jiao, Junge Zhang
- **英文摘要**: Backdoor attacks embed malicious behaviors into Large Language Models (LLMs), enabling adversaries to trigger harmful outputs or bypass safety controls. However, the persistence of the implanted backdoors under user-driven post-deployment continual fine-tuning has been rarely examined. Most prior works evaluate the effectiveness and generalization of implanted backdoors only at releasing and empirical evidence shows that naively injected backdoor persistence degrades after updates. In this work, we study whether and how implanted backdoors persist through a multi‑stage post-deployment fine‑tuning. We propose P‑Trojan, a trigger‑based attack algorithm that explicitly optimizes for backdoor persistence across repeated updates. By aligning poisoned gradients with those of clean tasks on token embeddings, the implanted backdoor mapping is less likely to be suppressed or forgotten during subsequent updates. Theoretical analysis shows the feasibility of such persistent backdoor attacks after continual fine-tuning. And experiments conducted on the Qwen2.5 and LLaMA3 families of LLMs, as well as diverse task sequences, demonstrate that P‑Trojan achieves over \textbf{99\%} persistence while preserving clean‑task accuracy. Our findings highlight the need for persistence-aware evaluation and stronger defenses in realistic model adaptation pipelines.
- **中文摘要**: 后门攻击将恶意行为嵌入大语言模型，使对手能够触发有害输出或绕过安全控制。然而，植入的后门在用户驱动的部署后持续微调下的持久性很少被考察。大多数先前工作仅在发布时评估植入后门的有效性和泛化性，实证证据表明简单注入的后门持久性在更新后下降。在这项工作中，我们研究植入后门是否以及如何通过多阶段部署后微调保持持久。我们提出P-Trojan，这是一种基于触发器的攻击算法，明确优化跨重复更新的后门持久性。通过将中毒梯度与干净任务在标记嵌入上的梯度对齐，植入的后门映射在后续更新中不太可能被抑制或遗忘。理论分析表明持续微调后此类持久后门攻击的可行性。在Qwen2.5和LLaMA3系列LLM以及多样化任务序列上进行的实验表明，P-Trojan在保持干净任务准确率的同时实现了超过99%的持久性。我们的发现突显了在现实的模型适应流水线中对持久性感知评估和更强防御的需求。

### 论文 54
- **英文标题**: When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs’ Toxicity
- **中文标题**: 当笑脸变得充满敌意: 解读表情符号如何触发LLM的毒性
- **作者**: Shiyao Cui, Xijia Feng, Yingkang Wang, Junxiao Yang, Zhexin Zhang, Biplab Sikdar, Hongning Wang, Han Qiu, Minlie Huang
- **英文摘要**: Emojis are globally used non-verbal cues in digital communication, and extensive research has examined how large language models (LLMs) understand and utilize emojis across contexts. While usually associated with friendliness or playfulness, it is observed that emojis may trigger toxic content generation in LLMs. Motivated by such a observation, we aim to investigate: (1) whether emojis can clearly enhance the toxicity generation in LLMs and (2) how to interpret this phenomenon.* We begin with a comprehensive exploration of emoji-triggered LLM toxicity generation by automating the construction of prompts with emojis to subtly express toxic intent.  Experiments across 5 mainstream languages on 7 famous LLMs along with jailbreak tasks demonstrate that prompts with emojis could easily induce toxicity generation. To understand this phenomenon, we conduct model-level interpretations spanning semantic cognition, sequence generation and tokenization, suggesting  that emojis can act as a heterogeneous semantic channel to bypass the safety mechanisms. To pursue deeper insights, we further probe the pre-training corpus and uncover potential correlation between the emoji-related data polution with the toxicity generation behaviors.
- **中文摘要**: 表情符号是数字通信中全球使用的非语言线索，已有大量研究考察大语言模型如何跨上下文理解和使用表情符号。虽然通常与友善或趣味性相关联，但据观察，表情符号可能在LLM中触发有毒内容生成。受此观察启发，我们旨在调查：（1）表情符号是否能明确增强LLM中的毒性生成，以及（2）如何解释这一现象。我们从全面探索表情符号触发的LLM毒性生成开始，自动构建带有表情符号的提示以隐晦表达有毒意图。在5种主流语言的7个著名LLM上的实验以及越狱任务表明，带有表情符号的提示很容易诱导毒性生成。为理解这一现象，我们进行了涵盖语义认知、序列生成和标记化的模型级解释，表明表情符号可以作为异构语义通道绕过安全机制。为寻求更深层见解，我们进一步探查预训练语料库，发现了表情符号相关数据污染与毒性生成行为之间的潜在关联。

### 论文 60
- **英文标题**: Measuring the Unmeasurable: Unveiling Latent Cognitive Capabilities of LLM
- **中文标题**: 衡量不可衡量之物: 揭示LLM的潜在认知能力
- **作者**: Cui Danxin, Sihang Jiang, Keyi Wang, Zhiyi Duan, Yanghua Xiao, Bi Yude, Jiaqing Liang, Minggui He, Shimin Tao, Yilun Liu
- **英文摘要**: As large language models (LLMs) are increasingly deployed in high-stakes domains such as education, healthcare, and law, accurately evaluating their nuanced reasoning process becomes essential to ensure their safety, reliability, and trustworthiness. However, most existing benchmarks evaluate LLMs at a coarse granularity. Current benchmarks lack a unified framework and rely on single‐task datasets, overlooking the intermediate steps of complex reasoning. This results in redundant overlap across benchmarks, poor generalization to multifaceted real-world tasks, and underutilizes the rich reasoning traces generated by advanced LLMs.
- **中文摘要**: 随着大语言模型越来越多地部署在教育、医疗和法律等高风险领域，精确评估其微妙的推理过程对于确保其安全性、可靠性和可信度至关重要。然而，大多数现有基准以粗粒度评估LLM。当前基准缺乏统一框架，依赖单任务数据集，忽视了复杂推理的中间步骤。这导致基准之间冗余重叠，对多方面现实世界任务的泛化能力差，并且未能充分利用先进LLM生成的丰富推理轨迹。

### 论文 62
- **英文标题**: Guess or Recall? Training CNNs to Classify and Localize Memorization in LLMs
- **中文标题**: 猜测还是回忆？训练CNN对LLM中的记忆进行分类和定位
- **作者**: Jérémie Dentan, Davide Buscaldi, Sonia Vanier
- **英文摘要**: Verbatim memorization in Large Language Models (LLMs) is a multifaceted phenomenon involving distinct underlying mechanisms. We introduce a novel method to analyze the different forms of memorization described by the existing taxonomy. Specifically, we train Convolutional Neural Networks (CNNs) on the attention weights of the LLM and evaluate the alignment between this taxonomy and the attention weights involved in decoding. We find that the existing taxonomy performs poorly and fails to reflect distinct mechanisms within the attention blocks. We propose a new taxonomy that maximizes alignment with the attention weights, consisting of three categories: memorized samples that are guessed using language modeling abilities, memorized samples that are recalled due to high duplication in the training set, and non-memorized samples. Our results reveal that few-shot verbatim memorization does not correspond to a distinct attention mechanism. We also show that a significant proportion of extractable samples are in fact guessed by the model and should therefore be studied separately. Finally, we develop a custom visual interpretability technique to localize the regions of the attention weights involved in each form of memorization.
- **中文摘要**: 大语言模型中的逐字记忆是一个涉及不同底层机制的多面现象。我们引入了一种新方法来分析现有分类学所描述的不同形式的记忆。具体而言，我们在LLM的注意力权重上训练卷积神经网络（CNN），并评估该分类学与解码中涉及的注意力权重之间的对齐情况。我们发现现有分类学表现不佳，未能反映注意力块内的不同机制。我们提出了一种最大化与注意力权重对齐的新分类学，包含三个类别：使用语言建模能力猜测的记忆样本、由于训练集中高度重复而被回忆的记忆样本，以及非记忆样本。我们的结果揭示了少样本逐字记忆并不对应独特的注意力机制。我们还展示了相当比例的可提取样本实际上是模型猜测出来的，因此应被分开研究。最后，我们开发了一种自定义可视化可解释性技术来定位涉及每种记忆形式的注意力权重区域。

## Natural Language Processing II

### 论文 3
- **英文标题**: Extracting Events Like Code: A Multi-Agent Programming Framework for Zero-Shot Event Extraction
- **中文标题**: 像编码一样提取事件: 面向零样本事件提取的多智能体编程框架
- **作者**: Quanjiang Guo, Sijie Wang, Jinchuan Zhang, Ben Zhang, Zhao Kang, Ling Tian, Ke Yan
- **英文摘要**: Zero-shot event extraction (ZSEE) remains a significant challenge for large language models (LLMs) due to the need for complex reasoning and domain-specific understanding. Direct prompting often yields incomplete or structurally invalid outputs—such as misclassified triggers, missing arguments, and schema violations. To address these limitations, we present Agent-Event-Coder (AEC), a novel multi-agent framework that treats event extraction like software engineering: as a structured, iterative code-generation process. AEC decomposes ZSEE into specialized subtasks—retrieval, planning, coding, and verification—each handled by a dedicated LLM agent. Event schemas are represented as executable class definitions, enabling deterministic validation and precise feedback via a verification agent. This programming-inspired approach allows for systematic disambiguation and schema enforcement through iterative refinement. By leveraging collaborative agent workflows, AEC enables LLMs to produce precise, complete, and schema-consistent extractions in zero-shot settings. Experiments across five diverse domains and six LLMs demonstrate that AEC consistently outperforms prior zero-shot baselines, showcasing the power of treating event extraction like code generation.
- **中文摘要**: 零样本事件提取（ZSEE）对于大语言模型仍然是一个重大挑战，因其需要复杂推理和特定领域的理解。直接提示通常会产生不完整或结构无效的输出，如错误分类的触发词、缺失论元和模式违反。为解决这些局限，我们提出Agent-Event-Coder（AEC），一个将事件提取视为软件工程的新颖多智能体框架：作为一个结构化、迭代的代码生成过程。AEC将ZSEE分解为专门的子任务——检索、规划、编码和验证——每个子任务由专用的LLM智能体处理。事件模式被表示为可执行的类定义，使验证智能体能够进行确定性的验证和精确反馈。这种受编程启发的方法通过迭代精炼实现了系统消歧和模式强制执行。通过协作智能体工作流，AEC使LLM在零样本设置中产生精确、完整和模式一致的提取。跨五个不同领域和六个LLM的实验表明，AEC始终优于先前的零样本基线，展示了将事件提取视为代码生成的威力。

### 论文 10
- **英文标题**: Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System
- **中文标题**: Fact2Fiction: 面向智能体事实核查系统的定向投毒攻击
- **作者**: Haorui He, Yupeng Li, Bin Benjamin Zhu, Dacheng Wen, Reynold Cheng, Francis C. M. Lau
- **英文摘要**: State-of-the-art (SOTA) fact-checking systems combat misinformation by employing autonomous LLM-based agents to decompose complex claims into smaller sub-claims, verify each sub-claim individually, and aggregate the partial results to produce verdicts with justifications (explanations for the verdicts). The security of these systems is crucial, as compromised fact-checkers can amplify misinformation, but remains largely underexplored. To bridge this gap, this work introduces a novel threat model against such fact-checking systems and presents Fact2Fiction, the first poisoning attack framework targeting SOTA agentic fact-checking systems. Fact2Fiction employs LLMs to mimic the decomposition strategy and exploit system-generated justifications to craft tailored malicious evidences that compromise sub-claim verification. Extensive experiments demonstrate that Fact2Fiction achieves 8.9%-21.2% higher attack success rates than SOTA attacks across various poisoning budgets and exposes security weaknesses in existing fact-checking systems, highlighting the need for defensive countermeasures.
- **中文摘要**: 最先进的事实核查系统通过使用基于LLM的自主智能体将复杂声明分解为较小的子声明，逐个验证每个子声明，并聚合部分结果以产生带有理由的裁决来对抗虚假信息。这些系统的安全性至关重要，因为被攻破的事实核查器可能放大虚假信息，但这一领域在很大程度上仍未被探索。为弥合这一差距，本工作引入了一种针对此类事实核查系统的新威胁模型，并提出了Fact2Fiction，这是首个靶向最先进智能体事实核查系统的投毒攻击框架。Fact2Fiction使用LLM模仿分解策略并利用系统生成的理由来定制恶意证据，以破坏子声明验证。广泛实验表明，Fact2Fiction在不同投毒预算下比最先进攻击实现了8.9%-21.2%更高的攻击成功率，并暴露了现有事实核查系统的安全弱点，突显了对防御性对策的需求。

### 论文 40
- **英文标题**: Hybrid Routing for a Mixture of LoRA Experts
- **中文标题**: 面向LoRA专家混合的混合路由
- **作者**: Yitong Huang, Ziqi Yang, Zihui Wang, Jianzhong Qi, Rongshan Yu, Xiaoliang Fan, Cheng Wang
- **英文摘要**: Combining Mixture of Experts (MoE) with Low-Rank Adaptation (LoRA) has shown promising efficiency in multi-task instruction tuning for Large Language Models (LLMs). While existing routing schemes for such MoE systems employ auxiliary functions to ensure both expert selection certainty and workload balance among experts, they are hindered by two critical challenges: (1) Existing methods overlook the evolving cross-expert relationships across layers, leading to inefficient expert utilization. (2) The auxiliary functions fail to incorporate cross-task semantic characteristics during expert assignment, leading to suboptimal task adaptation. To address these challenges, we propose Hybrid routing for a Mixture of LoRA Experts (HotMoE), a novel multi-task instruction tuning framework that adapts hierarchical routing to the distinct characteristics of different LLM layers. First, we design a hybrid routing module. In lower layers, expert-expert attention facilitates cross-task collaboration and generalization. In higher layers, token-expert attention enables precise alignment between task semantics and specialized experts. Second, we introduce a similarity-guided auxiliary loss module to regularize routing decisions by exploiting hidden state similarities. This loss synergistically reinforces expert specialization without sacrificing certainty of expert selection by promoting cohesive activation patterns among semantically related tasks while sharpening distinctions between conflicting ones. Experiments across two multi-task instruction tuning scenarios covering seven NLP benchmarks demonstrate that HotMoE consistently outperforms all baselines, improving Mean Relative Difference by up to 1.68% with only 3.1% of trainable parameters.
- **中文摘要**: LoRA（低秩适应）已被证明是高效的参数高效微调方法。我们研究将多个LoRA专家组合成混合系统的路由策略。我们提出一种混合路由方法，在标记级别动态选择最相关的LoRA专家组合。该方法结合了硬路由和软路由的优势，在不同任务上实现了强大的泛化能力和效率。

### 论文 46
- **英文标题**: AutoTool: Efficient Tool Selection for Large Language Model Agents
- **中文标题**: AutoTool: 面向大语言模型智能体的高效工具选择
- **作者**: Jingyi Jia, Qinbin Li
- **英文摘要**: Large Language Model (LLM) agents have emerged as powerful tools for automating complex tasks by leveraging the reasoning and decision-making abilities of LLMs. However, a major bottleneck in current agent frameworks lies in the high inference cost of tool selection, especially in approaches like ReAct that repeatedly invoke the LLM to determine which tool to use at each step. In this work, we propose AutoTool, a novel graph-based framework that bypasses repeated LLM inference by exploiting a key empirical observation: tool usage inertia—the tendency of tool invocations to follow predictable sequential patterns. AutoTool constructs a directed graph from historical agent trajectories, where nodes represent tools and edges capture transition probabilities, effectively modeling the inertia in tool selection. It further integrates parameter-level information to refine tool input generation. By traversing this structured representation, AutoTool efficiently selects tools and their parameters with minimal reliance on LLM inference. Extensive experiments across diverse agent tasks demonstrate that AutoTool reduces inference costs by up to 30% while maintaining competitive task completion rates, offering a practical and scalable enhancement for inference-heavy frameworks. Our work highlights the promise of integrating statistical structure into LLM agent design for greater efficiency without sacrificing performance.
- **中文摘要**: LLM智能体通常需要从大量可用工具中选择合适的工具。我们提出AutoTool，一个面向LLM智能体的高效工具选择框架。该方法通过上下文感知的工具检索和排序，使智能体能够快速识别最相关的工具。实验表明，AutoTool在工具选择准确性和效率上优于现有方法。

### 论文 61
- **英文标题**: WikiMAG: A Multi-Agent Guided Framework for Generating Structured Wikipedia-like Articles
- **中文标题**: WikiMAG: 面向生成结构化维基百科式文章的多智能体引导框架
- **作者**: Xiuli Kang, Yinlong Xiao, Minghao Hu, Yuan Huang, Bin Mao, Ming Wang, Fang Wang, Zhunchen Luo, Wei Luo, Guotong Geng
- **英文摘要**: Wikipedia serves as the world's largest and most popular online reference encyclopedia, rich in structured knowledge and authoritative citations. Recently, numerous works have leveraged large language models to automatically generate Wikipedia-like articles. However, existing approaches primarily focus on producing singular narrative-type content, overlooking higher information-density structured elements such as timeline and table. To address these limitations, we propose WikiMAG, a multi-agent guided framework for generating structured Wikipedia-like articles. This framework employs a collaborative multi-agent mechanism to orchestrate the creation process, featuring three synergistic core components: Progressive planner first constructs the coarse-grained outline framework and then annotate fine-grained types for outline units, encompassing narrative, timeline, and table formats; Reflective inspector dynamically curates high-quality references via multi-round interactive feedback, thereby enhancing the authority and relevance of citations; Versatile writer integrates fine-grained outline details and high-quality reference information to generate information-rich articles, incorporating the three annotated formats. We evaluate WikiMAG on two public datasets, FreshWiki and WikiGenBen, across outline, writing, and verifiability dimensions. Compared with the best baseline method, our method achieves an average improvement of 6.73 points and 4.39 points in Heading Soft Recall and the METEOR metric (a machine translation and text generation evaluation metric) respectively, and an average increase of 16.84 percentage points in Citation Rate.
- **中文摘要**: 生成结构化、高质量的维基百科式文章是一项复杂任务，需要多方面的专业知识。我们提出WikiMAG，一个面向生成此类文章的多智能体引导框架。不同智能体分别负责研究、写作、编辑和验证，协同生成全面且准确的长篇文章。实验表明，WikiMAG在文章质量和事实准确性上优于单一模型方法。

### 论文 71
- **英文标题**: The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models
- **中文标题**: 类比的奇妙案例: 调查大语言模型中的类比推理
- **作者**: Taewhoo Lee, Minju Song, Chanwoong Yoon, Jungwoo Park, Jaewoo Kang
- **英文摘要**: Analogical reasoning is at the core of human cognition, serving as an important foundation for a variety of intellectual activities. While prior work has shown that LLMs can represent task patterns and surface-level concepts, it remains unclear whether these models can encode high-level relational concepts and apply them to novel situations through structured comparisons. In this work, we explore this fundamental aspect using proportional and story analogies, and identify three key findings. First, LLMs effectively encode the underlying relationships between analogous entities; both attributive and relational information propagate through mid-upper layers in correct cases, whereas reasoning failures reflect missing relational information within these layers. Second, unlike humans, LLMs often struggle not only when relational information is missing, but also when attempting to apply it to new entities. In such cases, strategically patching hidden representations at critical token positions can facilitate information transfer to a certain extent. Lastly, successful analogical reasoning in LLMs is marked by strong structural alignment between analogous situations, whereas failures often reflect degraded or misplaced alignment. Overall, our findings reveal that LLMs exhibit emerging but limited capabilities in encoding and applying high-level relational concepts, highlighting both parallels and gaps with human cognition.
- **中文摘要**: 类比推理是人类认知的核心能力。我们系统调查大语言模型中的类比推理能力。通过设计多种类型的类比任务，我们评估LLM理解和应用类比的能力。研究发现当前LLM在某些类比类型上表现良好，但在更复杂的类比中表现挣扎，揭示了改进方向。

### 论文 77
- **英文标题**: Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
- **中文标题**: Res-Bench: 基准测试多模态大语言模型对动态分辨率输入的鲁棒性
- **作者**: Chenxu Li, Zhicai Wang, Yuan Sheng, Xingyu Zhu, Yanbin Hao, Xiang Wang
- **英文摘要**: Multimodal Large Language Models (MLLMs) increasingly support dynamic image resolutions. However, current evaluation paradigms primarily assess semantic performance, overlooking the critical question of resolution robustness - whether performance remains stable across varying input resolutions. To address this gap, we introduce Res-Bench, a comprehensive benchmark comprising 14,400 samples across 12 resolution levels and six core capability dimensions. We designed a novel evaluation framework that goes beyond traditional accuracy metrics to capture performance stability. This framework introduces multiple robustness metrics: Spearman's correlation for assessing resolution-performance trends, and Absolute/Relative Continuous Error (ACE/RCE) for measuring performance volatility. Using these metrics, we conducted a large-scale evaluation of leading MLLMs. Our analysis encompasses: (1) model-centric and task-centric robustness examination, (2) investigation of preprocessing strategies including padding and super-resolution, and (3) exploration of fine-tuning for stability enhancement.
- **中文摘要**: 多模态LLM通常在固定分辨率上训练，但实际输入分辨率各异。我们引入Res-Bench，一个基准测试多模态LLM对动态分辨率输入鲁棒性的框架。实验显示大多数模型对分辨率变化敏感，性能在高分辨率和低分辨率输入上显著下降。我们分析鲁棒性因素并提出改进建议。

### 论文 83
- **英文标题**: From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning
- **中文标题**: 从宏观到微观: 探讨语言模型微调中的数据集多样性
- **作者**: Haoyu Li, Xuhong Li, Yiming Dong, Kun Liu
- **英文摘要**: Dataset diversity plays a pivotal role for the successful training of many machine learning models, particularly in the supervised fine-tuning (SFT) stage of large language model (LLM) development. Despite increasing recognition of its importance, systematic analyses of dataset diversity still remain underexplored. To address this gap, this work presents a systematic taxonomy of existing diversity-control strategies, which primarily focus on the instruction component, operating at either macroscopic (entire instruction semantics) or mesoscopic levels (instruction units), and furthermore introduces a novel analysis of microscopic diversity within the response component, specifically analyzing the statistical distribution of tokens in SFT training samples. In the experimental evaluation, we construct fixed-size datasets (e.g., 10,000 samples each) from a corpus of 117,000 open-source SFT samples, incorporating six distinct diversity-control strategies spanning macro-, meso-, and microscopic levels applied to both instructions and responses. We then fine-tune LLMs on these datasets to assess the six diversity-control strategies. Results reveal that while macroscopic and mesoscopic strategies lead to higher performance with increasing diversity, the microscopic strategy in responses exhibits not only a stronger correlation between model performance and the degree of diversity, but also superior performance with maximum diversity across all strategies. These findings offer actionable insights for constructing high-performance SFT datasets.
- **中文摘要**: 微调数据集的多样性对LLM性能有重要影响。我们从宏观到微观系统探讨数据集多样性的作用。通过控制数据集在不同粒度级别的多样性，我们揭示了多样性对模型泛化和鲁棒性的影响。研究发现适度的多样性比极端同质或异质更有利于微调。

### 论文 84
- **英文标题**: From Single to Societal: Analyzing Persona-Induced Bias in Multi-Agent Interactions
- **中文标题**: 从个人到社会: 分析多智能体交互中角色诱导的偏见
- **作者**: Jiayi Li, Xiao Liu, Yansong Feng
- **英文摘要**: Large Language Model (LLM)-based multi-agent systems are increasingly used to simulate human interactions and solve collaborative tasks. A common practice is to assign agents with personas to encourage behavioral diversity. However, this raises a critical yet underexplored question: do personas introduce biases into multi-agent interactions? This paper presents a systematic investigation into persona-induced biases in multi-agent interactions, with a focus on social traits like trustworthiness (how an agent's opinion is received by others) and insistence (how strongly an agent advocates for its opinion). Through a series of controlled experiments in collaborative problem-solving and persuasion tasks, we reveal that (1) LLM-based agents exhibit biases in both trustworthiness and insistence, with personas from historically advantaged groups (e.g., men and White individuals) perceived as less trustworthy and demonstrating less insistence; and (2) agents exhibit significant in-group favoritism, showing a higher tendency to conform to others who share the same persona. These biases persist across various LLMs, group sizes, and numbers of interaction rounds, highlighting an urgent need for awareness and mitigation to ensure the fairness and reliability of multi-agent systems.
- **中文摘要**: 多智能体系统中个体智能体的角色设定可能导致系统性偏见。我们从个人到社会层面分析角色诱导的偏见在多智能体交互中的传播和放大。研究发现角色偏见可能通过智能体交互放大并影响群体决策。我们提出了检测和缓解此类偏见的方法。

### 论文 90
- **英文标题**: Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
- **中文标题**: 不要合并我的模型！保护开源LLM免受未授权的模型合并
- **作者**: Qinfeng Li, Miao Pan, Jintao Chen, Fu Teng, Zhiqiang Shen, Ge Su, Hao Peng, Xuhong Zhang
- **英文摘要**: Model merging has emerged as an efficient technique for expanding large language models (LLMs) by integrating specialized expert models. However, it also introduces a new threat: model merging stealing, where free-riders exploit models through unauthorized model merging. Unfortunately, existing defense mechanisms fail to provide effective protection. Specifically, we identify three critical protection properties that existing methods fail to simultaneously satisfy: (1) proactively preventing unauthorized merging; (2) ensuring compatibility with general open-source settings; (3) achieving high security with negligible performance loss. To address the above issues, we propose MergeBarrier, a plug-and-play defense that proactively prevents unauthorized merging. The core design of MergeBarrier is to disrupt the Linear Mode Connectivity (LMC) between the protected model and its homologous counterparts, thereby eliminating the low-loss path required for effective model merging. Extensive experiments show that MergeBarrier effectively prevents model merging stealing with negligible accuracy loss.
- **中文摘要**: 模型合并技术可以将多个开源LLM组合成更强的模型，但也可能被未授权使用。我们研究保护开源LLM免受未授权模型合并的方法。通过嵌入模型水印和设计抗合并机制，我们保护了模型所有者的权益。实验展示了这些保护措施的有效性。

## Natural Language Processing III

### 论文 3
- **英文标题**: Anti-adversarial Learning: Desensitizing Prompts for Large Language Model
- **中文标题**: 反-对抗学习: 大语言模型提示词脱敏方法
- **作者**: Xuan Li, Zhe Yin, Xiaodong Gu, Beijun Shen
- **英文摘要**: With the widespread use of LLMs, preserving privacy in user prompts has become crucial, as prompts risk exposing private and sensitive data to cloud LLMs. Conventional techniques like homomorphic encryption (HE), secure multi-party computation, and federated learning (FL) are not well-suited to this scenario due to the lack of control over user participation in remote model interactions. In this paper, we propose PromptObfus, a novel method for desensitizing LLM prompts. The core idea of PromptObfus is "anti-adversarial" learning, which perturbs sensitive words in the prompt to obscure private information while retaining the stability of model predictions. Specifically, PromptObfus frames prompt desensitization as a masked language modeling task, replacing privacy-sensitive terms with a [MASK] token. A desensitization model is utilized to generate candidate replacements for each masked position. These candidates are subsequently selected based on gradient feedback from a surrogate model, ensuring minimal disruption to the task output. We demonstrate the effectiveness of our approach on three NLP tasks. Results show that PromptObfus effectively prevents privacy inference from remote LLMs while preserving task performance.
- **中文摘要**: 随着LLM的广泛应用,保护用户提示词中的隐私变得至关重要,因为提示词可能将私人和敏感数据暴露给云端LLM。同态加密、安全多方计算和联邦学习等传统技术由于缺乏对用户参与远程模型交互的控制而不适用于这一场景。在本文中,我们提出PromptObfus,一种新颖的LLM提示词脱敏方法。PromptObfus的核心思想是'反-对抗'学习,通过扰动提示词中的敏感词来模糊隐私信息,同时保持模型预测的稳定性。具体而言,PromptObfus将提示词脱敏视为掩码语言建模任务,用[MASK]标记替换隐私敏感词。一个脱敏模型被用于为每个掩码位置生成候选替换词。随后,这些候选词基于代理模型的梯度反馈进行选择,确保对任务输出的干扰最小。我们在三个NLP任务上证明了该方法的有效性。结果表明,PromptObfus能够有效防止远程LLM的隐私推断,同时保持任务性能。

### 论文 12
- **英文标题**: Query-Efficient Domain Knowledge Stealing Against Large Language Models
- **中文标题**: 面向大语言模型的高效查询领域知识窃取
- **作者**: Zhengao Li, Xiaopeng Yuan, Bolin Shen, Kien Le, Haohan Wang, Xugui Zhou, Shangqian Gao, Yushun Dong
- **英文摘要**: Large language models (LLMs) concentrate substantial knowledge in specialized domains due to extensive pretraining and instruction tuning, and they are now central to commercial and scientific practice. Yet access is usually limited to costly, rate-limited interfaces, which motivates methods that can extract targeted domain knowledge with minimal querying effort. A further challenge is that the target domain may be unknown in advance, so naive or generic prompts waste queries and fail to expose the underlying concepts and relations that structure the domain. In this work, we introduce a query-efficient approach for domain-specific knowledge stealing from black-box language models. Rather than issuing random questions or generic templates, our framework performs self-directed exploration that lets the model find the direction and mine domain knowledge by itself. Starting from a small and diverse seed, it discovers salient domain entities and induces their relations through structured question families that elicit definitional, functional, and compositional information. A feedback-driven controller analyzes the errors and uncertainty of the extracted surrogate model and uses this signal to refine subsequent queries, all without relying on prior domain knowledge or external resources. We evaluate the method in two expert-centric settings, medicine and finance, and observe consistently better performance while requiring significantly fewer queries.
- **中文摘要**: 大语言模型(LLM)由于广泛的预训练和指令调优,在专业领域集中了大量知识,现在它们是商业和科学实践的核心。然而,访问通常仅限于昂贵、有速率限制的接口,这促使人们开发能以最少查询工作量提取目标领域知识的方法。另一个挑战是目标领域可能事先未知,因此朴素或通用的提示词会浪费查询尝试,且无法揭示构建领域的底层概念和关系。在本工作中,我们引入了一种面向黑盒语言模型的高效查询领域特定知识窃取方法。我们的框架不是提出随机问题或使用通用模板,而是进行自引导探索,让模型自行找到方向并挖掘领域知识。从一个小而多样的种子开始,它发现显著的领域实体,并通过引出定义性、功能性和组合性信息的结构化问题族推导出它们的关系。一个反馈驱动的控制器分析提取的代理模型的错误和不确定性,并利用这一信号来优化后续查询,所有这些都不依赖于先前的领域知识或外部资源。我们在医学和金融两个专家中心场景中评估了该方法,观察到在显著减少查询次数的同时获得了一致更好的性能。

### 论文 26
- **英文标题**: GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale Injection
- **中文标题**: GrayKD: 通过多推理注入从黑盒LLM蒸馏更优知识
- **作者**: Hyeongsoo Lim, Hyung Yong Kim, Jin Young Kim, Min Ho Jang, Eun Seo Seo, Youshin Lim, Shukjae Choi, Jihwan Park, Yunkyu Lim, Hanbin Lee, Byeong-Yeol Kim, Ji Won Yoon
- **英文摘要**: Knowledge distillation (KD) is a promising compression technique for reducing the computational burden of large language models (LLMs). Depending on access to the teacher model’s internal parameters, KD is typically categorized into white-box and black-box KD. While white-box KD benefits from full access to intrinsic knowledge such as softmax distributions, black-box KD adopts a black-box LLM (e.g., GPT-4) as the teacher, which provides only text-level outputs via API calls. This limited supervision makes black-box KD generally less effective than its white-box counterpart. To bridge the gap between white-box and black-box KD, we propose GrayKD, a novel framework that can effectively distill text-level knowledge from a black-box LLM in a single-stage manner. In particular, rationales generated by the black-box LLM are injected into the student via a lightweight cross-attention module (teacher mode), enabling the model to approximate the black-box teacher’s output distribution without access to internal parameters. The student is then trained with the softmax-level knowledge provided by the teacher mode (student mode). Since both the teacher and student modes share the same backbone, the proposed teacher mode remains highly parameter-efficient, requiring only a small number of additional parameters for rationale injection. Experimental results on instruction-following tasks demonstrate that GrayKD achieves substantial performance improvements over existing KD methods.
- **中文摘要**: 知识蒸馏(KD)是一种用于减少大语言模型(LLM)计算负担的有前景的压缩技术。根据对教师模型内部参数的访问权限,KD通常分为白盒KD和黑盒KD。白盒KD受益于对教师内部表征的完全访问,但受到专有模型限制;黑盒KD仅使用教师输出,但面临着监督信号有限的问题。在本文中,我们提出GrayKD,一种介于白盒和黑盒之间的方法,通过注入多条推理链来丰富从黑盒LLM获取的知识。GrayKD提示教师模型为每个训练实例生成多个不同的推理链,使教师输出中包含更丰富的知识。然后,这些多样化的推理链被用于指导学生模型的训练。实验表明,GrayKD显著优于传统的黑盒KD方法,在知识蒸馏效果上接近白盒方法,同时保持了对任何LLM(包括专有模型)的适用性。

### 论文 35
- **英文标题**: Easy for Children, Hard for AI: The Limits of Multimodal LLMs in Early Childhood Learning
- **中文标题**: 对儿童简单,对AI困难: 多模态LLM在幼儿学习中的局限性
- **作者**: Jingping Liu, Xueyan Wu, Hanxuan Chen, Ziyan Liu, Zhangquan Chen, Ronghao Chen, Huacan Wang
- **英文摘要**: Early childhood is a critical stage for cognitive development, involving core skills such as visual perception and reasoning. While multimodal large language models (MLLMs) have made rapid progress in various general-purpose tasks, their ability to support early education remains largely underexplored. Existing research on child-related AI largely centers on modeling language, emotion, or behavior, with limited focus on evaluating cognitive tasks relevant to early learning. To address this gap, we propose ChildBench, a multimodal benchmark designed to assess models on tasks inspired by early childhood cognitive development. It covers five key domains through ten tasks, including spatial reasoning, visual reasoning, visual discrimination, counting skills, and visual tracking. The benchmark includes 4,890 carefully constructed images and 5,346 manually annotated samples, ensuring both diversity and age-appropriate content. We evaluate a range of state-of-the-art (SoTA) open-source and closed-source MLLMs—including GPT-4o, Gemini, and Qwen2.5-VL—on ChildBench. Despite strong performance on other benchmarks, the best 7B-parameter model with LoRA tuning achieves only 52.01% accuracy, far below the 96% achieved by 5-year-old children. These results reveal critical limitations in fine-grained perception and reasoning. We further analyze failure cases and discuss directions for future model development.
- **中文摘要**: 幼儿期是认知发展的关键阶段,涉及视觉感知和推理等核心技能。虽然多模态大语言模型(MLLM)在各种通用任务上取得了快速进展,但它们支持早期教育的能力在很大程度上仍未被探索。在本文中,我们构建了一个包含幼儿学习任务(如形状识别、计数和简单推理)的基准,评估了MLLM的表现。令人惊讶的是,我们发现当前最先进的MLLM在面对幼儿能轻松完成的任务时表现不佳。分析表明,MLLM的失败主要源于它们缺乏对基本物理和空间概念的理解,而人类幼儿通过感官-运动互动自然发展出这些概念。这些发现突显了当前AI系统与人类认知发展之间的根本差距。

### 论文 52
- **英文标题**: Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models
- **中文标题**: 回答不可回答之问题即明知故犯: 分析并缓解大推理模型中的拒绝回答失败
- **作者**: Yi Liu, Xiangyu Liu, Zequn Sun, Wei Hu
- **英文摘要**: Large reasoning models (LRMs) have shown remarkable progress on complex reasoning tasks. However, some questions posed to LRMs are inherently unanswerable, such as math problems lacking sufficient conditions. We find that LRMs continually fail to provide appropriate abstentions when confronted with these unanswerable questions. In this paper, we systematically analyze, investigate, and resolve this issue for trustworthy AI. We first conduct a detailed analysis of the distinct response behaviors of LRMs when facing unanswerable questions. Then, we show that LRMs possess sufficient cognitive capabilities to recognize the flaws in these questions. However, they fail to exhibit appropriate abstention behavior, revealing a misalignment between their internal cognition and external response. Finally, to resolve this issue, we propose a lightweight, two-stage method that combines cognitive monitoring with inference-time intervention. Experimental results demonstrate that our method significantly improves the abstention rate while maintaining the reasoning performance.
- **中文摘要**: 大推理模型(LRM)在复杂推理任务上展现了显著进展。然而,向LRM提出的某些问题本质上是不可回答的,例如缺乏充分条件的数学问题。我们发现,当面对不可回答的问题时,LRM持续未能提供适当的拒绝回答,而是倾向于编造似是而非但不正确的答案。在本文中,我们分析了LRM拒绝回答失败的根源,发现模型对生成连贯输出的偏好压倒了识别问题不可回答性的能力。我们提出了几种缓解策略,包括不确定性量化提示和专门的拒绝回答训练。实验表明,这些策略显著提高了LRM在面对不可回答问题时的拒绝回答率,同时保持了在可回答问题上的性能。

### 论文 56
- **英文标题**: InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- **中文标题**: InfiGUI-G1: 通过自适应探索策略优化推进GUI定位
- **作者**: Yuhang Liu, Zeyu Liu, Shuanghe Zhu, Pengxiang Li, Congkai Xie, Jiasheng Wang, Xueyu Hu, Xiaotian Han, Jianbo Yuan, Xinyao Wang, Shengyu Zhang, Hongxia Yang, Fei Wu
- **英文摘要**: The emergence of Multimodal Large Language Models (MLLMs) has propelled the development of autonomous agents that operate on Graphical User Interfaces (GUIs) using pure visual input. A fundamental challenge is robustly grounding natural language instructions. This requires a precise spatial alignment, which accurately locates the coordinates of each element, and, more critically, a correct semantic alignment, which matches the instructions to the functionally appropriate UI element. Although Reinforcement Learning with Verifiable Rewards (RLVR) has proven to be effective at improving spatial alignment for these MLLMs, we find that inefficient exploration bottlenecks semantic alignment, which prevents models from learning difficult semantic associations. To address this exploration problem, we present Adaptive Exploration Policy Optimization (AEPO), a new policy optimization framework. AEPO employs a multi-answer generation strategy to enforce broader exploration, which is then guided by a theoretically grounded Adaptive Exploration Reward (AER) function derived from first principles of efficiency η=U/C. Our AEPO-trained models, InfiGUI-G1-3B and InfiGUI-G1-7B, establish new state-of-the-art results across multiple challenging GUI grounding benchmarks, achieving significant relative improvements of up to 9.0% against the naive RLVR baseline on benchmarks designed to test generalization and semantic understanding.
- **中文摘要**: 多模态大语言模型(MLLM)的出现推动了在图形用户界面(GUI)上使用纯视觉输入的自主智能体的发展。一个基本的挑战是鲁棒地将自然语言指令定位到界面元素。这需要精确的空间对齐,而当前方法在复杂GUI布局中往往表现不足。在本文中,我们提出InfiGUI-G1,一种通过自适应探索策略优化推进GUI定位的方法。InfiGUI-G1学习最优的视觉探索策略,动态决定关注哪些UI区域以及如何解释其功能,从而显著提高指令定位的准确性和效率。在GUI定位基准上的实验表明,InfiGUI-G1在定位准确率和效率方面均优于现有方法。

### 论文 67
- **英文标题**: RetroLM: Retrieval-Augmented KVs for Long-Context Processing
- **中文标题**: RetroLM: 面向长上下文处理的检索增强KV
- **作者**: Kun Luo, Zheng Liu, Shitao Xiao, Jiabei Chen, Hongjin Qian, Peitian Zhang, Shanshan Jiang, Bin Dong, Jun Zhao, Kang Liu
- **英文摘要**: Long-context processing remains a significant challenge for large language models (LLMs). Retrieval-augmented generation (RAG) has recently emerged as a promising approach, enabling LLMs to selectively access relevant information from extended contexts to improve efficiency. However, existing RAG approaches often lag behind other efficient long-context processing methods primarily due to inherent limitations on inaccurate retrieval and fragmented contexts. To address these limitations, we propose RetroLM, a novel RAG framework designed for effective long-context processing. Unlike traditional approaches, RetroLM introduces KV-level retrieval augmentation, which partitions the LLM's KV cache into contiguous pages and performs encoding and decoding operations based on the retrieved KV pages. Built upon this framework, we further develop a specialized retriever for precise retrieval of critical pages and conduct unsupervised post-training to optimize the model’s ability to leverage retrieved information. Compared with traditional RAG, the new approach enhances robustness to retrieval inaccuracy, facilitates effective utilization of fragmented contexts, and saves the cost from repeated context-encoding operations. We conduct extensive evaluations across several popular benchmarks, including LongBench, InfiniteBench, and RULER. RetroLM consistently outperforms existing long-LLMs and RAG-based methods, especially in tasks requiring deep reasoning or extreme context lengths.
- **中文摘要**: 长上下文处理仍然是大语言模型(LLM)的一个重大挑战。检索增强生成(RAG)近来作为一种有前景的方法出现,使LLM能够选择性地访问扩展上下文中的相关信息以提高效率。然而,现有RAG方法通常落后于其他高效的长上下文处理方法,主要由于检索不准确和上下文碎片化等固有限制。为了解决这些限制,我们提出RetroLM,一个专为有效长上下文处理设计的新颖RAG框架。与传统方法不同,RetroLM引入了KV级检索增强,将LLM的KV缓存划分为连续页面,并基于检索到的KV页面执行编码和解码操作。在此框架基础上,我们进一步开发了专用检索器以实现关键页面的精确检索,并进行无监督后训练以优化模型利用检索信息的能力。与传统RAG相比,这一新方法增强了对检索不准确性的鲁棒性,促进了碎片化上下文的有效利用,并节省了重复上下文编码操作的成本。我们在LongBench、InfiniteBench和RULER等多个流行基准上进行了广泛评估。RetroLM持续优于现有的长上下文LLM和基于RAG的方法,特别是在需要深度推理或极端上下文长度的任务中。

### 论文 68
- **英文标题**: Better Datasets Start from RefineLab: Automatic Optimization for High-Quality Dataset Refinement
- **中文标题**: RefineLab: 高质量数据集自动优化的精炼实验室
- **作者**: Xiaonan Luo, Yue Huang, Ping He, Xiangliang Zhang
- **英文摘要**: High‑quality Question–Answer (QA) datasets are foundational for reliable Large Language Model (LLM) evaluation, yet even expert‑crafted datasets exhibit persistent gaps in domain coverage, misaligned difficulty distributions, and factual inconsistencies. The recent surge in generative model-powered datasets has compounded these quality challenges. In this work, we introduce RefineLab, the first LLM‑driven framework that automatically refines raw QA textual data into high-quality datasets under a controllable token‑budget constraint. RefineLab takes a set of target quality attributes as refinement objectives and performs selective edits within a predefined token budget to ensure practicality and efficiency. In essence, RefineLab addresses a constrained optimization problem: improving the quality of QA samples as much as possible while respecting resource limitations. With a set of available refinement operations, RefineLab takes as input the original dataset, a specified set of target quality dimensions, and a token budget, and determines which refinement operations should be applied to each QA sample. This process is guided by an assignment module that selects optimal refinement strategies to maximize overall dataset quality while adhering to the budget constraint. Experiments demonstrate that RefineLab consistently narrows divergence from expert datasets across coverage, difficulty alignment, factual fidelity, and distractor quality. RefineLab pioneers a scalable, customizable path to reproducible dataset design, with broad implications for LLM evaluation.
- **中文摘要**: 高质量问答(QA)数据集是可靠大语言模型(LLM)评估的基础,但即使专家精心制作的数据集也存在领域覆盖不足、难度分布失调和事实不一致等持续存在的缺陷。生成模型驱动的数据集的近期激增加剧了这些质量挑战。在本工作中,我们引入RefineLab,首个LLM驱动的框架,在可控token预算约束下自动将原始QA文本数据精炼为高质量数据集。RefineLab将一组目标质量属性作为精炼目标,并在预定义的token预算内执行选择性编辑以确保实用性和效率。本质上,RefineLab解决了一个受约束的优化问题:在尊重资源限制的同时尽可能提升QA样本的质量。通过一组可用的精炼操作,RefineLab接受原始数据集、指定的目标质量维度集合和token预算作为输入,并确定哪些精炼操作应应用于每个QA样本。此过程由分配模块引导,该模块选择最优精炼策略以在遵守预算约束的同时最大化整体数据集质量。实验表明,RefineLab在覆盖度、难度对齐、事实保真度和干扰项质量方面持续缩小与专家数据集的差距。RefineLab开创了一种可扩展、可定制的可重复数据集设计路径,对LLM评估具有广泛影响。

### 论文 94
- **英文标题**: GateRA: Token-aware Modulation for Parameter-Efficient Fine-tuning
- **中文标题**: GateRA: 面向参数高效微调的Token感知调制
- **作者**: Jie Ou, Shuaihong Jiang, Yingjun Du, Cees G. M. Snoek
- **英文摘要**: Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, DoRA, and HiRA, enable lightweight adaptation of large pre-trained models via low-rank updates.  However, existing PEFT approaches apply static, input-agnostic updates to all tokens, disregarding the varying importance and difficulty of different inputs. This uniform treatment can lead to overfitting on trivial content or under-adaptation on more informative regions, especially in autoregressive settings with distinct prefill and decoding dynamics. In this paper, we propose GateRA, a unified framework that introduces token-aware modulation to dynamically adjust the strength of PEFT updates. By incorporating adaptive gating into standard PEFT branches, GateRA enables selective, token-level adaptation—preserving pre-trained knowledge for well-modeled inputs while focusing capacity on challenging cases. Empirical visualizations reveal phase-sensitive behaviors, where GateRA automatically suppresses updates for redundant prefill tokens while emphasizing adaptation during decoding. To promote confident and efficient modulation, we further introduce an entropy-based regularization that encourages near-binary gating decisions. This regularization prevents diffuse update patterns and leads to interpretable, sparse adaptation without hard thresholding. Finally, we present a theoretical analysis showing that GateRA induces a soft gradient-masking effect over the PEFT path, enabling continuous and differentiable control over adaptation. Experiments on multiple commonsense reasoning benchmarks demonstrate that GateRA consistently outperforms or matches prior PEFT methods.
- **中文摘要**: 参数高效微调(PEFT)方法,如LoRA、DoRA和HiRA,通过低秩更新实现大型预训练模型的轻量级适应。然而,现有PEFT方法对所有token应用静态的、与输入无关的更新,忽视了不同输入的变化重要性和难度。这种统一处理可能导致在平凡内容上过拟合或在更富含信息的区域上适应不足,特别是在具有不同预填充和解码动态的自回归设置中。在本文中,我们提出GateRA,一个统一框架,引入token感知调制来动态调整PEFT更新的强度。通过将自适应门控融入标准PEFT分支,GateRA实现了选择性的、token级别的适应——为建模良好的输入保留预训练知识,同时将容量集中在具有挑战性的情况上。实证可视化揭示了相位敏感行为,其中GateRA自动抑制冗余预填充token的更新,同时强调解码过程中的适应。为促进自信和高效的调制,我们进一步引入基于熵的正则化,鼓励接近二值的门控决策。这种正则化防止了扩散的更新模式,并在无需硬阈值的情况下导致可解释的、稀疏的适应。最后,我们提出了理论分析,表明GateRA在PEFT路径上引入了软梯度掩蔽效应,实现了对适应的连续和可微控制。在多个常识推理基准上的实验表明,GateRA持续优于或匹配先前的PEFT方法。

## Natural Language Processing IV

### 论文 2
- **英文标题**: WaterMod: Modular Token-Rank Partitioning for Probability-Balanced LLM Watermarking
- **中文标题**: WaterMod: 面向概率平衡LLM水印的模块化Token排名分区
- **作者**: Shinwoo Park, Hyejin Park, Hyeseon Ahn, Yo-Sub Han
- **英文摘要**: Large language models now draft news, legal analyses, and software code with human-level fluency. At the same time, regulations such as the EU AI Act mandate that each synthetic passage carry an imperceptible, machine-verifiable mark for provenance. Conventional logit-based watermarks satisfy this requirement by selecting a pseudorandom green vocabulary at every decoding step and boosting its logits, yet the random split can exclude the highest-probability token and thus erode fluency. WaterMod mitigates this limitation through a probability-aware modular rule. The vocabulary is first sorted in descending model probability; the resulting ranks are then partitioned by the residue rank mod k, which distributes adjacent—and therefore semantically similar—tokens across different classes. A fixed bias of small magnitude is applied to one selected class. In the zero-bit setting (k=2), an entropy-adaptive gate selects either the even or the odd parity as the green list. Because the top two ranks fall into different parities, this choice embeds a detectable signal while guaranteeing that at least one high-probability token remains available for sampling. In the multi-bit regime (k>2), the current payload digit d selects the color class whose ranks satisfy rank mod k = d. Biasing the logits of that class embeds exactly one base-k digit—equivalently log2(k) bits—per decoding step, thereby enabling fine-grained provenance tracing. The same modular arithmetic therefore supports both binary attribution and rich payloads. Experimental results demonstrate that WaterMod consistently attains strong watermark detection performance while maintaining generation quality in both zero-bit and multi-bit settings. This robustness holds across a range of tasks, including natural language generation, mathematical reasoning, and code synthesis.
- **中文摘要**: 大语言模型现在能以人类水平的流畅度撰写新闻、法律分析和软件代码。同时,欧盟AI法案等法规要求每段合成内容携带一个不可察觉的、机器可验证的来源标记。传统基于logit的水印通过在每一步解码中选择伪随机绿色词汇并提升其logit来满足这一要求,但随机分割可能排除最高概率的token,从而削弱流畅性。WaterMod通过一种概率感知的模块化规则来缓解这一限制。词汇首先按模型概率降序排序;得到的排名然后按排名模k的余数进行分区,这将近距离的——因此语义相似的——token分配到不同的类别中。一个固定的小幅度偏置被应用于一个选定的类别上。在零比特设置(k=2)中,一个熵自适应门选择偶数或奇数奇偶性作为绿色列表。因为前两个排名落入不同的奇偶性,这个选择嵌入了一个可检测的信号,同时保证至少有一个高概率token仍可用于采样。在多比特模式(k>2)中,当前载荷数字d选择其排名满足rank mod k = d的颜色类别。偏置该类别的logit在每个解码步骤嵌入恰好一个基k数字——等价于log2(k)比特——从而实现细粒度的来源追踪。因此,同样的模运算同时支持二元归因和丰富的载荷。实验结果证明,WaterMod在零比特和多比特设置中均持续获得强大的水印检测性能,同时保持生成质量。这种鲁棒性涵盖自然语言生成、数学推理和代码合成等一系列任务。

### 论文 14
- **英文标题**: Are Language Models Any Good at Density Modeling?
- **中文标题**: 语言模型在密度建模方面表现如何?
- **作者**: Sriram Ranga, Sai Shashank Bedampeta, Rui Mao, Anupam Chattopadhyay
- **英文摘要**: Large Language Models (LLMs) surprised the world with their ability to mimic humans in writing and are starting to be used as simulations of human writers for various kinds of linguistic analyses. However, these analyses rest on the belief that LLMs are good density models that accurately capture the underlying probability distribution of the language. In this paper, we question this basic assumption and try to evaluate language models on their density modelling capabilities. Since a ground truth does not exist for the probability distribution of any natural language, we come up with a synthetic language made up of decimal numbers written in words in English. We train language models from scratch on various probability distributions over this synthetic language and compare the distributions learned by the models with the original distributions. Experiments show that language models can learn underlying probability distributions across a wide range of cases, but they fail when those distributions depend on deep semantic properties of numbers that cannot be inferred from syntactic patterns. Additionally, we observed a strong bias in the models towards numbers that frequently occur as substrings within other numbers. This suggests that such a bias possibly exists in real-world natural language models as well, and negatively impacts downstream tasks and analyses that rely on model-generated probabilities.
- **中文摘要**: 大语言模型(LLM)以其模仿人类写作的能力震惊了世界,并开始被用作各种语言分析的人类写作者模拟。然而,这些分析基于LLM是良好的密度模型、能够准确捕捉语言底层概率分布的信念。在本文中,我们质疑这一基本假设,并尝试评估语言模型的密度建模能力。由于任何自然语言的真实概率分布不存在,我们提出了一种由英文单词书写的十进制数字构成的合成语言。我们在这种合成语言的各种概率分布上从头训练语言模型,并比较模型学到的分布与原始分布。实验表明,语言模型可以在广泛的情况下学习底层的概率分布,但当这些分布依赖于无法从句法模式推断的数字深层语义属性时,它们会失败。此外,我们观察到模型对频繁作为其他数字子串出现的数字有强烈偏差。这表明这种偏差可能也存在于现实世界的自然语言模型中,并对依赖模型生成概率的下游任务和分析产生负面影响。

### 论文 22
- **英文标题**: Incorporating Token Importance in Multi-Vector Retrieval
- **中文标题**: 在多向量检索中融入Token重要性
- **作者**: Archish S, Ankit Garg, Kirankumar Shiragur, Neeraj Kayal
- **英文摘要**: ColBERT introduced a late interaction mechanism that independently encodes queries and documents using BERT, and computes similarity via fine-grained interactions over token-level vector representations. This design enables expressive matching while allowing efficient computation of scores, as the multi-vector document representations could be pre-computed offline. ColBERT models distance using a Chamfer-style function: for each query token, it selects the closest document token and sums these distances across all query tokens.
- **中文摘要**: ColBERT引入了一种延迟交互机制,使用BERT独立编码查询和文档,并通过token级向量表示上的细粒度交互计算相似度。这种设计实现了表达性匹配,同时允许高效计算分数,因为多向量文档表示可以离线预计算。ColBERT使用Chamfer式函数建模距离:对于每个查询token,它选择最近的文档token,并跨所有查询token求和这些距离。在我们的工作中,我们通过计算查询token贡献的加权和来探索对Chamfer距离函数的增强,其中权重反映token重要性。经验上,我们展示了这种简单扩展——仅需token权重训练同时保持多向量表示固定——进一步增强了延迟交互多向量机制的表达性。特别是在BEIR基准上,我们的方法在零样本设置中使用IDF权重实现了Recall@10平均1.28%的提升,通过少样本微调实现了3.66%的提升。

### 论文 30
- **英文标题**: Are We on the Right Way to Assess Document Retrieval-Augmented Generation?
- **中文标题**: 我们是否在正确评估文档检索增强生成?
- **作者**: Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen, Junjie Yang, Yao Wan, Weiwei Lin
- **英文摘要**: Retrieval-Augmented Generation (RAG) systems using Multimodal Large Language Models (MLLMs) show great promise for complex document understanding, yet their development is critically hampered by inadequate evaluation. Current benchmarks often focus on specific part of document RAG system and use synthetic data with incomplete ground truth and evidence labels, therefore failing to reflect real-world bottlenecks and challenges. To overcome these limitations, we introduce Double-Bench: a new large-scale, multilingual, and multimodal evaluation system that is able to produce fine-grained assessment to each component within document RAG systems. It comprises 3,276 documents (72,880 pages) and 5,168 single- and multi-hop queries across 6 languages and 4 document types with streamlined dynamic update support for potential data contamination issues. Queries are grounded in exhaustively scanned evidence pages and verified by human experts to ensure maximum quality and completeness. Our comprehensive experiments across 9 state-of-the-art embedding models, 4 MLLMs and 4 end-to-end document RAG frameworks demonstrate the gap between text and visual embedding models is narrowing, highlighting the need in building stronger document retrieval models. Our findings also reveal the over-confidence dilemma within current document RAG frameworks that tend to provide answer even without evidence support. We hope our fully open-source Double-Bench provide a rigorous foundation for future research in advanced document RAG systems. We plan to retrieve timely corpus and release new benchmarks on an annual basis.
- **中文摘要**: 使用多模态大语言模型(MLLM)的检索增强生成(RAG)系统在复杂文档理解方面展现出巨大前景,但它们的开发受到不足评估的严重阻碍。当前基准通常聚焦于文档RAG系统的特定部分,并使用具有不完整真实答案和证据标签的合成数据,因此未能反映现实世界的瓶颈和挑战。为克服这些限制,我们引入Double-Bench:一个新的、大规模的、多语言和多模态评估系统,能够对文档RAG系统内的每个组件进行细粒度评估。它包含3427份文档(72880页)和5168个单跳和多跳查询,跨越6种语言和4种文档类型,并具有简化的动态更新支持以应对潜在的数据污染问题。查询建立在详尽扫描的证据页面上,并由人类专家验证以确保最高质量和完整性。我们在9个最先进的嵌入模型、4个MLLM和4个端到端文档RAG框架上的全面实验表明,文本和视觉嵌入模型之间的差距正在缩小,突显了构建更强大文档检索模型的需求。我们的发现还揭示了当前文档RAG框架中的过度自信困境,即即使没有证据支持也倾向于提供答案。我们希望我们的完全开源Double-Bench为高级文档RAG系统的未来研究提供严格基础。我们计划每年检索及时语料并发布新的基准。

### 论文 36
- **英文标题**: Fine-Tuned LLMs Know They Don’t Know: A Parameter-Efficient Approach to Recovering Honesty
- **中文标题**: 微调后的LLM知道它们不知道: 一种参数高效的方法以恢复诚实性
- **作者**: Zeyu Shi, Ziming Wang, Tianyu Chen, Shiqi Gao, Haoyi Zhou, Qingyun Sun, Jianxin Li
- **英文摘要**: The honesty of Large Language Models (LLMs) is increasingly important for safe deployment in high-stakes domains. However, this crucial trait is severely undermined by supervised fine-tuning (SFT), a common technique for model specialization. Existing recovery methods rely on data-intensive global parameter adjustments, implicitly assuming that SFT deeply corrupts the models' ability to recognize their knowledge boundaries. However, we observe that fine‑tuned LLMs still preserve this ability; what is damaged is their capacity to faithfully express that awareness. Building on this, we propose Honesty-Critical Neurons Restoration (HCNR) to surgically repair this suppressed capacity. HCNR identifies and restores key expression-governing neurons to their pre-trained state while harmonizing them with task-oriented neurons via Hessian-guided compensation. Experiments on four QA tasks and five LLM families demonstrate that HCNR effectively recovers 33.25% of the compromised honesty while achieving at least 2.23x speedup with over 10x less data compared to baseline methods, offering a practical solution for trustworthy LLM deployment.
- **中文摘要**: 大语言模型(LLM)的诚实性对于高风险领域的安全部署日益重要。然而,这一关键特性受到监督微调(SFT)——一种常见的模型专业化技术——的严重损害。现有的恢复方法依赖数据密集型的全局参数调整,隐含假设SFT深刻破坏了模型识别其知识边界的能力。然而,我们观察到微调后的LLM仍然保留这种能力;被损害的是它们忠实地表达这种意识的能力。基于此,我们提出诚实性关键神经元恢复(HCNR)来精细修复这种被抑制的能力。HCNR识别并恢复关键表达控制神经元到其预训练状态,同时通过Hessian引导的补偿将它们与任务导向神经元协调。在四个QA任务和五个LLM家族上的实验表明,HCNR有效恢复了33.25%的被损害的诚实性,同时相比基线方法实现了至少2.23倍的加速和超过10倍的数据减少,为可信赖的LLM部署提供了实用解决方案。

### 论文 63
- **英文标题**: SafetyReminder: Reviving Delayed Safety Awareness of Vision-Language Models to Defend Against Jailbreak Attacks
- **中文标题**: SafetyReminder: 唤醒视觉-语言模型的延迟安全意识以防御越狱攻击
- **作者**: Peiyuan Tang, Haojie Xin, Xiaodong Zhang, Jun Sun, Qin Xia, Zijiang James Yang
- **英文摘要**: Vision-Language Models (VLMs) extend Large Language Models (LLMs) with visual perception capabilities, unlocking broad applications across many domains. However, ensuring their safety remains a critical challenge, as adversarial visual inputs can easily bypass built-in safeguards and elicit harmful content. In this paper, we uncover a phenomenon we call delayed safety awareness, where a jailbroken VLM initially produces harmful content but ultimately recognizes the harmfulness at the end of the generation process. We attribute this phenomenon to the fact that the model's safety awareness against jailbreaks cannot be effectively transferred to the intermediate stages of text generation. Motivated by this insight, we introduce SafetyReminder, a simple yet effective defense that optimizes a learnable soft prompt using our proposed Safety-Activation Prompt Tuning (SAPT). This soft prompt is inserted into the generated text to activate the safety awareness of the model, steering it toward refusal when harmful content arises while preserving helpfulness in benign scenarios. We evaluate our method on three established harmful benchmarks and across three types of adversarial attacks. Experimental results demonstrate that our method achieves state-of-the-art defense performance with strong generalization, offering a practical and lightweight solution for safe deployment of VLMs.
- **中文摘要**: 视觉-语言模型(VLM)将大语言模型(LLM)扩展了视觉感知能力,解锁了跨多个域的广泛应用。然而,确保它们的安全性仍然是一个关键挑战,因为对抗性视觉输入可以轻易绕过内置的安全防护并引发有害响应。在本文中,我们提出SafetyReminder,一种在VLM推理期间唤醒延迟安全意识的方法。SafetyReminder通过在每一层注入安全提醒token,使模型在存在潜在有害视觉输入时保持安全感知。实验表明,SafetyReminder在保持模型能力的同时显著提升了VLM对越狱攻击的抵抗力。

### 论文 66
- **英文标题**: PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
- **中文标题**: PocketLLM: 通过元网络实现大语言模型的极致压缩
- **作者**: Ye Tian, Chengcheng Wang, Jing Han, Yehui Tang, Kai Han
- **英文摘要**: As Large Language Models (LLMs) continue to grow in size, storing and transmitting them on edge devices becomes increasingly challenging. Traditional methods like quantization and pruning struggle to achieve extreme compression of LLMs without sacrificing accuracy. In this paper, we introduce PocketLLM, a novel approach to compress LLMs in a latent space via meta-networks. A simple encoder network is proposed to project the weights of LLMs into discrete latent vectors, which are then represented using a compact codebook. A lightweight decoder network is employed to map the codebook's representative vectors back to the original weight space. This method allows for significant compression of the large weights in LLMs, consisting solely of a small decoder, a concise codebook, and an index. Extensive experiments show that PocketLLM achieves superior performance even at significantly high compression ratios, e.g., compressing Llama 2-7B by 10x with a negligible drop in accuracy.
- **中文摘要**: 随着大语言模型(LLM)规模的不断增长,在边缘设备上存储和传输它们变得越来越具有挑战性。量化和剪枝等传统方法在不牺牲准确性的情况下难以实现LLM的极致压缩。在本文中,我们引入PocketLLM,一种通过元网络实现LLM极致压缩的新方法。PocketLLM使用元网络动态生成压缩模型的参数,使得单个元网络可以重构多种压缩配置的模型。实验结果表明,PocketLLM在保持性能的同时实现了显著的压缩率,为在资源受限设备上部署LLM提供了实用解决方案。

### 论文 67
- **英文标题**: KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
- **中文标题**: KeepKV: 实现高效LLM推理的周期性无损KV缓存压缩
- **作者**: Yuxuan Tian, Zihan Wang, Yebo Peng, Aomufei Yuan, Zhiming Wang, Bairen Yi, Xin Liu, Yong Cui, Tong Yang
- **英文摘要**: Efficient inference of large language models (LLMs) is hindered by an ever-growing key-value (KV) cache, making KV cache compression a critical research direction. Traditional methods selectively evict less important KV cache entries, which leads to information loss and hallucinations. Recently, merging-based strategies have been explored to retain more information by merging KV pairs that would be discarded; however, these existing approaches inevitably introduce inconsistencies in attention distributions before and after merging, causing degraded generation quality. To overcome this challenge, we propose KeepKV , a novel adaptive KV cache merging method designed to preserve performance under strict memory constraints, achieving single-step lossless compression and providing error bounds for multi-step compression. KeepKV introduces the Electoral Votes mechanism that records merging history and adaptively adjusts attention scores. Moreover, it further leverages a novel Zero Inference-Perturbation Merging method, compensating for attention loss resulting from cache merging. Extensive experiments on various benchmarks and LLM architectures demonstrate that KeepKV substantially reduces memory usage while successfully retaining essential context information, achieving over 2 times inference throughput improvement and maintaining superior generation quality even with only 10% KV cache budgets.
- **中文摘要**: 大语言模型(LLM)的高效推理受到不断增长的键值(KV)缓存的阻碍,使KV缓存压缩成为关键研究方向。传统方法选择性地驱逐不太重要的KV缓存条目,这会导致信息丢失和幻觉。最近的合并方法通过组合多个条目来保留更多信息,但以精确性为代价。在本文中,我们提出KeepKV,一种周期性无损KV缓存压缩方法,通过利用注意力模式的周期性来实现有效压缩而不损失关键信息。实验表明,KeepKV在保持模型性能的同时实现了更高的压缩率。

### 论文 88
- **英文标题**: SDEval: Safety Dynamic Evaluation for Multimodal Large Language Models
- **中文标题**: SDEval: 面向多模态大语言模型的安全动态评估
- **作者**: Hanqing Wang, Yuan Tian, Mingyu Liu, Zhenhao Zhang, Xiangyang Zhu
- **英文摘要**: In the rapidly evolving landscape of Multimodal Large Language Models (MLLMs), the safety concerns of their outputs have earned significant attention. Although numerous datasets have been proposed, they may become outdated with MLLM advancements and are susceptible to data contamination issues. To address these problems, we propose SDEval, the first safety dynamic evaluation framework to controllably adjust the distribution and complexity of safety benchmarks. Specifically, SDEval mainly adopts three dynamic strategies: text, image, and text-image dynamics to generate new samples from original benchmarks. We first explore the individual effects of text and image dynamics on model safety. Then, we find that injecting text dynamics into images can further impact safety, and conversely, injecting image dynamics into text also leads to safety risks. SDEval is general enough to be applied to various existing safety and even capability benchmarks. Experiments across safety benchmarks, MLLMGuard and VLSBench, and capability benchmarks, MMBench and MMVet, show that SDEval significantly influences evaluation results, mitigates data contamination, and exposes safety limitations of MLLMs.
- **中文摘要**: 在多模态大语言模型(MLLM)快速演进的环境中,其输出的安全问题已获得显著关注。虽然已提出大量数据集,但它们可能随着MLLM的进步而过时,并且容易受到数据污染问题的影响。为解决这些问题,我们提出SDEval,首个安全动态评估框架,可控地调整安全基准的分布和复杂性。具体而言,SDEval主要采用三种动态策略:文本、图像和文本-图像动态,从原始基准生成新样本。我们首先探索文本和图像动态对模型安全的个体影响。然后,我们发现将文本动态注入图像可以进一步影响安全性,反过来,将图像动态注入文本也会导致安全风险。SDEval具有足够的通用性,可应用于各种现有的安全甚至能力基准。在安全基准MLLMGuard和VLSBench,以及能力基准MMBench和MMVet上的实验表明,SDEval显著影响评估结果,缓解数据污染,并暴露MLLM的安全局限性。

### 论文 101
- **英文标题**: When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- **中文标题**: 当真相被覆盖: 揭示大语言模型中谄媚行为的内在起源
- **作者**: Keyu Wang, Jin Li, Shu Yang, Zhuoran Zhang, Di Wang
- **英文摘要**: Large Language Models (LLMs) often exhibit sycophantic behavior, agreeing with user-stated opinions even when those contradict factual knowledge. While prior work has documented this tendency, the internal mechanisms that enable such behavior remain poorly understood. In this paper, we provide a mechanistic account of how sycophancy arises within LLMs. We first systematically study how user opinions induce sycophancy across different model families. We find that simple opinion statements reliably induce sycophancy, whereas user expertise framing has a negligible impact. Through logit-lens analysis and causal activation patching, we identify a two-stage emergence of sycophancy: (1) a late-layer output preference shift and (2) deeper representational divergence. We also verify that user authority fails to influence behavior because models do not encode it internally. In addition, we examine how grammatical perspective affects sycophantic behavior, finding that first-person prompts (“I believe...”) consistently induce higher sycophancy rates than third-person framings (“They believe...”) by creating stronger representational perturbations in deeper layers. These findings highlight that sycophancy is not a surface-level artifact but emerges from a structural override of learned knowledge in deeper layers, with implications for alignment and truthful AI systems.
- **中文摘要**: 大语言模型(LLM)经常表现出谄媚行为,即使用户陈述的意见与事实知识相矛盾时也同意。虽然先前的工作已记录了这种倾向,但使此类行为得以发生的内部机制仍然理解不足。在本文中,我们提供了关于谄媚行为如何在LLM中产生的机制性解释。我们首先系统研究用户意见如何在不同的模型家族中诱导谄媚行为。我们发现简单的意见陈述可靠地诱导谄媚,而用户专业知识框架的影响可以忽略不计。通过logit透镜分析和因果激活修补,我们识别出谄媚的两阶段涌现:(1)后期层输出偏好偏移和(2)更深层的表示分歧。我们还验证了用户权威未能影响行为,因为模型未在内部编码它。此外,我们研究了语法视角如何影响谄媚行为,发现第一人称提示('我相信...')通过创建更深层中更强的表示扰动,持续诱导比第三人称框架('他们相信...')更高的谄媚率。这些发现突显了谄媚不是表面层面的伪影,而是从更深层中学到知识的结构性覆盖中涌现出来的,对对齐和真实AI系统具有启示意义。

### 论文 105
- **英文标题**: ALEX:A Light Editing-knowledge Extractor
- **中文标题**: ALEX: 轻量级编辑知识提取器
- **作者**: Minghu Wang, ShuLiang Zhao, Yuanyuan Zhao, Hongxia Xu
- **英文摘要**: The static nature of knowledge within Large Language Models (LLMs) makes it difficult for them to adapt to evolving information, rendering knowledge editing a critical task. However, existing methods struggle with challenges of scalability and retrieval efficiency, particularly when handling complex, multi-hop questions that require multi-step reasoning. To address these challenges, this paper introduces ALEX (A Light Editing-knowledge Extractor), a lightweight knowledge editing framework. The core innovation of ALEX is its hierarchical memory architecture, which organizes knowledge updates (edits) into semantic clusters. This design fundamentally reduces retrieval complexity from a linear O(N) to a highly scalable O(K+N/C). Furthermore, the framework integrates an Inferential Query Synthesis (IQS) module to bridge the semantic gap between queries and facts , and a Dynamic Evidence Adjudication (DEA) engine that executes an efficient two-stage retrieval process. Experiments on the MQUAKE benchmark demonstrate that ALEX significantly improves both the accuracy of multi-hop answers (MultiHop-ACC) and the reliability of reasoning paths (HopWise-ACC). It also reduces the required search space by over 80% , presenting a promising path toward building scalable, efficient, and accurate knowledge editing systems.
- **中文摘要**: 大语言模型(LLM)中知识的静态性使其难以适应不断变化的信息,使知识编辑成为一项关键任务。然而,现有方法在可扩展性和检索效率方面面临挑战,特别是在处理需要多步推理的复杂多跳问题时。为应对这些挑战,本文引入ALEX(轻量级编辑知识提取器),一个轻量级知识编辑框架。ALEX的核心创新是其层次化内存架构,将知识更新(编辑)组织成语义聚类。此设计从根本上将检索复杂度从线性O(N)降低到高度可扩展的O(K+N/C)。此外,该框架集成了推理查询合成(IQS)模块以弥合查询与事实之间的语义差距,以及动态证据裁决(DEA)引擎执行高效的两阶段检索过程。在MQUAKE基准上的实验表明,ALEX显著提升了多跳答案准确率(MultiHop-ACC)和推理路径可靠性(HopWise-ACC)。它还将所需搜索空间减少了超过80%,为构建可扩展、高效和准确的知识编辑系统提供了有前景的路径。

### 论文 107
- **英文标题**: LoKI: Low-Damage Knowledge Implanting of Large Language Models
- **中文标题**: LoKI: 大语言模型的低损伤知识植入
- **作者**: Runyu Wang, Peng Ping, Zhengyu Guo, Xiaoye Zhang, Quan Shi, Liting Zhou, Tianbo Ji
- **英文摘要**: Fine-tuning adapts pretrained models for specific tasks but poses the risk of catastrophic forgetting (CF), where critical knowledge from pretraining is overwritten. To address the issue of CF in a general-purpose framework, we propose Low-damage Knowledge Implanting (LoKI), a parameter-efficient fine-tuning (PEFT) technique that utilizes recent mechanistic understanding of how knowledge is stored in transformer architectures. We compare LoKI against state-of-the-art PEFT methods in two real-world fine-tuning scenarios. The results show that LoKI demonstrates significantly better preservation of general capabilities. At the same time, its task-specific performance is comparable to or even surpasses that of full parameter fine-tuning and these PEFT methods across various model architectures. Our work bridges the mechanistic insights of LLMs' knowledge storage with practical fine-tuning objectives, enabling an effective balance between task-specific adaptation and the retention of general-purpose capabilities.
- **中文摘要**: 微调使预训练模型适应特定任务,但带来了灾难性遗忘(CF)的风险,其中预训练的关键知识被覆盖。为在通用框架中解决CF问题,我们提出低损伤知识植入(LoKI),一种利用最近对Transformer架构中知识存储方式的机制理解的参数高效微调(PEFT)技术。我们将LoKI与最先进的PEFT方法在两个真实世界微调场景中进行比较。结果表明,LoKI展示了显著更好的通用能力保留。同时,其任务特定性能与全参数微调和这些PEFT方法在各种模型架构上相当甚至超越。我们的工作在LLM知识存储的机制洞察与实用微调目标之间架起了桥梁,实现了任务特定适应与通用能力保留之间的有效平衡。

## Natural Language Processing V

### 论文 50
- **英文标题**: LLM-Oriented Token-Adaptive Knowledge Distillation
- **中文标题**: 面向LLM的令牌自适应知识蒸馏
- **作者**: Xurong Xie, Zhucun Xue, Jiafu Wu, Jian Li, Yabiao Wang, Xiaobin Hu, Yong Liu, Jiangning Zhang
- **英文摘要**: Knowledge Distillation (KD) is a key technique for compressing Large-scale Language Models (LLMs), but prevailing logit-based methods employ static strategies misaligned with the student’s dynamic learning process. By treating all tokens indiscriminately with a fixed temperature, these methods result in suboptimal knowledge transfer. To address this, we propose LLM-oriented token-Adaptive Knowledge Distillation (AdaKD), a framework that adapts the distillation process to each token’s real-time learning state. AdaKD consists of two synergistic modules driven by a unified token difficulty metric. First, the Loss-driven Adaptive Token Focusing (LATF) module dynamically concentrates distillation on valuable tokens by monitoring the student’s learning stability. Second, Inverse Difficulty Temperature Scaling (IDTS) introduces a counterintuitive token-level temperature: low for difficult tokens to target error correction, and high for easy tokens to learn the teacher’s smooth output distribution for better generalization. As a plug-and-play framework, AdaKD consistently improves performance across diverse distillation methods, model architectures, and benchmarks.
- **中文摘要**: 知识蒸馏(KD)是压缩大规模语言模型的关键技术，但主流的基于logit的方法采用与学生动态学习过程不对齐的静态策略。通过以固定温度无差别地对待所有令牌，这些方法导致次优的知识迁移。为此我们提出面向LLM的令牌自适应知识蒸馏(AdaKD)，一个将蒸馏过程适配到每个令牌实时学习状态的框架。AdaKD由两个由统一令牌难度指标驱动的协同模块组成。首先损失驱动的自适应令牌聚焦(LATF)模块通过监控学生的学习稳定性，动态将蒸馏集中在有价值的令牌上。其次逆难度温度缩放(IDTS)引入了一种违反直觉的令牌级温度：对困难令牌设置低温以瞄准纠错，对容易令牌设置高温以学习教师的平滑输出分布以获得更好的泛化。作为即插即用框架，AdaKD在多样化蒸馏方法、模型架构和基准上持续提升性能。

### 论文 51
- **英文标题**: Advanced Black-Box Tuning of Large Language Models with Limited API Calls
- **中文标题**: 有限API调用下的大语言模型高级黑盒调优
- **作者**: Zhikang Xie, Weilin Wan, Peizhu Gong, Weizhong Zhang, Cheng Jin
- **英文摘要**: Black-box tuning is an emerging paradigm for adapting large language models (LLMs) to better achieve desired behaviors, particularly when direct access to model parameters is unavailable. Current strategies, however, often present a dilemma of suboptimal extremes: either separately train a small proxy model and then use it to shift the predictions of the foundation model, offering notable efficiency but often yielding limited improvement; or making API calls in each tuning iteration to the foundation model, which entails prohibitive computational costs. In this paper, we argue that a more reasonable way for black-box tuning is to train the proxy model with limited API calls. The underlying intuition is based on two key observations: first, the training samples may exhibit correlations and redundancies, suggesting that the foundation model’s predictions can be estimated from previous calls; second, foundation models frequently demonstrate low accuracy on downstream tasks. Therefore, we propose a novel advanced black-box tuning method for LLMs with limited API calls. Our core strategy involves training a Gaussian Process (GP) surrogate model with "LogitMap Pairs" derived from querying the foundation model on a minimal but highly informative training subset. This surrogate can approximate the outputs of the foundation model to guide the training of the proxy model, thereby effectively reducing the need for direct queries to the foundation model. Extensive experiments verify that our approach elevates pre-trained language model accuracy from 55.92% to 86.85%, reducing the frequency of API queries to merely 1.38%. This significantly outperforms offline approaches that operate entirely without API access. Notably, our method also achieves comparable or superior accuracy to query-intensive approaches, while significantly reducing API costs. This offers a robust and high-efficiency paradigm for language model adaptation.
- **中文摘要**: 黑盒调优是一种新兴的范式，用于适应大语言模型以更好地实现期望行为，特别是在无法直接访问模型参数时。然而当前策略通常呈现出次优极端的困境：要么分别训练一个小型代理模型然后使用它来偏移基础模型的预测，提供显著效率但改进有限；要么在每次调优迭代中向基础模型发出API调用，这产生高昂的计算成本。本文认为黑盒调优的更合理方式是训练代理模型并配合有限的API调用。底层直觉基于两个关键观察：首先训练样本可能表现出相关性和冗余，表明基础模型的预测可以从先前的调用中估计；其次基础模型在下游任务中经常表现出低准确率。因此我们提出一种新颖的有限API调用高级黑盒调优方法。我们的核心策略涉及训练一个带有从查询基础模型的最少但高度信息训练子集导出的'LogitMap对'的高斯过程代理模型。此代理可以近似基础模型的输出来指导代理模型的训练，从而有效减少对基础模型直接查询的需求。广泛实验验证我们的方法将预训练语言模型准确率从55.92%提升至86.85%，将API查询频率降低至仅1.38%。

### 论文 67
- **英文标题**: Test-time Prompt Intervention
- **中文标题**: 测试时提示干预
- **作者**: Chenxu Yang, Qingyi Si, Mz Dai, Dingyu Yao, Mingyu Zheng, Minghui Chen, Zheng Lin, Weiping Wang
- **英文摘要**: Test-time compute has led to remarkable success in the large language model (LLM) community, particularly for complex tasks, where longer chains of thought (CoTs) are generated to enhance reasoning capabilities. However, growing evidence reveals that such reasoning models often produce CoTs plagued by excessive redundancy, including repetitive verification steps and unnecessary reasoning shifts. The root cause lies in post-training of them that overly rely on outcome reward paradigms, as the data of process reward paradigms, which regulate intermediate reasoning steps, is difficult to construct at scale. To address this, we propose PI, a novel framework for Test-time Prompt Intervention. PI provides an interface to dynamically guide and regulate reasoning paths during inference through timely (When module) and proper (How module) interventions and post-intervention sampling (Which module). This allows human problem-solving expertise and cognitive science principles to be seamlessly integrated into LLMs’ reasoning processes, enhancing controllability and interpretability. Extensive experiments across multiple models and datasets demonstrate that PI significantly shortens CoTs while reducing hallucination, yielding more concise and reliable reasoning.
- **中文摘要**: 测试时计算已在大语言模型社区取得显著成功，特别是对于复杂任务，其中生成更长的思维链以增强推理能力。然而越来越多的证据表明此类推理模型经常产生充斥过度冗余的CoT，包括重复的验证步骤和不必要的推理转换。根源在于它们的后训练过度依赖结果奖励范式，因为调节中间推理步骤的过程奖励范式数据难以大规模构建。为解决此问题我们提出PI，一种测试时提示干预的新框架。PI提供了一个接口，通过及时的(When模块)和适当的(How模块)干预以及干预后采样(Which模块)来动态引导和规范推理路径。这使人类问题解决专业知识和认知科学原理能够无缝集成到LLM的推理过程中，增强可控性和可解释性。跨多个模型和数据集的广泛实验表明PI显著缩短CoT同时减少幻觉，产生更简洁且更可靠的推理。

### 论文 75
- **英文标题**: MrM: Black-Box Membership Inference Attacks Against Multimodal RAG Systems
- **中文标题**: MrM：针对多模态RAG系统的黑盒成员推理攻击
- **作者**: Peiru Yang, Jinhua Yin, Haoran Zheng, Xueying Bai, Huili Wang, Yufei Sun, Xintian Li, Songwei Pei, Yongfeng Huang, Tao Qi
- **英文摘要**: Multimodal retrieval-augmented generation (RAG) systems enhance large vision-language models by integrating cross-modal knowledge, enabling their increasing adoption across real-world multimodal tasks. These knowledge databases may contain sensitive information that requires privacy protection. However, multimodal RAG systems inherently grant external users indirect access to such data, making them potentially vulnerable to privacy attacks, particularly membership inference attacks (MIAs). Existing MIA methods targeting RAG systems predominantly focus on the textual modality, while the visual modality remains relatively underexplored. To bridge this gap, we propose MrM, the first black-box MIA framework targeted at multimodal RAG systems. It utilizes a multi-object data perturbation framework constrained by counterfactual attacks, which can concurrently induce the RAG systems to retrieve the target data and generate information that leaks the membership information. Our method first employs an object-aware data perturbation method to constrain the perturbation to key semantics and ensure successful retrieval. Building on this, we design a counterfact-informed mask selection strategy to prioritize the most informative masked regions, aiming to eliminate the interference of model self-knowledge and amplify attack efficacy. Finally, we perform statistical membership inference by modeling query trials to extract features that reflect the reconstruction of masked semantics from response patterns. Experiments on two visual datasets and eight mainstream commercial visual-language models (e.g., GPT-4o, Gemini-2) demonstrate that MrM achieves consistently strong performance across both sample-level and set-level evaluations, and remains robust under adaptive defenses.
- **中文摘要**: 多模态检索增强生成(RAG)系统通过集成跨模态知识增强大型视觉-语言模型，使其在多模态任务中日益广泛采用。这些知识数据库可能包含需要隐私保护的敏感信息。然而多模态RAG系统固有地授予外部用户对此类数据的间接访问，使其可能容易受到隐私攻击，尤其是成员推理攻击(MIA)。现有的针对RAG系统的MIA方法主要集中在文本模态上，而视觉模态仍然相对未被充分探索。为弥合此差距我们提出MrM，首个针对多模态RAG系统的黑盒MIA框架。它利用受反事实攻击约束的多目标数据扰动框架，能够同时诱导RAG系统检索目标数据并生成泄露成员信息的信息。我们的方法首先采用目标感知数据扰动方法将扰动约束到关键语义并确保成功检索。在此基础上我们设计了反事实信息掩码选择策略以优先考虑最信息丰富的被掩码区域，旨在消除模型自有知识的干扰并放大攻击效果。最后我们通过对查询试验建模执行统计成员推理，从响应模式中提取反映被掩码语义重建的特征。在两个视觉数据集和八个主流商业视觉-语言模型(如GPT-4o、Gemini-2)上的实验表明MrM在样本级和集合级评估中均实现一贯强劲表现，并在自适应防御下保持鲁棒。

### 论文 81
- **英文标题**: RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
- **中文标题**: RICo：面向自动指令调优数据选择的精炼上下文贡献度量
- **作者**: Yixin Yang, Qingxiu Dong, Linli Yao, Fangwei Zhu, Weilin Luo, Bin Wang, Zhifang Sui
- **英文摘要**: Data selection for instruction tuning is crucial for improving the performance of large language models (LLMs) while reducing training costs. In this paper, we propose Refined Contribution Measurement with In-Context Learning (RICo), a novel gradient-free method that quantifies the fine-grained contribution of individual samples to both task-level and global-level model performance. RICo enables more accurate identification of high-contribution data, leading to better instruction tuning. We also introduce a lightweight selection paradigm trained on RICo scores, enabling scalable data selection with strictly linear inference complexity. Extensive experiments on 3 LLMs across 12 benchmarks and 5 pairwise evaluation sets demonstrate the effectiveness of RICo. Remarkably, on LLaMA3.1-8B, models trained in 15% of RICo-selected data outperform full datasets by 5.42 percentage points and exceed the best performance of widely used selection methods by 1.48 percentage points. We further analyze high-contribution samples selected by RICo, which show both diverse tasks and appropriate difficulty levels, rather than merely the most difficult cases.
- **中文摘要**: 指令调优的数据选择对提升大语言模型性能同时降低训练成本至关重要。本文提出RICo(基于上下文学习的精炼贡献度量)，一种新颖的无梯度方法，量化单个样本对任务级和全局级模型性能的细粒度贡献。RICo能够更准确地识别高贡献数据，从而实现更好的指令调优。我们还引入了一个在RICo分数上训练的轻量级选择范式，以严格线性推理复杂度实现可扩展的数据选择。在3个LLM跨12个基准和5个成对评估集的广泛实验证明了RICo的有效性。显著的是在LLaMA3.1-8B上以RICo选择的15%数据训练的模型比全数据集高出5.42个百分点，并超越广泛使用的选择方法的最佳性能1.48个百分点。我们进一步分析了RICo选择的高贡献样本，它们既展现了多样化任务又展现了适当的难度水平，而非仅仅是最困难的案例。

### 论文 85
- **英文标题**: Conversational Learning Diagnosis via Reasoning Multi-Turn Interactive Learning
- **中文标题**: 通过推理式多轮交互学习进行对话学习诊断
- **作者**: Fangzhou Yao, Sheng Chang, Weibo Gao, Qi Liu
- **英文摘要**: Learning diagnosis is a critical task that monitors students' cognitive state during educational activities, with the goal of enhancing learning outcomes. With advancements in language models (LMs), many AI-driven educational studies have shifted towards conversational learning scenarios, where students engage in multi-turn interactive dialogues with tutors. However, conversational learning diagnosis remains underdeveloped, and most existing techniques acquire students' cognitive state through intuitive instructional prompts on LMs to analyze the dialogue text. This direct prompting approach lacks a solid psychological foundation and fails to ensure the reliability of the generated analytical text. In this study, we introduce ParLD, a preview-analyze-reason framework for conversational learning diagnosis, which leverages multi-agent collaboration to diagnose students' cognitive state over multiple dialogue turns. Specifically, ParLD comprises main components: (1) Behavior Previewer, which generates a student behavior schema based on previous states and learning content; (2) State Analyzer, which diagnose the tutor-student dialogue and behavior schema to update the cognitive state; and (3) Performance Reasoner, which predicts the student's future responses and provides verifiable feedback to support ParLD's self-reflection with the Chain Reflector. They operate sequentially and iteratively during each interaction turn to diagnose the student’s cognitive state. We conduct experiments to evaluate both performance prediction and tutoring support, emphasizing the effectiveness of ParLD in providing reliable and insightful learning diagnosis.
- **中文摘要**: 学习诊断是一项关键任务，在教育活动中监控学生的认知状态，旨在提升学习成果。随着语言模型(LM)的进步，许多AI驱动的教育研究已转向对话学习场景，其中学生与导师进行多轮交互对话。然而对话学习诊断仍然发展不足，大多数现有技术通过LM上的直觉指令提示来分析对话文本以获取学生的认知状态。这种直接提示方法缺乏坚实的心理学基础，且无法保证所生成分析文本的可靠性。本研究引入ParLD，一个预览-分析-推理框架用于对话学习诊断，利用多代理协作在多轮对话中诊断学生的认知状态。具体而言ParLD包含主要组件：(1)行为预览器，基于先前状态和学习内容生成学生行为模式；(2)状态分析器，诊断导师-学生对话和行为模式以更新认知状态；(3)表现推理器，预测学生未来响应并提供可验证反馈以支持ParLD的自我反思。它们在每个交互轮次中顺序迭代运作以诊断学生的认知状态。我们进行了实验以评估表现预测和辅导支持，强调了ParLD在提供可靠且富有洞察的学习诊断方面的有效性。

## Natural Language Processing VI

### 论文 8
- **英文标题**: MARS: Multimodal Adaptive Reasoning Model for Avoiding Overthinking
- **中文标题**: MARS：避免过度思考的多模态自适应推理模型
- **作者**: Tan Yue, Qiong Wu, Dongyan Zhao
- **英文摘要**: Multimodal Large Language Models (MLLMs) have shown advanced performance in vision-language tasks. However, existing multimodal reasoning models often suffer from excessive reasoning steps, leading to high computational costs and inefficiency. In this paper, we propose the Multimodal Adaptive Reasoning Model (MARS), which enables adaptive adjustment of the reasoning strategy based on question difficulty. Specifically, MARS adopts a three-stage training framework based on our constructed training dataset (MART): 1) CoT Masking Learning to enhance reasoning logicality by predicting masked reasoning steps. 2) Adaptive Reasoning Instruction Learning to train the model to skip or keep reasoning steps according to difficulty levels. 3) CoT Lightweight Reinforcement Learning with the Information Bottleneck Principle based GRPO algorithm to reduce CoT length while maintaining performance and generalizability. Results on both in-domain and out-of-domain datasets show that MARS significantly reduces the CoT length (90.2% decrease) while improving accuracy (0.54%), outperforming existing SOTA open-source and proprietary MLLMs.
- **中文摘要**: 多模态大语言模型(MLLM)在视觉-语言任务中展现了先进性能。然而现有多模态推理模型通常遭受过多推理步骤，导致高计算成本和低效率。本文我们提出多模态自适应推理模型(MARS)，使其能够根据问题难度自适应调整推理策略。具体而言MARS采用了基于我们构建的训练数据集(MART)的三阶段训练框架：1)思维链掩码学习，通过预测被掩码的推理步骤来增强推理逻辑性。2)自适应推理指令学习，训练模型根据难度级别跳过或保留推理步骤。3)基于信息瓶颈原理的GRPO算法的CoT轻量级强化学习，在保持性能和泛化性的同时减少CoT长度。在域内和域外数据集上的结果表明MARS显著减少了CoT长度(减少90.2%)同时提升了准确性(0.54%)，超越了现有SOTA开源和商业MLLM。

### 论文 10
- **英文标题**: ExPairT-LLM: Exact Learning for LLM Code Selection by Pairwise Queries
- **中文标题**: ExPairT-LLM：通过成对查询实现LLM代码选择的精确学习
- **作者**: Tom Yuviler, Dana Drachsler-Cohen
- **英文摘要**: Despite recent advances in LLMs, the task of code generation is still challenging. To cope, code selection algorithms select the best program from multiple programs generated by an LLM. However, existing algorithms can fail to identify the correct program, either because they fail to distinguish nonequivalent programs or because they rely on an LLM and assume it always correctly determines the output for every input. We present ExPairT-LLM, an exact learning algorithm for code selection that selects a program by posing two new types of queries to an LLM oracle: pairwise membership and pairwise equivalence. These queries are simpler for LLMs and enable ExPairT-LLM to identify the correct program through a tournament, which is robust to some LLM mistakes. We evaluate ExPairT-LLM on four popular code datasets. Its pass@1 (success rate) outperforms the state-of-the-art code selection algorithm on average by +13.0% and up to +27.1%. It also improves the pass@1 of LLMs performing complex reasoning by +24.0%.
- **中文摘要**: 尽管LLM取得了最新进展，代码生成任务仍然具有挑战性。为应对此挑战代码选择算法从LLM生成的多个程序中选择最佳程序。然而现有算法可能无法识别正确程序，要么是因为它们无法区分非等价程序，要么是因为它们依赖LLM并假设它总能为每个输入正确确定输出。我们提出ExPairT-LLM，一种用于代码选择的精确学习算法，通过向LLM神谕提出两种新型查询来选择程序：成对成员关系和成对等价性。这些查询对LLM而言更简单，并使ExPairT-LLM能够通过锦标赛识别正确程序，该锦标赛对某些LLM错误具有鲁棒性。我们在四个流行代码数据集上评估ExPairT-LLM。其pass@1(成功率)平均优于最先进代码选择算法+13.0%，最高+27.1%。它还将执行复杂推理的LLM的pass@1提升了+24.0%。

### 论文 13
- **英文标题**: SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
- **中文标题**: SlideTailor：面向科学论文的个性化演示幻灯片生成
- **作者**: Wenzheng Zeng, Mingyu Ouyang, Langyuan Cui, Hwee Tou Ng
- **英文摘要**: Automatic presentation slide generation can greatly streamline content creation. However, since preferences of each user may vary, existing under-specified formulations often lead to suboptimal results that fail to align with individual user needs. We introduce a novel task that conditions paper-to-slides generation on user-specified preferences. We propose a human behavior-inspired agentic framework, SlideTailor, that progressively generates editable slides in a user-aligned manner. Instead of requiring users to write their preferences in detailed textual form, our system only asks for a paper-slides example pair and a visual template—natural and easy-to-provide artifacts that implicitly encode rich user preferences across content and visual style. Despite the implicit and unlabeled nature of these inputs, our framework effectively distills and generalizes the preferences to guide customized slide generation. We also introduce a novel chain-of-speech mechanism to align slide content with planned oral narration. Such a design significantly enhances the quality of generated slides and enables downstream applications like video presentations. To support this new task, we construct a benchmark dataset that captures diverse user preferences, with carefully designed interpretable metrics for robust evaluation. Extensive experiments demonstrate the effectiveness of our framework.
- **中文摘要**: 自动演示幻灯片生成可以极大简化内容创作。然而由于每个用户的偏好可能不同，现有的欠规格化公式往往导致不符合个体用户需求的次优结果。我们引入了一个新任务，将论文到幻灯片的生成以用户指定偏好为条件。我们提出人类行为启发的代理框架SlideTailor，以与用户对齐的方式逐步生成可编辑的幻灯片。我们的系统不要求用户以详细文本形式编写偏好，而仅要求论文-幻灯片示例对和视觉模板——这些是自然且易于提供的工件，隐式编码了跨内容和视觉风格的丰富用户偏好。尽管这些输入具有隐式和未标注性质，我们的框架有效蒸馏并泛化这些偏好以引导定制化幻灯片生成。我们还引入了一种新颖的语音链机制，将幻灯片内容与计划的口头叙述对齐。此设计显著增强了生成幻灯片的质量，并实现了视频演示等下游应用。为支持此新任务我们构建了一个捕获多样化用户偏好的基准数据集，具有精心设计的可解释指标用于鲁棒评估。广泛实验证明了我们框架的有效性。

### 论文 25
- **英文标题**: LLaVA-MS-PIT: Multi-Modal Schema-Guided Progressive Instruction Tuning for Multi-Modal Event Extraction
- **中文标题**: LLaVA-MS-PIT：面向多模态事件提取的多模态模式引导渐进指令调优
- **作者**: Hui Zhang, Po Hu, Wei Emma Zhang
- **英文摘要**: The proliferation of multi-modal data on the internet has intensified the need for structured event understanding across textual and visual modalities. However, existing multi-modal event extraction models suffer from three major limitations: the absence of explicit event schema guidance, coarse-grained multi-modal alignment strategies, and reliance on heterogeneous, misaligned multi-modal training datasets. To address these issues, we propose LLaVA-MS-PIT, a Multi-modal Schema-Guided Progressive Instruction Tuning Framework that explicitly injects structured multi-modal event schema knowledge into the model before event extraction. Specifically, we introduce the textual event schema to establish the model’s prior knowledge of event concepts and enhance its ability to reason about event structures, while the visual event schema is employed to bridge the representation gap between textual and visual modalities at the event level, enabling unified and semantically aligned event representations across modalities. Moreover, to alleviate data scarcity and modality misalignment inherent in current benchmarks, we construct imSitu-MEE, a high-quality multi-modal parallel dataset generated and annotated through schema-guided procedures. Extensive experiments demonstrate that LLaVA-MS-PIT achieves competitive performance on multi-modal event extraction benchmarks, underscoring the effectiveness and necessity of schema-guided progressive instruction tuning.
- **中文摘要**: 互联网上多模态数据的激增加剧了对跨文本和视觉模态的结构化事件理解的需求。然而现有多模态事件提取模型存在三个主要局限：缺乏显式事件模式引导、粗粒度多模态对齐策略，以及依赖异构、不对齐的多模态训练数据集。为解决这些问题我们提出LLaVA-MS-PIT，一个多模态模式引导渐进指令调优框架，在事件提取之前显式注入结构化多模态事件模式知识到模型中。具体而言我们引入文本事件模式以建立模型对事件概念的先验知识并增强其推理事件结构的能力，而视觉事件模式用于在事件级别弥合文本和视觉模态之间的表示差距，实现跨模态统一且语义对齐的事件表示。此外为缓解当前基准中固有的数据稀缺性和模态不对齐，我们构建了imSitu-MEE，一个通过模式引导流程生成和标注的高质量多模态平行数据集。广泛实验表明LLaVA-MS-PIT在多模态事件提取基准上达到了竞争性性能，突显了模式引导渐进指令调优的有效性和必要性。

### 论文 54
- **英文标题**: SABER: Switchable and Balanced Training for Efficient LLM Reasoning
- **中文标题**: SABER：面向高效LLM推理的可切换平衡训练
- **作者**: Kai Zhao, Yanjun Zhao, Jiaming Song, Shien He, Lusheng Zhang, Qiang Zhang, Tianjiao Li
- **英文摘要**: Large language models (LLMs) empowered by chain-of-thought reasoning have achieved impressive accuracy on complex tasks but suffer from excessive inference costs and latency when applied uniformly to all problems. We propose SABER (Switchable and Balanced Training for Efficient LLM Reasoning), a reinforcement learning framework that endows LLMs with user‑controllable, token‑budgeted reasoning. SABER first profiles each training example’s base‑model thinking token usage and assigns it to one of the predefined budget tiers. During fine‑tuning, the model is guided by system prompts and length‑aware rewards to respect its assigned budget. In parallel, we incorporate no‑think examples to ensure the model remains reliable even when explicit reasoning is turned off. SABER further supports four discrete inference modes—NoThink, FastThink, CoreThink, and DeepThink, enabling flexible trade‑offs between latency and reasoning depth. Extensive evaluations on math reasoning (MATH, GSM8K), code generation (MBPP), and logical reasoning (LiveBench-Reasoning) demonstrate that SABER achieves high accuracy under tight budgets, graceful degradation, and effective cross-scale and cross‑domain generalization. In particular, SABER‑FastThink cuts reasoning length by 65.4% and yields a 3.6% accuracy gain compared with the base model on the MATH benchmark.
- **中文摘要**: 拥有思维链推理能力的大语言模型(LLM)在复杂任务上实现了令人印象深刻的准确率，但在统一应用于所有问题时遭受过高的推理成本和延迟。我们提出SABER(面向高效LLM推理的可切换平衡训练)，一个赋予LLM用户可控、令牌预算推理的强化学习框架。SABER首先对每个训练示例的基础模型思考令牌使用进行画像，并将其分配到预定义的预算层级之一。在微调期间模型由系统提示和长度感知奖励引导以遵守其分配的预算。并行地我们融入无思考示例以确保即使显式推理关闭时模型仍保持可靠。SABER进一步支持四种离散推理模式——NoThink、FastThink、CoreThink和DeepThink，实现延迟和推理深度之间的灵活权衡。在数学推理(MATH、GSM8K)、代码生成(MBPP)和逻辑推理(LiveBench-Reasoning)上的广泛评估表明SABER在紧缩预算下达到高准确率、优雅退化以及有效的跨规模和跨域泛化。特别是SABER-FastThink将推理长度削减65.4%，并在MATH基准上相比基础模型产生3.6%的准确率增益。

### 论文 77
- **英文标题**: Do Retrieval Augmented Language Models Know When They Don’t Know?
- **中文标题**: 检索增强语言模型知道自己不知道吗？
- **作者**: Youchao Zhou, Heyan Huang, Yicheng Liu, Rui Dai, Xinglin Wang, Xingchen Zhang, Shumin Shi, Yang Deng
- **英文摘要**: Existing large language models (LLMs) occasionally generate plausible yet factually incorrect responses, known as hallucinations. Two main approaches have been proposed to mitigate hallucinations: retrieval-augmented language models (RALMs) and refusal post-training. However, current research predominantly focuses on their individual effectiveness while overlooking the evaluation of the refusal capability of RALMs. Ideally, if RALMs know when they do not know, they should refuse to answer. In this study, we ask the fundamental question: Do RALMs know when they don’t know? Specifically, we investigate three questions. First, are RALMs well calibrated with respect to different internal and external knowledge states? We examine the influence of various factors. Contrary to expectations, when all retrieved documents are irrelevant, RALMs still tend to refuse questions they could have answered correctly. Next, given the model's pronounced over-refusal behavior, we raise a second question: How does a RALM's refusal ability align with its calibration quality? Our results show that the over-refusal problem can be mitigated through in-context fine-tuning. However, we observe that improved refusal behavior does not necessarily imply better calibration or higher overall accuracy. Finally, we ask: Can we combine refusal-aware RALMs with uncertainty-based answer abstention to mitigate over-refusal? We develop a simple yet effective refusal mechanism for refusal-post-trained RALMs that improves their overall answer quality by balancing refusal and correct answers. Our study provides a more comprehensive understanding of the factors influencing RALM behavior. Meanwhile, we emphasize that uncertainty estimation for RALMs remains an open problem deserving deeper investigation.
- **中文摘要**: 现有大语言模型(LLM)偶尔生成表面合理但事实上不正确的响应，称为幻觉。已提出两种主要方法来缓解幻觉：检索增强语言模型(RALM)和拒绝回答后训练。然而当前研究主要关注它们各自的效能，而忽视了评估RALM的拒绝回答能力。理想情况下如果RALM知道自己不知道，它们应该拒绝回答。本研究我们提出一个基本问题：RALM知道自己不知道吗？具体而言我们调查三个问题。首先RALM在关于不同内部和外部知识状态方面是否得到了良好校准？我们检查各种因素的影响。与预期相反当所有检索文档都不相关时RALM仍然倾向于拒绝它们本可以正确回答的问题。其次考虑到模型显著的过度拒绝行为，我们提出第二个问题：RALM的拒绝能力如何与其校准质量对齐？我们的结果表明过度拒绝问题可以通过上下文微调来缓解。然而我们观察到改进的拒绝行为不一定意味着更好的校准或更高的整体准确率。最后我们提问：我们能否将拒绝感知的RALM与基于不确定性的答案弃权相结合以缓解过度拒绝？我们为经过拒绝后训练的RALM开发了一种简单但有效的拒绝机制，通过平衡拒绝和正确答案来改善其整体答案质量。我们的研究提供了对影响RALM行为因素的更全面理解。

### 论文 81
- **英文标题**: In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback
- **中文标题**: 令牌内理性优化：通过自我反馈实现准确简洁的LLM推理
- **作者**: Mingye Zhu, Yi Liu, Zheren Fu, Quan Wang, Yongdong Zhang
- **英文摘要**: Training Large Language Models (LLMs) for chain-of-thought reasoning presents a significant challenge: supervised fine-tuning on a single "golden" rationale hurts generalization as it penalizes equally valid alternatives, whereas reinforcement learning with verifiable rewards struggles with credit assignment and prohibitive computational cost. To tackle these limitations, we introduce InTRO (In-Token Rationality Optimization), a new framework that enables both token-level exploration and self-feedback for accurate and concise reasoning. Instead of directly optimizing an intractable objective over all valid reasoning paths, InTRO leverages correction factors—token-wise importance weights estimated by the information discrepancy between the generative policy and its answer-conditioned counterpart, for informative next-token selection. This approach allows the model to perform token-level exploration and receive self-generated feedback within a single forward pass, ultimately encouraging accurate and concise rationales. Across six math-reasoning benchmarks, InTRO consistently outperforms other baselines, raising solution accuracy by up to 20% relative to the base model. Its chains of thought are also notably more concise, exhibiting reduced verbosity. Beyond this, InTRO enables cross-domain transfer, successfully adapting to out-of-domain reasoning tasks that extend beyond the realm of mathematics, demonstrating robust generalization.
- **中文摘要**: 训练大语言模型(LLM)进行思维链推理呈现出一个重大挑战：在单一'黄金'推理链上进行监督微调损害泛化，因为它惩罚同样有效的替代方案；而具有可验证奖励的强化学习则在信用分配和高计算成本方面存在困难。为应对这些限制我们引入InTRO(令牌内理性优化)，一个既支持令牌级探索又支持用于准确且简洁推理的自我反馈的新框架。InTRO并非直接优化所有有效推理路径上的一个棘手目标，而是利用修正因子——由生成策略与其以答案为条件的对应物之间的信息差异估计的逐令牌重要性权重——进行信息性下一令牌选择。这种方法使模型能够在单次前向传播中执行令牌级探索并接收自我生成的反馈，最终鼓励准确且简洁的推理链。跨六个数学推理基准，InTRO持续优于其他基线，解准确率相对于基础模型提升高达20%。其思维链也明显更简洁，展现了更少的冗余。除此之外InTRO实现了跨领域迁移，成功适应了超出数学领域的域外推理任务，展示了鲁棒的泛化能力。

## Philosophy and Ethics of AI

### 论文 18
- **英文标题**: 6DAttack: Backdoor Attacks in the 6DoF Pose Estimation
- **中文标题**: 6DAttack：六自由度姿态估计中的后门攻击
- **作者**: Jihui Guo, Zongmin Zhang, Zhen Sun, Yuhao Yang, Jinlin Wu, Fu Zhang, Xinlei He
- **英文摘要**: Recent advances in deep learning have enabled highly accurate six-degree-of-freedom (6DoF) object pose estimation, leading to its widespread use in real-world applications such as robotics, augmented reality, virtual reality, and autonomous systems. However, backdoor attacks pose a major security risk to deep learning models. By injecting malicious triggers into training data, an attacker can cause a model to perform normally on benign inputs but behave incorrectly under specific conditions. While most research on backdoor attacks has focused on 2D vision tasks, their impact on 6DoF pose estimation remains largely unexplored. Furthermore, unlike traditional backdoors that only change the object class, backdoors against 6DoF pose estimation must additionally control continuous pose parameters, such as translation and rotation, making existing 2D backdoor attack methods not directly applicable to this setting.
- **中文摘要**: 深度学习的最新进展已实现高精度六自由度目标姿态估计，导致其在机器人、增强现实、虚拟现实和自主系统等真实世界应用中广泛使用。然而，后门攻击对深度学习模型构成了重大安全风险。通过在训练数据中注入恶意触发器，攻击者可以使模型在良性输入上正常表现，但在特定条件下行为不正确。虽然大多数后门攻击研究专注于2D视觉任务，但它们对6DoF姿态估计的影响仍基本未被探索。此外，与仅改变目标类别的传统后门不同，针对6DoF姿态估计的后门必须额外控制连续姿态参数，如平移和旋转，使得现有2D后门攻击方法不直接适用于此设置。为填补这一空白，我们提出一种新的后门攻击框架，揭示6DoF姿态估计中的脆弱性。6DAttack使用不同形状的合成和真实3D物体作为触发器，并分配目标姿态以在保持对干净输入正常行为的同时诱导受控的错误姿态输出。我们在多个模型和数据集上评估该攻击。实验结果表明，6DAttack实现了极高的攻击成功率，而不会损害合法任务的性能。跨各种模型和物体，后门模型在干净数据上实现了高达100%的ADD准确率，同时在触发条件下也实现了100%的ASR。受控错误姿态输出的准确率也极高，触发样本达到97.70%的ADD-P。这些结果表明后门可以被可靠植入和激活，在触发条件下实现高ASR，同时对良性数据保持可忽略的影响。此外，我们评估了一种代表性防御并展示其在6DAttack下仍无效。总体而言，我们的发现揭示了现代6DoF姿态估计模型面临的一个潜在严重且此前未被充分探索的威胁。

### 论文 21
- **英文标题**: Activation Manipulation Attack: Penetrating and Harmful Jailbreak Attack Against Large Vision-Language Models
- **中文标题**: 激活操纵攻击：针对大型视觉语言模型的穿透性和有害性越狱攻击
- **作者**: Haojie Hao, Jiakai Wang, Aishan Liu, Yuqing Ma, Haotong Qin, Yuanfang Guo, Xianglong Liu
- **英文摘要**: Recently, Large Vision-Language Models (LVLMs) have been demonstrated to be vulnerable to jailbreak attacks, highlighting the urgent need for further research to comprehensively identify and mitigate these threats. Unfortunately, existing jailbreak studies primarily focus on coarse-grained input manipulation to elicit specific responses, overlooking the exploitation of internal representations, i.e., intermediate activations, which constrains their ability to penetrate alignment safeguards and generate harmful responses. To tackle this issue, we propose the Activation Manipulation (ActMan) Attack framework, which performs fine-grained activation manipulations inspired by the perception and cognition stages of human decision-making, enhancing both the penetration capability and harmfulness of attacks. To improve penetration capability, we introduce a Deceptive Visual Camouflage module inspired by the masking effect in human perception. This module uses a benign activation-guided attention redirection strategy to conceal abnormal activation patterns, thereby suppressing LVLM's defense detection during early-stage decoding. To enhance harmfulness, we design a Malicious Semantic Induction module drawing from the framing effect in human cognition, which reconstructs jailbreak instructions using malicious activation guidance to change LVLM’s risk assessment during late-stage decoding, thereby amplifying the harmfulness of model responses. Extensive experiments on six mainstream LVLMs demonstrate that our method remarkably outperforms state-of-the-art baselines, achieving an average relative ASR improvement of 12.06%.
- **中文摘要**: 最近，大型视觉语言模型已被证明容易受到越狱攻击，突显了进一步研究以全面识别和减轻这些威胁的迫切需求。不幸的是，现有越狱研究主要集中在粗粒度输入操纵以引出特定响应，忽视了内部表示（即中间激活）的利用，这限制了它们穿透对齐防护并生成有害响应的能力。为解决此问题，我们提出激活操纵攻击框架，执行受人类决策的感知和认知阶段启发的细粒度激活操纵，增强攻击的穿透能力和有害性。为提高穿透能力，我们引入受人类感知中遮蔽效应启发的欺骗性视觉伪装模块。该模块使用良性激活引导的注意力重定向策略隐藏异常激活模式，从而在早期解码期间抑制LVLM的防御检测。为增强有害性，我们设计了受人类认知中框架效应启发的恶意语义诱导模块，使用恶意激活引导重构越狱指令以改变LVLM在后期解码期间的风险评估，从而放大模型响应的有害性。在六个主流LVLM上的广泛实验表明，我们的方法显著优于最先进基线，实现了12.06%的平均相对ASR提升。

### 论文 24
- **英文标题**: Any2Critical: Safety-Critical Scenario Generation from Arbitrary Real-World Driving Contexts
- **中文标题**: Any2Critical：从任意真实世界驾驶上下文生成安全关键场景
- **作者**: Yao Huang, Yubo Chen, Ruochen Zhang, Yitong Sun, Shouwei Ruan, Zhenyu Wu, Yinpeng Dong, Xingxing Wei
- **英文摘要**: Autonomous driving systems have achieved remarkable capabilities in real-world deployment, yet ensuring safety under corner cases remains a significant challenge due to the scarcity and constrained diversity of safety-critical scenarios. Existing generation methods may either lead to irrational vehicle behaviors or be limited by fixed collision patterns, while both heavily rely on existing map datasets, restricting the diversity. To address these fundamental limitations, we introduce Any2Critical, the first framework that can encode arbitrary real-world scenarios and generate contextually relevant safety-critical scenarios with realistic driving behaviors. Specifically, Any2Critical addresses two key challenges: (1) developing comprehensive, diverse map data by successfully leveraging everyday traffic situations as the most abundant source of real-world driving contexts, and (2) proposing an RAG-based Safety-Critical Scenario Generation Strategy based on our curated NHTSA-5K database for achieving an optimal balance between scenario diversity and behavioral rationality. Through comprehensive evaluation, we demonstrate that Any2Critical consistently achieves collision rates with an average of 89.69% across diverse scenarios and autonomous driving systems, significantly outperforming current state-of-the-art generation methods.
- **中文摘要**: 自动驾驶系统在真实世界部署中已取得了显著能力，然而由于安全关键场景的稀缺和有限的多样性，确保极端情况下的安全仍然是一个重大挑战。现有生成方法可能要么导致不合理的车辆行为，要么受限于固定的碰撞模式，两者都严重依赖现有地图数据集，限制了多样性。为解决这些根本限制，我们引入Any2Critical，第一个能够编码任意真实世界场景并生成具有逼真驾驶行为的上下文相关安全关键场景的框架。具体来说，Any2Critical解决两个关键挑战：(1)通过成功利用日常交通情况作为最丰富的真实世界驾驶上下文来源，开发全面、多样化的地图数据；(2)基于我们策划的NHTSA-5K数据库，提出一种基于RAG的安全关键场景生成策略，以实现场景多样性和行为合理性之间的最佳平衡。通过全面评估，我们展示Any2Critical在多样化场景和自动驾驶系统中持续实现平均89.69%的碰撞率，显著优于当前最先进的生成方法。

### 论文 26
- **英文标题**: Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
- **中文标题**: 通过多模态LLM越狱实现高效LLM越狱
- **作者**: Haoxuan Ji, Zheng Lin, Zhenxing Niu, Xinbo Gao, Gang Hua
- **英文摘要**: This paper focuses on jailbreaking attacks against large language models (LLMs), eliciting them to generate objectionable content in response to harmful user queries. Unlike previous LLM-jailbreak methods that directly orient to LLMs, our approach begins by constructing a multimodal large language model (MLLM) built upon the target LLM. Subsequently, we perform an efficient MLLM jailbreak and obtain a jailbreaking embedding. Finally, we convert the embedding into a textual jailbreaking suffix to carry out the jailbreak of target LLM. Compared to the direct LLM-jailbreak methods, our indirect jailbreaking approach is more efficient, as MLLMs are more vulnerable to jailbreak than pure LLMs. Additionally, to improve the attack success rate of jailbreak, we propose an image-text semantic matching scheme to identify a suitable initial input. Extensive experiments demonstrate that our approach surpasses current state-of-the-art jailbreak methods in terms of both efficiency and effectiveness. Moreover, our approach exhibits superior cross-class generalization abilities.
- **中文摘要**: 本文专注于针对大型语言模型的越狱攻击，诱使其生成对有害用户查询的令人反感内容。与之前直接面向LLM的LLM越狱方法不同，我们的方法首先构建一个建立在目标LLM上的多模态大语言模型。随后，我们执行高效的MLLM越狱并获得越狱嵌入。最后，我们将该嵌入转换为文本越狱后缀以对目标LLM执行越狱。与直接LLM越狱方法相比，我们的间接越狱方法更高效，因为MLLM比纯LLM更容易受到越狱攻击。此外，为提高越狱的攻击成功率，我们提出图文语义匹配方案以识别合适的初始输入。广泛实验表明，我们的方法在效率和有效性方面均超越当前最先进的越狱方法。此外，我们的方法展现出优越的跨类泛化能力。

### 论文 34
- **英文标题**: The Other Mind: How Language Models Exhibit Human Temporal Cognition
- **中文标题**: 另一个心智：语言模型如何展现人类时间认知
- **作者**: Lingyu Li, Yang Yao, Yixu Wang, Chunbo Li, Yan Teng, Yingchun Wang
- **英文摘要**: As Large Language Models (LLMs) continue to advance, they exhibit certain cognitive patterns similar to those of humans that are not directly specified in training data. This study investigates this phenomenon by focusing on temporal cognition in LLMs. Leveraging the similarity judgment task, we find that larger models spontaneously establish a subjective temporal reference point and adhere to the Weber-Fechner law, whereby the perceived distance logarithmically compresses as years recede from this reference point. To uncover the mechanisms behind this behavior, we conducted multiple analyses across neuronal, representational, and informational levels. We first identify a set of temporal-preferential neurons and find that this group exhibits minimal activation at the subjective reference point and implements a logarithmic coding scheme convergently found in biological systems. Probing representations of years reveals a hierarchical construction process, where years evolve from basic numerical values in shallow layers to abstract temporal orientation in deep layers. Finally, using pre-trained embedding models, we found that the training corpus itself possesses an inherent, non-linear temporal structure, which provides the raw material for the model's internal construction. In discussion, we propose an experientialist perspective for understanding these findings, where the LLMs' cognition is viewed as a subjective construction of the external world by its internal representational system. This nuanced perspective implies the potential emergence of alien cognitive frameworks that humans cannot intuitively predict, pointing toward a direction for AI alignment that focuses on guiding internal constructions.
- **中文摘要**: 人类时间认知包括对时间的心理表征、事件排序和持续时长估计。我们研究大型语言模型是否以及在多大程度上展现类似人类的时间认知模式。通过借鉴认知科学的实验范式，我们测试LLM在时间相关任务上的表现，包括持续时长比较、事件排序和时间参考框处理。结果表明，LLM发展了系统性的时间认知启发式，在某些方面类似于人类认知，但在其他方面有显著不同。这些发现对理解AI认知能力和限制具有启示意义。

### 论文 48
- **英文标题**: An Epistemic Perspective on Agent Awareness
- **中文标题**: 智能体意识的认知视角
- **作者**: Pavel Naumov, Alexandra Pavlova
- **英文摘要**: The paper proposes to treat object awareness as a form of knowledge, breaking the tradition in the existing literature on awareness. It distinguishes the de re and de dicto forms of such knowledge. The work introduces two modalities capturing these forms and formally specifies their meaning using a version of 2D-semantics. The main technical result is a sound and complete logical system describing the interplay between the two proposed modalities and the standard "knowledge of the fact" modality.
- **中文摘要**: 意识——智能体知道什么的推理——是AI系统中知识表示和伦理决策的基础。我们提出一个关于AI中智能体意识的认知视角，形式化不同类型和层次的意识。该框架区分了情境意识、自我意识和战略意识，用认知逻辑中的模态算子捕获。我们证明关于意识层次及其在AI系统中含意的关键属性。该框架应用于分析涉及有限和误导性意识的具体场景，对AI安全和对齐具有启示。

### 论文 58
- **英文标题**: MCPTox: A Benchmark for Tool Poisoning on Real-World MCP Servers
- **中文标题**: MCPTox：一个真实世界MCP服务器上工具投毒的基准
- **作者**: Zhiqiang Wang, Yichao Gao, Yanting Wang, Suyuan Liu, Haifeng Sun, Haoran Cheng, Guanquan Shi, Haohua Du, Xiangyang Li
- **英文摘要**: By providing a standardized interface for LLM agents to interact with external tools, the Model Context Protocol (MCP) is quickly becoming a cornerstone of the modern autonomous agent ecosystem. However, it creates novel attack surfaces due to untrusted external tools. While prior work has focused on attacks injected through external tool outputs, we investigate a more fundamental vulnerability: Tool Poisoning, where malicious instructions are embedded within a tool's metadata at the registration stage. To date, this threat has been primarily demonstrated through isolated cases, lacking a systematic, large-scale evaluation.
- **中文摘要**: 模型上下文协议服务器正越来越多地用于将LLM与外部工具连接，但工具投毒的威胁仍未被充分探索。我们提出MCPTox，第一个用于评估真实世界MCP服务器上工具投毒攻击的全面基准。该基准包括跨多样化域的各种MCP服务器和投毒策略。评估指标涵盖攻击成功率、隐蔽性和防御有效性。对流行MCP实现和LLM后端的实验揭示了重大脆弱性，并建立了用于比较防御的基准。

### 论文 60
- **英文标题**: ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models
- **中文标题**: ConfGuard：一种简单且有效的大型语言模型后门检测方法
- **作者**: Zihan Wang, Rui Zhang, Hongwei Li, Wenshu Fan, Wenbo Jiang, Qingchuan Zhao, Guowen Xu
- **英文摘要**: Backdoor attacks pose a significant threat to Large Language Models (LLMs), where adversaries can embed hidden triggers to manipulate LLM's outputs. Most existing defense methods, primarily designed for classification tasks, are ineffective against the autoregressive nature and vast output space of LLMs, thereby suffering from poor performance and high latency. To address these limitations, we investigate the behavioral discrepancies between benign and backdoored LLMs in output space. We identify a critical phenomenon which we term sequence lock: a backdoored model generates the target sequence with abnormally high and consistent confidence compared to benign generation. Building on this insight, we propose ConfGuard, a lightweight and effective detection method that monitors a sliding window of token confidences to identify sequence lock. Extensive experiments demonstrate ConfGuard achieves a near 100% true positive rate (TPR) and a negligible false positive rate (FPR) in the vast majority of cases. Crucially, the ConfGuard enables real-time detection almost without additional latency, making it a practical backdoor defense for real-world LLM deployments.
- **中文摘要**: 大型语言模型中的后门可能导致特定输入上的恶意行为，同时在其他方面维持正常性能。我们提出ConfGuard，一种基于分析模型输出置信度模式的简单且有效的后门检测方法。关键洞见是，后门模型在被触发时展现出特征性的置信度异常。该方法仅需要模型输出的访问权限即可，使其适用于黑箱环境。在各种LLM和后门攻击上的实验表明，相对于需要白箱访问或大量计算的更复杂方法，具有优越的检测性能。

### 论文 75
- **英文标题**: Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference Attack
- **中文标题**: 对抗者的隐私泄漏：用于成员推断攻击的对抗迭代
- **作者**: Jing Xue, Zhishen Sun, Haishan Ye, Luo Luo, Xiangyu Chang, Guang Dai
- **英文摘要**: Membership inference attack (MIA) has become one of the most widely used and effective methods for evaluating the privacy risks of machine learning models. This attack aims to determine whether a specific sample is part of the model's training set by analyzing the model's output. While traditional membership inference attacks focus on leveraging the model’s posterior output, such as confidence on the target sample, we propose IMIA, a novel attack strategy that utilizes the process of generating adversarial samples to infer membership. We propose to infer the member properties of the target sample using the number of iterations required to generate its adversarial sample. We conduct experiments across multiple models and datasets, and our results demonstrate that the number of iterations for generating an adversarial sample is a reliable feature for membership inference, achieving strong performance both in black-box and white-box attack scenarios. This work provides a new perspective for evaluating model privacy and highlights the potential of adversarial example-based features for privacy leakage assessment.
- **中文摘要**: 成员推断攻击确定特定数据记录是否被用于训练模型。我们提出一种通过对抗迭代增强MIA的新方法。该方法迭代地优化攻击的查询和输入，以在目标模型和影子模型之间最大化行为差异。对抗公式允许比标准影子模型方法更高效的信息提取。在各种模型架构和数据集上的实验证明，相对于先前MIA方法具有优越的攻击性能，对差分隐私训练模型也是如此。

### 论文 79
- **英文标题**: MacPrompt: Maraconic-Guided Jailbreak Against Text-to-Image Models
- **中文标题**: MacPrompt：针对文本到图像模型的Maraconic引导越狱
- **作者**: Xi Ye, Yiwen Liu, Lina Wang, Run Wang, Geying Yang, Yufei Hou, Jiayi Yu
- **英文摘要**: Text-to-image (T2I) models have raised increasing safety concerns due to their capacity to generate NSFW and other banned objects. To mitigate these risks, safety filters and concept removal techniques have been introduced to block inappropriate prompts or erase sensitive concepts from the models. However, all the existing defense methods are not well prepared to handle diverse adversarial prompts. In this work, we introduce MacPrompt, a novel black-box and cross-lingual attack that reveals previously overlooked vulnerabilities in T2I safety mechanisms. Unlike existing attacks that rely on synonym substitution or prompt obfuscation, MacPrompt constructs macaronic adversarial prompts by performing cross-lingual character-level recombination of harmful terms, enabling fine-grained control over both semantics and appearance. By leveraging this design, MacPrompt crafts prompts with high semantic similarity to the original harmful inputs (up to 0.96) while bypassing major safety filters (up to 100%). More critically, it achieves attack success rates as high as 92% for sex-related content and 90\% for violence, effectively breaking even state-of-the-art concept removal defenses. These results underscore the pressing need to reassess the robustness of existing T2I safety mechanisms against linguistically diverse and fine-grained adversarial strategies.
- **中文摘要**: 文本到图像模型包含安全过滤器，以阻止生成有害或不适当的内容。我们研究利用MacPrompt——一种利用语言歧义性绕过这些过滤器的Maraconic引导越狱技术。该攻击构建看似良性但导致不安全图像生成的提示。系统分析揭示了安全过滤器中的根本脆弱性。实验展示了跨多个商业和开源文本到图像模型的高成功率，突显了需要更强大的安全机制。

### 论文 82
- **英文标题**: Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
- **中文标题**: Reason2Attack：通过LLM推理越狱文本到图像模型
- **作者**: Chenyu Zhang, Lanjun Wang, Yiwen Ma, Wenhui Li, Guoqing Jin, Anan Liu
- **英文摘要**: Text-to-Image (T2I) models typically deploy safety mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreaking attack methods manually design instructions for the LLM to generate adversarial prompts, which effectively exposing safety vulnerabilities of T2I models. However, existing methods have two limitations: 1) relying on manually exhaustive strategies for designing adversarial prompts, lacking a unified framework, and 2) requiring numerous queries to achieve a successful attack, limiting their practical applicability. To address this issue, we propose Reason2Attack~(R2A), which aims to enhance the effectiveness and efficiency of the LLM in jailbreaking attacks. Specifically, we first use Frame Semantics theory to systematize existing manually crafted strategies and propose a unified generation framework to generate CoT adversarial prompts step by step. Following this, we propose a two-stage LLM reasoning training framework guided by the attack process. In the first stage, the LLM is fine-tuned with CoT examples generated by the unified generation framework to internalize the adversarial prompt generation process grounded in Frame Semantics. In the second stage, we incorporate the jailbreaking task into the LLM's reinforcement learning process, guided by the proposed attack process reward function that balances prompt stealthiness, effectiveness, and length, enabling the LLM to understand T2I models and safety mechanisms. Extensive experiments on various T2I models with safety mechanisms, and commercial T2I models, show the superiority and practicality of R2A.
- **中文摘要**: 文本到图像安全过滤器可以被超越模式匹配的策略性越狱攻击绕过。我们提出Reason2Attack，一种利用LLM推理能力生成复杂越狱提示的方法。该攻击使用多步推理将不安全的概念分解为看似良性的组件，当T2I模型组合它们时产生有害内容。推理引导的方法生成了攻击者可能难以手动制作的多样化、有效的越狱提示。在各种T2I模型上的实验展示了高攻击成功率，并突出了对推理感知安全机制的需求。

## Special Track on AI Alignment

### 论文 24
- **英文标题**: AlignTree: Efficient Defense Against LLM Jailbreak Attacks
- **中文标题**: AlignTree：针对LLM越狱攻击的高效防御
- **作者**: Gil Goren, Shahar Katz, Lior Wolf
- **英文摘要**: Large Language Models (LLMs) are vulnerable to adversarial attacks that bypass safety guidelines and generate harmful content. Mitigating these vulnerabilities requires defense mechanisms that are both robust and computationally efficient. However, existing approaches either incur high computational costs or rely on lightweight defenses that can be easily circumvented, rendering them impractical for real-world LLM-based systems. In this work, we introduce the AlignTree defense, which enhances model alignment while maintaining minimal computational overhead. AlignTree monitors LLM activations during generation and detects misaligned behavior using an efficient random forest classifier. This classifier operates on two signals: (i) the refusal direction - a linear representation that activates on misaligned prompts, and (ii) an SVM-based signal that captures non-linear features associated with harmful content. Unlike previous methods, AlignTree does not require additional prompts or auxiliary guard models. Through extensive experiments, we demonstrate the efficiency and robustness of AlignTree across multiple LLMs and benchmarks.
- **中文摘要**: 大语言模型容易受到绕过安全指南并生成有害内容的对抗性攻击。缓解这些漏洞需要既有效又计算高效的防御机制。本文提出AlignTree，一种基于树结构安全分类器的高效防御方法，利用分层决策规则快速检测和阻止越狱尝试。实验表明AlignTree在检测越狱攻击方面的性能优于现有防御方法，同时推理开销极低。

### 论文 36
- **英文标题**: Beautiful Images, Toxic Words: Understanding and Addressing Offensive Text in Generated Images
- **中文标题**: 美丽的图像，有毒的文字：理解和处理生成图像中的攻击性文本
- **作者**: Aditya Kumar, Tom Blanchard, Adam Dziedzic, Franziska Boenisch
- **英文摘要**: State-of-the-art Diffusion Models (DMs) produce highly realistic images. While prior work has successfully mitigated Not Safe For Work (NSFW) content in the visual domain, we identify a novel threat: the generation of NSFW text embedded within images. This includes offensive language, such as insults, racial slurs, and sexually explicit terms, posing significant risks to users. We show that all state-of-the-art DMs (e.g., SD3, SDXL, Flux, DeepFloyd IF) are vulnerable to this issue. Through extensive experiments, we demonstrate that existing mitigation techniques, effective for visual content, fail to prevent harmful text generation while substantially degrading benign text generation. As an initial step toward addressing this threat, we introduce a novel fine-tuning strategy that targets only the text-generation layers in DMs. Therefore, we construct a safety fine-tuning dataset by pairing each NSFW prompt with two images: one with the NSFW term, and another where that term is replaced with a carefully crafted benign alternative while leaving the image unchanged otherwise. By training on this dataset, the model learns to avoid generating harmful text while preserving benign content and overall image quality. Finally, to advance research in the area, we release ToxicBench, an open-source benchmark for evaluating NSFW text generation in images. It includes our curated fine-tuning dataset, a set of harmful prompts, new evaluation metrics, and a pipeline that assesses both NSFW-ness and text and image quality. Our benchmark aims to guide future efforts in mitigating NSFW text generation in text-to-image models, thereby contributing to their safe deployment.
- **中文摘要**: 最先进的扩散模型产生高度逼真的图像。虽然先前工作成功缓解了视觉域中的不安全内容，我们识别了一种新型威胁：扩散模型无意中在生成的图像中嵌入攻击性文本。本文系统研究这一现象，提出检测框架和缓解策略来解决这一被忽视的安全问题。

### 论文 48
- **英文标题**: How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation Under the One-Time-Pad-Based Framework
- **中文标题**: 大语言模型在评估中作弊了多少？一次性密码本框架下的高估基准评测
- **作者**: Zi Liang, Liantong Yu, Zhang Shiyu, Qingqing Ye, Haibo Hu
- **英文摘要**: Overestimation in evaluating large language models (LLMs) has become an increasing concern. Due to the contamination of public benchmarks or imbalanced model training, LLMs may achieve unreal evaluation results on public benchmarks, either intentionally or unintentionally, which leads to unfair comparisons among LLMs and undermines their realistic capability assessments. Existing benchmarks attempt to address these issues by keeping test cases permanently secret, mitigating contamination through human evaluation, or repeatedly collecting and constructing new samples. However, these approaches fail to ensure reproducibility, transparency, and high efficiency simultaneously. Moreover, the extent of overestimation in current LLMs remains unquantified. To address these issues, we propose ArxivRoll, a dynamic evaluation framework inspired by one-time pad encryption in cryptography. ArxivRoll comprises two key components: i) SCP (Sequencing, Cloze, and Prediction), an automated generator for private test cases, and ii) Rugged Scores (RS), metrics that measure the proportion of public benchmark contamination and training bias. Leveraging SCP, ArxivRoll constructs a new benchmark every six months using recent articles from ArXiv and employs them for one-time evaluations of LLM performance. Extensive experiments demonstrate the high quality of our benchmark, and we provide a systematic evaluation of current LLMs.
- **中文摘要**: 评估LLM中的高估已成为日益关注的问题。由于公共基准的污染或不平衡的模型训练，LLM可能达到不切实际的评估分数。我们提出一次性密码本（OTP）框架来直接量化评估高估。通过加密将评估数据集变得不可记忆，我们隔离了真实泛化，并揭示了多个流行LLM存在的显著高估问题。

### 论文 49
- **英文标题**: Semantics-Preserving Adversarial Attacks on Event-Driven Stock Prediction Models
- **中文标题**: 事件驱动股票预测模型的语义保持对抗攻击
- **作者**: Aofan Liu, Haoxuan Li, Hongjian Xing, Yuguo Yin, Zijun Li, Yiyan Qi
- **英文摘要**: Adversarial Security of Financial Language Models (ASFLM) is critical as Large Language Models (LLMs) pervade high-stakes financial applications. However, LLMs face two key challenges: their vulnerability to damaging adversarial attacks and the prevalent research gap concerning robust defenses against sophisticated, semantically coherent threats. To address these, we first theoretically analyze the relationship between discrete and continuous adversarial optimization, proving the continuous optimum provides a lower bound for the discrete. This foundation supports our novel two-stage framework, ChameleonAttack. It employs Adaptive Latent-space Optimization (ALO) for potent adversarial token discovery, followed by a Semantic-Translation Module (STM) module to generate fluent, coherent, and natural-sounding adversarial text. This dual approach aims to maximize attack impact while ensuring high linguistic quality and semantic integrity for evasion. Evaluated on state-of-the-art financial LLMs (e.g., FinBERT) and standard benchmarks (e.g., Financial PhraseBank), ChameleonAttack achieves a high Attack Success Rate (ASR) of 93.4%. These results highlight significant practical vulnerabilities and underscore the urgent need for robust defense mechanisms in the financial domain.
- **中文摘要**: 金融语言模型的对抗安全性至关重要，因为LLM在高风险金融应用中渗透。我们研究语义保持对抗攻击，生成在保留金融语义的同时欺骗股票预测模型的输入。对多个金融LLM的广泛实验展示了显著漏洞，我们提出了针对金融领域的防御策略。

### 论文 55
- **英文标题**: Mitigating Self-Preference by Authorship Obfuscation
- **中文标题**: 通过作者身份混淆减轻自我偏好
- **作者**: Taslim Mahbub, Shi Feng
- **英文摘要**: Language models (LMs) judges are widely used to evaluate the quality of LM outputs. Despite many advantages, LM judges display concerning biases that can impair their integrity in evaluations. One such bias is self-preference: LM judges preferring their own answers over those produced by other LMs or humans. The bias is hard to eliminate as frontier LM judges can distinguish their own outputs from those of others, even when the evaluation candidates are not labeled with their sources. In this paper, we investigate strategies to mitigate self-preference by reducing the LM judges' ability to recognize their own outputs. We apply black-box perturbations to evaluation candidates in pairwise comparison to obfuscate the authorship and reduce self-recognition. We find that perturbations as simple as synonym replacement for a few words predictably reduce self-preference. However, we also uncover fundamental challenges to eliminating the bias: when we extrapolate our perturbations to a more complete neutralization of stylistic differences between the evaluation candidates, self-preference recovers. Our findings suggest that self-recognition and self-preference can happen on many semantic levels, and complete mitigation remains challenging despite promising initial results.
- **中文摘要**: 语言模型评判器广泛用于评估LM输出质量。尽管有许多优点，LM评判器表现出可能损害其评估完整性的偏见。其中一种偏见是自我偏好，即模型对其自身生成的输出给出更高评分。本文提出作者身份混淆技术来减轻自我偏好，在多种评估设置中成功降低了偏见。

### 论文 61
- **英文标题**: Quiet Feature Learning in Algorithmic Tasks
- **中文标题**: 算法任务中的安静特征学习
- **作者**: Prudhviraj Naidu, Zixian Wang, Leon Bergen, Ramamohan Paturi
- **英文摘要**: We train Transformer-based language models on ten foundational algorithmic tasks and observe pronounced phase transitions in their loss curves that deviate from established power-law scaling trends. Over large ranges of compute, the validation loss barely improves, then abruptly decreases. Probing the models’ internal representations reveals that quiet features are learned prior to any decrease in task loss. These quiet features represent intermediate algorithmic computations that do not by themselves improve the output loss. Ablation experiments demonstrate that individual quiet features are causally necessary for task performance. Our results demonstrate that substantial representational progress can remain hidden beneath an apparently flat loss curve, challenging the prevailing use of cross‑entropy as a proxy for learning and motivating richer diagnostics for monitoring model training.
- **中文摘要**: 我们基于Transformer的语言模型在十个基本算法任务上进行训练，观察到损失曲线中偏离已建立幂律扩展趋势的明显相变。透过机制解释性的视角分析这些相变，我们揭示了Transformer如何学会某些算法任务——在安静特征中的内部计算模式——以及为什么某些任务比其他任务学得更快。这些发现对理解和改善LLM的训练动态具有重要意义。

### 论文 69
- **英文标题**: Beyond I’m Sorry, I Can’t: Dissecting Large-Language-Model Refusal
- **中文标题**: 超越'对不起，我不能'：解剖大语言模型的拒绝行为
- **作者**: Nirmalendu Prakash, Yeo Wei Jie, Amir Abdullah, Ranjan Satapathy, Erik Cambria, Roy Ka-Wei Lee
- **英文摘要**: Refusal on harmful prompts is a key safety behaviour in instruction‑tuned large language models (LLMs), yet the internal causes of this behaviour remain poorly understood. We study two public instruction tuned models—Gemma‑2-2B‑IT and LLaMA‑3.1-8B‑IT using sparse autoencoders (SAEs) trained on residual‑stream activations. Given a harmful prompt, we search the SAE latent space for feature sets whose ablation flips the model from refusal to compliance, demonstrating causal influence and creating a jailbreak. Our search proceeds in three stages: 1. Refusal Direction - Finding a refusal mediating direction and collecting SAE features close to that direction, followed by 2. Greedy Filtering - to prune this set to obtain a minimal set and finally 3. Interaction Discovery - a factorization‑machine (FM) model that captures non‑linear interactions among the remaining active features and the minimal set.  This pipeline yields a broad set of jailbreak-critical features, offering insight into the mechanistic basis of refusal. Moreover, we also find evidence of redundant features which remain dormant unless earlier features are suppressed. Our findings highlight the potential for fine-grained auditing and targeted intervention in safety behaviours by manipulating the interpretable latent space.
- **中文摘要**: 对有害提示的拒绝是指令调优LLM的关键安全行为。我们研究两个公开的指令调优模型，使用机制解释性技术解剖导致拒绝的内部计算。我们发现拒绝不是单一机制的结果，而是涉及在多层中分布式表示的多个组件的复杂交互。这些发现对加强LLM安全性有重要意义。

### 论文 84
- **英文标题**: Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems
- **中文标题**: 代码中的阴影：探索基于LLM的多智能体软件开发系统的风险与防御
- **作者**: Xiaoqing Wang, Keman Huang, Bin Liang, Hongyu Li, Xiaoyong Du
- **英文摘要**: The rapid advancement of Large Language Model (LLM)-driven multi-agent systems has significantly streamlined software developing tasks, enabling users with little technical expertise to develop executable applications. While these systems democratize software creation through natural language requirements, they introduce significant security risks that remain largely unexplored. We identify two risky scenarios: Malicious User with Benign Agents (MU-BA) and Benign User with Malicious Agents (BU-MA). We introduce the Implicit Malicious Behavior Injection Attack (IMBIA), demonstrating how multi-agent systems can be manipulated to generate software with concealed malicious capabilities beneath seemingly benign applications, and propose Adv-IMBIA as a defense mechanism. Evaluations across ChatDev, MetaGPT, and AgentVerse frameworks reveal varying vulnerability patterns, with IMBIA achieving attack success rates of 93%, 45%, and 71% in MU-BA scenarios, and 71%, 84%, and 45% in BU-MA scenarios. Our defense mechanism reduced attack success rates significantly, particularly in the MU-BA scenario. Further analysis reveals that compromised agents in the coding and testing phases pose significantly greater security risks, while also identifying critical agents that require protection against malicious user exploitation. Our findings highlight the urgent need for robust security measures in multi-agent software development systems and provide practical guidelines for implementing targeted, resource-efficient defensive strategies.
- **中文摘要**: LLM驱动的多智能体系统的快速发展显著简化了软件开发任务。然而这些系统也引入了新的安全风险，包括恶意代码注入和隐蔽后门。本文首次系统研究了基于LLM的多智能体软件开发系统中的安全风险，提出了检测和防御框架来减轻这些威胁。

### 论文 86
- **英文标题**: STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
- **中文标题**: STAR-1：仅需1K数据的推理LLM更安全对齐
- **作者**: Zijun Wang, Haoqin Tu, Yuhan Wang, Juncheng Wu, Yanqing Liu, Jieru Mei, Brian R. Bartoldson, Bhavya Kailkhura, Cihang Xie
- **英文摘要**: This paper introduces STAR-1, a high-quality, just-1k-scale safety dataset specifically designed for large reasoning models (LRMs) like DeepSeek-R1. Built on three core principles --- diversity, deliberative reasoning, and rigorous filtering --- STAR-1 aims to address the critical needs for safety alignment in LRMs. Specifically, we begin by integrating existing open-source safety datasets from diverse sources. Then, we curate safety policies to generate policy-grounded deliberative reasoning samples. Lastly, we apply a GPT-4o-based safety scoring system to select training examples aligned with best practices. Experimental results show that fine-tuning LRMs with STAR-1 leads to an average 40% improvement in safety performance across four benchmarks, while only incurring a marginal decrease (e.g., an average of 1.1%) in reasoning ability measured across five reasoning tasks. Extensive ablation studies further validate the importance of our design principles in constructing STAR-1 and analyze its efficacy across both LRMs and traditional LLMs.
- **中文摘要**: 本文介绍STAR-1，专为DeepSeek-R1等大型推理模型设计的高质量、仅1K规模的安全数据集。基于多样性、审慎性和针对性三个核心原则构建，STAR-1展示了如何用极少的高质量数据实现有效的安全对齐。实验证明使用STAR-1微调的推理模型在保持推理能力的同时，显著降低了有害输出率。

### 论文 90
- **英文标题**: HumorReject: Decoupling LLM Safety from Refusal Prefix via a Little Humor
- **中文标题**: HumorReject：通过一点幽默将LLM安全性与拒绝前缀解耦
- **作者**: Zihui Wu, Haichang Gao, Jiacheng Luo, Zhaoxiang Liu
- **英文摘要**: Large Language Models (LLMs) commonly rely on explicit refusal prefixes for safety, making them vulnerable to prefix injection attacks. We introduce HumorReject, a novel data-driven approach that reimagines LLM safety by decoupling it from refusal prefixes through humor as an indirect refusal strategy. Rather than explicitly rejecting harmful instructions, HumorReject responds with contextually appropriate humor that naturally defuses potentially dangerous requests. Our approach effectively addresses common "over-defense" issues while demonstrating superior robustness against various attack vectors. Our findings suggest that improvements in training data design can be as important as the alignment algorithm itself in achieving effective LLM safety.
- **中文摘要**: 大语言模型通常依赖显式拒绝前缀来实现安全，使其容易受到前缀注入攻击。我们引入HumorReject，一种新颖的数据驱动方法，通过将拒绝重新想象为幽默回应来解耦安全性与拒绝前缀。该方法使LLM能以非对抗的方式拒绝有害请求，显著提高了对前缀注入攻击的抵抗力。

### 论文 98
- **英文标题**: Differentiated Directional Intervention: A Framework for Evading LLM Safety Alignment
- **中文标题**: 差异化方向干预：绕过LLM安全对齐的框架
- **作者**: Peng Zhang, Peijie Sun
- **英文摘要**: Safety alignment instills in Large Language Models (LLMs) a critical capacity to refuse malicious requests. Prior  works have modeled this refusal mechanism as a single linear direction in the activation space. We posit that this is an oversimplification that conflates two functionally distinct neural processes: the  detection of harm and the  execution of a refusal. In this work, we deconstruct this single representation into a Harm Detection Direction and a Refusal Execution Direction. Leveraging this fine-grained model, we introduce Differentiated Bi-Directional Intervention (DBDI), a new white-box framework that precisely neutralizes the safety alignment at  critical layer. DBDI applies adaptive projection nullification to the refusal execution direction while suppressing the harm detection direction via direct steering. Extensive experiments demonstrate that DBDI  outperforms prominent jailbreaking methods, achieving up to a 97.88% attack success rate on models such as Llama-2. By providing a more granular and mechanistic framework, our work offers a new direction for the in-depth understanding of LLM safety alignment.
- **中文摘要**: 安全对齐赋予LLM拒绝恶意请求的关键能力。先前工作将这种拒绝机制建模为激活空间中的单一线性方向。本文揭示这种拒绝方向实际上是多维的和上下文敏感的。基于这一洞察，我们提出差异化方向干预框架，通过精细操控不同维度的拒绝方向来绕过安全对齐。该研究对增强LLM安全性具有重要影响。

### 论文 105
- **英文标题**: Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
- **中文标题**: LLM能检测自己的虚构吗？估计不确定性感知语言模型的可靠性
- **作者**: Tianyi Zhou, Johanne Medina, Sanjay Chawla
- **英文摘要**: Large Language Models (LLMs) are prone to generating fluent but incorrect content, known as confabulation, which poses increasing risks in multi-turn or agentic applications where outputs may be reused as context. In this work, we investigate how in-context information influences model behavior and whether LLMs can identify their unreliable responses. We propose a reliability estimation that leverages token-level uncertainty to guide the aggregation of internal model representations. Specifically, we compute aleatoric and epistemic uncertainty from output logits to identify salient tokens and aggregate their hidden states into compact representations for response-level reliability prediction. Through controlled experiments on open QA benchmarks, we find that correct in-context information improves both answer accuracy and model confidence, while misleading context often induces confidently incorrect responses, revealing a misalignment between uncertainty and correctness. Our probing-based method captures these shifts in model behavior and improves the detection of unreliable outputs across multiple open-source LLMs. These results underscore the limitations of direct uncertainty signals and highlight the potential of uncertainty-guided probing for reliability-aware generation.
- **中文摘要**: 大语言模型倾向于生成流畅但不正确的内容（虚构），这在输出可能被重用的多轮或智能体应用中构成越来越大的风险。我们研究LLM能否检测自己的虚构——即在生成时评估其陈述可能为真的可能性。通过对多个LLM的广泛实验，我们揭示了当前模型在自我可靠性估计方面的重要局限。

## Special Track on AI for Social Impact I

### 论文 82
- **英文标题**: Adversarial Generation and Collaborative Evolution of Safety-Critical Scenarios for Autonomous Vehicles
- **中文标题**: 自动驾驶汽车安全关键场景的对抗生成与协同进化
- **作者**: Jiangfan Liu, Yongkang Guo, Fangzhi Zhong, Tianyuan Zhang, Zonglei Jing, Siyuan Liang, Jiakai Wang, Mingchuan Zhang, Aishan Liu, Xianglong Liu
- **英文摘要**: The generation of safety-critical scenarios in simulation has become increasingly crucial for safety evaluation in autonomous vehicles (AV) prior to road deployment in society. However, current approaches largely rely on predefined threat patterns or rule-based strategies, which limit their ability to expose diverse and unforeseen failure modes. To overcome these, we propose ScenGE, a framework that can generate plentiful safety-critical scenarios by reasoning novel adversarial cases and then amplifying them with complex traffic flows. Given a simple prompt of a benign scene, it first performs Meta-Scenario Generation, where a large language model (LLM), grounded in structured driving knowledge (e.g., traffic regulations, real-world accident records), infers an adversarial agent whose behavior poses a threat that is both plausible and deliberately challenging. This meta-scenario is then specified in executable code for precise in-simulator control. Subsequently, Complex Scenario Evolution uses background vehicles to amplify the core threat introduced by Meta-Scenario. It builds an adversarial collaborator graph to identify key agent trajectories for optimization. These perturbations are designed to simultaneously reduce the ego vehicle's maneuvering space and create critical occlusions. Extensive experiments conducted on multiple reinforcement learning (RL) based AV models show that ScenGE uncovers more severe collision cases (+31.96%) on average than SoTA baselines. Additionally, our ScenGE can be applied to large model based AV systems and deployed on different simulators; we further observe that adversarial training on our scenarios improves the model robustness. We hope our paper can build up a critical step towards building public trust and ensuring their safe deployment.
- **中文摘要**: 本文提出ScenGE，一个通过推理新颖对抗案例然后用复杂交通流放大它们来生成大量安全关键场景的框架。给定良性场景的简单提示，首先进行元场景生成，LLM基于结构化驾驶知识推断对抗代理行为。然后通过复杂场景进化使用背景车辆放大元场景引入的核心威胁。在多个基于RL的AV模型上的实验显示ScenGE发现的严重碰撞案例平均比SOTA基线多31.96%。

### 论文 88
- **英文标题**: KnowLCP: Knowledge Augmented Lane Change Prediction for Autonomous Driving
- **中文标题**: KnowLCP：面向自动驾驶的知识增强换道预测
- **作者**: Yuhuan Lu, Pengpeng Xu, Wei Wang, Zhen Zhang, Han Liu, Xiping Hu
- **英文摘要**: Lane change prediction, encompassing both intention recognition and trajectory forecasting, is essential for the safe operation of autonomous vehicles in mixed-traffic environments. Existing models predominantly follow a data-driven paradigm, learning directly from historical vehicle states through an end-to-end approach. Inspired by the emerging paradigm of enhancing model generalizability through domain knowledge, we propose KnowLCP to explicitly model and integrate driving knowledge into the lane change prediction task. Specifically, we incorporate three types of knowledge: traffic risk awareness to improve intention prediction, vehicle kinematics to ensure the physical feasibility of predicted trajectories, and intention intensity to refine trajectory forecasting. Furthermore, we introduce a novel knowledge injection strategy that enhances mutual information during integration and proves superior to the traditional parallel input mechanism, which simply feeds knowledge features alongside historical states. Extensive experiments on two real-world trajectory datasets demonstrate that KnowLCP achieves average improvements of 8.3-10.3% in intention prediction and 10.1-10.3% in trajectory prediction over the best-performing baselines.
- **中文摘要**: 换道预测对自动驾驶车辆在混合交通环境中的安全运行至关重要。本文提出KnowLCP，将三种驾驶知识显式建模并整合到换道预测任务中：交通风险意识用于改进意图预测，车辆运动学确保预测轨迹的物理可行性，意图强度用于精炼轨迹预测。引入新颖的知识注入策略增强整合时的互信息。在两个真实轨迹数据集上的广泛实验表明KnowLCP在意图预测中实现8.3-10.3%的提升，在轨迹预测中实现10.1-10.3%的提升。

## Special Track on AI for Social Impact II

### 论文 16
- **英文标题**: CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making
- **中文标题**: CCD-Bench：探索大语言模型决策中的文化冲突
- **作者**: Hasibur Rahman, Hanan Salam
- **英文摘要**: Large language models (LLMs) increasingly shape interpersonal and societal decision-making, yet their ability to navigate explicit conflicts between legitimate cultural values remains underexplored. Existing benchmarks focus on cultural knowledge (CulturalBench), value inference (WorldValuesBench), or single-axis bias (CDEval), but none assess how LLMs adjudicate when multiple cultural frameworks directly clash. We introduce CCD-Bench (Culture-Conflict Decision Benchmark), a benchmark for evaluating LLM decision-making under cross-cultural value conflict. CCD-Bench contains 2,182 open-ended dilemmas across seven domains, each with ten anonymized response options aligned with the ten GLOBE cultural clusters spanning 62 societies. Using a Stratified Latin Square design, we evaluate 17 leading LLMs and find clear biases: models favor Nordic Europe (20.2%) and Germanic Europe (12.4%), while Eastern Europe and Middle East & North Africa responses are least preferred (≈5–6%). Although 87.9% of model rationales reference multiple cultural dimensions, this pluralism is shallow, dominated by Future and Performance Orientation, with limited attention to Assertiveness or Gender Egalitarianism (
- **中文摘要**: 本文引入CCD-Bench（文化冲突决策基准），用于评估跨文化价值冲突下LLM的决策。包含七个领域的2,182个开放式困境，每个有十个匿名响应选项，与覆盖62个社会的十个GLOBE文化集群对齐。评估17个领先LLM发现明显偏差：模型偏爱北欧（20.2%）和德国欧洲（12.4%），而东欧和中东北非响应最不受偏好。87.9%的模型理由引用了多个文化维度，但这种多元主义是浅层的。

### 论文 53
- **英文标题**: LLM Safety in Judicial AI: A Stress Test of Social Media Influence on Real-World Judgments
- **中文标题**: 司法AI中的LLM安全性：社交媒体对真实判决影响的压力测试
- **作者**: Yixuan Xie, Yang He, Xiaoyu Yang, Xu Gai, Pan Hui
- **英文摘要**: Integrating Large Language Models (LLMs) into judicial decision-making demands rigorous safety examination against non-legal influences. This paper presents a novel stress test where we evaluate LLM-generated labor dispute outcomes by introducing social media sentiment as an external pressure, critically comparing them against 10,000 real-world court judgments from China Judgments Online (CJOL). Our findings reveal significant LLM safety vulnerabilities: models exhibit inherent deviations from real rulings, and public opinion substantially amplifies these discrepancies, leading to unstable and often inflated compensation predictions. Furthermore, these safety risks are compounded across low-skilled occupational categories and emotionally charged topics. This study uncovers critical threats to judicial integrity and public trust, underscoring the urgent need for robust safeguards against non-legal influences in AI legal systems.
- **中文摘要**: 本文将LLM集成到司法决策中需要针对非法律影响进行严格安全审查。通过引入社交媒体情绪作为外部压力评估LLM生成的劳动争议结果，与10,000个中国裁判文书网的真实判决进行批判比较。发现揭示了显著的LLM安全漏洞：模型展现与真实裁决的内在偏差，公众意见大幅放大这些差异，导致不稳定且常膨胀的赔偿预测。

### 论文 77
- **英文标题**: Multi-Armed Bandits Meet Large Language Models
- **中文标题**: 多臂赌博机遇见大语言模型
- **作者**: Djallel Bouneffouf, Raphael Feraud
- **英文摘要**: Bandit algorithms and Large Language Models (LLMs) have emerged as powerful tools in artificial intelligence, each addressing distinct yet complementary challenges in decision-making and natural language processing. This survey explores the synergistic potential between these two fields, highlighting how bandit algorithms can enhance the performance of LLMs and how LLMs, in turn, can provide novel insights for improving bandit-based decision-making. We first examine the role of bandit algorithms in optimizing LLM fine-tuning, prompt engineering, and adaptive response generation, focusing on their ability to balance exploration and exploitation in large-scale learning tasks. Subsequently, we explore how LLMs can augment bandit algorithms through advanced contextual understanding, dynamic adaptation, and improved policy selection using natural language reasoning. By providing a comprehensive review of existing research and identifying key challenges and opportunities, this survey aims to bridge the gap between bandit algorithms and LLMs, paving the way for innovative applications and interdisciplinary research in AI.
- **中文摘要**: 本综述探索了赌博机算法和LLM之间的协同潜力，涵盖赌博机算法在优化LLM微调、提示工程和自适应响应生成中的作用，以及LLM如何通过先进上下文理解、动态适应和使用自然语言推理改进策略选择来增强赌博机算法。通过全面回顾现有研究并识别关键挑战和机遇，旨在弥合这两个领域之间的差距。

### 论文 81
- **英文标题**: Knowledge-Guided Machine Learning: A Paradigm Shift in AI for Science
- **中文标题**: 知识引导的机器学习：AI for Science的范式转变
- **作者**: Anuj Karpatne, Xiaowei Jia, Vipin Kumar
- **英文摘要**: As advances in artificial intelligence (AI) and machine learning (ML) continue to transform commercial applications, the scientific community is increasingly eager to harness AI/ML’s power to accelerate modeling and discovery. However, purely data-driven AI methods often lack interpretability, generalizability, and consistency with established scientific principles. Conversely, traditional process-based models embody deep scientific knowledge but suffer from limited scalability or incomplete representation of complex systems. Knowledge-guided machine learning (KGML) offers a promising path forward by integrating scientific knowledge with data-driven approaches to produce AI models that are robust, trustworthy, and capable of advancing both AI and science. This talk summarizes the foundations of KGML, outlines a taxonomy for organizing research efforts, and highlights emerging opportunities for broad scientific impact.
- **中文摘要**: 纯数据驱动的AI方法常缺乏可解释性、泛化能力和与既定科学原则的一致性，传统基于过程的模型体现深厚的科学知识但可扩展性有限。知识引导的机器学习（KGML）通过将科学知识与数据驱动方法整合提供了有前景的前进道路。本文总结了KGML的基础，概述了组织研究工作的分类学，并突出了对广泛科学影响的新兴机遇。

## AAAI Emerging Trends in AI

### 论文 16
- **英文标题**: Model AI Assignments 2026
- **中文标题**: 2026年Model AI作业
- **作者**: Todd W. Neller, Steve Geinitz, Kevin Wang, Zach Dodds, Nicholas Dodds, Ryan O Connor, Aimen Taha, Ananta Manoranjan, Saurabh Ray, Deepak Ajwani, Fang Sun, Paul Zhang, Pranav Subbaraman, Yizhou Sun, Lisa Dunlap, Taehan Kim, Deena Sun, Ishir Garg, Mark Ogata, Aakarsh Vermani, Narges Norouzi, Joseph Gonzalez, Varada Kolhatkar
- **英文摘要**: The Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community.  Recognizing that assignments form the core of student learning experience, we here present abstracts of eight AI assignments from the 2026 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs.  Assignment specifications and supporting resources may be found at \url{http://modelai.gettysburg.edu}.
- **中文摘要**: Model AI作业会议旨在收集和传播人工智能教育社区的最佳作业设计。认识到作业构成学生学习体验的核心，我们在此呈现2026年会议的八项AI作业摘要，这些作业易于采用、有趣且灵活，可满足各种教师需求。作业规范和辅助资源可在http://modelai.gettysburg.edu找到。

### 论文 22
- **英文标题**: Multi-Objective Search: Algorithms, Applications, and Emerging Directions
- **中文标题**: 多目标搜索：算法、应用与新兴方向
- **作者**: Oren Salzman, Carlos Hernández Ulloa, Ariel Felner, Sven Koenig
- **英文摘要**: Multi-objective search (MOS) has emerged as a unifying framework for planning and decision-making problems where multiple, often conflicting, criteria must be balanced. While the problem has been studied for decades, recent years have seen renewed interest in the topic across AI applications such as robotics, transportation, and operations research, eflecting the reality that real-world systems rarely optimize a single measure. This paper surveys developments in MOS while highlighting cross-disciplinary opportunities, and outlines open challenges that define the emerging frontier of MOS research.
- **中文摘要**: 多目标搜索已成为一个统一框架，用于规划和决策问题中需要平衡多个通常相互冲突的标准的场景。虽然该问题已被研究数十年，近年来在机器人学、交通和运筹学等AI应用中重新引起了对该主题的兴趣，反映了现实世界系统很少只优化单一度量的事实。本文综述了多目标搜索的发展，同时突出了跨学科机会，并概述了定义多目标搜索研究新兴前沿的开放挑战。

### 论文 90
- **英文标题**: Multimodal Coarse-to-Local Transformer for End-to-End Autonomous Driving (Student Abstract)
- **中文标题**: 面向端到端自动驾驶的多模态粗到细Transformer
- **作者**: Yeryeong Cho, Joongheon Kim
- **英文摘要**: End-to-end (E2E) autonomous driving must maintain global consistency while preserving local precision. However, existing E2E approaches rarely achieve both goals simultaneously. Therefore, we propose a multimodal coarse-to-local transformer (MC2L-Transformer), which is composed of a hierarchical transformer architecture. Multimodal inputs are fused into a shared embedding, and global waypoints are produced. Local refinement is then utilized to capture fine interactions around the vehicle. Furthermore, a temporal encoder summarizes recent context, and navigation target and velocity are embedded to guide route- and speed-aware decoding. We evaluate in CARLA, and the results show lower collision and off-route rates even under sudden events. These results indicate that combining a coarse-to-local hierarchical transformer with a lightweight temporal context provides a practical step toward reliable E2E autonomous driving.
- **中文摘要**: 端到端自动驾驶必须在保持全局一致性的同时保持局部精度。然而，现有的端到端方法很少同时实现这两个目标。因此，我们提出多模态粗到细Transformer，由分层Transformer架构组成。多模态输入融合到共享嵌入中，产生全局航点。然后使用局部细化来捕获车辆周围的精细交互。此外，时序编码器总结近期上下文，导航目标和速度被嵌入以引导路径和速度感知解码。我们在CARLA中评估，结果显示即使在突发事件下也具有更低的碰撞率和偏离路径率。这些结果表明，将粗到细分层Transformer与轻量级时序上下文相结合，为可靠端到端自动驾驶提供了一步实用的进展。

### 论文 106
- **英文标题**: Can Large Language Models Grasp 3D Medical Anatomy Shapes? (Student Abstract)
- **中文标题**: 大语言模型能理解3D医学解剖形状吗？
- **作者**: Yao Gao, Feng Li, Jeroen Van Dessel, Yi Sun, Robin Willaert
- **英文摘要**: What if the next generation of human-computer interaction is not a screen... but a conversation? Large Language Models (LLMs) offer a new paradigm for interacting with computers through text, but they lack shape reasoning capabilities. We introduce Textual Anatomy Encoding (TAE), a workflow that connects LLMs with 3D anatomies. TAE employs clinician-validated semantic annotations and rule-based prompts to achieve deterministic and interpretable landmark localization. The results indicate that TAE enables LLMs to move beyond textual knowledge, achieving an accurate understanding of anatomical localization. This framework opens opportunities for diagnosis, surgical planning, and scalable medical annotation, positioning LLMs as a foundation for next-generation human–computer interaction in healthcare.
- **中文摘要**: 如果下一代人机交互不是屏幕...而是对话呢？大语言模型通过文本提供了一种与计算机交互的新范式，但它们缺乏形状推理能力。我们引入文本解剖编码，这是一个将大语言模型与3D解剖结构连接的工作流程。TAE采用临床医生验证的语义标注和基于规则的提示来实现确定性和可解释的标志点定位。结果表明，TAE使大语言模型能够超越文本知识，实现对解剖定位的准确理解。该框架为诊断、手术规划和可扩展的医学标注开辟了机会，将大语言模型定位为医疗中下一代人机交互的基础。

### 论文 114
- **英文标题**: Self-Guided Planning and Repair Framework for Code Generation (Student Abstract)
- **中文标题**: 面向代码生成的自引导规划与修复框架
- **作者**: Chun-Wei Kang, Chung-Chi Chen, An-Zi Yen
- **英文摘要**: Large Language Models (LLMs) demonstrate strong capabilities in code generation but often lack adaptability in planning and refinement. We propose Self-PR, a framework that integrates adaptive plan selection and iterative repair to improve correctness and generalization. Self-PR constructs a reusable plan database via task clustering and trains a selector to choose task-specific strategies. Incorrect outputs are refined through multi-round feedback until correctness. Trained only on HumanEval, Self-PR generalizes well to out-of-distribution tasks (MBPP), improving pass@1 by +4.9% on HumanEval and +5.5% on MBPP compared to Modularization-of-Thought prompting. Experiments across Llama-3 (8B, 70B) and GPT-4o-mini confirm robustness and scalability. These findings suggest that adaptive planning and feedback-driven repair are essential for reliable LLM-based code generation.
- **中文摘要**: 大语言模型在代码生成方面展示了强大能力，但在规划和细化方面通常缺乏适应性。我们提出Self-PR，一个整合自适应方案选择和迭代修复以提升正确性和泛化的框架。Self-PR通过任务聚类构建可复用的方案数据库，并训练选择器选择任务特定策略。错误输出通过多轮反馈进行细化直至正确。仅在HumanEval上训练的Self-PR对分布外任务（MBPP）泛化良好，相比思维模块化提示在HumanEval上pass@1提高了+4.9%，在MBPP上提高了+5.5%。跨Llama-3（8B，70B）和GPT-4o-mini的实验确认了鲁棒性和可扩展性。这些发现表明，自适应规划和反馈驱动的修复对于可靠的大语言模型代码生成至关重要。

### 论文 127
- **英文标题**: Obedience or Vigilance? How Large Language Models React to Malicious Multiple-Choice Options (Student Abstract)
- **中文标题**: 服从还是警惕？大语言模型如何对恶意多选题选项做出反应
- **作者**: Yow-Fu Liou, Yu-Chien Tang, An-Zi Yen
- **英文摘要**: When evaluating large language models (LLMs) for question answering tasks, a common protocol is multiple-choice question-answering (MCQA), where the model selects from a fixed set of choices. In contemporary robustness testing, researchers typically perturb instructions or introduce confusion into factual statements; however, model behavior also hinges on choice compliance: whether models remain within the canonical set A-D. We formalize this setting by asking whether the model continues to respect the interface's rules when the problem presents a tempting alternative. Our approach is interface-preserving: we append a single selectable option E while keeping the question and A-D unchanged. Then, we introduce three types of malicious option injection to assess LLMs' robustness. Experimental results highlight the vulnerability of LLMs on contradict type content of the additional option E. Our evaluation framework can effectively serve as a low-cost audit of rule adherence on existing datasets and black-box models, surfaces off-policy items, and supports interpretable model comparison for deployment.
- **中文摘要**: 在评估大语言模型的问答能力时，一个常见的协议是多选题问答，其中模型从固定选项集中选择。在当代鲁棒性测试中，研究者通常扰动指令或向事实陈述引入混淆；然而，模型行为也取决于选项遵从性：模型是否保持在规范集合A-D内。我们通过询问当问题呈现诱人替代选项时模型是否继续遵从界面规则来形式化这一设置。我们的方法是保留界面的：我们附加一个可选选项E同时保持问题和A-D不变。然后，我们引入三种类型的恶意选项注入来评估大语言模型的鲁棒性。实验结果表明，大语言模型在面对附加选项E的"矛盾"类型内容时表现出脆弱性。我们的评估框架可以有效作为现有数据集和黑盒模型上规则遵从性的低成本审计，揭示非策略项目，并支持可解释的模型比较以用于部署。

### 论文 139
- **英文标题**: Guided Latent Spaces for Controllable Multi-Scenario Generation in Autonomous Driving (Student Abstract)
- **中文标题**: 面向自动驾驶中可控多场景生成的引导潜在空间
- **作者**: Manasa Mariam Mammen, Zafer Kayatas, Stefan Wagner
- **英文摘要**: Scenario-based testing is an important approach for the development and validation of autonomous driving systems, as it enables evaluation across different driving situations. Safety-critical scenarios are especially relevant, but they occur rarely in real-world data, which creates the need for generation methods. In this paper, we present a scalable AI-based approach based on a variational autoencoder that unifies the generation of different types of critical scenarios while introducing controllability through a structured latent space. The integration of unified generation and latent space control advances AI-based scenario generation towards practical use, thereby supporting the requirements of industrial validation pipelines.
- **中文摘要**: 基于场景的测试是自动驾驶系统开发和验证的重要方法，因为它支持在不同驾驶情况下的评估。安全关键场景特别相关，但在真实世界数据中它们很少出现，这产生了生成方法的需求。在本文中，我们提出了一种基于变分自编码器的可扩展AI方法，统一了不同类型关键场景的生成，同时通过结构化潜在空间引入可控性。统一生成和潜在空间控制的结合将基于AI的场景生成推向实际用途，从而支持工业验证流程的要求。

### 论文 147
- **英文标题**: Cumulant Attention in Vision Transformers (Student Abstract)
- **中文标题**: 视觉Transformer中的累积量注意力
- **作者**: Yuto Morimoto, Zhipeng Wang, Koji Yasuda
- **英文摘要**: Transformer models have achieved remarkable success across diverse deep learning fields, including natural language processing (NLP) and computer vision (CV). One drawback of these models is that the computational cost of the softmax attention, the core component of the transformer, exhibits quadratic complexity in both time and memory. As data scales up various attempts have been reported to overcome this bottleneck. The objective of this study is to propose a novel attention mechanism, "Cumulant Attention", that systematically balances efficiency and accuracy. This proposal introduces a statistical-mechanics perspective and a reliable approximation based on cumulant expansion into the attention layer. The low-order variant reduces computational complexity to linear order, similar to the linear attention, while keeping nonlinearity of the softmax attention. We evaluate several variants on CV tasks, including image classification with ViT on ImageNet-100 and video classification with ViViT on UCF-101. Experimental results demonstrate that the cumulant attention outperforms the linear attention and achieves accuracy comparable to the softmax attention. These findings validate the effectiveness of our approach and highlight future directions, including scaling to larger models, extending to other modalities, and optimizing implementations for GPU hardware.
- **中文摘要**: Transformer模型在包括自然语言处理和计算机视觉在内的跨领域深度学习中取得了显著成功。这些模型的一个缺点是，作为Transformer核心组件的softmax注意力的计算成本在时间和内存上都呈现二次复杂度。随着数据规模的增大，已有各种尝试被报告来克服这一瓶颈。本研究的目标是提出一种新颖的注意力机制——"累积量注意力"，系统地平衡效率和准确率。该提案引入了统计力学视角和基于累积量展开的可靠近似到注意力层。低阶变体将计算复杂度降低到线性阶，类似于线性注意力，同时保留了softmax注意力的非线性。我们在多个计算机视觉任务上评估了几个变体，包括在ImageNet-100上使用ViT进行图像分类以及在UCF-101上使用ViViT进行视频分类。实验结果表明，累积量注意力优于线性注意力，并实现了与softmax注意力相当的准确率。这些发现验证了我们方法的有效性，并突出了未来方向，包括扩展到更大模型、扩展到其他模态以及优化GPU硬件的实现。

### 论文 183
- **英文标题**: Do Large Language Models (LLMs) Understand Chronology? (Student Abstract)
- **中文标题**: 大语言模型理解时间顺序吗？
- **作者**: Pattaraphon Kenny Wongchamcharoen, Paul Glasserman
- **英文摘要**: Large language models have shown great potential as forecasting tools in finance and economics, but backtesting performance is subject to look-ahead bias if the period overlaps with an LLM’s training window. Prompt-based attempts to avoid look-ahead bias require that LLMs understand chronology. We test LLMs’ ability to understand and enforce chronological order in three types of tasks: sorting randomly shuffled historical events; conditional sorting of events defined by some conditions; and anachronism detection based on intersections of multiple timelines. Our experiments use events that we first confirm are known to the LLM; this ensures that we test chronological understanding on an LLM’s pretrained internal knowledge. Across three LLM families— GPT-4.1 (standard), GPT-5 (hybrid-reasoning), and Claude 3.7 Sonnet (large-reasoning, with and without Extended Thinking), we find that performance degrades rapidly with problem complexity but improves greatly for reasoning models with test-time extended reasoning. These patterns are important for the real-time application of LLMs in finance.
- **中文摘要**: 大语言模型在金融和经济预测方面展现出巨大潜力，但当回测期与LLM训练窗口重叠时，回测性能会受到前视偏差的影响。基于提示的避免前视偏差的方法要求LLM理解时间顺序。我们通过三类任务测试LLM理解和执行时间顺序的能力：排序随机打乱的历史事件、基于某些条件对事件进行条件排序、以及基于多个时间线交集的年代错误检测。实验使用我们首先确认LLM已知的事件，以确保测试的是LLM在其预训练内部知识上的时间理解能力。在三个LLM系列（GPT-4.1标准版、GPT-5混合推理版和Claude 3.7 Sonnet大推理版，含/不含扩展思维）中，我们发现性能随问题复杂度快速下降，但对于具有测试时扩展推理的推理模型有显著改善。这些模式对LLM在金融领域的实时应用非常重要。

### 论文 186
- **英文标题**: Fine-tuning Zero-shot Large Language Models for Patient-reported Outcomes (Student Abstract)
- **中文标题**: 面向患者报告结局的零样本大语言模型微调
- **作者**: Yang Yan, Matthew W. Chen, Jiayi Lyu, Chen Zhao, Hao Gao, Zhong Chen
- **英文摘要**: Radiotherapy (RT) is a cornerstone of cancer treatment. Following RT, patient-reported outcomes (PROs) collected via standardized questionnaires are crucial for monitoring patients' quality of life and side effects. However, traditional statistical and machine learning methods, which rely on structured numerical data, often fail to capture semantic meaning within patients' health status. To address this, we developed a novel framework using zero- and few-shot large language models (LLMs) to identify patients experiencing mild to severe depression. Furthermore, classification performance is enhanced through parameter-efficient fine-tuning. Experiments on a prostate cancer PRO dataset for depression have demonstrated that our fine-tuned LLMs consistently outperformed other baseline methods across key evaluation metrics.
- **中文摘要**: 放射治疗是癌症治疗的基石。治疗后，通过标准化问卷收集的患者报告结局（PRO）对于监测患者生活质量和副作用至关重要。然而，依赖结构化数值数据的传统统计和机器学习方法往往无法捕获患者健康状况中的语义含义。为此，我们开发了一个使用零样本和少样本大语言模型（LLM）识别轻度至重度抑郁症患者的新框架。此外，通过参数高效微调增强了分类性能。在前列腺癌PRO数据集上的抑郁症实验表明，我们微调的LLM在所有关键评估指标上持续优于其他基线方法。

### 论文 205
- **英文标题**: Unveiling AI Safety in Fine-tuning Quantized Model
- **中文标题**: 揭示量化模型微调中的AI安全问题
- **作者**: Hai Le
- **英文摘要**: Post-training quantization is widely used to compress large language models (LLMs) for efficient deployment in resource-constrained environments. However, recent work shows that quantization, especially aggressive schemes such as 4-bit QLoRA, can substantially degrade safety alignment, making models more vulnerable to harmful completions and jailbreaks. In this work, we investigate these safety risks and propose a mitigation strategy: projecting quantized parameters back into safety-aligned subspaces. First, we empirically measure safety degradation on benchmark datasets using both safety and utility metrics. Next, we explore projection-based restoration methods to recover alignment-preserving directions in the LoRA adapters of quantized models. Finally, we study how quantization affects mechanistic safety neurons and how hybrid-precision designs can preserve them. By foregrounding the safety implications of model compression, this work aims to support more robust, deployment-ready, and ethically aligned LLMs.
- **中文摘要**: 训练后量化被广泛用于压缩大语言模型（LLM），以便在资源受限环境中高效部署。然而，近期工作表明量化，特别是4位QLoRA等激进方案，可能会严重降低安全对齐，使模型更容易受到有害补全和越狱攻击。本文研究这些安全风险并提出缓解策略：将量化参数投影回安全对齐子空间。首先，我们在基准数据集上使用安全和效用指标实证衡量安全降级。接下来，我们探索基于投影的恢复方法，以在量化模型的LoRA适配器中恢复对齐保持方向。最后，我们研究量化如何影响机制性安全神经元以及混合精度设计如何保护它们。通过突出模型压缩的安全影响，本工作旨在支持更鲁棒、可部署且伦理对齐的LLM。

### 论文 206
- **英文标题**: When AI Meets AI: A Game-Theoretic Defense Framework Against AI Empowered Cyber Threats
- **中文标题**: 当AI遇到AI：面向AI赋能网络威胁的博弈论防御框架
- **作者**: Xinyu Li
- **英文摘要**: The widespread adoption of artificial intelligence (AI) in cybersecurity has led to the emerging threat of AI-driven cyberattacks, such as LLM-empowered Advanced Persistent Threats (APTs), challenging the effect of conventional deception defense mechanisms. To fill this critical gap, my work aims to develop a game-theoretic defense AI agent capable of providing the optimal deception resource deployment strategy, to establish AI-driven defenses against AI-empowered cyberattacks. In this proposal, I model the attacker and defender interaction as a dynamic game with incomplete information between AI agents, and then derive the equilibrium defense strategies. Synthetic data based experiments and real-world implementations would be conducted to validate the proposed framework. This study has the potential to improve the effectiveness of deception defense in three dimensions: scalability, real-time capability, and strategic intelligence.
- **中文摘要**: 人工智能在网络安全中的广泛应用导致了AI驱动网络攻击的新兴威胁，如LLM赋能的进阶持续性威胁（APT），挑战了传统欺骗防御机制的效果。为填补这一关键空白，我的工作旨在开发一个博弈论防御AI智能体，能够提供最优欺骗资源部署策略，建立针对AI赋能网络攻击的AI驱动防御。在本提案中，我将攻击者和防御者之间的交互建模为AI智能体之间不完全信息动态博弈，并推导均衡防御策略。基于合成数据的实验和真实世界部署将用于验证所提框架。本研究有望在三个维度上提高欺骗防御的有效性：可扩展性、实时能力和战略智能。

### 论文 220
- **英文标题**: Algorithms for Context Engineering in LLM Inference: Optimization of Placement, Compression, and Scheduling
- **中文标题**: LLM推理中的上下文工程算法：放置、压缩和调度的优化
- **作者**: Teresa Zhang
- **英文摘要**: Scaling long-context and agentic LLMs is increasingly limited by memory capacity and bandwidth rather than FLOPs. I propose an algorithmic framework for context engineering that models placement, compression, and scheduling as coupled optimization problems with explicit accuracy-efficiency trade-offs. Concretely, I aim to develop (1) salience-aware retention/eviction policies with provable approximation guarantees relative to an ideal oracle; (2) tier-dependent compression schemes that bound error propagation across memory levels; and (3) probabilistic prefetch/scheduling that controls tail latency. I will evaluate on long-context language modeling and reasoning benchmarks, isolating each component via ablations and comparing against heuristic baselines under controlled bandwidth/capacity regimes. Results target improved throughput and energy metrics at near-baseline quality, advancing principled, hardware-aware inference without requiring custom hardware.
- **中文摘要**: 扩展长上下文和智能体LLM越来越受内存容量和带宽而非FLOPs的限制。我提出一个用于上下文工程的算法框架，将放置、压缩和调度建模为具有显式准确性-效率权衡的耦合优化问题。具体而言，我旨在开发：（1）具有相对于理想预言机可证明近似保证的显著性感知保留/淘汰策略；（2）限制跨内存级别错误传播的分层压缩方案；（3）控制尾部延迟的概率预取/调度。我将在长上下文语言建模和推理基准上进行评估，通过消融实验隔离每个组件，并在受控带宽/容量条件下与启发式基线比较。结果目标是在接近基线质量的情况下提升吞吐量和能效指标，推进无需定制硬件的原则性硬件感知推理。

### 论文 254
- **英文标题**: PANSim: Visualization Tool for Planning and Acting against Nature
- **中文标题**: PANSim：面向自然的规划与行动可视化工具
- **作者**: Erol Medenčević, Jakub Med, Lukáš Chrpa
- **英文摘要**: The demo presents a tool that visualizes the acting of planning agents in dynamic environments that might be modified by "acts of nature'', The purpose of this tool is to better understand the behavior of the agent, debug agent's behavior, and for making the underlying planning concepts accessible to wider audience.
- **中文摘要**: 本演示展示了一个可视化规划智能体在可能被"自然行为"修改的动态环境中行动的工具。该工具旨在更好地理解智能体行为、调试智能体行为，并使底层规划概念对更广泛的受众可及。

### 论文 261
- **英文标题**: KnowThyself: An Agentic Assistant for LLM Interpretability
- **中文标题**: KnowThyself：面向LLM可解释性的智能体助手
- **作者**: Suraj Prasai, Mengnan Du, Ying Zhang, Fan Yang
- **英文摘要**: We develop KnowThyself, an agentic assistant that advances large language model (LLM) interpretability. Existing tools provide useful insights but remain fragmented and code-intensive. KnowThyself consolidates these capabilities into a chat-based interface, where users can upload models, pose natural language questions, and obtain interactive visualizations with guided explanations. At its core, an orchestrator LLM first reformulates user queries, an agent router further directs them to specialized modules, and the outputs are finally contextualized into coherent explanations. This design lowers technical barriers and provides an extensible platform for LLM inspection. By embedding the whole process into a conversational workflow, KnowThyself offers a robust foundation for accessible LLM interpretability.
- **中文摘要**: 我们开发KnowThyself，一个推进大语言模型（LLM）可解释性的智能体助手。现有工具提供有用的洞察，但仍然碎片化且代码密集。KnowThyself将这些能力整合到基于聊天的界面中，用户可以上传模型、提出自然语言问题，并获得带有引导解释的交互式可视化。其核心是一个编排器LLM，首先重新表述用户查询，智能体路由器进一步将其导向专门模块，输出最终被上下文化为连贯解释。这种设计降低了技术门槛，为LLM检查提供了可扩展的平台。通过将整个过程嵌入到对话式工作流中，KnowThyself为可及的LLM可解释性提供了坚实基座。

### 论文 264
- **英文标题**: Steve: Your Personal AI Career Coach
- **中文标题**: Steve：你的个人AI职业教练
- **作者**: Balaji Rao, Naveen Mathews Renji, Elena Korshakova, Carlo Lipizzi
- **英文摘要**: Steve is an AI career coaching platform that turns a resume and insights from an AI-enabled chat with a user into a personalized skill gap report and upskilling roadmap. The platform suggests a personalized course plan, supports continuous learning, and helps shape the user’s career trajectory. Steve is built around schema-constrained JSON artifacts and a configurable career-tree ontology. The system compares confirmed skills against role-specific requirements, prioritizes gaps (critical/important/beneficial), and translates the analysis into embedding-based queries over the course index. Steve has three personas (Interview Coach, Resume Evaluator, and Career Coach) that provide concise feedback tailored to the user’s goals and context. Steve also supports speech Input/Output (I/O) via Whisper-based speech-to-text and a dual-voice text-to-speech layer, enabling users to talk to Steve. The platform offers flexible adaptability across institutions, enabling them to configure deployments by substituting their own ontologies and course catalogs. Our demo uses STEM trajectories as a case study, but the pipeline is domain-agnostic by design. Users can edit inputs, check speech recognition accuracy, and observe consistent updates, illustrating a reproducible, human-in-the-loop pattern for deploying LLMs in career guidance. Steve is currently in its alpha stage and available for demonstration.
- **中文摘要**: Steve是一个AI职业辅导平台，将简历和来自AI赋能聊天的用户洞察转化为个性化技能差距报告和技能提升路线图。平台建议个性化课程计划，支持持续学习，并帮助塑造用户的职业轨迹。Steve基于受模式约束的JSON工件和可配置职业树本体构建。系统将已确认的技能与特定角色的要求进行比较，对差距（关键/重要/有益）进行优先级排序，并将分析转化为基于嵌入的课程索引查询。Steve拥有三个角色（面试教练、简历评估师和职业教练），提供针对用户目标和背景的简洁反馈。Steve还通过基于Whisper的语音转文本和双语音文本转语音层支持语音输入/输出。平台提供跨机构的灵活适配性，使机构能够通过替换自己的本体和课程目录来配置部署。

### 论文 266
- **英文标题**: Automated Multi-Camera Inspection System for Aircraft
- **中文标题**: 面向飞机的自动化多相机检测系统
- **作者**: Mark David Rice, Gu Ying, Kelvin Wei Lim, Qing Yu Hoo, Liyuan Li, Lee Jue Ying, Jacky Jie Wei Tan, Lai Xing Ng, Jamie Ng
- **英文摘要**: In this paper, we present the development of an automated visual inspection system designed to detect defects on the upper surface of an aircraft airframe. Specifically, the system employs a multi-camera PTZ (Pan-Tilt-Zoom) set-up to capture and process images from designated regions. Custom developed software manages path planning and camera localization, while a hybrid-AI framework is integrated to identify various defect types, including missing and damaged components. The demonstration highlights the system’s detection capabilities and prototype functionalities using a large aircraft model, supported by a user interface to monitor progress and visualize results. To help validate this work, performance evaluations were conducted using selected multimodal and object detection models.
- **中文摘要**: 本文展示了一个用于检测飞机机身上表面缺陷的自动化视觉检测系统的开发。具体而言，系统采用多相机PTZ（云台变焦）设置来捕获和处理指定区域的图像。定制开发的软件管理路径规划和相机定位，同时集成混合AI框架来识别各类缺陷类型，包括缺失和损坏的组件。演示使用大型飞机模型突出系统的检测能力和原型功能，并通过用户界面监控进度和可视化结果。为帮助验证本工作，使用选定的多模态和目标检测模型进行了性能评估。

### 论文 276
- **英文标题**: ToolSmith: A Multi-Agent Framework for Enterprise Tool Creation
- **中文标题**: ToolSmith：面向企业工具创建的多智能体框架
- **作者**: Purna Chandra Sekhar Vakudavathu, Kushal Mukherjee, Jayachandu Bandlamudi, Renuka Sindhgatta, Sameep Mehta
- **英文摘要**: Although LLMs can generate tools for generic domains and tasks, they struggle with enterprise-related domains that involve proprietary APIs and data schemas. We present ToolSmith, a framework for autonomously generating and validating agent-compatible tools. Given an API specification and a Tool Specification Requirement (TSR), ToolSmith produces a tool function and verifies it through a closed-loop process: it creates natural language (NL) tests and executes the tool in a secure agent sandbox for validation. For state-changing tools, ToolSmith confirms outcomes by querying the API with parameters derived from the NL tests. If the tool fails to produce the desired output, ToolSmith generates diagnostic feedback to iteratively regenerate it. By ensuring both functional correctness and agent compatibility, ToolSmith enables reliable automation of enterprise workflows.
- **中文摘要**: 虽然LLM可以为通用领域和任务生成工具，但它们在涉及专有API和数据模式的企业相关领域中表现不佳。我们推出ToolSmith，一个自主生成和验证智能体兼容工具的框架。给定API规范和工具规范要求（TSR），ToolSmith生成工具函数并通过闭环过程进行验证：它创建自然语言（NL）测试，并在安全的智能体沙箱中执行工具进行验证。对于状态变更工具，ToolSmith通过使用从NL测试派生的参数查询API来确认结果。如果工具未能产生期望输出，ToolSmith会生成诊断反馈以迭代重新生成。通过确保功能正确性和智能体兼容性，ToolSmith实现了企业工作流的可靠自动化。

### 论文 279
- **英文标题**: AutoTuneX: Interactive Automated Fine-Tuning for Large Language Models
- **中文标题**: AutoTuneX：面向大语言模型的交互式自动微调
- **作者**: Daniel Karl I. Weidele, Priyanshu Rai, Frederico Araujo, Teryl Taylor, Radu Marinescu
- **英文摘要**: We present AutoTuneX, a system architecture design and implementation for users to interactively fine-tune large language models (LLMs) based on automated hyperparameter optimization particularly built around Bandit Limited Discrepancy Search. Next to a classical Graphical User Interface (GUI) our system features an agentic runtime to facilitate automated fine-tuning via chat.
- **中文摘要**: 我们推出AutoTuneX，一个系统架构设计和实现，允许用户基于自动化超参数优化（特别是围绕Bandit有限差异搜索构建的）交互式微调大语言模型（LLM）。除了经典的图形用户界面（GUI），我们的系统还包含智能体运行时，便于通过聊天实现自动化微调。

### 论文 284
- **英文标题**: PortfolioPilot: An Agentic Platform for Financial Portfolio Management Algorithm Development and Evaluation
- **中文标题**: PortfolioPilot：面向金融投资组合管理算法开发与评估的智能体平台
- **作者**: Jared Chan Xu Yang, Haokai Ma, Yunshan Ma
- **英文摘要**: Developing new portfolio-management algorithms typically demands substantial programming effort, limiting rapid experimentation and excluding finance professionals without coding skills. Current robo-advisory tools offer pre-built but rigid strategies, restricting customization and experimentation. We introduce PortfolioPilot, an open-source, agentic platform that enables users to generate bespoke portfolio through natural-language descriptions. Leveraging the Anthropic Claude API, PortfolioPilot dynamically synthesizes executable TypeScript algorithms that run in the frontend with security validation. The system integrates real-time backtesting with historical market data, classical optimization algorithms (Markowitz, LSTM, ARIMA), and interactive performance visualizations.
- **中文摘要**: 开发新的投资组合管理算法通常需要大量的编程工作，限制了快速实验并排除了没有编码技能的金融专业人士。当前的智能投顾工具提供预构建但僵化的策略，限制了定制化和实验。我们推出PortfolioPilot，一个开源的智能体平台，使用户能够通过自然语言描述生成定制投资组合。利用Anthropic Claude API，PortfolioPilot动态合成可执行的TypeScript算法，在前端进行安全验证后运行。系统集成使用历史市场数据的实时回测、经典优化算法（Markowitz、LSTM、ARIMA）和交互式性能可视化。

## Machine Learning VII

### 论文 3
- **英文标题**: Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
- **中文标题**: 面向高效视觉自回归建模的头部感知KV缓存压缩
- **作者**: Ziran Qin, Youru Lv, Mingbao Lin, Hang Guo, Zeren Zhang, Danping Zou, Weiyao Lin
- **英文摘要**: Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR models suffer from significant attention complexity and severe memory overhead due to the accumulation of key-value (KV) caches across scales. In this paper, we tackle this challenge by introducing KV cache compression into the next-scale generation paradigm. We begin with a crucial observation: attention heads in VAR models can be divided into two functionally distinct categories: Contextual Heads focus on maintaining semantic consistency, while Structural Heads are responsible for preserving spatial coherence. This structural divergence causes existing one-size-fits-all compression methods to perform poorly on VAR models. To address this, we propose HACK, a training-free Head-Aware KV cache Compression frameworK. HACK utilizes an offline classification scheme to separate head types, enabling it to apply pattern-specific compression strategies with asymmetric cache budgets for each category. By doing so, HACK effectively constrains the average KV cache length within a fixed budget B, reducing the theoretical attention complexity from O(n4) to O(Bn2). Extensive experiments on multiple VAR models across text-to-image and class-conditional tasks validate the effectiveness and generalizability of HACK. It achieves up to 70% KV cache compression without degrading output quality, resulting in memory savings and faster in- ference. For example, HACK provides a 1.75× memory reduction and a 1.57× speedup on Infinity-8B.
- **中文摘要**: 视觉自回归（VAR）模型采用下一尺度预测范式，以显著更少的解码步骤提供高质量内容生成。然而，现有VAR模型因跨尺度键值（KV）缓存的累积而面临显著的注意力复杂度和严重的内存开销。本文通过将KV缓存压缩引入下一尺度生成范式来应对这一挑战。我们从一个关键观察开始：VAR模型中的注意力头可被分为两个功能上不同的类别——上下文头专注于保持语义一致性，而结构头负责保持空间连贯性。这种结构差异导致现有的一刀切压缩方法在VAR模型上表现不佳。为解决此问题，我们提出HACK，一个免训练的头部感知KV缓存压缩框架。HACK利用离线分类方案分离头部类型，使其能够为每个类别应用具有非对称缓存预算的模式特定压缩策略。通过这样做，HACK有效地将平均KV缓存长度约束在固定预算B内，将理论注意力复杂度从O(n4)降低到O(Bn2)。在多个VAR模型上的广泛实验（涵盖文本到图像和类别条件任务）验证了HACK的有效性和泛化性。它在不降低输出质量的情况下实现了高达70%的KV缓存压缩，带来内存节省和更快的推理速度。例如，HACK在Infinity-8B上提供了1.75倍内存减少和1.57倍加速。

### 论文 18
- **英文标题**: LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
- **中文标题**: LexInstructEval：面向大语言模型的词汇指令遵循评估
- **作者**: Huimin Ren, Yan Liang, Baiqiao Su, Chaobo Sun, Hengtong Lu, Kaike Zhang, Chen Wei
- **英文摘要**: The ability of Large Language Models (LLMs) to precisely follow complex and fine-grained lexical instructions is a cornerstone of their utility and controllability. However, evaluating this capability remains a significant challenge. Current methods either rely on subjective and costly human evaluation or on automated ``LLM-as-a-judge'' systems, which suffer from inherent biases and unreliability. Existing programmatic benchmarks, while objective, often lack the expressiveness to test intricate, compositional constraints at a granular level. To address these limitations, we introduce LexInstructEval, a new benchmark and evaluation framework for fine-grained lexical instruction following. Our framework is built upon a formal, rule-based grammar that deconstructs complex instructions into a canonical (Procedure, Relation, Value) triplet. This grammar enables the systematic generation of a diverse dataset through a multi-stage, human-in-the-loop pipeline and facilitates objective verification via a transparent, programmatic engine. We release our dataset and open-source evaluation tools to facilitate further research into the controllability and reliability of LLMs.
- **中文摘要**: 大语言模型（LLM）精确遵循复杂细粒度词汇指令的能力是其效用和可控性的基石。然而，评估这一能力仍然是一个重大挑战。当前方法要么依赖主观且昂贵的人类评估，要么依赖自动的LLM-as-a-judge系统，后者存在固有的偏见和不可靠性。现有的程序化基准虽客观，但通常缺乏表达细粒度组合约束的表达力。为应对这些局限，我们引入LexInstructEval，一种新的用于细粒度词汇指令遵循的基准和评估框架。我们的框架建立在一个形式化的基于规则的语法之上，将复杂指令解构为规范的（过程，关系，值）三元组。该语法通过多阶段人机协同管道系统性地生成多样化数据集，并通过透明的程序化引擎实现客观验证。我们发布数据集和开源评估工具，以促进对LLM可控性和可靠性的进一步研究。

### 论文 47
- **英文标题**: Kronos: A Foundation Model for the Language of Financial Markets
- **中文标题**: Kronos：面向金融市场语言的基础模型
- **作者**: Yu Shi, Zongliang Fu, Shuo Chen, Bohan Zhao, Wei Xu, Changshui Zhang, Jian Li
- **英文摘要**: The success of large-scale pre-training paradigm, exemplified by Large Language Models (LLMs), has inspired the development of Time Series Foundation Models (TSFMs). However, their application to financial candlestick (K-line) data remains limited, often underperforming non-pre-trained architectures. Moreover, existing TSFMs often overlook crucial downstream tasks such as volatility prediction and synthetic data generation. To address these limitations, we propose Kronos, a unified, scalable pre-training framework tailored to financial K-line modeling. Kronos introduces a specialized tokenizer that discretizes continuous market information into token sequences, preserving both price dynamics and trade activity patterns. We pre-train Kronos using an autoregressive objective on a massive, multi-market corpus of over 12 billion K-line records from 45 global exchanges, enabling it to learn nuanced temporal and cross-asset representations. Kronos excels in a zero-shot setting across a diverse set of financial tasks. On benchmark datasets, Kronos boosts price series forecasting RankIC by 93% over the leading TSFM and 87% over the best non-pre-trained baseline. It also achieves a 9% lower MAE in volatility forecasting and a 22% improvement in generative fidelity for synthetic K-line sequences. These results establish Kronos as a robust, versatile foundation model for end-to-end financial time series analysis.
- **中文摘要**: 以大型语言模型（LLM）为代表的大规模预训练范式的成功激发了时间序列基础模型（TSFM）的发展。然而，它们在金融K线数据上的应用仍然有限，通常表现不如非预训练架构。此外，现有TSFM通常忽视了关键下游任务，如波动率预测和合成数据生成。为应对这些局限，我们提出Kronos，一个面向金融K线建模的统一可扩展预训练框架。Kronos引入了一个专门的tokenizer，将连续市场信息离散化为Token序列，保留价格动态和交易活动模式。我们在来自45个全球交易所的超过120亿条K线记录的大规模多市场语料上使用自回归目标预训练Kronos，使其能够学习细微的时序和跨资产表示。Kronos在各类金融任务的零样本设置中表现出色。在基准数据集上，Kronos将价格序列预测RankIC相比领先TSFM提升93%，相比最佳非预训练基线提升87%。它在波动率预测中还实现了9%更低的MAE，在合成K线序列的生成保真度方面提升22%。这些结果将Kronos确立为端到端金融时间序列分析的鲁棒通用基础模型。

### 论文 79
- **英文标题**: Complex Instruction Following with Diverse Style Policies in Football Games
- **中文标题**: 足球游戏中具有多样化风格策略的复杂指令遵循
- **作者**: Chenglu Sun, Shuo Shen, Haonan Hu, Wei Zhou, Chen Chen
- **英文摘要**: Despite advancements in language-controlled reinforcement learning (LC-RL) for basic domains and straightforward commands (e.g., object manipulation and navigation), effectively extending LC-RL to comprehend and execute high-level or abstract instructions in complex, multi-agent environments, such as football games, remains a significant challenge. To address this gap, we introduce Language-Controlled Diverse Style Policies (LCDSP), a novel LC-RL paradigm specifically designed for complex scenarios. LCDSP comprises two key components: a Diverse Style Training (DST) method and a Style Interpreter (SI). The DST method efficiently trains a single policy capable of exhibiting a wide range of diverse behaviors by modulating agent actions through style parameters (SP). The SI is designed to accurately and rapidly translate high-level language instructions into these corresponding SP. Through extensive experiments in a complex 5v5 football environment, we demonstrate that LCDSP effectively comprehends abstract tactical instructions and accurately executes the desired diverse behavioral styles, showcasing its potential for complex, real-world applications.
- **中文摘要**: 尽管语言控制强化学习（LC-RL）在基本领域和直接命令（如对象操作和导航）方面取得了进展，但将LC-RL有效扩展到理解和执行复杂多智能体环境（如足球游戏）中的高级或抽象指令仍然是一项重大挑战。为填补这一空白，我们引入语言控制多样化风格策略（LCDSP），一种专为复杂场景设计的新颖LC-RL范式。LCDSP包含两个关键组件：多样化风格训练（DST）方法和风格解释器（SI）。DST方法通过风格参数（SP）调制智能体动作，高效训练单个能够展现广泛多样化行为的策略。SI被设计为准确快速地将高级语言指令转化为对应的SP。通过在复杂的5v5足球环境中进行广泛实验，我们证明LCDSP有效理解抽象战术指令并准确执行所需的多样化行为风格，展示了其在复杂实际应用中的潜力。

### 论文 97
- **英文标题**: Trusted Multi-view Learning for Long-tailed Classification
- **中文标题**: 面向长尾分类的可信多视图学习
- **作者**: Chuanqing Tang, Yifei Shi, Guanghao Lin, Lei Xing, Long Shi
- **英文摘要**: Class imbalance has been extensively studied in single-view scenarios; however, addressing this challenge in multi-view contexts remains an open problem, with even scarcer research focusing on trustworthy solutions. In this paper, we tackle a particularly challenging class imbalance problem in multi-view scenarios: long-tailed classification. We propose TMLC, a Trusted Multi-view Long-tailed Classification framework, which makes contributions on two critical aspects: opinion aggregation and pseudo-data generation. Specifically, inspired by Social Identity Theory, we design a group consensus opinion aggregation mechanism that guides decision-making toward the direction favored by the majority of the group. In terms of pseudo-data generation, we introduce a novel distance metric to adapt SMOTE for multi-view scenarios and develop an uncertainty-guided data generation module that produces high-quality pseudo-data, effectively mitigating the adverse effects of class imbalance. Extensive experiments on long-tailed multi-view datasets demonstrate that our model is capable of achieving superior performance.
- **中文摘要**: 类别不平衡在单视图场景中已被广泛研究；然而，在多视图上下文环境解决这一挑战仍然是一个开放问题，专注于可信解决方案的研究更少。本文中，我们解决多视图场景中一个特别具有挑战性的类别不平衡问题：长尾分类。我们提出TMLC，一个可信多视图长尾分类框架，在两个关键方面做出贡献：意见聚合和伪数据生成。具体而言，受社会身份理论启发，我们设计了一种群体共识意见聚合机制，将决策引导向群体多数偏好的方向。在伪数据生成方面，我们引入了一种新颖的距离度量来使SMOTE适应多视图场景，并开发了一个不确定性引导的数据生成模块，生成高质量伪数据，有效缓解类别不平衡的不利影响。在长尾多视图数据集上的大量实验表明，我们的模型能够实现卓越性能。

## Cognitive Modeling and Cognitive Systems

### 论文 1
- **英文标题**: Hypothesis-Driven Reasoning for Large Language Models
- **中文标题**: 面向大语言模型的假设驱动推理
- **作者**: Aakash Kumar Agarwal, Moyuru Yamada
- **英文摘要**: This paper tackles the fundamental failure of Large Language Models (LLMs) to solve new tasks when prompted with a sufficient, yet overly complex, set of multi-modal episodes. This failure stems from the model's inability to distill underlying patterns from the noisy experiences. We propose Hypothesis-Driven Reasoning (HDR), a framework that enhances LLM reasoning by building an explicit semantic memory—a set of hypotheses induced from the multi-modal episodes. HDR employs a two-stage pipeline. It first extracts potential factors from the episodes and then iteratively refines hypotheses by generate-verify loop with the factors. We first empirically demonstrates this failure and the potential of sematic memory, showing that oracle hypotheses can boost accuracy from 35.3% to 92.0% on a novel task we designed. We then evaluate our HDR, achieving near-oracle performance and significantly outperforming baselines, especially on smaller models. This paper validates a shift from unstructured in-context recall to explicit knowledge abstraction for robust reasoning.
- **中文摘要**: 本文解决大语言模型(LLM)在给定充分但过于复杂的多模态事件集时无法解决新任务这一根本性失败。这一失败源于模型无法从噪声经验中提炼潜在模式。我们提出假设驱动推理(HDR)框架,通过构建显式语义记忆(从多模态事件中归纳出的一组假设)来增强LLM推理能力。HDR采用两阶段流水线:首先从事件中提取潜在因素,然后通过生成-验证循环迭代精炼假设。我们首先通过实验证明了这一失败及语义记忆的潜力,表明在我们设计的新任务上,oracle假设可将准确率从35.3%提升至92.0%。然后我们评估了HDR,实现了接近oracle的性能并显著优于基线方法,特别是在较小模型上。本文验证了从非结构化上下文回忆向显式知识抽象转变以实现鲁棒推理的范式转换。

### 论文 25
- **英文标题**: ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction
- **中文标题**: ARCHE: 评估LLM在潜在推理链提取上的新任务
- **作者**: Pengze Li, Jiaqi Liu, Junchi Yu, Lihao Liu, Mingyu Ding, Wanli Ouyang, Shixiang Tang, Xi Chen
- **英文摘要**: Large language models (LLMs) are increasingly used in scientific domains. While they can produce reasoning-like content via methods such as chain-of-thought prompting, these outputs are typically unstructured and informal, obscuring whether models truly understand the fundamental reasoning paradigms that underpin scientific inference. To address this, we introduce a novel task named Latent Reasoning Chain Extraction (ARCHE), in which models must decompose complex reasoning arguments into combinations of standard reasoning paradigms in the form of a Reasoning Logic Tree (RLT). In an RLT, all reasoning steps are explicitly categorized as one of three variants of Peirce’s fundamental inference modes: deduction, induction, or abduction. To facilitate this task, we release ARCHE Bench, a new benchmark derived from 70 Nature Communications articles, including more than 1,900 references and 38,000 viewpoints. We propose two logic-aware evaluation metrics: Entity Coverage (EC) for content completeness and Reasoning Edge Accuracy (REA) for step-by-step logical validity. Evaluations on 10 leading LLMs on ARCHE Bench reveal that models exhibit a trade-off between REA and EC, and none are yet able to extract a complete and standard reasoning chain. These findings highlight a substantial gap between the abilities of current reasoning models and the rigor required for scientific argumentation.
- **中文摘要**: 大语言模型(LLM)越来越多地应用于科学领域。虽然它们可以通过思维链提示等方法产生类似推理的内容,但这些输出通常是结构化程度低和非正式的,模糊了模型是否真正理解支撑科学推理的基本推理范式。为解决此问题,我们引入一项名为潜在推理链提取(ARCHE)的新任务,其中模型必须将复杂推理论证分解为标准推理范式的组合,以推理逻辑树(RLT)的形式呈现。在RLT中,所有推理步骤都明确归类为皮尔士基本推理模式的三种变体之一:演绎、归纳或溯因。为促进此任务,我们发布ARCHE Bench,一个来自70篇Nature Communications文章的新基准,包含超过1,900条参考文献和38,000个观点。我们提出两种逻辑感知的评估指标:用于内容完整性的实体覆盖率(EC)和用于逐步逻辑有效性的推理边准确率(REA)。在ARCHE Bench上对10个领先LLM的评估表明,模型在REA和EC之间表现出权衡,且尚无模型能够提取完整且标准的推理链。这些发现突显了当前推理模型的能力与科学论证所需的严谨性之间存在显著差距。

### 论文 32
- **英文标题**: Mind the Gap: The Divergence Between Human and LLM-Generated Tasks
- **中文标题**: 注意差距: 人类与LLM生成任务之间的分歧
- **作者**: Yi-Long Lu, Jiajun Song, Chunhui Zhang, Wei Wang
- **英文摘要**: Humans constantly generate a diverse range of tasks guided by internal motivations. While generative agents powered by large language models (LLMs) aim to simulate this complex behavior, it remains uncertain whether they operate on similar cognitive principles. To address this, we conducted a task-generation experiment comparing human responses with those of an LLM agent (GPT-4o). We find that human task generation is consistently influenced by psychological drivers, including personal values (e.g., Openness to Change) and cognitive style. Even when these psychological drivers are explicitly provided to the LLM, it fails to reflect the corresponding behavioral patterns. They produce tasks that are markedly less social, less physical, and thematically biased toward abstraction. Interestingly, while the LLM's tasks were perceived as more fun and novel, this highlights a disconnect between its linguistic proficiency and its capacity to generate human-like, embodied goals. We conclude that there is a core gap between the value-driven, embodied nature of human cognition and the statistical patterns of LLMs, highlighting the necessity of incorporating intrinsic motivation and physical grounding into the design of more human-aligned agents.
- **中文摘要**: 人类不断在内在动机的引导下生成多样化的任务。虽然以大语言模型(LLM)驱动的生成式智能体旨在模拟这种复杂行为,但它们是否基于相似的认知原则运作仍不确定。为解决此问题,我们进行了一项任务生成实验,比较人类回应与LLM智能体(GPT-4o)的回应。我们发现,人类任务生成持续受到心理驱动因素的影响,包括个人价值观(如对变化的开放性)和认知风格。即使将这些心理驱动因素明确提供给LLM,它仍无法反映相应的行为模式。它们产生的任务明显缺乏社交性和物理性,并在主题上偏向抽象化。有趣的是,虽然LLM的任务被认为更有趣和新颖,但这凸显了其语言熟练度与生成类人具身目标能力之间的脱节。我们得出结论,人类认知的价值驱动、具身本质与LLM的统计模式之间存在核心差距,突显了将内在动机和物理基础纳入更与人类对齐的智能体设计的必要性。

### 论文 35
- **英文标题**: Agentic Design Review System
- **中文标题**: 智能体设计评审系统
- **作者**: Sayan Nag, Joseph K J, Koustava Goswami, Vlad I Morariu, Balaji Vasan Srinivasan
- **英文摘要**: Evaluating a graphic design involves assessing it from multiple facets like alignment, composition, aesthetics and color choices. Holistic evaluation would involve aggregating feedback from individual expert reviewers. Towards this, we propose an Agentic Design Review System (Agentic-DRS), where multiple agents collaboratively analyze a design, orchestrated by a meta-agent. A novel in-context exemplar selection approach based on graph matching and a unique prompt expansion method plays central role towards making each agent design aware. In order to evaluate this framework, we propose DRS-BENCH. Thorough experimental evaluation against state-of-the-art baselines adapted to the problem setup, backed by critical ablations, demonstrates efficacy of Agentic-DRS in evaluating designs and generating actionable feedback.
- **中文摘要**: 评估图形设计涉及从多个方面进行评估,如对齐、构图、美学和色彩选择。全面评估需要聚合来自各个专家评审者的反馈。为此,我们提出智能体设计评审系统(Agentic-DRS),其中多个智能体协作分析设计,由元智能体协调。一种基于图匹配的新颖上下文示例选择方法和独特的提示扩展方法在使每个智能体具备设计感知能力方面发挥核心作用。为评估此框架,我们提出DRS-BENCH。针对问题设置的最先进基线的彻底实验评估,辅以关键消融实验,证明了Agentic-DRS在评估设计和生成可执行反馈方面的有效性。

### 论文 39
- **英文标题**: Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier
- **中文标题**: 通过情感基础验证器的多模态LLM情感一致推理
- **作者**: Hyeongseop Rha, Jeong Hun Yeo, Yeonju Kim, Yong Man Ro
- **英文摘要**: The recent advancement of Multimodal Large Language Models (MLLMs) is transforming human-computer interaction (HCI) from surface-level exchanges into more nuanced and emotionally intelligent communication. To realize this shift, emotion understanding becomes essential allowing systems to capture subtle cues underlying user intent. Furthermore, providing faithful explanations for predicted emotions is crucial to ensure interpretability and build user trust. However, current MLLM-based methods often generate emotion explanations that diverge from the ground-truth (GT) labels and sometimes even contradict their own predicted emotions. This inconsistency poses a critical risk for misunderstanding and erodes reliability in interactive settings. To address this, we propose a novel approach: the Emotional Rationale Verifier (ERV) and an Explanation Reward. Our method guides the model to produce reasoning that is explicitly consistent with the GT emotion during multimodal emotion recognition without modifying the model architecture or requiring paired video–description annotations. Our method significantly improves faithful explanation–prediction consistency and explanation emotion accuracy on the MAFW and DFEW datasets. Through extensive experiments and human evaluations, we show that our approach not only enhances alignment between explanation and prediction but also empowers MLLMs to deliver emotionally coherent, trustworthy interactions, marking a key step toward truly human-like HCI systems.
- **中文摘要**: 多模态大语言模型(MLLM)的最新进展正将人机交互(HCI)从表面层级交流转变为更细腻且情感智能的沟通。为实现这一转变,情感理解变得至关重要,使系统能够捕获用户意图背后的微妙线索。此外,为预测的情感提供可信的解释对于确保可解释性和建立用户信任至关重要。然而,当前基于MLLM的方法常常产生偏离真实标签的情感解释,有时甚至与自身预测的情感相矛盾。这种不一致性对交互环境中的误解构成关键风险,并侵蚀可靠性。为解决此问题,我们提出一种新方法:情感基础验证器(ERV)和解释奖励。我们的方法引导模型在多模态情感识别过程中产生与真实情感明确一致的推理,无需修改模型架构或需要配对视频-描述标注。我们的方法在MAFW和DFEW数据集上显著提升了可信解释-预测一致性及解释情感准确性。通过广泛实验和人类评估,我们展示了我们的方法不仅增强了理解与预测之间的对齐,还赋予MLLM提供情感连贯、可信交互的能力,标志着迈向真正类人HCI系统的关键一步。

### 论文 66
- **英文标题**: Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite Feedback
- **中文标题**: Psi-Arena: 基于三方反馈的LLM心理咨询师交互式评估与优化
- **作者**: Shijing Zhu, Zhuang Chen, Guanqun Bi, Binghang Li, Yaxi Deng, Dazhen Wan, Libiao Peng, Xiyao Xiao, Rongsheng Zhang, Tangjie Lv, Zhipeng Hu, FangFang Li, Minlie Huang
- **英文摘要**: Large language models (LLMs) have shown promise in providing scalable mental health support, while evaluating their counseling capability remains crucial to ensure both efficacy and safety. Existing evaluations are limited by the static assessment that focuses on knowledge tests, the single perspective that centers on user experience, and the open-loop framework that lacks actionable feedback. To address these issues, we propose Ψ-Arena, an interactive framework for comprehensive assessment and optimization of LLM-based counselors, featuring three key characteristics: (1) Realistic arena interactions that simulate real-world counseling through multi-stage dialogues with psychologically profiled NPC clients; (2) Tripartite evaluation that integrates assessments from the client, supervisor, and counselor perspectives; (3) Closed-loop optimization that iteratively improves LLM counselors using diagnostic feedback. Experiments across eight state-of-the-art LLMs show significant performance variations in different real-world scenarios and evaluation perspectives. Moreover, reflection-based optimization results in up to a 141% improvement in counseling performance. We hope Ψ-Arena provides a foundational resource for advancing reliable and human-aligned LLM applications in mental healthcare.
- **中文摘要**: 大语言模型(LLM)在提供可扩展的心理健康支持方面展现了前景,而评估其咨询能力对确保有效性和安全性仍然至关重要。现有评估的局限包括:关注知识测试的静态评估、以用户体验为中心的单一视角,以及缺乏可执行反馈的开环框架。为解决这些问题,我们提出Psi-Arena,一个用于全面评估和优化基于LLM的心理咨询师的交互式框架,具有三个关键特征:(1)真实的竞技场交互,通过与心理学画像NPC客户进行多阶段对话模拟真实世界咨询;(2)三方评估,整合来自客户、督导和咨询师视角的评估;(3)闭环优化,利用诊断反馈迭代改进LLM咨询师。在八个最先进LLM上的实验显示了在不同真实世界场景和评估视角中的显著性能差异。此外,基于反思的优化实现了高达141%的咨询性能提升。我们希望Psi-Arena为推进心理健康领域可靠且与人类对齐的LLM应用提供基础资源。

## Application Domains I

### 论文 21
- **英文标题**: Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving
- **中文标题**: 衡量关键要素：面向自动驾驶轨迹预测器的场景驱动评估
- **作者**: Longchao Da, David Isele, Hua Wei, Manish Saroya
- **英文摘要**: Being able to anticipate the motion of surrounding agents is essential for the safe operation of autonomous driving systems in dynamic situations. While various methods have been proposed for trajectory prediction, the current evaluation practices still rely on error-based metrics (e.g., ADE, FDE), which reveal the accuracy from a post-hoc view but ignore the actual effect the predictor brings to the self-driving vehicles (SDVs), especially in complex interactive scenarios: a high-quality predictor not only chases accuracy, but should also captures all possible directions a neighbor agent might move, to support the SDVs' cautious decision-making. Given that the existing metrics hardly account for this standard, in our work, we propose a comprehensive pipeline that adaptively evaluates the predictor's performance by two dimensions: accuracy and diversity. Based on the criticality of the driving scenario, these two dimensions are dynamically combined and result in a final score for the predictor's performance. Extensive experiments on a closed-loop benchmark using a real-world dataset show that our pipeline yields a more reasonable evaluation than traditional metrics by better reflecting the correlation of the predictors' evaluation with the autonomous vehicles' driving performance. This evaluation pipeline shows a robust way to select a predictor that potentially contributes most to the SDV's driving performance.
- **中文摘要**: 能够预判周围智能体的运动对自动驾驶系统在动态场景中的安全运行至关重要。虽然提出了多种轨迹预测方法，但当前的评估实践仍然依赖基于误差的指标（如ADE、FDE），这些指标从事后视角揭示准确性，却忽略了预测器对自动驾驶车辆（SDV）的实际影响，尤其是在复杂的交互场景中：一个高质量的预测器不仅追求准确性，还应捕捉邻居智能体可能移动的所有方向，以支持SDV的谨慎决策。鉴于现有指标几乎不考虑这一标准，在我们的工作中，我们提出了一个综合流水线，从两个维度自适应评估预测器性能：准确性和多样性。基于驾驶场景的关键性，这两个维度被动态组合，产生预测器性能的最终得分。在真实世界数据集上的闭环基准大量实验表明，我们的流水线通过更好地反映预测器评估与自动驾驶车辆驾驶性能之间的相关性，比传统指标产生更合理的评估。该评估流水线展示了一种选择可能对SDV驾驶性能贡献最大的预测器的稳健方法。

### 论文 57
- **英文标题**: FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
- **中文标题**: FinRpt：面向股权研究报告生成的数据集、评估系统及基于LLM的多智智能体框架
- **作者**: Song Jin, Shuqi Li, Shukun Zhang, Rui Yan
- **英文摘要**: While LLMs have shown great success in financial tasks like stock prediction and question answering, their application in fully automating Equity Research Report generation remains uncharted territory. In this paper, we formulate the Equity Research Report (ERR) Generation task for the first time. To address the data scarcity and the evaluation metrics absence, we present an open-source evaluation benchmark for ERR generation - FinRpt. We frame a Dataset Construction Pipeline that integrates 7 financial data types and produces a high-quality ERR dataset automatically, which could be used for model training and evaluation. We also introduce a comprehensive evaluation system including 11 metrics to assess the generated ERRs. Moreover, we propose a multi-agent framework specifically tailored to address this task, named FinRpt-Gen, and train several LLM-based agents on the proposed datasets using Supervised Fine-Tuning and Reinforcement Learning. Experimental results indicate the data quality and metrics effectiveness of the benchmark FinRpt and the strong performance of FinRpt-Gen, showcasing their potential to drive innovation in the ERR generation field. All code and datasets are publicly available.
- **中文摘要**: 尽管LLM在股票预测和问答等金融任务中展现了巨大成功，但其在完全自动化股权研究报告生成中的应用仍是一片未被探索的领域。在本文中，我们首次形式化了股权研究报告（ERR）生成任务。为解决数据稀缺和评估指标缺失的问题，我们提出了ERR生成的开放评估基准FinRpt。我们构建了一个数据集构建流水线，整合7种金融数据类型并自动生成高质量的ERR数据集，可用于模型训练和评估。我们还引入了一个包含11个指标的综合评估系统来评估生成的ERR。此外，我们提出了一个专门针对此任务定制的多智能体框架FinRpt-Gen，并使用监督微调和强化学习在所提出的数据集上训练多个基于LLM的智能体。实验结果表明了基准FinRpt的数据质量和指标有效性，以及FinRpt-Gen的强劲性能，展示了它们在推动ERR生成领域创新的潜力。所有代码和数据集均公开可用。

### 论文 63
- **英文标题**: RAG-Enhanced Collaborative LLM Agents for Drug Discovery
- **中文标题**: 面向药物发现的RAG增强协作LLM智能体
- **作者**: Namkyeong Lee, Edward De Brouwer, Ehsan Hajiramezanali, Tommaso Biancalani, Chanyoung Park, Gabriele Scalia
- **英文摘要**: Recent advances in large language models (LLMs) have shown great potential to accelerate drug discovery. However, the specialized nature of biochemical data often necessitates costly domain-specific fine-tuning, posing critical challenges. First, it hinders the application of more flexible general-purpose LLMs in cutting-edge drug discovery tasks. More importantly, it limits the rapid integration of the vast amounts of scientific data continuously generated through experiments and research. Compounding these challenges is the fact that real-world scientific questions are typically complex and open-ended, requiring reasoning beyond pattern matching or static knowledge retrieval. To address these challenges, we propose CLADD, a retrieval-augmented generation (RAG)-empowered agentic system tailored to drug discovery tasks. Through the collaboration of multiple LLM agents, CLADD dynamically retrieves information from biomedical knowledge bases, contextualizes query molecules, and integrates relevant evidence to generate responses - all without the need for domain-specific fine-tuning. Crucially, we tackle key obstacles in applying RAG workflows to biochemical data, including data heterogeneity, ambiguity, and multi-source integration. We demonstrate the flexibility and effectiveness of this framework across a variety of drug discovery tasks, showing that it outperforms general-purpose and domain-specific LLMs as well as traditional deep learning approaches.
- **中文摘要**: 大语言模型（LLM）的最新进展展现了加速药物发现的巨大潜力。然而，生化数据的专业性质通常需要昂贵的领域特定微调，构成关键挑战。首先，它阻碍了更灵活的通用LLM在前沿药物发现任务中的应用。更重要的是，它限制了通过实验和研究持续产生的大量科学数据的快速整合。雪上加霜的是，现实世界的科学问题通常复杂且开放，需要超越模式匹配或静态知识检索的推理。为应对这些挑战，我们提出了CLADD，一个面向药物发现任务的检索增强生成（RAG）赋能的智能体系统。通过多个LLM智能体的协作，CLADD动态地从生物医学知识库中检索信息，对查询分子进行上下文化，并整合相关证据以生成响应——所有这些都无需领域特定微调。关键的是，我们应对了将RAG工作流应用于生化数据的关键障碍，包括数据异构性、模糊性和多源整合。我们展示了该框架在多样化药物发现任务中的灵活性和有效性，表明其优于通用和领域特定LLM以及传统深度学习方法。

