When Adaptation Hurts: Connecting Representational Drift to OOD Failures in MedS

📄 When Adaptation Hurts: Connecting Representational Drift to OOD Failures in MedSAM Fine-Tuning

Foundation models for medical image segmentation, like prompt-based MedSAM, generalize well across domains and modalities, often in zero or few-shot setups. However, their performance depends on the quality of prompts and the adaptation of the models to custom datasets. This work systematically exam

📄 View on arXiv 📥 PDF

Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Unde

📄 Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Understanding

Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supervised methods. However, deploying these models in real-world scenarios is severely hindered by their inability to efficiently handle streaming RGB-D

📄 View on arXiv 📥 PDF

Specification Portability Across LLM Development Agents: Cross-Agent Compatibili

📄 Specification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software Migration

This paper investigates cross-agent specification portability using Oracle-to-PostgreSQL migration as a controlled software transformation task. The study combines two experimental stages. First, a specification-first migration pipeline was evaluated on 1,006 PL/SQL files, of which 623 were successf

📄 View on arXiv 📥 PDF

Toward Vision Language Model-based Assessment of Clinical Quality and Usability

📄 Toward Vision Language Model-based Assessment of Clinical Quality and Usability of LGE-MR Images for Cardiac Ablation Planning

LGE cardiac MRI is widely used for left atrial fibrosis assessment and ablation planning in atrial fibrillation patients as knowledge of fibrotic tissue regions identified from LGE-MRI is critical for catheter ablation. Often, poor quality images used during ablation planning can cause mis-localizat

📄 View on arXiv 📥 PDF

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery

📄 Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. Reference-based metrics such as BLEU and ROUGE fail to capture semantic correctness, while expert h

📄 View on arXiv 📥 PDF

Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models f

📄 Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI

Federated fine-tuning enables large language models to adapt on edge devices without centralizing private data, but practical deployments must address hardware instability and adversarial update corruption together. Thermally constrained clients may throttle, slow local training, or delay synchronou

📄 View on arXiv 📥 PDF

Personalized Privacy Control in LLMs via Attention Head Intervention

The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms. However, acceptable disclosure boundaries may vary across users even within

📄 View on arXiv 📥 PDF

VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for

📄 VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation

We propose VT-MUSE, a Multimodal Unified SEquential representation learning framework for visuotactilemanipulation. Existing approaches often encode visual and tactile observations independently before fusion, limiting their ability to capture fine-grained cross-modal dependencies. Moreover, most me

📄 View on arXiv 📥 PDF

Consistency Models for Fast MRI Reconstruction Using Regularization by Denoising

Diffusion models (DMs) have emerged as powerful generative priors for MRI reconstruction with promising results. Yet DM-based methods require extensive iterative refinement, limiting their practical deployment. Consistency models (CMs) provide a compelling alternative, aiming to map out the diffusio

📄 View on arXiv 📥 PDF

Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored. This setting requires models to jointly understand spatial layouts, interpret

📄 View on arXiv 📥 PDF

HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Work

📄 HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads

Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communication patterns complicate efficient hardware acceleration. Chiplet-based architectures offer a scalable solution by integrating specialized compute and memory units.

📄 View on arXiv 📥 PDF

Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering

Accurate and responsible medical question answering (QA) is important in healthcare, where complex cases require factual knowledge and nuanced reasoning. Existing medical QA systems, typically based on single-agent architectures and static retrieval, often lack adaptability, persistent memory, and s

📄 View on arXiv 📥 PDF

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accelerating and extending costly human evaluation. However, most work treats a judge as a static artifact, evaluating it onc

📄 View on arXiv 📥 PDF

Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering

Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question-related key fr

📄 View on arXiv 📥 PDF

TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missi

📄 TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs

Large video-language models (LVLMs) have shown remarkable performance on multimodal tasks like multimodal emotion recognition (ER) in the wild. ER is inherently multimodal, requiring a joint understanding of facial expressions, vocalizations, language, biosignals, and gestures. However, real-world d

📄 View on arXiv 📥 PDF

A Comprehensive Review of Large Language Models for Nanophotonics: From Surrogat

📄 A Comprehensive Review of Large Language Models for Nanophotonics: From Surrogate Modeling to Autonomous Design

Metasurfaces have revolutionized the development of photonic devices by enabling unprecedented precision in light manipulation. However, their design processes are often constrained by computationally expensive simulations and complex high-dimensional design spaces. Although deep learning has accele

📄 View on arXiv 📥 PDF

Do Large Language Models Play Six Degrees of Separation? Measuring Topological C

📄 Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds

Large Language Models (LLMs) demonstrate remarkable multi-hop reasoning capabilities over long contexts, yet the internal mechanisms enabling these distant cognitive leaps remain poorly understood. Traditional attention-based interpretability often fails to capture true semantic proximity due to rou

📄 View on arXiv 📥 PDF

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hiera

📄 HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it chal

📄 View on arXiv 📥 PDF

AISA: AI Safety Assistant Framework for Continuous Improvement of Highway Constr

📄 AISA: AI Safety Assistant Framework for Continuous Improvement of Highway Construction

Job Safety Analysis (JSA) and pre-task planning can benefit from prior incident records, yet historical accident data is often stored as unstructured narratives that are difficult to consult at the point of planning. A novel framework centered on large language models (LLMs) for highway construction

📄 View on arXiv 📥 PDF

Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted

📄 Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents

Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge-poisoning attacks, in which misinformation in retrieved documents can influence the model's final outputs. Notably, an LLM may correctly d

📄 View on arXiv 📥 PDF