Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models
arXiv:2609.22161v1 Announce Type: new Abstract: Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types differentially shape model
An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users
arXiv:2609.22277v1 Announce Type: new Abstract: Visual impairment affects over 2.2 billion people worldwide, yet conventional white canes cannot detect elevated hazards or provide semantic environmental context. Existing AI-assisted navigat
PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation
arXiv:2609.22353v1 Announce Type: new Abstract: Mobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints alone cannot describe. On resource constrained platforms, converting imperfect visual pre
Social Influence and the Allocation of Scientific Attention in AI Populations
arXiv:2609.22408v1 Announce Type: new Abstract: AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human
Learning 3D biophysical cell properties from 2D images and cell-population statistics
arXiv:2609.22410v1 Announce Type: new Abstract: Inferring 3D cellular properties from 2D microscopy is difficult when a reference instrument reports only population statistics rather than labels for individual cells. Here we develop a popul
Goal-driven Variant Categorization
arXiv:2609.22475v1 Announce Type: new Abstract: Process discovery rarely yields a single coherent process structure. For analysis, a common step is to cluster process variants based on structural similarity and then assign business meaning
Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
arXiv:2609.22478v1 Announce Type: new Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whethe
The Wisdom of Artificial Deliberative Crowds
arXiv:2609.22497v1 Announce Type: new Abstract: The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually attributed to the independence of estimate
Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and
IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law
arXiv:2609.22529v1 Announce Type: new Abstract: International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its texts remain without token-level named ent
EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability
arXiv:2609.22537v1 Announce Type: new Abstract: Enterprise AI assistants must produce responses that are verifiable and traceable to source evidence. However, retrieval augmented generation (RAG) over heterogeneous enterprise data can suffe
AutoGym: Blueprint-First Generation of Verifiable Agent Gyms
arXiv:2609.22592v1 Announce Type: new Abstract: Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a verifier that reliably distinguishes success
MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
arXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated resp
GaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment
arXiv:2609.22619v1 Announce Type: new Abstract: Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation sessions, yet objective 3D measurement remains confined to specialized motion-capture labo
Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation
arXiv:2609.22620v1 Announce Type: new Abstract: Scanned mail, uploaded PDFs, and consolidated attachments often arrive as page streams that must be split into individual documents before downstream classification, extraction, or routing. Ze
Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA
arXiv:2609.22628v1 Announce Type: new Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the prompt, jud
Self-Organizing Agent Teams Learn to Reason Together
arXiv:2609.22682v1 Announce Type: new Abstract: Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor c
Generative Embodied Multiple Behavior Control Systems for Human-like Agents
arXiv:2609.22691v1 Announce Type: new Abstract: An enduring and richly elaborated dichotomy in cognitive neuroscience is that of human behavior control mechanisms, divided into habitual versus goal-directed. While existing human-like agent
Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review
arXiv:2609.22694v1 Announce Type: new Abstract: Public crash databases increasingly support automated safety analysis, but crash severity prediction remains difficult to translate into public-sector decision workflows when models are evalua
A Survey on the Linear Representation Hypothesis
arXiv:2609.22695v1 Announce Type: new Abstract: The term "linear representation hypothesis" (LRH) has appeared across diverse subfields of artificial intelligence, neuroscience, and cognitive science. But previous works have not consistentl
Building Trustworthy Mental Health Benchmarks on Bluesky: A Validation-Aware Weak-Supervision Framework
arXiv:2609.22696v1 Announce Type: new Abstract: Decentralized social media platforms create new opportunities and challenges for computational mental health research because data access, moderation, labeling, and deployment responsibilities
Hapi: A Multivariable Land-Surface Transformer for Medium-Range Hydrological Forecasting at Continental Scale
arXiv:2609.22702v1 Announce Type: new Abstract: Accurate flood forecasts several days in advance are essential for flood control, water-resource management, and emergency response. Producing them at high resolution over a continental domain
Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems
arXiv:2609.22712v1 Announce Type: new Abstract: Agentic AI systems built on large language models can plan over multiple steps, use external tools, retain information in memory, and coordinate with other agents. These capabilities make them
ProcessLight: Process Supervision for Large Language Model Based Traffic Signal Control
arXiv:2609.22746v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently been introduced into traffic signal control (TSC) as decision agents due to their strengths in human-readable reasoning generation. Yet, existing LLM
CTSpinoPelvic1K: spine, pelvis, ribs and femora in one coordinate frame, annotated for lumbosacral transitional anatomy
arXiv:2609.22760v1 Announce Type: new Abstract: Purpose: A vertebra at the lumbosacral junction is named by counting caudally from C2 on whole-spine imaging, but a lumbar case is planned on lumbar-only imaging (T12 to S1), without C2. Abdom
DVA-Neurons: Design and Verification of Adaptive LIF Neurons: From Single-Neuron Dynamics to Multi-Neuron Spiking Networks
arXiv:2609.22775v1 Announce Type: new Abstract: Spiking Neural Networks (SNNs) offer a promising path toward ultra-low-power artificial intelligence inference by emulating the event-driven computation of biological neurons. However, two cha
From Research Frontier to Laboratory Bench: Design of a Four-Tier Experimental Teaching System for Multimodal Medical Image Intelligent Diagnosis
arXiv:2609.22790v1 Announce Type: new Abstract: Undergraduate programmes in intelligent medical engineering are expanding, yet laboratory curricula lag behind the multimodal, long-tailed, and distributionally shifting realities of clinical
ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures
arXiv:2609.22878v1 Announce Type: new Abstract: Large language model code generation benchmarks primarily evaluate well-resourced languages like Python and Java, where models benefit from abundant training data. They provide limited evidenc
When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used
arXiv:2609.22910v1 Announce Type: new Abstract: Vision-language agents that crop and zoom are trained with rewards that credit a successful tool call, yet a successful call does not show that the model needed to look or used the pixels it r
Beyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language Model
arXiv:2609.22939v1 Announce Type: new Abstract: Long-context models read a novel the way a person reads a printout: one token after another, in narrative order, with the whole history competing for a fixed budget of attention. A detective d
AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows
arXiv:2609.22951v1 Announce Type: new Abstract: Enterprise agentic systems that route every trajectory step to a frontier model waste 60-80% of their inference budget on subtasks that smaller models handle equally well. Existing routing sol
A Compact Stance-Indexed Anterior-Posterior COP Representation for Parkinson's Disease Classification from Plantar VGRF
arXiv:2609.22956v1 Announce Type: new Abstract: Parkinson's disease alters gait and bilateral coordination, but machine-learning performance also depends on how continuous gait signals are represented. This study investigates whether preser
R-GEAN: Regimen-Guided Edit Action Network for Within-Admission Medication Change Prediction
arXiv:2609.22959v1 Announce Type: new Abstract: The medications prescribed to a patient often change during a hospital admission as clinicians start, stop, or continue therapies. We study whether models can predict which medication classes
OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling
arXiv:2609.22987v1 Announce Type: new Abstract: Automated operations research (OR) modeling requires LLMs to translate natural-language decision problems into correct mathematical programs. Existing methods can improve individual formulatio
PINNForge: Execution-Grounded Evolutionary Design of Physics-Informed Neural Networks for PDE Solving via Large Language Models
arXiv:2609.23023v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) require coordinated choices over network representation, sampling, loss construction, and optimization, while effective configurations often vary subst
Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
arXiv:2609.23038v1 Announce Type: new Abstract: Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transition
Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees
arXiv:2609.23043v1 Announce Type: new Abstract: Large Language Models (LLMs) enable open-ended dialogue in interactive games, but their non-deterministic outputs make it difficult to preserve authorial control, factual consistency, and the
LazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic Programs
arXiv:2609.23058v1 Announce Type: new Abstract: Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for agent-authored programs organized around
FireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire Dynamics
arXiv:2609.23064v1 Announce Type: new Abstract: Understanding the physical world requires more than object recognition, scene description, and short-term visual prediction, as real-world physical systems involve multiple continuous fields,
Tutoring Large Language Models to be Domain-adaptive, Precise and Safe
arXiv:2609.23071v1 Announce Type: new Abstract: This thesis proposes a framework for "responsible intelligence" to address AI's critical challenges in safety, ethics, and cultural sensitivity. It advances three core areas: First, it improve
Event Signature Transfer: Model-Agnostic Forecast Scenario Construction from Historical Events
arXiv:2609.23074v1 Announce Type: new Abstract: Forecasters often know an event is imminent but not the shape, size, or timing of its effect. We introduce Event Signature Transfer (EST), a training-free, model-agnostic operator that turns a
From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving
arXiv:2609.23130v1 Announce Type: new Abstract: Large language model (LLM) inference is evolving from an engine-local optimization problem into a distributed control problem involving reusable state, phase placement, heterogeneous accelerat
CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine
arXiv:2609.23142v1 Announce Type: new Abstract: Building gameplay features in a game engine requires more than code, as code that compiles and runs does not necessarily implement the requested gameplay. We introduce CraftBenchUE, an evaluat
Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation
arXiv:2609.23201v1 Announce Type: new Abstract: Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretable only in relation to the system tested,
Expansion Counts under Standard A* Tie-Breaking Strategies on the Final Plateau
arXiv:2609.23293v1 Announce Type: new Abstract: In the A* search algorithm, the tie-breaking strategies for nodes with the same $f$-value determines which states A* expands on the final $f$-layer. For nine standard tie-breaking strategies,
TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents
arXiv:2609.23363v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have led to the emergence of coding agents capable of performing complex engineering tasks, including register-transfer level (RTL) design and o
Leaky-integrator reconstruction: taming error accumulation in recursive differenced time-series forecasting
arXiv:2609.23378v1 Announce Type: new Abstract: We introduce leaky-integrator reconstruction, a training-free method that cures the error accumulation of recursive differenced forecasting. Our first contribution is diagnostic: predicting on
AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction
arXiv:2609.23512v1 Announce Type: new Abstract: Large language model agents are typically deployed with predefined configurations, although the required model capability, context, tools, permissions, memory, and computational resources can
Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment
arXiv:2609.23640v1 Announce Type: new Abstract: Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the resp
PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI
arXiv:2609.23695v1 Announce Type: new Abstract: Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must perceive, reason, plan, and act in dynami
ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents
arXiv:2609.23735v2 Announce Type: new Abstract: Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim assessment. Most existing systems, however,
On Probabilistic Inference Through Parametric Tensor Decomposition in Base Tensor Networks
arXiv:2609.23774v1 Announce Type: new Abstract: Probabilistic inference is generally only tractable in low-treewidth graphical models, limiting its effective applicability in high-treewidth settings. Many existing methods improve efficiency
Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows
arXiv:2609.23790v1 Announce Type: new Abstract: Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected tokens are billed as input tokens at the same
WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks
arXiv:2609.23806v1 Announce Type: new Abstract: Many knowledge-work benchmarks are constructed around individual tasks, with the context needed for each task selected together with or after the task has been specified. This design measures
Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models
arXiv:2609.23860v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) commonly reuse visual encoders pretrained with CLIP, although the features of these ViTs are ultimately consumed by autoregressive LLMs. We refer to th
Explainable Recommendations at Scale: LLM Rationales for YouTube Music Artist Discovery
arXiv:2609.23877v1 Announce Type: new Abstract: Modern music streaming platforms face a persistent tradeoff: exploiting familiar content versus driving the exploration of novel items. While users frequently desire discovery, they hesitate t
Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer
arXiv:2609.23917v1 Announce Type: new Abstract: Chess involves complex reasoning in a deterministic environment, which makes it a useful setting for studying the mechanisms of computation inside transformers. The Maia-3 chess transformer ta
Echo State Network (ESN) for Signal Recovery in RF-Impaired IBFD MIMO Systems
arXiv:2609.23945v1 Announce Type: new Abstract: In-band full-duplex (IBFD) multiple-input multiple-output (MIMO) systems enable simultaneous transmission and reception on the same frequency band, improving spectral efficiency for next-gener
Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control
arXiv:2609.23953v1 Announce Type: new Abstract: AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful. We ask what tha
Divergent strategies and convergent outcomes in autonomous materials discovery
arXiv:2609.23957v1 Announce Type: new Abstract: Scientific agents are mostly evaluated on whether they complete tasks or recover known results; we instead study variation across repeated open-ended campaigns. Sixteen separately initialized
UniK: Universal Knowledge Perception for Digital and Physical AI
arXiv:2609.23971v1 Announce Type: new Abstract: Two transformative classes of AI systems are reshaping how organizations operate: \textit{digital AI}, which reasons over enterprise knowledge to power chatbots and agent workflows; and \texti
LEAP-NBV: Lightweight Edge Active-Perception for Foundation-Model Next-Best-View Planning
arXiv:2609.23974v1 Announce Type: new Abstract: Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through visual perception. A representative example
Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
arXiv:2609.23986v1 Announce Type: new Abstract: Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expen
ACLArena: Agent Continue Learning in Multi-stage Post-training
arXiv:2609.23989v1 Announce Type: new Abstract: Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of training. Yet there is currently no well-es
FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering
arXiv:2609.24002v1 Announce Type: new Abstract: Large language model agents increasingly answer financial questions by searching regulatory filings. Such questions are often deceptively under-specified: Meta Platforms' "operating income" is
Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure
arXiv:2609.24012v2 Announce Type: new Abstract: Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We
Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech
arXiv:2609.24016v1 Announce Type: new Abstract: Commercial large language models are increasingly deployed across African fintech infrastructure for fraud detection and customer communication, yet no Nigerian or African continental regulato
Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search
arXiv:2609.24025v1 Announce Type: new Abstract: We present a method for synthesizing reactive character behaviors for continuous games as compact, human-readable programs. Game AI practice still relies heavily on manually authored behavior
Structured Decomposition for Reliable LLM-Generated Access Control Policies
arXiv:2609.24036v1 Announce Type: new Abstract: This paper presents an LLM-based system that translates natural-language access control policies (NLACPs) into executable Rego code for Open Policy Agent (OPA). It provides a modular, end-to-e
Representation-guided in-context learning for medical image interpretation with multimodal large language models
arXiv:2609.24057v1 Announce Type: new Abstract: Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuni
Incremental Consistency Execution for Autonomous Intelligent Systems
arXiv:2609.24090v1 Announce Type: new Abstract: Long-horizon autonomous intelligent systems rely on heterogeneous components such as large language models, databases, external APIs, and rule engines, while their external states continuously
DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents
arXiv:2609.24092v1 Announce Type: new Abstract: Real-world document processing systems rely on rigid, predefined schemas, yet critical target fields often lack direct visual counterparts on the page. Extracting these implicit values require
When More Evidence Hurts: Publication-Bias Drift and Principled Stopping for Biomedical Causal Search
arXiv:2609.24101v1 Announce Type: new Abstract: Automated biomedical evidence synthesis depends on retrieving published studies, but the biomedical literature is systematically skewed toward positive findings. Deeper retrieval can therefore
EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation
arXiv:2609.24115v1 Announce Type: new Abstract: Tool-calling LLM agents are increasingly deployed in enterprise applications. However, effective evaluation and optimization require high-quality, diverse task datasets that are often difficul
Self-Healing Harness for Runtime Oversight of Agent Self-Modification
arXiv:2609.24130v1 Announce Type: new Abstract: LLM agents can change their own future behavior, raising a basic control question of which self-generated changes should be allowed to persist. We formulate this as admission control for self-
APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction
arXiv:2609.24165v1 Announce Type: new Abstract: Synchrotron data reduction, detector calibration followed by azimuthal integration of terabyte-scale diffraction series, is a multi-step, expert-bound bottleneck that increasingly limits the s
CREDO: Variance-Guided Rubric Evolution for Replay-Corrected Credit Assignment
arXiv:2609.24174v1 Announce Type: new Abstract: Long-horizon language agents receive sparse terminal feedback, while intermediate rubrics provide structured but potentially misspecified assessments of progress. In resettable training enviro
LIMIT: Less Is More for Instruction Tuning in Text-to-SQL
arXiv:2609.24186v1 Announce Type: new Abstract: Large language models have achieved remarkable progress on Text-to-SQL through reasoning-enhanced fine-tuning, yet existing approaches predominantly rely on massive instruction corpora under t
SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision--Language Models
arXiv:2609.24198v1 Announce Type: new Abstract: This paper describes the SKstars submission to SHROOM-Visions 2026, a shared task on fine-grained hallucination detection in large vision-language model outputs. The task requires systems to i
Recovering Lost Details: Multi-Scale Frequency Compensation for Long-Term Time Series Forecasting
arXiv:2609.24229v1 Announce Type: new Abstract: Long-term time series forecasting has made significant progress by leveraging multi-scale information to capture hierarchical temporal patterns and model long-range dependencies. However, temp
Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation Enforcement
arXiv:2609.24243v1 Announce Type: new Abstract: Reinforcement learning (RL) improves reasoning in vision-language models (VLMs) but can induce chain-of-thought (CoT) obfuscation: an operational, non-intentional outcome where task reward or
Unsupervised Brain Anomaly Detection as a Bayesian Inverse Problem with Diffusion Prior
arXiv:2609.24265v1 Announce Type: new Abstract: Unsupervised anomaly detection (UAD) aims to localize abnormal regions in medical scans without pixel-level annotations. A typical strategy seeks to reconstruct a pseudo-healthy image that pre
How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation
arXiv:2609.24277v1 Announce Type: new Abstract: GUI agents predict click coordinates as digit-token sequences, but standard text-LLM confidence estimation methods rank correct clicks from wrong ones only weakly. GUI-specific alternatives us
When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain
arXiv:2609.24290v1 Announce Type: new Abstract: Instruction-tuned LLMs faced with underspecified queries often commit to a single interpretation rather than ask for clarification, producing confidently wrong answers. In our experiments, pro
Brain-Token Learning: Microstate-Based Tokenization and Multi-Scale Interaction for Long-Horizon EEG Sequence Modeling
arXiv:2609.24324v1 Announce Type: new Abstract: Electroencephalography (EEG) provides a non-invasive window into dynamic brain activity, yet modeling long-horizon EEG sequences remains challenging due to their high temporal complexity, subs
LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning
arXiv:2609.24346v1 Announce Type: new Abstract: Graph Retrieval-Augmented Generation (GraphRAG) has remarkably enhanced large language models on complex reasoning by leveraging structured entity topologies. However, existing frameworks heav
Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs
arXiv:2609.24352v1 Announce Type: new Abstract: Large language models (LLMs), when acting as agents, are expected to take observed data in context, infer the latent state space underlying the world, and leverage it for downstream prediction
VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning
arXiv:2609.24362v1 Announce Type: new Abstract: Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language models to vision-language models (VLMs) in
Predicting Postprandial Glycemic Response from Meal Images, Clinical Variables, and Gut Microbiome Information
arXiv:2609.24453v1 Announce Type: new Abstract: Predicting postprandial glycemic response (PPGR) is fundamental to personalized nutrition and type 2 diabetes management, yet existing approaches typically rely on manually reported dietary in
Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards
arXiv:2609.24480v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the convergent, evidence-driven task of inferring a
Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging
arXiv:2609.24517v1 Announce Type: new Abstract: Model merging aims to combine multiple fine-tuned models derived from a common pretrained model into a single multi-task model without additional joint training. Recent spectral merging method
The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
arXiv:2609.24555v1 Announce Type: new Abstract: We introduce the Endless Exam, a benchmark for measuring mathematical progress from today's models toward artificial superintelligence through fourteen parameterised construction families. Eac
Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis
arXiv:2609.24620v1 Announce Type: new Abstract: Answering epidemiological questions from real-world clinical data requires medical coding, schema-aware SQL, and validation of implicit choices about populations, denominators, and time. We pr
Custom Named Entity Recognition and Topic Classification for Global Health Publications
arXiv:2609.24625v1 Announce Type: new Abstract: How should natural language processing models be selected and adapted for global health literature in environments where annotated data and computational resources are limited? This thesis inv
DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security
arXiv:2609.24662v1 Announce Type: new Abstract: LLM-based agents increasingly operate in environments where they interact with users, tools, and external systems. Yet most security evaluations assume passive users and static control, ignori
Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents
arXiv:2609.24663v1 Announce Type: new Abstract: Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated th
TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction
arXiv:2609.24677v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cannot reveal whether correct answers reflect
World State Generator
arXiv:2609.24744v1 Announce Type: new Abstract: Language agents solve complex tasks through plans and actions. A single step the world refuses puts the goal out of reach, and what the agent does next decides the task. Prompted planners fail
Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents
arXiv:2609.24755v1 Announce Type: new Abstract: Autonomous AI agents are increasingly deployed in areas where wrong decisions are hard to reverse. This paper examines schema mismatch: the condition in which an agent operates within an inter
Construting Reverse Thinking: Developing Large Language Models' Reverse Thingking Ability
arXiv:2609.24760v1 Announce Type: new Abstract: When facing complex problems, humans tend to try various ideas for different issues. Human thinking patterns exhibit remarkable flexibility in adapting to diverse scenarios. GPT-o1, GPT-o3, an