news.jao.life

A calm feed of what's actually happening — without the tracking.

Privacy-first by design. This site shows headlines and short snippets only. Every story opens a preview page first, then links out to the original publisher with no referrer. No cookies. No analytics. No third-party requests. Read the full promise →

Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models

arXiv:2609.22161v1 Announce Type: new Abstract: Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types differentially shape model

An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users

arXiv:2609.22277v1 Announce Type: new Abstract: Visual impairment affects over 2.2 billion people worldwide, yet conventional white canes cannot detect elevated hazards or provide semantic environmental context. Existing AI-assisted navigat

PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation

arXiv:2609.22353v1 Announce Type: new Abstract: Mobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints alone cannot describe. On resource constrained platforms, converting imperfect visual pre

Social Influence and the Allocation of Scientific Attention in AI Populations

arXiv:2609.22408v1 Announce Type: new Abstract: AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human

Learning 3D biophysical cell properties from 2D images and cell-population statistics

arXiv:2609.22410v1 Announce Type: new Abstract: Inferring 3D cellular properties from 2D microscopy is difficult when a reference instrument reports only population statistics rather than labels for individual cells. Here we develop a popul

Goal-driven Variant Categorization

arXiv:2609.22475v1 Announce Type: new Abstract: Process discovery rarely yields a single coherent process structure. For analysis, a common step is to cluster process variants based on structural similarity and then assign business meaning

Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation

arXiv:2609.22478v1 Announce Type: new Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whethe

The Wisdom of Artificial Deliberative Crowds

arXiv:2609.22497v1 Announce Type: new Abstract: The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually attributed to the independence of estimate

Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and

IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

arXiv:2609.22529v1 Announce Type: new Abstract: International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its texts remain without token-level named ent

EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability

arXiv:2609.22537v1 Announce Type: new Abstract: Enterprise AI assistants must produce responses that are verifiable and traceable to source evidence. However, retrieval augmented generation (RAG) over heterogeneous enterprise data can suffe

AutoGym: Blueprint-First Generation of Verifiable Agent Gyms

arXiv:2609.22592v1 Announce Type: new Abstract: Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a verifier that reliably distinguishes success

MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators

arXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated resp

GaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment

arXiv:2609.22619v1 Announce Type: new Abstract: Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation sessions, yet objective 3D measurement remains confined to specialized motion-capture labo

Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation

arXiv:2609.22620v1 Announce Type: new Abstract: Scanned mail, uploaded PDFs, and consolidated attachments often arrive as page streams that must be split into individual documents before downstream classification, extraction, or routing. Ze

Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA

arXiv:2609.22628v1 Announce Type: new Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the prompt, jud

Self-Organizing Agent Teams Learn to Reason Together

arXiv:2609.22682v1 Announce Type: new Abstract: Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor c

Generative Embodied Multiple Behavior Control Systems for Human-like Agents

arXiv:2609.22691v1 Announce Type: new Abstract: An enduring and richly elaborated dichotomy in cognitive neuroscience is that of human behavior control mechanisms, divided into habitual versus goal-directed. While existing human-like agent

Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review

arXiv:2609.22694v1 Announce Type: new Abstract: Public crash databases increasingly support automated safety analysis, but crash severity prediction remains difficult to translate into public-sector decision workflows when models are evalua

A Survey on the Linear Representation Hypothesis

arXiv:2609.22695v1 Announce Type: new Abstract: The term "linear representation hypothesis" (LRH) has appeared across diverse subfields of artificial intelligence, neuroscience, and cognitive science. But previous works have not consistentl

Building Trustworthy Mental Health Benchmarks on Bluesky: A Validation-Aware Weak-Supervision Framework

arXiv:2609.22696v1 Announce Type: new Abstract: Decentralized social media platforms create new opportunities and challenges for computational mental health research because data access, moderation, labeling, and deployment responsibilities

Hapi: A Multivariable Land-Surface Transformer for Medium-Range Hydrological Forecasting at Continental Scale

arXiv:2609.22702v1 Announce Type: new Abstract: Accurate flood forecasts several days in advance are essential for flood control, water-resource management, and emergency response. Producing them at high resolution over a continental domain

Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

arXiv:2609.22712v1 Announce Type: new Abstract: Agentic AI systems built on large language models can plan over multiple steps, use external tools, retain information in memory, and coordinate with other agents. These capabilities make them

ProcessLight: Process Supervision for Large Language Model Based Traffic Signal Control

arXiv:2609.22746v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently been introduced into traffic signal control (TSC) as decision agents due to their strengths in human-readable reasoning generation. Yet, existing LLM

CTSpinoPelvic1K: spine, pelvis, ribs and femora in one coordinate frame, annotated for lumbosacral transitional anatomy

arXiv:2609.22760v1 Announce Type: new Abstract: Purpose: A vertebra at the lumbosacral junction is named by counting caudally from C2 on whole-spine imaging, but a lumbar case is planned on lumbar-only imaging (T12 to S1), without C2. Abdom

DVA-Neurons: Design and Verification of Adaptive LIF Neurons: From Single-Neuron Dynamics to Multi-Neuron Spiking Networks

arXiv:2609.22775v1 Announce Type: new Abstract: Spiking Neural Networks (SNNs) offer a promising path toward ultra-low-power artificial intelligence inference by emulating the event-driven computation of biological neurons. However, two cha

From Research Frontier to Laboratory Bench: Design of a Four-Tier Experimental Teaching System for Multimodal Medical Image Intelligent Diagnosis

arXiv:2609.22790v1 Announce Type: new Abstract: Undergraduate programmes in intelligent medical engineering are expanding, yet laboratory curricula lag behind the multimodal, long-tailed, and distributionally shifting realities of clinical

ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures

arXiv:2609.22878v1 Announce Type: new Abstract: Large language model code generation benchmarks primarily evaluate well-resourced languages like Python and Java, where models benefit from abundant training data. They provide limited evidenc

When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used

arXiv:2609.22910v1 Announce Type: new Abstract: Vision-language agents that crop and zoom are trained with rewards that credit a successful tool call, yet a successful call does not show that the model needed to look or used the pixels it r

Beyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language Model

arXiv:2609.22939v1 Announce Type: new Abstract: Long-context models read a novel the way a person reads a printout: one token after another, in narrative order, with the whole history competing for a fixed budget of attention. A detective d

AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows

arXiv:2609.22951v1 Announce Type: new Abstract: Enterprise agentic systems that route every trajectory step to a frontier model waste 60-80% of their inference budget on subtasks that smaller models handle equally well. Existing routing sol

A Compact Stance-Indexed Anterior-Posterior COP Representation for Parkinson's Disease Classification from Plantar VGRF

arXiv:2609.22956v1 Announce Type: new Abstract: Parkinson's disease alters gait and bilateral coordination, but machine-learning performance also depends on how continuous gait signals are represented. This study investigates whether preser

R-GEAN: Regimen-Guided Edit Action Network for Within-Admission Medication Change Prediction

arXiv:2609.22959v1 Announce Type: new Abstract: The medications prescribed to a patient often change during a hospital admission as clinicians start, stop, or continue therapies. We study whether models can predict which medication classes

OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling

arXiv:2609.22987v1 Announce Type: new Abstract: Automated operations research (OR) modeling requires LLMs to translate natural-language decision problems into correct mathematical programs. Existing methods can improve individual formulatio

PINNForge: Execution-Grounded Evolutionary Design of Physics-Informed Neural Networks for PDE Solving via Large Language Models

arXiv:2609.23023v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) require coordinated choices over network representation, sampling, loss construction, and optimization, while effective configurations often vary subst

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

arXiv:2609.23038v1 Announce Type: new Abstract: Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transition

Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees

arXiv:2609.23043v1 Announce Type: new Abstract: Large Language Models (LLMs) enable open-ended dialogue in interactive games, but their non-deterministic outputs make it difficult to preserve authorial control, factual consistency, and the

LazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic Programs

arXiv:2609.23058v1 Announce Type: new Abstract: Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for agent-authored programs organized around

FireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire Dynamics

arXiv:2609.23064v1 Announce Type: new Abstract: Understanding the physical world requires more than object recognition, scene description, and short-term visual prediction, as real-world physical systems involve multiple continuous fields,

Tutoring Large Language Models to be Domain-adaptive, Precise and Safe

arXiv:2609.23071v1 Announce Type: new Abstract: This thesis proposes a framework for "responsible intelligence" to address AI's critical challenges in safety, ethics, and cultural sensitivity. It advances three core areas: First, it improve

Event Signature Transfer: Model-Agnostic Forecast Scenario Construction from Historical Events

arXiv:2609.23074v1 Announce Type: new Abstract: Forecasters often know an event is imminent but not the shape, size, or timing of its effect. We introduce Event Signature Transfer (EST), a training-free, model-agnostic operator that turns a

From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving

arXiv:2609.23130v1 Announce Type: new Abstract: Large language model (LLM) inference is evolving from an engine-local optimization problem into a distributed control problem involving reusable state, phase placement, heterogeneous accelerat

CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine

arXiv:2609.23142v1 Announce Type: new Abstract: Building gameplay features in a game engine requires more than code, as code that compiles and runs does not necessarily implement the requested gameplay. We introduce CraftBenchUE, an evaluat

Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation

arXiv:2609.23201v1 Announce Type: new Abstract: Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretable only in relation to the system tested,

Expansion Counts under Standard A* Tie-Breaking Strategies on the Final Plateau

arXiv:2609.23293v1 Announce Type: new Abstract: In the A* search algorithm, the tie-breaking strategies for nodes with the same $f$-value determines which states A* expands on the final $f$-layer. For nine standard tie-breaking strategies,

TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents

arXiv:2609.23363v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have led to the emergence of coding agents capable of performing complex engineering tasks, including register-transfer level (RTL) design and o

Leaky-integrator reconstruction: taming error accumulation in recursive differenced time-series forecasting

arXiv:2609.23378v1 Announce Type: new Abstract: We introduce leaky-integrator reconstruction, a training-free method that cures the error accumulation of recursive differenced forecasting. Our first contribution is diagnostic: predicting on

AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction

arXiv:2609.23512v1 Announce Type: new Abstract: Large language model agents are typically deployed with predefined configurations, although the required model capability, context, tools, permissions, memory, and computational resources can

Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment

arXiv:2609.23640v1 Announce Type: new Abstract: Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the resp

PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI

arXiv:2609.23695v1 Announce Type: new Abstract: Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must perceive, reason, plan, and act in dynami

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

arXiv:2609.23735v2 Announce Type: new Abstract: Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim assessment. Most existing systems, however,

On Probabilistic Inference Through Parametric Tensor Decomposition in Base Tensor Networks

arXiv:2609.23774v1 Announce Type: new Abstract: Probabilistic inference is generally only tractable in low-treewidth graphical models, limiting its effective applicability in high-treewidth settings. Many existing methods improve efficiency

Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows

arXiv:2609.23790v1 Announce Type: new Abstract: Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected tokens are billed as input tokens at the same

WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks

arXiv:2609.23806v1 Announce Type: new Abstract: Many knowledge-work benchmarks are constructed around individual tasks, with the context needed for each task selected together with or after the task has been specified. This design measures

Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models

arXiv:2609.23860v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) commonly reuse visual encoders pretrained with CLIP, although the features of these ViTs are ultimately consumed by autoregressive LLMs. We refer to th

Explainable Recommendations at Scale: LLM Rationales for YouTube Music Artist Discovery

arXiv:2609.23877v1 Announce Type: new Abstract: Modern music streaming platforms face a persistent tradeoff: exploiting familiar content versus driving the exploration of novel items. While users frequently desire discovery, they hesitate t

Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer

arXiv:2609.23917v1 Announce Type: new Abstract: Chess involves complex reasoning in a deterministic environment, which makes it a useful setting for studying the mechanisms of computation inside transformers. The Maia-3 chess transformer ta

Echo State Network (ESN) for Signal Recovery in RF-Impaired IBFD MIMO Systems

arXiv:2609.23945v1 Announce Type: new Abstract: In-band full-duplex (IBFD) multiple-input multiple-output (MIMO) systems enable simultaneous transmission and reception on the same frequency band, improving spectral efficiency for next-gener

Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control

arXiv:2609.23953v1 Announce Type: new Abstract: AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful. We ask what tha

Divergent strategies and convergent outcomes in autonomous materials discovery

arXiv:2609.23957v1 Announce Type: new Abstract: Scientific agents are mostly evaluated on whether they complete tasks or recover known results; we instead study variation across repeated open-ended campaigns. Sixteen separately initialized

UniK: Universal Knowledge Perception for Digital and Physical AI

arXiv:2609.23971v1 Announce Type: new Abstract: Two transformative classes of AI systems are reshaping how organizations operate: \textit{digital AI}, which reasons over enterprise knowledge to power chatbots and agent workflows; and \texti

LEAP-NBV: Lightweight Edge Active-Perception for Foundation-Model Next-Best-View Planning

arXiv:2609.23974v1 Announce Type: new Abstract: Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through visual perception. A representative example

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

arXiv:2609.23986v1 Announce Type: new Abstract: Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expen

ACLArena: Agent Continue Learning in Multi-stage Post-training

arXiv:2609.23989v1 Announce Type: new Abstract: Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of training. Yet there is currently no well-es

FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering

arXiv:2609.24002v1 Announce Type: new Abstract: Large language model agents increasingly answer financial questions by searching regulatory filings. Such questions are often deceptively under-specified: Meta Platforms' "operating income" is

Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure

arXiv:2609.24012v2 Announce Type: new Abstract: Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We

Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech

arXiv:2609.24016v1 Announce Type: new Abstract: Commercial large language models are increasingly deployed across African fintech infrastructure for fraud detection and customer communication, yet no Nigerian or African continental regulato

Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search

arXiv:2609.24025v1 Announce Type: new Abstract: We present a method for synthesizing reactive character behaviors for continuous games as compact, human-readable programs. Game AI practice still relies heavily on manually authored behavior

Structured Decomposition for Reliable LLM-Generated Access Control Policies

arXiv:2609.24036v1 Announce Type: new Abstract: This paper presents an LLM-based system that translates natural-language access control policies (NLACPs) into executable Rego code for Open Policy Agent (OPA). It provides a modular, end-to-e

Representation-guided in-context learning for medical image interpretation with multimodal large language models

arXiv:2609.24057v1 Announce Type: new Abstract: Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuni

Incremental Consistency Execution for Autonomous Intelligent Systems

arXiv:2609.24090v1 Announce Type: new Abstract: Long-horizon autonomous intelligent systems rely on heterogeneous components such as large language models, databases, external APIs, and rule engines, while their external states continuously

DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents

arXiv:2609.24092v1 Announce Type: new Abstract: Real-world document processing systems rely on rigid, predefined schemas, yet critical target fields often lack direct visual counterparts on the page. Extracting these implicit values require

When More Evidence Hurts: Publication-Bias Drift and Principled Stopping for Biomedical Causal Search

arXiv:2609.24101v1 Announce Type: new Abstract: Automated biomedical evidence synthesis depends on retrieving published studies, but the biomedical literature is systematically skewed toward positive findings. Deeper retrieval can therefore

EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation

arXiv:2609.24115v1 Announce Type: new Abstract: Tool-calling LLM agents are increasingly deployed in enterprise applications. However, effective evaluation and optimization require high-quality, diverse task datasets that are often difficul

Self-Healing Harness for Runtime Oversight of Agent Self-Modification

arXiv:2609.24130v1 Announce Type: new Abstract: LLM agents can change their own future behavior, raising a basic control question of which self-generated changes should be allowed to persist. We formulate this as admission control for self-

APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction

arXiv:2609.24165v1 Announce Type: new Abstract: Synchrotron data reduction, detector calibration followed by azimuthal integration of terabyte-scale diffraction series, is a multi-step, expert-bound bottleneck that increasingly limits the s

CREDO: Variance-Guided Rubric Evolution for Replay-Corrected Credit Assignment

arXiv:2609.24174v1 Announce Type: new Abstract: Long-horizon language agents receive sparse terminal feedback, while intermediate rubrics provide structured but potentially misspecified assessments of progress. In resettable training enviro

LIMIT: Less Is More for Instruction Tuning in Text-to-SQL

arXiv:2609.24186v1 Announce Type: new Abstract: Large language models have achieved remarkable progress on Text-to-SQL through reasoning-enhanced fine-tuning, yet existing approaches predominantly rely on massive instruction corpora under t

SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision--Language Models

arXiv:2609.24198v1 Announce Type: new Abstract: This paper describes the SKstars submission to SHROOM-Visions 2026, a shared task on fine-grained hallucination detection in large vision-language model outputs. The task requires systems to i

Recovering Lost Details: Multi-Scale Frequency Compensation for Long-Term Time Series Forecasting

arXiv:2609.24229v1 Announce Type: new Abstract: Long-term time series forecasting has made significant progress by leveraging multi-scale information to capture hierarchical temporal patterns and model long-range dependencies. However, temp

Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation Enforcement

arXiv:2609.24243v1 Announce Type: new Abstract: Reinforcement learning (RL) improves reasoning in vision-language models (VLMs) but can induce chain-of-thought (CoT) obfuscation: an operational, non-intentional outcome where task reward or

Unsupervised Brain Anomaly Detection as a Bayesian Inverse Problem with Diffusion Prior

arXiv:2609.24265v1 Announce Type: new Abstract: Unsupervised anomaly detection (UAD) aims to localize abnormal regions in medical scans without pixel-level annotations. A typical strategy seeks to reconstruct a pseudo-healthy image that pre

How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation

arXiv:2609.24277v1 Announce Type: new Abstract: GUI agents predict click coordinates as digit-token sequences, but standard text-LLM confidence estimation methods rank correct clicks from wrong ones only weakly. GUI-specific alternatives us

When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain

arXiv:2609.24290v1 Announce Type: new Abstract: Instruction-tuned LLMs faced with underspecified queries often commit to a single interpretation rather than ask for clarification, producing confidently wrong answers. In our experiments, pro

Brain-Token Learning: Microstate-Based Tokenization and Multi-Scale Interaction for Long-Horizon EEG Sequence Modeling

arXiv:2609.24324v1 Announce Type: new Abstract: Electroencephalography (EEG) provides a non-invasive window into dynamic brain activity, yet modeling long-horizon EEG sequences remains challenging due to their high temporal complexity, subs

LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning

arXiv:2609.24346v1 Announce Type: new Abstract: Graph Retrieval-Augmented Generation (GraphRAG) has remarkably enhanced large language models on complex reasoning by leveraging structured entity topologies. However, existing frameworks heav

Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs

arXiv:2609.24352v1 Announce Type: new Abstract: Large language models (LLMs), when acting as agents, are expected to take observed data in context, infer the latent state space underlying the world, and leverage it for downstream prediction

VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning

arXiv:2609.24362v1 Announce Type: new Abstract: Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language models to vision-language models (VLMs) in

Predicting Postprandial Glycemic Response from Meal Images, Clinical Variables, and Gut Microbiome Information

arXiv:2609.24453v1 Announce Type: new Abstract: Predicting postprandial glycemic response (PPGR) is fundamental to personalized nutrition and type 2 diabetes management, yet existing approaches typically rely on manually reported dietary in

Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards

arXiv:2609.24480v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the convergent, evidence-driven task of inferring a

Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging

arXiv:2609.24517v1 Announce Type: new Abstract: Model merging aims to combine multiple fine-tuned models derived from a common pretrained model into a single multi-task model without additional joint training. Recent spectral merging method

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence

arXiv:2609.24555v1 Announce Type: new Abstract: We introduce the Endless Exam, a benchmark for measuring mathematical progress from today's models toward artificial superintelligence through fourteen parameterised construction families. Eac

Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis

arXiv:2609.24620v1 Announce Type: new Abstract: Answering epidemiological questions from real-world clinical data requires medical coding, schema-aware SQL, and validation of implicit choices about populations, denominators, and time. We pr

Custom Named Entity Recognition and Topic Classification for Global Health Publications

arXiv:2609.24625v1 Announce Type: new Abstract: How should natural language processing models be selected and adapted for global health literature in environments where annotated data and computational resources are limited? This thesis inv

DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security

arXiv:2609.24662v1 Announce Type: new Abstract: LLM-based agents increasingly operate in environments where they interact with users, tools, and external systems. Yet most security evaluations assume passive users and static control, ignori

Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

arXiv:2609.24663v1 Announce Type: new Abstract: Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated th

TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction

arXiv:2609.24677v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cannot reveal whether correct answers reflect

World State Generator

arXiv:2609.24744v1 Announce Type: new Abstract: Language agents solve complex tasks through plans and actions. A single step the world refuses puts the goal out of reach, and what the agent does next decides the task. Prompted planners fail

Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents

arXiv:2609.24755v1 Announce Type: new Abstract: Autonomous AI agents are increasingly deployed in areas where wrong decisions are hard to reverse. This paper examines schema mismatch: the condition in which an agent operates within an inter

Construting Reverse Thinking: Developing Large Language Models' Reverse Thingking Ability

arXiv:2609.24760v1 Announce Type: new Abstract: When facing complex problems, humans tend to try various ideas for different issues. Human thinking patterns exhibit remarkable flexibility in adapting to diverse scenarios. GPT-o1, GPT-o3, an