arXiv:2608.11905v1 Announce Type: new Abstract: In many practical applications of generative AI systems, from tax rules to airline baggage allowance,…
Tag: cs.AI updates on arXiv.org
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
arXiv:2608.11888v1 Announce Type: new Abstract: Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape…
Proportional Analogies on Probability Distributions via Bayesian Updating
arXiv:2608.11724v1 Announce Type: new Abstract: Analogies are quaternary relations of the form “A is to B as C is to D”. Among the various formalizations…
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting
arXiv:2608.11692v1 Announce Type: new Abstract: Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which…
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
arXiv:2608.11727v1 Announce Type: new Abstract: When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing…
HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry
arXiv:2608.11768v1 Announce Type: new Abstract: The adaptive neuro-fuzzy inference system (ANFIS) is an interpretable reasoning framework capable of…
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning
arXiv:2608.11705v1 Announce Type: new Abstract: Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the…
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication
arXiv:2608.11676v1 Announce Type: new Abstract: Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can…
Making AI-Generated Feedback Matter: From Provision to Student Enactment
arXiv:2608.11625v1 Announce Type: new Abstract: Feedback processes strongly influence student learning, yet their educational value depends on addressing…
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents
arXiv:2608.11683v1 Announce Type: new Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the…
