arXiv:2609.07943v1 Announce Type: new Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the…
Tag: cs.AI updates on arXiv.org
PRIMUS: Identity, Governance, and Verification for Multi-Agent Federations
arXiv:2609.07910v1 Announce Type: new Abstract: Multi-agent federations need governance that answers three questions under adversarial conditions: who…
FrogNano: Training a 4B Coding Agent via Online Task Synthesis
arXiv:2609.07925v2 Announce Type: new Abstract: We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently and…
Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Serving-Time Weight Compression
arXiv:2609.07901v1 Announce Type: new Abstract: Weight quantization largely determines the economics of serving open-weight LLMs. Its costs are usually…
xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
arXiv:2609.07784v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only…
Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging
arXiv:2609.07803v1 Announce Type: new Abstract: Model pruning is widely used to compress deep neural networks, reducing memory and computational…
Do Large Language Models Know What They Don’t Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty
arXiv:2609.07879v1 Announce Type: new Abstract: Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question…
When Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AI Systems
arXiv:2609.07741v1 Announce Type: new Abstract: Persistent AI assistants are intended to extend human attention, memory, and coordination across changing…
What Does an LLM-Agent Leaderboard Rank Actually Compare?
arXiv:2609.07785v1 Announce Type: new Abstract: An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent.…
Aegix Pulse: A Traceable Three-Stage Architecture for Personalized Content Generation and Context-Preserving Revision
arXiv:2609.07672v1 Announce Type: new Abstract: Production content-generation systems must integrate a user’s immediate task, long-term brand identity,…
