arXiv:2609.23806v1 Announce Type: new Abstract: Many knowledge-work benchmarks are constructed around individual tasks, with the context needed for each…
Tag: AI
ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents
arXiv:2609.23735v2 Announce Type: new Abstract: Scientific agents support a range of literature-based research tasks, such as retrieval, question…
Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models
arXiv:2609.23860v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) commonly reuse visual encoders pretrained with CLIP, although the…
Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows
arXiv:2609.23790v1 Announce Type: new Abstract: Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects…
TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents
arXiv:2609.23363v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have led to the emergence of coding agents capable of…
Leaky-integrator reconstruction: taming error accumulation in recursive differenced time-series forecasting
arXiv:2609.23378v1 Announce Type: new Abstract: We introduce leaky-integrator reconstruction, a training-free method that cures the error accumulation of…
Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment
arXiv:2609.23640v1 Announce Type: new Abstract: Human-feedback alignment has made language models useful assistants and is commonly described as aligning…
PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI
arXiv:2609.23695v1 Announce Type: new Abstract: Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as…
AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction
arXiv:2609.23512v1 Announce Type: new Abstract: Large language model agents are typically deployed with predefined configurations, although the required…
Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation
arXiv:2609.23201v1 Announce Type: new Abstract: Benchmark scores increasingly influence the development, marketing, and selection of large language models…
