arXiv:2608.21097v1 Announce Type: new Abstract: LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and…
Tag: AI
XPENG Robotics Raises $900M+ First Round at $6.3B+ Valuation
XPENG’s robotics business has entered share purchase agreements with investors to raise more than US$900 million at a post-money valuation above US$6.3…
Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda
arXiv:2608.21107v1 Announce Type: new Abstract: Large Language Models (LLMs) are moving from code completion toward repository-scale agents that retrieve…
Kids outlearn AI—and we still don’t know why
People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world…
ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models
arXiv:2608.21100v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) extend model capabilities beyond text, they also make…
Socialized Division and Collaboration: Rethinking Class-Incremental Learning under Optimization Conflicts
arXiv:2608.21044v1 Announce Type: new Abstract: Class-incremental learning is commonly instantiated as a single-model paradigm, where a unified model…
Don’t Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents
arXiv:2608.21027v1 Announce Type: new Abstract: LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use,…
Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models
arXiv:2608.20975v1 Announce Type: new Abstract: Effective social interaction requires agents to translate mental state inferences into coordinated…
When AI Reads Between the Lines: OCR vs. VLMs
Can machines truly understand documents, or have they simply become more effective at extracting information from them? With traditional OCR, an error can…
Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance
arXiv:2608.21036v1 Announce Type: new Abstract: The transport of dangerous goods by sea is a high-consequence activity governed by the International…
