arXiv:2510.09011v4 Announce Type: replace Abstract: In our deployed travel-planning service, most users give minimal inputs or free-form requests rather…
Tag: cs.AI updates on arXiv.org
Behavioral Coherence: A Method for Sensitive-Domain LLM Evaluation
arXiv:2512.13142v5 Announce Type: replace Abstract: People use LLMs for reproductive-health questions, including abortion-related support. A response can…
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
arXiv:2510.07731v4 Announce Type: replace Abstract: Organic reaction mechanisms describe the step-wise elementary processes by which reactants transform…
Architectural Design, Not Only Model Intelligence, Governs Multi-Agent LLM Performance
arXiv:2602.03128v2 Announce Type: replace Abstract: Multi-agent LLM frameworks are data-intensive systems that govern how agents orchestrate tasks, manage…
ERCPMP-Gx: Endoscopic Image and Video Dataset for Morphological, Histopathological, and Genomic Characterization of Colorectal Polyposis
arXiv:2609.20815v1 Announce Type: cross Abstract: Hereditary polyposis syndromes can be precursor lesions to colorectal cancer and are associated with a…
Paint-Anything: Unified Any-Color Control for Image Generation and Editing
arXiv:2609.20816v1 Announce Type: cross Abstract: Professional design requires any-color control: the ability to specify an object’s target color with any…
Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision
arXiv:2609.20820v1 Announce Type: cross Abstract: Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As…
FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
arXiv:2609.20817v1 Announce Type: cross Abstract: Modeling articulated objects from sparse monocular views is challenging because each observation reveals…
Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation
arXiv:2609.20822v1 Announce Type: cross Abstract: Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the…
Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
arXiv:2609.20779v1 Announce Type: cross Abstract: Safety evaluations for large language models rely on surface-form classifiers that report declining harm…
