arXiv:2609.05439v1 Announce Type: new Abstract: Current evaluation methods for large language models are coarse-grained and decoupled from generation,…
Tag: cs.AI updates on arXiv.org
AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents
arXiv:2609.05446v1 Announce Type: new Abstract: We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model…
Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
arXiv:2609.05437v1 Announce Type: new Abstract: Previous AI alignment efforts have focused primarily on first-order social norms — teaching models what…
Damage-Aware Bandit Pruning for Vision and Language Transformers
arXiv:2609.05448v1 Announce Type: new Abstract: Structured post-training pruning of transformers requires selecting complete functional units whose…
Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points
arXiv:2412.11194v3 Announce Type: replace-cross Abstract: Security vulnerabilities in software can have severe consequences; however, manual vulnerability…
AI-Powered CPS-Enabled Vulnerable-User-Aware Urban Transportation Digital Twin: Methods and Applications
arXiv:2501.10396v4 Announce Type: replace-cross Abstract: We present methods and applications for the development of digital twins (DT) for urban traffic…
Measuring proximity to standard planes during fetal brain ultrasound scanning
arXiv:2404.07124v2 Announce Type: replace-cross Abstract: This paper presents a pipeline designed to bring ultrasound (US) plane pose estimation closer to…
An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders
arXiv:2406.02465v2 Announce Type: replace-cross Abstract: Can pretrained models generalize to new datasets without any retraining? We deploy pretrained…
Hyperedge Anomaly Detection with Hypergraph Neural Network
arXiv:2412.05641v2 Announce Type: replace-cross Abstract: Hypergraph is a data structure that enables us to model higher-order associations among data…
Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
arXiv:2608.31082v2 Announce Type: replace Abstract: Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings,…
