arXiv:2608.24263v1 Announce Type: new Abstract: Change data synthesis provides a cost-effective solution for expanding training data and improving the…
Category: AI
SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction
arXiv:2608.24252v1 Announce Type: new Abstract: LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful…
STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation
arXiv:2608.24237v1 Announce Type: new Abstract: Longitudinal radiology report generation (LRRG) requires identifying both current findings and their…
Spineart’s PERLA TL App Gains FDA Clearance for Robotic Spine Surgery
Spineart and eCential Robotics received FDA 510(k) clearance on August 26, 2026 for Spineart’s PERLA TL application to run on eCential’s Op.n navigation…
Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks
arXiv:2608.24258v1 Announce Type: new Abstract: AI systems are increasingly evaluated for legally accountable settings, where correct outputs must also be…
Fastino Releases GLiNER2.5: A Boundary-Prediction Architecture That Removes Span Enumeration From Information Extraction
Fastino released GLiNER2.5, replacing span enumeration with boundary prediction so entity width no longer costs compute. Three Apache 2.0 checkpoints ship…
Matched Excess-Outranker Regularization for Candidate-Set Interference in Continual Knowledge Graph Embedding
arXiv:2608.24273v1 Announce Type: new Abstract: Continual knowledge graph embedding updates entity and relation representations as a graph grows. Existing…
Trump bought SpaceX shares two weeks after blockbuster IPO
The president bought when the stock was in the mid-$150 range. SpaceX finished trading on Monday back at its IPO price of $135.
AI models flub these intelligence tests. Can you fare any better?
Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic…
Evaluating Multiple LLM Generations with Validated Task Coverage
arXiv:2608.24228v1 Announce Type: new Abstract: Many LLM applications are most useful when they provide several candidate outputs for comparison,…
