13 posts published in the last hour
- 05:32ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
- 05:32The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?
- 05:32SNOMED CT Concept Recommendation from Masked Clinical Context
- 05:32Roundtables: Could AI really kill us all?
- 05:32Do Frontier Models Seek Safety Evidence Before Acting?
- 05:32Meta now lets AI agents handle the boring parts of WhatsApp Business setup
- 05:32OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning
- 05:03FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment
- 05:03Imitation Learning for Autonomous Driving in CARLA
- 05:03Learning Heterogeneous Preferences
- 05:03SAGE: Governed Artifact Generation from Enterprise Guidelines
- 05:03Iceland-based Treble raises $18 million for its voice simulation platform
- 05:03A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning
