11 posts published in the last hour
- 08:32What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models
- 08:32Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents
- 08:32Visual Compliance via Executable Safety Rule Entailment
- 08:32Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts
- 08:32Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition
- 08:04REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement
- 08:04Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI
- 08:04Building Trust in Artificial Intelligence: A Necessity for Railway Applications
- 08:03BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs
- 08:03OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training
- 08:03Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum
