San Francisco’s housing market is in trouble again.
Tag: AI
INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation
arXiv:2603.21607v2 Announce Type: replace Abstract: While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it…
Social World Models
arXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about…
CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
arXiv:2608.07460v1 Announce Type: cross Abstract: While post-training improves the capabilities of large language models (LLMs), it generally lowers their…
“LLM Agent Performance” Is Not a Single Evaluation Target
arXiv:2602.03238v3 Announce Type: replace Abstract: LLM agent benchmark scores are shaped not only by the model but also by the agent harness,…
Boundary Density Likelihood for Direct Event-Time Supervision
arXiv:2408.12792v2 Announce Type: replace Abstract: Event detection turns long recordings into a sparse set of ranked timestamps. Yet many sequence models…
As AI-led attacks multiply, OpenAI launches a new cyber model
OpenAI is expanding its AI cybersecurity defense program Daybreak, and rolling out a new cyber-trained AI model with it.
Serious Games: Human-AI Interaction, Evolution, and Coevolution
arXiv:2505.16388v3 Announce Type: replace Abstract: The serious games between humans and AI have only just begun. Evolutionary Game Theory (EGT) models…
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
arXiv:2608.07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making…
Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools
arXiv:2608.07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational,…