arXiv:2505.11326v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown remarkable progress in offline tasks such as image…
From In-Silico to Wet-Lab: Evaluating AI Protein Design Performance
In this tutorial, we analyze Anthropic’s 1,440 AI-designed protein binder dataset to benchmark 10 leading structure predictors. Discover how target…
HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement
arXiv:2505.15740v2 Announce Type: replace-cross Abstract: Formal methods play a crucial role in ensuring the reliability of critical systems through…
OpenAI researcher warns ultrafast AI could leave security teams in the dust
An OpenAI researcher warns that state-of-the-art AI models running 50 times faster could infiltrate systems before human teams can react. Simple…
Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation
arXiv:2506.21599v5 Announce Type: replace-cross Abstract: Advancing large language models (LLMs) for the next point-of-interest (POI) recommendation task…
CollaFuse: Collaborative Diffusion Models
arXiv:2406.14429v5 Announce Type: replace-cross Abstract: In the landscape of generative artificial intelligence, diffusion-based models have emerged as a…
The BS-meter: Detecting Politics and Labour through ChatGPT’s Language
arXiv:2411.15129v3 Announce Type: replace-cross Abstract: What can we learn about language from studying how it is used by ChatGPT and other large…
Why Travel Needs Layered AI Adoption, Not a Race to Autonomy
AI discussions in travel often tend to center on how the technology can rapidly transform the sector and deliver high-impact outcomes. The reality,…
Recurrent Reinforcement Learning with Memoroids
arXiv:2402.09900v4 Announce Type: replace-cross Abstract: Memory models such as Recurrent Neural Networks (RNNs) and Transformers address Partially…
Adam Gross, Co-Founder and CEO of HarmonEyes – Interview Series
Adam Gross, Co-Founder and CEO of HarmonEyes, is a serial entrepreneur with nearly three decades of experience building and scaling technology-driven…
Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
arXiv:2504.05216v5 Announce Type: replace-cross Abstract: Dense retrieval is a crucial task in Information Retrieval (IR), serving as the basis for…
Yardstik Raises $30M Series B as AI Reshapes the Future of Workforce Trust
Yardstik has raised $30 million in Series B funding as the workforce technology company looks to expand beyond traditional background checks and tackle a…
Communication styles and reader preferences of LLM- and human-authored COVID-19 information explanations: a case study
arXiv:2505.08143v2 Announce Type: replace-cross Abstract: With the wide adoption of large language models (LLMs) in information assistance, it is…
Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents
arXiv:2608.22963v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where…
ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation
arXiv:2608.22559v2 Announce Type: replace Abstract: Rubrics aim to make language-model evaluation transparent by decomposing response quality into…
Account Consistency from Gameplay Traces: Same-Player Verification in Counter-Strike 2
arXiv:2608.24893v2 Announce Type: replace Abstract: In competitive first-person shooter (FPS) games such as Counter-Strike 2 (CS2), account-integrity…
Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement
arXiv:2608.23098v2 Announce Type: replace Abstract: Classical Chinese poetry composition has long valued Tuiqiao, the iterative refinement of words,…
Our decision on Cursor following its acquisition by SpaceX
Our decision to wind down our contract providing OpenAI models to Cursor following its acquisition by SpaceX.
