arXiv:2608.18104v1 Announce Type: new Abstract: Large language model (LLM)-based agents are increasingly becoming self-evolving systems that persist…
Tag: AI
Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges
arXiv:2608.18080v1 Announce Type: new Abstract: We present a review on the applications of large language models (LLMs) in health, e.g., social media…
Position: Profiling Game Worlds by Transition Complexity
arXiv:2608.18079v1 Announce Type: new Abstract: Game world modeling (GWM) and reinforcement learning (RL) are often confounded because research papers…
Etched’s valuation doubles to $21B in a month
Jane Street has installed Etched’s first shipped AI cluster system, and was so impressed, it led another massive round, the startup says.
Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models
arXiv:2608.18086v1 Announce Type: new Abstract: The growth of open-weight foundation models (OWFMs) has prompted the AI community to re-evaluate…
Improve contract search accuracy with auto-generated filters in Amazon Bedrock
In this post, we describe how AIDA works at a high level and how it helps address these challenges — grounding users in the right contracts, under the…
Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions
arXiv:2608.18078v1 Announce Type: new Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to…
How Jumio built a real-time feature store on AWS
Learn how Jumio built a centralized, real-time feature store on AWS with Amazon SageMaker Feature Store, Amazon Managed Service for Apache Flink, and…
Position: Behavioral Systems Require Behavioral Tests
arXiv:2608.18081v1 Announce Type: new Abstract: Artificial agentic systems increasingly operate as behavioral systems by interacting with dynamic…
SOD: Step-wise On-policy Distillation for Small Language Model Agents
arXiv:2605.07725v3 Announce Type: replace-cross Abstract: Tool-integrated reasoning (TIR) is difficult to scale to small language models due to…
