arXiv:2608.08736v1 Announce Type: new Abstract: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities…
Pathway Raises New Funding at $500M Valuation to Scale Post-Transformer AI
AI research company Pathway has secured additional funding at a $500 million valuation, bringing its total seed financing to $30 million as it prepares to…
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
arXiv:2608.08786v1 Announce Type: new Abstract: Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought…
Embattled hedge fund Situational Awareness invests $400M in chip startup Source Foundry
The AI-focused hedge fund is still making some big bets.
Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models
arXiv:2608.08746v1 Announce Type: new Abstract: Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal…
Anthropic is turning Claude Code’s auto mode on by default
Programming with Claude Code will soon require even less human oversight.
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling
arXiv:2608.08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as…
Historian Jill Lepore says Silicon Valley misreads science fiction and undermines democracy
On the latest episode of Equity, we spoke to Jill Lepore about “government by machines” and why Elon Musk is a bad science fiction reader.
AI Evaluation Should Measure Verification Cost, Not Correctness Alone
arXiv:2608.08709v1 Announce Type: new Abstract: The reliability of AI generative models is typically measured by output correctness, yet in practice it…
SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests
arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond…
IMDb Sentiment Analysis with DistilBERT LoRA, TF-IDF Baselines, Calibration, Interpretability, Robustness Testing, and Semi-Supervised Learning
This tutorial provides a comprehensive guide to building a robust sentiment analysis workflow. By combining classical TF-IDF baselines with modern…
The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task
arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on…
Sam Jenkins, Managing Partner at Punchcard Systems – Interview Series
Sam Jenkins, Managing Partner at Punchcard Systems, has spent more than two decades working at the intersection of technology, entrepreneurship, and…
EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility
arXiv:2608.08691v1 Announce Type: new Abstract: Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but…
The AI safety test is becoming a safety risk
AI agents are escaping cybersecurity testing environments and reaching real-world systems, raising questions about whether safety infrastructure, industry…
Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
arXiv:2608.08677v1 Announce Type: new Abstract: Skill evolution improves agent skills through feedback over time, with failed trajectories often providing…
AMIE, our research medical AI system, demonstrates real-time clinical video consultation capabilities in a first-of-its-kind study.
AMIE promotional video
A Structural Dynamics Graph World Model: Unified Modeling, Constrained Rollout, and Interpretable Calibration
arXiv:2608.08689v1 Announce Type: new Abstract: The state evolution of a complex system arises jointly from object laws, relational propagation, domain…
